Nemotron-3 Diarization (GGUF)
GGUF conversions of NVIDIA's Nemotron-3 Diarization (Streaming Sortformer v3) for transcribe.cpp, the runtime behind speaker labels in Glimpse. The model answers "who spoke when" for up to 8 speakers at 10 ms resolution. It produces no text.
| File | Size | Notes |
|---|---|---|
nemotron-3-diarization-Q8_0.gguf |
106 MB | Recommended. Same accuracy as F32 on AMI |
nemotron-3-diarization-F16.gguf |
199 MB | |
nemotron-3-diarization-F32.gguf |
397 MB | Reference precision |
Results
AMI meeting corpus, IHM test set: 16 meetings, 9.06 hours. Scored against NVIDIA's forced-alignment RTTMs with a 0 s collar, overlapped speech included, and a plain 0.5 threshold at 10 ms with no post-processing. Every row went through the same pipeline.
| Model | Lookahead | DER | JER | Missed | False alarm | Speaker confusion |
|---|---|---|---|---|---|---|
| NVIDIA reference (HF Transformers, fp32) | 30.4 s | 9.22% | 12.89% | 4.68% | 3.67% | 0.88% |
| This port, F32 | 30.4 s | 9.23% | 12.84% | 4.68% | 3.67% | 0.88% |
| This port, F16 | 30.4 s | 9.23% | 12.84% | 4.68% | 3.67% | 0.88% |
| This port, Q8_0 | 30.4 s | 9.23% | 12.85% | 4.79% | 3.58% | 0.86% |
| This port, Q8_0 | 1.04 s | 9.48% | 13.02% | 5.02% | 3.51% | 0.95% |
| This port, Q8_0 | 0.64 s | 9.64% | 13.16% | 5.13% | 3.51% | 1.01% |
| Sortformer v2.1 Q8_0 (previous diarizer) | 30.4 s | 15.99% | 21.20% | 7.42% | 4.97% | 3.60% |
Compared with Sortformer v2.1, errors drop by 42% and speaker confusion by 76%, and the 1-second streaming mode beats v2.1's offline mode.
Throughput on an Apple M2 Pro (offline mode, times faster than real time):
| Model | Metal, 10 min meeting | CPU, 10 min meeting |
|---|---|---|
| Nemotron-3 Q8_0 | 339x | 59x |
| Nemotron-3 F16 | 335x | 48x |
| Sortformer v2.1 Q8_0 | 105x | 44x |
Fidelity to the reference
On a 12-second two-speaker clip, every stage (mel, embedder, encoder layers, projection, upsampler, speaker probabilities) matches the Transformers reference within 2e-4. The final probabilities match within 6e-6 for every latency preset, including runs that force speaker-cache compression. Over the 9 AMI hours, the Metal build and the reference disagree on 0.03% of 10 ms decisions. These come from floating-point order flipping a near-tie in the discrete cache compression, and DER is unchanged.
Usage
transcribe-cli -m nemotron-3-diarization-Q8_0.gguf meeting.wav
Speaker turns are read through transcribe_n_speaker_segments /
transcribe_get_speaker_segment. Speaker ids are 1-based, in order of first appearance.
The latency preset comes from the Sortformer run extension (transcribe/sortformer.h):
| Preset | Lookahead |
|---|---|
DEFAULT / VERY_HIGH_LATENCY |
30.4 s |
LOW_LATENCY |
1.04 s |
VERY_LOW_LATENCY |
0.64 s |
ULTRA_LOW_LATENCY |
0.32 s |
In Rust, through Glimpse-Speech:
let turns = glimpse_speech::diarization::diarize(&model_path, &samples, sample_rate, true)?;
Input is 16 kHz mono; other rates are resampled.
Provenance
Converted from nvidia/Nemotron-3-Diarization at revision
a435e9867d79e789e90053f9b6d6834053af564a (model.safetensors, fp32) with
scripts/convert-nemotron3_diar.py, then quantized with transcribe-quantize. The learned
speaker-cache silence embedding stays F32 in every file.
| File | SHA-256 |
|---|---|
nemotron-3-diarization-Q8_0.gguf |
877ff9e77e829e30158528cfaabf56188fca11d05349e09821b3033a97d31688 |
nemotron-3-diarization-F16.gguf |
5513da21cc39fc3ab5a36bd945324aeb013369b63172849b4ef174686e15f27c |
nemotron-3-diarization-F32.gguf |
46ce1267731375c7c366b67b5ce36c7a41bca45ab43cbedd98a0debffc4b85eb |
Glimpse
This model labels speakers in Glimpse, a free, open-source dictation app for Mac and Windows, entirely on your device. The source is on GitHub.
License
The model and its weights are NVIDIA's, released under the OpenMDW License Agreement 1.1, which permits commercial use. This repository only converts the original checkpoint. It is not an official NVIDIA release. See NVIDIA's model card for training data, intended use, and bias, safety and privacy notes.
- Downloads last month
- 142
8-bit
16-bit
32-bit
Model tree for Glimpse-Dictation/Nemotron-3-Diarization-gguf
Base model
nvidia/Nemotron-3-Diarization