Nemotron-3 Diarization (GGUF)

GGUF conversions of NVIDIA's Nemotron-3 Diarization (Streaming Sortformer v3) for transcribe.cpp, the runtime behind speaker labels in Glimpse. The model answers "who spoke when" for up to 8 speakers at 10 ms resolution. It produces no text.

File Size Notes
nemotron-3-diarization-Q8_0.gguf 106 MB Recommended. Same accuracy as F32 on AMI
nemotron-3-diarization-F16.gguf 199 MB
nemotron-3-diarization-F32.gguf 397 MB Reference precision

Results

AMI meeting corpus, IHM test set: 16 meetings, 9.06 hours. Scored against NVIDIA's forced-alignment RTTMs with a 0 s collar, overlapped speech included, and a plain 0.5 threshold at 10 ms with no post-processing. Every row went through the same pipeline.

Model Lookahead DER JER Missed False alarm Speaker confusion
NVIDIA reference (HF Transformers, fp32) 30.4 s 9.22% 12.89% 4.68% 3.67% 0.88%
This port, F32 30.4 s 9.23% 12.84% 4.68% 3.67% 0.88%
This port, F16 30.4 s 9.23% 12.84% 4.68% 3.67% 0.88%
This port, Q8_0 30.4 s 9.23% 12.85% 4.79% 3.58% 0.86%
This port, Q8_0 1.04 s 9.48% 13.02% 5.02% 3.51% 0.95%
This port, Q8_0 0.64 s 9.64% 13.16% 5.13% 3.51% 1.01%
Sortformer v2.1 Q8_0 (previous diarizer) 30.4 s 15.99% 21.20% 7.42% 4.97% 3.60%

Compared with Sortformer v2.1, errors drop by 42% and speaker confusion by 76%, and the 1-second streaming mode beats v2.1's offline mode.

Throughput on an Apple M2 Pro (offline mode, times faster than real time):

Model Metal, 10 min meeting CPU, 10 min meeting
Nemotron-3 Q8_0 339x 59x
Nemotron-3 F16 335x 48x
Sortformer v2.1 Q8_0 105x 44x

Fidelity to the reference

On a 12-second two-speaker clip, every stage (mel, embedder, encoder layers, projection, upsampler, speaker probabilities) matches the Transformers reference within 2e-4. The final probabilities match within 6e-6 for every latency preset, including runs that force speaker-cache compression. Over the 9 AMI hours, the Metal build and the reference disagree on 0.03% of 10 ms decisions. These come from floating-point order flipping a near-tie in the discrete cache compression, and DER is unchanged.

Usage

transcribe-cli -m nemotron-3-diarization-Q8_0.gguf meeting.wav

Speaker turns are read through transcribe_n_speaker_segments / transcribe_get_speaker_segment. Speaker ids are 1-based, in order of first appearance. The latency preset comes from the Sortformer run extension (transcribe/sortformer.h):

Preset Lookahead
DEFAULT / VERY_HIGH_LATENCY 30.4 s
LOW_LATENCY 1.04 s
VERY_LOW_LATENCY 0.64 s
ULTRA_LOW_LATENCY 0.32 s

In Rust, through Glimpse-Speech:

let turns = glimpse_speech::diarization::diarize(&model_path, &samples, sample_rate, true)?;

Input is 16 kHz mono; other rates are resampled.

Provenance

Converted from nvidia/Nemotron-3-Diarization at revision a435e9867d79e789e90053f9b6d6834053af564a (model.safetensors, fp32) with scripts/convert-nemotron3_diar.py, then quantized with transcribe-quantize. The learned speaker-cache silence embedding stays F32 in every file.

File SHA-256
nemotron-3-diarization-Q8_0.gguf 877ff9e77e829e30158528cfaabf56188fca11d05349e09821b3033a97d31688
nemotron-3-diarization-F16.gguf 5513da21cc39fc3ab5a36bd945324aeb013369b63172849b4ef174686e15f27c
nemotron-3-diarization-F32.gguf 46ce1267731375c7c366b67b5ce36c7a41bca45ab43cbedd98a0debffc4b85eb

Glimpse

This model labels speakers in Glimpse, a free, open-source dictation app for Mac and Windows, entirely on your device. The source is on GitHub.

License

The model and its weights are NVIDIA's, released under the OpenMDW License Agreement 1.1, which permits commercial use. This repository only converts the original checkpoint. It is not an official NVIDIA release. See NVIDIA's model card for training data, intended use, and bias, safety and privacy notes.

Downloads last month
142
GGUF
Model size
99.2M params
Architecture
nemotron3_diar
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Glimpse-Dictation/Nemotron-3-Diarization-gguf

Quantized
(22)
this model