Nemotron 3 Diarization β ONNX streaming step
An ONNX export of one streaming step of NVIDIA's
Nemotron 3 Diarization
(revision f667ed73aee57d40cc39428eb768b4fd87a0a29e), in fp32, for running the
model outside NeMo β here, from Rust through ONNX Runtime with DirectML on
Windows. It is the diarization model of the Tishina transcription app.
The graph is not a whole diarizer. It is forward_for_export: one chunk of mel
features plus the speaker cache and FIFO go in, per-frame speaker
probabilities and the chunk's embeddings come out. The mel frontend and the
speaker-cache update (SortformerModules.streaming_update_async and its cache
compression) run outside the graph and have to be reimplemented around it.
Files
| file | what it is |
|---|---|
nemotron3_step.onnx |
one streaming step, fp32, opset 17, fixed shapes for the offline profile |
mel_filterbank.f32 |
the checkpoint's mel filterbank, 128 Γ 257, row-major little-endian f32 |
learnable_sil_emb.f32 |
the checkpoint's learned silence embedding, 512 little-endian f32 |
LICENSE |
the OpenMDW License Agreement, version 1.1 |
Why fp32, and how it was made
NVIDIA publishes the NeMo checkpoint (Nemotron-3-Diarization.nemo) in
bfloat16, and the same weights in fp32 only as the transformers
model.safetensors. The two name their tensors differently, so the NeMo state
dict was rebuilt in fp32 by content rather than by name: each bf16 tensor was
replaced by the one fp32 tensor that rounds to it bit for bit. The fused
w_qkv projections were matched a third at a time; the Hann window and mel
filterbank were recomputed from the checkpoint's config in fp32. Two blocks
have no fp32 copy because inference never reads them β the frozen
hidden_to_spks layer and the training-only activity_head β and keep their
bf16 values; they are absent from the exported graph. Every one of the 417
fp32 tensors found its place, and the rebuilt checkpoint loads in NeMo with
strict=True.
The step was then exported with torch.onnx.export. NeMo's own exporter does
not handle the transformer encoder, and FlexAttention was replaced by the
equivalent dense attention for tracing (outputs agree to 5e-6).
Tensors
Offline profile from the model card: speaker cache 264, FIFO 40, chunk 340, right context 40, update period 300 (input buffer 30.4 s).
| input | shape |
|---|---|
chunk |
1 Γ 3040 Γ 128 β log-mel frames for left context + chunk + right context |
chunk_lengths |
1 |
spkcache |
1 Γ 264 Γ 512 |
spkcache_lengths |
1 |
fifo |
1 Γ 40 Γ 512 |
fifo_lengths |
1 |
| output | shape |
|---|---|
preds |
1 Γ 684 Γ 8 β speaker probabilities per 80 ms frame over cache + FIFO + chunk |
preds_hr |
1 Γ 5472 Γ 8 β the same at 10 ms |
chunk_embs |
1 Γ 380 Γ 512 β the chunk's pre-encoded embeddings, for the cache update |
chunk_emb_lengths |
1 |
encoded_lengths |
1 |
Short final chunks are zero-padded to 3040 frames with their true length in
chunk_lengths.
Mel frontend
16 kHz mono input, 0.97 pre-emphasis, 512-point FFT, 25 ms Hann window, 10 ms
hop, power spectrum through mel_filterbank.f32, natural log with a 2^-24
guard, no per-feature normalisation, frame count padded to a multiple of 16.
The checkpoint is in streaming mode, so NeMo does not scale the recording by
its peak before feature extraction; doing so shifts every feature.
Agreement with NeMo
Measured against NeMo running the same fp32 checkpoint in async mode with
async_pad_to_max, which is the computation a fixed-shape graph reproduces:
- per step, on captured real inputs: speaker probabilities within 2.7e-4;
- whole recordings through a Rust runtime built on this graph: seven of eight synthetic multi-speaker files frame-identical, eight AMI meetings within 0.004 DER of NeMo, a 45-minute Russian meeting within 0.006.
An fp16 export was measured as well and is not published here: on AMI it scored the same against the reference labels, but on one long meeting it moved 35 seconds of speech between speakers relative to fp32, with no reference to say which was right.
License and attribution
Nemotron 3 Diarization is by NVIDIA Corporation and is released under the
OpenMDW License Agreement, version 1.1; its
text is in LICENSE, and it applies to these files as a conversion of that
model. Any redistribution has to keep the agreement and these notices of
origin. See the original
model card for the
model's training data, evaluation, intended use and limitations β all of which
apply here unchanged.
Model tree for beshkenadze/nemotron-3-diarization-onnx
Base model
nvidia/Nemotron-3-Diarization