Nemotron 3 Diarization β€” ONNX streaming step

An ONNX export of one streaming step of NVIDIA's Nemotron 3 Diarization (revision f667ed73aee57d40cc39428eb768b4fd87a0a29e), in fp32, for running the model outside NeMo β€” here, from Rust through ONNX Runtime with DirectML on Windows. It is the diarization model of the Tishina transcription app.

The graph is not a whole diarizer. It is forward_for_export: one chunk of mel features plus the speaker cache and FIFO go in, per-frame speaker probabilities and the chunk's embeddings come out. The mel frontend and the speaker-cache update (SortformerModules.streaming_update_async and its cache compression) run outside the graph and have to be reimplemented around it.

Files

file what it is
nemotron3_step.onnx one streaming step, fp32, opset 17, fixed shapes for the offline profile
mel_filterbank.f32 the checkpoint's mel filterbank, 128 Γ— 257, row-major little-endian f32
learnable_sil_emb.f32 the checkpoint's learned silence embedding, 512 little-endian f32
LICENSE the OpenMDW License Agreement, version 1.1

Why fp32, and how it was made

NVIDIA publishes the NeMo checkpoint (Nemotron-3-Diarization.nemo) in bfloat16, and the same weights in fp32 only as the transformers model.safetensors. The two name their tensors differently, so the NeMo state dict was rebuilt in fp32 by content rather than by name: each bf16 tensor was replaced by the one fp32 tensor that rounds to it bit for bit. The fused w_qkv projections were matched a third at a time; the Hann window and mel filterbank were recomputed from the checkpoint's config in fp32. Two blocks have no fp32 copy because inference never reads them β€” the frozen hidden_to_spks layer and the training-only activity_head β€” and keep their bf16 values; they are absent from the exported graph. Every one of the 417 fp32 tensors found its place, and the rebuilt checkpoint loads in NeMo with strict=True.

The step was then exported with torch.onnx.export. NeMo's own exporter does not handle the transformer encoder, and FlexAttention was replaced by the equivalent dense attention for tracing (outputs agree to 5e-6).

Tensors

Offline profile from the model card: speaker cache 264, FIFO 40, chunk 340, right context 40, update period 300 (input buffer 30.4 s).

input shape
chunk 1 Γ— 3040 Γ— 128 β€” log-mel frames for left context + chunk + right context
chunk_lengths 1
spkcache 1 Γ— 264 Γ— 512
spkcache_lengths 1
fifo 1 Γ— 40 Γ— 512
fifo_lengths 1
output shape
preds 1 Γ— 684 Γ— 8 β€” speaker probabilities per 80 ms frame over cache + FIFO + chunk
preds_hr 1 Γ— 5472 Γ— 8 β€” the same at 10 ms
chunk_embs 1 Γ— 380 Γ— 512 β€” the chunk's pre-encoded embeddings, for the cache update
chunk_emb_lengths 1
encoded_lengths 1

Short final chunks are zero-padded to 3040 frames with their true length in chunk_lengths.

Mel frontend

16 kHz mono input, 0.97 pre-emphasis, 512-point FFT, 25 ms Hann window, 10 ms hop, power spectrum through mel_filterbank.f32, natural log with a 2^-24 guard, no per-feature normalisation, frame count padded to a multiple of 16. The checkpoint is in streaming mode, so NeMo does not scale the recording by its peak before feature extraction; doing so shifts every feature.

Agreement with NeMo

Measured against NeMo running the same fp32 checkpoint in async mode with async_pad_to_max, which is the computation a fixed-shape graph reproduces:

  • per step, on captured real inputs: speaker probabilities within 2.7e-4;
  • whole recordings through a Rust runtime built on this graph: seven of eight synthetic multi-speaker files frame-identical, eight AMI meetings within 0.004 DER of NeMo, a 45-minute Russian meeting within 0.006.

An fp16 export was measured as well and is not published here: on AMI it scored the same against the reference labels, but on one long meeting it moved 35 seconds of speech between speakers relative to fp32, with no reference to say which was right.

License and attribution

Nemotron 3 Diarization is by NVIDIA Corporation and is released under the OpenMDW License Agreement, version 1.1; its text is in LICENSE, and it applies to these files as a conversion of that model. Any redistribution has to keep the agreement and these notices of origin. See the original model card for the model's training data, evaluation, intended use and limitations β€” all of which apply here unchanged.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for beshkenadze/nemotron-3-diarization-onnx

Quantized
(26)
this model