Nemotron-3-Diarization GGUF

Speaker diarization model converted from nvidia/Nemotron-3-Diarization to GGUF format for parakeet.cpp.

Model description

Nemotron-3-Diarization is a 31-layer pre-LN RoPE Transformer encoder with a subpixel Conv1D upsampling speaker head. It takes mel spectrograms as input and outputs per-frame speaker activity probabilities for up to 8 speakers.

  • Encoder: 31-layer pre-LN RoPE Transformer (d_model=512, 8 heads, head_dim=64, ff_dim=2048)
  • Subsampling: FeatureStacking 8x
  • Speaker head: encoder_proj (512->192) + subpixel Conv1D 8x upsampling + sigmoid (192->8)
  • Input: 128-bin mel spectrograms, 16kHz, 10ms frame resolution
  • Output: 8 speaker tracks at 10ms resolution (upsampled from encoder rate)

Parity vs PyTorch reference

Metric Value
probs max_diff 0.0097 (F16 noise)
probs mean_diff 0.000118
segments (F16) 583 (ours) vs 581 (reference)
segments (Q8_0) 547

Available files

File Size Notes
nemotron-3-diarization-f16.gguf 192 MB F16 linear weights, F32 biases/norms
nemotron-3-diarization-q8_0.gguf 104 MB Q8_0 quantized linear weights

Only linear weight tensors are quantized. All biases, layer norms, embeddings, and the mel filterbank/window are kept in F32 for numerical stability.

Usage

Standalone diarization

# Build parakeet.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

# Run diarization on an audio file
./build/examples/cli/diarize nemotron-3-diarization-f16.gguf audio.wav

Output is JSON:

{
  "speakers": 8,
  "segments": [
    {"speaker": 0, "start": 0.07, "end": 0.08},
    {"speaker": 1, "start": 30.02, "end": 30.03},
    ...
  ]
}

Speaker-attributed ASR (SAS)

Combine with an ASR model (e.g. parakeet-tdt-0.6b-v3) to get speaker-attributed transcripts:

#include "parakeet_capi.h"

parakeet_ctx* asr  = parakeet_capi_load("parakeet-tdt-0.6b-v3.q8_0.gguf");
parakeet_ctx* diar = parakeet_capi_load("nemotron-3-diarization-f16.gguf");

int n = 0;
parakeet_sas_result* results = parakeet_capi_transcribe_and_diarize(
    asr, diar, samples, n_samples, sample_rate, &n);

for (int i = 0; i < n; i++) {
    printf("[speaker %d] %s (%.2f-%.2f)\n",
           results[i].speaker, results[i].text,
           results[i].start, results[i].end);
    parakeet_capi_free_string(results[i].text);
}
parakeet_capi_free_sas_results(results);

Or get JSON output:

char* json = parakeet_capi_transcribe_and_diarize_json(
    asr, diar, samples, n_samples, sample_rate);
// {"speakers":8,"utterances":[...],"words":[...]}
parakeet_capi_free_string(json);

Conversion

To convert from the original .nemo checkpoint:

python3 scripts/convert_parakeet_to_gguf.py \
  --model nvidia/Nemotron-3-Diarization \
  --output nemotron-3-diarization-f16.gguf \
  --dtype f16

Limitations

  • Offline only: Streaming diarization (AOSC speaker cache) is not yet implemented. The model processes the full audio at once.
  • No mel sharing with ASR: The diarization model uses 128 mel bins while most ASR models use 80, so mel spectrograms are computed separately for each model in the SAS path.
  • Q8_0 accuracy: The Q8_0 variant produces ~6% fewer segments than F16 due to quantization noise in the sigmoid speaker head. For maximum accuracy, use F16.

License

NVIDIA Open Model License. See license link.

Downloads last month
41
GGUF
Model size
99.3M params
Architecture
parakeet
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mudler/Nemotron-3-Diarization-GGUF

Quantized
(22)
this model