Nemotron-3-Diarization GGUF
Speaker diarization model converted from nvidia/Nemotron-3-Diarization to GGUF format for parakeet.cpp.
Model description
Nemotron-3-Diarization is a 31-layer pre-LN RoPE Transformer encoder with a subpixel Conv1D upsampling speaker head. It takes mel spectrograms as input and outputs per-frame speaker activity probabilities for up to 8 speakers.
- Encoder: 31-layer pre-LN RoPE Transformer (d_model=512, 8 heads, head_dim=64, ff_dim=2048)
- Subsampling: FeatureStacking 8x
- Speaker head: encoder_proj (512->192) + subpixel Conv1D 8x upsampling + sigmoid (192->8)
- Input: 128-bin mel spectrograms, 16kHz, 10ms frame resolution
- Output: 8 speaker tracks at 10ms resolution (upsampled from encoder rate)
Parity vs PyTorch reference
| Metric | Value |
|---|---|
| probs max_diff | 0.0097 (F16 noise) |
| probs mean_diff | 0.000118 |
| segments (F16) | 583 (ours) vs 581 (reference) |
| segments (Q8_0) | 547 |
Available files
| File | Size | Notes |
|---|---|---|
nemotron-3-diarization-f16.gguf |
192 MB | F16 linear weights, F32 biases/norms |
nemotron-3-diarization-q8_0.gguf |
104 MB | Q8_0 quantized linear weights |
Only linear weight tensors are quantized. All biases, layer norms, embeddings, and the mel filterbank/window are kept in F32 for numerical stability.
Usage
Standalone diarization
# Build parakeet.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
# Run diarization on an audio file
./build/examples/cli/diarize nemotron-3-diarization-f16.gguf audio.wav
Output is JSON:
{
"speakers": 8,
"segments": [
{"speaker": 0, "start": 0.07, "end": 0.08},
{"speaker": 1, "start": 30.02, "end": 30.03},
...
]
}
Speaker-attributed ASR (SAS)
Combine with an ASR model (e.g. parakeet-tdt-0.6b-v3) to get speaker-attributed transcripts:
#include "parakeet_capi.h"
parakeet_ctx* asr = parakeet_capi_load("parakeet-tdt-0.6b-v3.q8_0.gguf");
parakeet_ctx* diar = parakeet_capi_load("nemotron-3-diarization-f16.gguf");
int n = 0;
parakeet_sas_result* results = parakeet_capi_transcribe_and_diarize(
asr, diar, samples, n_samples, sample_rate, &n);
for (int i = 0; i < n; i++) {
printf("[speaker %d] %s (%.2f-%.2f)\n",
results[i].speaker, results[i].text,
results[i].start, results[i].end);
parakeet_capi_free_string(results[i].text);
}
parakeet_capi_free_sas_results(results);
Or get JSON output:
char* json = parakeet_capi_transcribe_and_diarize_json(
asr, diar, samples, n_samples, sample_rate);
// {"speakers":8,"utterances":[...],"words":[...]}
parakeet_capi_free_string(json);
Conversion
To convert from the original .nemo checkpoint:
python3 scripts/convert_parakeet_to_gguf.py \
--model nvidia/Nemotron-3-Diarization \
--output nemotron-3-diarization-f16.gguf \
--dtype f16
Limitations
- Offline only: Streaming diarization (AOSC speaker cache) is not yet implemented. The model processes the full audio at once.
- No mel sharing with ASR: The diarization model uses 128 mel bins while most ASR models use 80, so mel spectrograms are computed separately for each model in the SAS path.
- Q8_0 accuracy: The Q8_0 variant produces ~6% fewer segments than F16 due to quantization noise in the sigmoid speaker head. For maximum accuracy, use F16.
License
NVIDIA Open Model License. See license link.
- Downloads last month
- 41
Hardware compatibility
Log In to add your hardware
8-bit
16-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for mudler/Nemotron-3-Diarization-GGUF
Base model
nvidia/Nemotron-3-Diarization