Nemotron 3 Diarization is NVIDIA’s open-weight speaker-diarization model for separating up to eight voices in live or recorded audio. Released on September 23, 2026, the approximately 100-million-parameter model handles overlapping speech and can feed speaker timestamps into an automatic speech recognition system, producing labeled transcripts for meetings, calls, interviews, and podcasts.
What does Nemotron 3 Diarization actually do?
Nemotron 3 Diarization determines when each person speaks and assigns anonymous labels to as many as eight speakers, including during overlapping speech. The 2026 model processes live streams or completed recordings, returning timestamps at configurable 10-millisecond intervals rather than identifying participants by name or transcribing their words.
Speaker diarization is an audio-processing task that divides speech by speaker and answers the question “who spoke when?” NVIDIA’s model names channels according to the order in which voices first appear. The first detected participant becomes the first channel, but the system doesn’t know that person’s identity.
Pairing the diarizer with automatic speech recognition, or ASR, turns those time ranges into a speaker-attributed transcript. A meeting application can therefore associate recognized words with Speaker 1 or Speaker 2, while a call-analysis system can separate an agent’s speech from a customer’s.
The distinction matters. Diarization error rate measures incorrect speaker segmentation, missed speech, false speech detections, and speaker confusion; it isn’t a word-accuracy score. Audio quality also remains consequential, so developers preparing compressed recordings should understand how bitrate affects usable audio quality before blaming the model for every mistake.
How accurate is Nemotron 3 Diarization?
Nemotron 3 Diarization ranked first among 12 systems in VoiceArena’s initial 2026 Diarization-Bench, recording a 14.72% diarization error rate. The runner-up, DiariZen WavLM Large s80 MD v2, scored 19.34% on the same 139 English conversations and roughly 22 hours of evaluated audio.
| System or metric | 2026 result | What it indicates |
|---|---|---|
| Nemotron 3 Diarization DER | 14.72% | Total diarization error |
| DiariZen WavLM Large s80 MD v2 DER | 19.34% | Second-place system’s total error |
| Nemotron JER | 25.53% | Per-speaker overlap error |
| Exact speaker-count accuracy | 83.45% | Recordings with the correct speaker count |
| Missed speech | 3.63% | Speech not detected |
| False alarms | 8.88% | Nonspeech marked as speech |
| Speaker confusion | 2.21% | Speech assigned to the wrong speaker |
The headline comparison amounts to a 4.62-percentage-point reduction in DER. Calculated against DiariZen’s 19.34% result, NVIDIA’s 14.72% score is about 23.9% lower. That’s a meaningful lead, although one benchmark shouldn’t be treated as a universal ranking.
VoiceArena included overlapping speech, system-generated speech-activity detection, and a zero-millisecond scoring collar in its 2026 test. Its published 95% confidence interval for the NVIDIA result was approximately 13.4% to 16.1%, showing the uncertainty around a score drawn from a finite test set.
English conversations dominated the initial benchmark, so the result doesn’t establish equivalent performance for every language, room, microphone, or speaking style. Argmax separately reported in 2026 that the model achieved the lowest average DER in its six-system OpenBench comparison across 11 datasets and led five datasets individually. Useful corroboration, yes. A guarantee for your recordings, no.
How much latency does live speaker tracking add?
Nemotron 3 Diarization accepts buffers ranging from an experimental 80 milliseconds to 30.4 seconds, according to NVIDIA’s 2026 model card. NVIDIA recommends at least 0.32 seconds for streaming, but shorter buffers return decisions sooner at the cost of higher diarization error on the company’s NOTSOFAR1-MHM evaluation.
NVIDIA reported 6.77% DER with a 30.4-second buffer in 2026. Error rose to 7.70% at 1.04 seconds and 8.65% at 0.32 seconds. Moving from the longest tested setting to the recommended low-latency setting therefore increased DER by 1.88 percentage points, or about 27.8% relative to the 6.77% baseline.
You should choose the buffer around the product experience, not the smallest number on the configuration screen. Live captions and agent-assistance software may justify a 0.32-second buffer, whereas an uploaded interview can tolerate a 30.4-second window in exchange for better segmentation.
The input must be mono audio sampled at 16 kHz. Chunked inference removes a fixed recording-duration limit, and the output resolution can be configured in 10-millisecond increments. Low output granularity doesn’t erase the buffer trade-off: timestamps may be precise while the system still needs more surrounding audio to decide who spoke.
Where does the model fit in a practical speech stack?
Nemotron 3 Diarization fits between audio ingestion and transcript assembly in a 2026 speech stack. The model separates speakers first or alongside ASR, after which an application aligns recognized words with speaker channels for meeting notes, call analytics, interviews, podcast editing, or private transcription.
NVIDIA’s NeMo-Speech.cpp runtime supports the diarizer by itself or combined with ASR. The project added day-zero support in 2026 for live microphone transcription, local GGUF execution, an eight-speaker preset, 10-millisecond output frames, and optional Q8_0 conversion.
A sensible deployment review covers the following practical checks:
- Convert incoming audio to mono 16-kHz input and test the effect of the source codec.
- Select a buffer using measured latency and DER on your own conversations.
- Combine diarization with ASR when the product requires words as well as speaker timing.
- Map anonymous channels to known names only through a separate, consent-aware workflow.
- Test meetings with interruptions, silence, background audio, and more than eight participants.
Argmax Pro SDK 3 added real-time support and speaker-attributed transcription for up to eight participants in 2026. Argmax also introduced a “pre-diarized” transcription API that separates speakers before recognition, claiming improved transcription for complicated overlapping conversations.
For sensitive recordings, local execution can reduce the need to send raw voices to an external service. Enterprises still need governance around binaries, model weights, transcript retention, and downstream agents; a review of shadow AI discovery platforms can help identify unmanaged speech tools, while agent-related enterprise security risks remain relevant when transcripts trigger automated actions.
What are the model’s biggest limitations?
Nemotron 3 Diarization cannot identify people by name, transcribe speech on its own, or reliably represent more than eight simultaneous speaker channels. NVIDIA’s 2026 model documentation warns that audio with more than eight speakers may lead to missed speech or assignments to an incorrect channel.
The eight-speaker ceiling is easy to overlook because many ordinary calls stay below it. A nine-person panel, rotating classroom discussion, or large hybrid meeting is an edge case with real consequences: the model doesn’t simply create a ninth label. Speech can disappear from the output or merge into an existing channel.
Anonymous labels also require careful product copy. Speaker 1 isn’t a verified identity, and channel order can change across separate recordings because labels follow first appearance. If you need names, you’ll need participant metadata, manual assignment, or a separate voice-identification system with appropriate consent and legal review.
NVIDIA describes the 2026 architecture as a 31-layer Transformer encoder using rotary position embeddings, 80-millisecond encoder frames, and a Conv1D layer that restores predictions to 10-millisecond resolution. NVIDIA also says training used roughly 10,000 hours of real conversations plus simulated mixtures with one to eight speakers. Those details help explain the compact model’s scope, but your production audio remains the decisive test.
The OpenMDW 1.1 license permits commercial and non-commercial use, according to the 2026 model card. Honestly, the release makes the most sense when you value local control, streaming support, or inspectable weights; a hosted transcription API may remain simpler if your team doesn’t want to operate an audio pipeline.
Nemotron 3 Diarization FAQ
Is Nemotron 3 Diarization open source?
Nemotron 3 Diarization is an open-weight model released under the OpenMDW 1.1 license in 2026. NVIDIA describes the weights as available for commercial and non-commercial use, but teams should review the license terms for their intended distribution and deployment.
Does Nemotron 3 Diarization transcribe audio?
Nemotron 3 Diarization doesn’t convert speech into words. The model supplies anonymous speaker labels and timestamps, which can be paired with an ASR model to create a speaker-attributed transcript.
Can Nemotron 3 Diarization run locally?
Nemotron 3 Diarization can run locally through NVIDIA’s NeMo-Speech.cpp runtime, which added GGUF execution and live microphone support in 2026. Community Core ML and MLX INT8 ports for Apple Silicon appeared on September 24, 2026, but those conversions weren’t official NVIDIA releases.
What audio format does Nemotron 3 Diarization require?
Nemotron 3 Diarization accepts mono audio sampled at 16 kHz, according to NVIDIA’s 2026 model card. Long recordings can be processed through chunked inference rather than a fixed maximum recording length.


