Google (Speaker, Voice and Language team) has released the DiarizationLM-Gemma-4-E4B-v1 model for diarization post-processing: it corrects speaker labels and utterance boundaries in completed ASR transcripts. The model is built on Gemma 4 E4B (4B parameters, half the size of the previous DiarizationLM-8b-Fisher-v2) and for the first time delivers statistically significant improvements not only on the Fisher and Callhome telephone dialogues, but also on 4–9-speaker ICSI and AMI meetings. Weights and code are open under Apache-2.0, and the quantized version weighs about 5.3 GB, so the "ASR + diarization + LLM correction" stack can be assembled even on a single desktop.

image
image

What happened

The model was fine-tuned using Locality-Preserving Oracle Supervision on 51,063 Fisher pairs and 20,762 Callhome/ICSI/AMI pairs with LoRA r=256 for 10,000 steps on 8× TPU v5p. The key recipe trick: edits are only allowed on short 1–5-word "backchannel" insertions, while utterances of 6+ words are preserved as acoustic anchors so that speaker identity doesn't drift in long monologues. This is a deliberate trade-off between boundary accuracy and attribution stability. On all four benchmarks, WDER and cpWER improvements are statistically significant (p<0.0001): on Fisher, WDER drops to 2.99% versus 3.28% for the 8B version (cpWER 17.62% versus 18.37%), on Callhome — to 4.92% versus 6.66%, with perfectly accurate speaker count (MAE=0). Weights are distributed in GGUF Q4_K_M (~5.3 GB) and safetensors (~16 GB) formats under the Apache-2.0 license.

Context

Post-processing diarization with a language model is not a new idea: previous DiarizationLM models from the same team already corrected transcription hypotheses with tags. The bottleneck was coverage: they were trained and evaluated primarily on two-speaker telephone corpora Fisher and Callhome, whereas a typical office meeting gathers four to nine participants, where utterances are interrupted by short insertions like "uh-huh" and "mm-hmm," and boundary errors multiply with each additional voice. Two metrics help understand the scale of improvements: WDER measures the share of speaker attribution errors, while cpWER — the overall transcription quality accounting for boundary shifts, i.e., showing the full effect of correction.

Why this matters for the industry

Mixed speaker attributions remain the main residual source of errors in multi-speaker meeting transcription, and LLM correction of this stage until recently yielded significant gains only on two-speaker phone calls. The new release shows significant improvement for the first time on 4–9-speaker ICSI and AMI meetings, and with a model half the size, meaning the approach scales to the typical office scenario without losing quality on phone conversations. The method itself is telling: quality improved thanks to a LoRA fine-tuning recipe and locality supervision, not scale — a rare case where a compact 4B model outperforms a larger predecessor. For companies, this turns the corrector into an addable batch stage of an existing pipeline: weights are open, the input format is simple, no ASR migration is needed, and the barrier to entry for accurate speaker attribution in self-hosted stacks drops sharply.

Why this matters for users

You can try the model today without a GPU farm: the quantized GGUF Q4_K_M build weighs about 5.3 GB and runs in llama.cpp-compatible runtimes, while the full weights in safetensors (~16 GB) work via Python — pip install diarizationlm — or through the transformers library. The exchange format is simple: a transcription hypothesis with tags is fed as input, and corrected label and boundary markup is produced as output. There is also a live HF Space demo DiarizationLM-GGUF where the pipeline can be tested before installation. A practical pilot plan is a low-risk A/B experiment: feed the current hypothesis of your diarization pipeline, get the corrected markup, and compare WDER on your own labeled sample against the baseline version without correction.

What is still unknown / limitations

Results were obtained on research corpora (telephone dialogues and meetings) with fixed upstream recognition quality, so transferability to other domains, languages, and live ASR engines has not yet been shown. In the available data, comparison is only provided with the previous model from the same team, and the position relative to other diarization approaches remains unclear. Moreover, this is a layer on top of a completed transcript: the model corrects speaker labels and utterance boundaries, but the quality of the words themselves is determined by upstream ASR, so claims of production readiness on arbitrary data should be verified on your own labeled data.

Sources

Author

Look at AI, editorial team