🤖 Google Releases Compact Diarization Correction Model
DiarizationLM-Gemma-4-E4B-v1 is now available for diarization post-processing: a Gemma 4 E4B-based model (4B parameters — half the size of the previous 8B version) corrects speaker labels and utterance boundaries in ASR transcripts, while utterances of 6+ words are preserved as acoustic anchors against identity drift.
🌍 Previous DiarizationLM models mainly helped in two-speaker telephone calls, while the new model delivers significant gains for the first time on ICSI and AMI meetings with 4–9 speakers: WDER of 2.99% on Fisher versus 3.28% for the 8B model, and 4.92% versus 6.66% on Callhome.
👤 Try it locally: the GGUF Q4_K_M (~5.3 GB) runs via llama.cpp, there is a Python package diarizationlm and a web demo, and weights and code are under Apache-2.0. Before commercial deployment, check the terms of the base Gemma model.
Source 1: https://huggingface.co/google/DiarizationLM-Gemma-4-E4B-v1 Source 2: https://github.com/google/speaker-id
