🩺 One in Three Notes from an AI Scribe Has a Confirmed Error
Researchers checked notes from three commercial ambient AI scribes—systems that automatically draft clinical records—across a set of 142 consultations. Errors were confirmed in 31.3% of cases: most often in allergy and medication data, in fabricated patient identification, and in an 'examination' that is impossible during a phone consultation. Adversarial verification yielded 618 findings.
🌍 The audit of deployed medical AI scribes strikes at vendors' argument that 'the doctor will check it anyway': the defect reached the doctor's signature, and omissions of allergies and medications lead to clinical incidents.
👤 The OmissionBench dataset on Hugging Face and the MIT-licensed code are open—the measurement can be replicated. Conclusion: auto-generated medical records should be read selectively, not signed off on trust.
Source 1: https://arxiv.org/abs/2608.31017
Source 2: https://huggingface.co/datasets/ComposoAI/OmissionBench
