Google Research and Google DeepMind have introduced AMIE (Video) — a medical AI system for real-time video consultations based on Gemini and Project Astra with a three-agent architecture that showed results on par with certified primary care physicians in a randomized study.


What happened
The AMIE (Articulate Medical Intelligence Explorer) system uses a three-agent asynchronous architecture: the Talker Agent conducts low-latency dialogue, the Planner Agent continuously updates the differential diagnosis, and the Perception Agent analyzes audiovisual signals — signs of distress, physical symptoms, and sound indicators. In a randomized OSCE study with 100 clinical scenarios across five organ systems, 15 actor-patients, and 300 consultations, AMIE (Video) achieved results comparable to the assessments of 10 certified primary care physicians. Actor-patients preferred the video mode to text chat. The study is published on arXiv under number 2608.09861, 40 authors, submitted on August 10, 2026.
Context
The system is built on the basis of Gemini and Project Astra — Google's key multimodal platforms. Before AMIE (Video), medical AI tools worked primarily in text format: chatbots received symptom descriptions and provided recommendations, but could not assess the patient's non-verbal signals. The three-agent architecture solves the fundamental conflict between response speed and depth of analysis, isolating the latency-critical dialogue stream from heavy clinical reasoning. AMIE is not Google's first medical project: previous versions of AMIE in text mode already demonstrated diagnostic accuracy, and AMIE (Video) represents an evolutionary step towards full-fledged video consultations.
Why this matters for the industry
The three-agent asynchronous architecture is becoming a reference pattern not only for medical AI, but also for any agent systems where simultaneous low dialogue latency and complex multimodal reasoning are required. Engineers can adapt the pattern — a dialogue agent plus a background planner plus a perception module — to their own products based on the Gemini API. Signal for the market: telemedicine through multimodal AI is ceasing to be an experimental area, which may trigger a reassessment of projects working only in text mode. The paper with a detailed description of the architecture is available on arXiv, allowing competing laboratories to begin replication right now.
Why this matters for users
AMIE (Video) is a clear example of the transition of multimodal AI from text chatbots to full-fledged video consultations, where the system sees and hears the patient. The paper on arXiv is open for reading, and architects and developers can study the approach in detail and begin applying similar patterns in their own projects. For end users, systematic progress in this area means that in the future, primary medical consultations via AI video calls may become a reality, especially in regions with a shortage of doctors.
What is still unknown / limitations
The study was conducted under highly controlled conditions with actor-patients and pre-written scenarios — this is not equivalent to real clinical practice, where patients formulate complaints in an unstructured way and symptoms may be atypical. The system documents limitations in fine anatomical accuracy, subtle affective nuances, and high-frequency movements. Before real clinical application, validation with real patients, safety protocols, and regulatory approval are necessary. There is no public API, and the system remains research-oriented.
Sources
- Google Research Blog — Advancing AMIE towards expert-level audio-visual clinical consultations
- arXiv 2608.09861 — Towards Expert-level Medical AI for Real-time Video Consultations
- The Keyword (Google Blog) — AMIE: Advancing medical AI for video consultations
Author
Look at AI, editorial team
