On September 4, 2026, the Ant Ling model team (inclusionAI, an open AGI initiative of Ant Group from the Alibaba ecosystem) released Ling-3.0-flash-Sante — a medical fine-tuned variant of Ling-3.0-flash. The model claims results on par with flagship models on medical benchmarks and is already available for free via hosted API on OpenRouter and Vercel AI Gateway.

What happened
Ling-3.0-flash-Sante is built on the base of Ling-3.0-flash: MoE architecture with 124 billion total parameters and 5.1 billion active per token, context of about 262k tokens (Vercel indicates 256K), max output 32K tokens. According to the benchmark table from the official announcement, the model scores 83.8 on DiagnosisArena-MCQ versus 81.9 for GPT-5.6 Sol, 78.4 for Kimi K3, and 76.2 for Gemini 3.6 Flash. On MedXpertQA-Text the result is more modest: 53.9 versus 60.2 for GPT-5.6 Sol, 62.4 for Gemini 3.6 Flash, and 53.5 for Kimi K3. On medical ethics (MedEthicAlign, AFUSAFE-MedSCE), the announcement claims results higher than Claude Opus 4.8 and slightly lower than GPT-5.6 Sol. The model is available only as a hosted API: free on OpenRouter (identifier inclusionai/ling-3.0-flash-sante:free, provider Novita AI) and in Vercel AI Gateway, where a promo is active until October 4, 2026. Weights are not published: on Hugging Face the model returns 401.
Context
inclusionAI is the model team of Ant Group in the Alibaba ecosystem, releasing the Ling line of models. The base Ling-3.0-flash is available on Hugging Face, and previously the team released Ling-3.0-flash-Fin — a financial fine-tuned variant under the MIT license on the same base. Ling-3.0-flash-Sante continues this pattern: domain specialization of a general MoE base for a narrow area. All benchmark figures for the new model are provided in the chart of the official announcement; there is no technical report, description of the fine-tuning methodology, or independent reproduction of the results yet.
Why this matters for the industry
A free medical MoE model with benchmarks on par with flagship models on medical tasks lowers the barrier to entry for clinical applications: medical reasoning, checking the safety of drug therapy, and retrieval from an evidence base are available via standard API without self-hosting a 124-billion model. For startups, this is a free resource for prototyping clinical co-pilots and evaluating medical scenarios. At the same time, closed weights, metrics claimed by the team itself without a technical report, and the absence of an SLA limit the use of the model in production and the reproducibility of results.
Why this matters for users
The model can be tried right now for free: on OpenRouter under the identifier inclusionai/ling-3.0-flash-sante:free and in Vercel AI Gateway, where the promo is active until October 4, 2026. A context of about 262k tokens is convenient for long medical documents and clinical cases. To verify the claims, it is worth comparing the model on your own medical prompts with the base Ling-3.0-flash and with GPT-5.6 Sol.
What is still unknown / limitations
All metrics are claimed by the team itself and taken from the announcement chart: there is no technical report, no description of the evaluation protocol (dataset versions, prompts, inference parameters), and no independent reproduction. Without open weights, architectural claims cannot be verified independently. The benchmark picture is uneven: leadership on DiagnosisArena-MCQ coexists with a lag behind GPT-5.6 Sol and Gemini 3.6 Flash on MedXpertQA-Text, which may indicate narrow specialization. The model has no SLA, and pricing and access terms after the promo ends are not disclosed.
Sources
- Ant Ling — official announcement of Ling-3.0-flash-Sante (X/Twitter, September 4, 2026)
- OpenRouter — Ling 3.0 Flash Sante (free): specification, provider Novita AI, price 0
- Vercel AI Gateway — Ling 3.0 Flash Sante (Free): MoE 124B/5.1B, 256K context, promo until 04.10.2026
Author
Look at AI, editorial team
