🚀 InclusionAI (Ant Group's lab) has introduced Ling-3.0-flash — an agentic MoE model with 124B parameters.

The model utilizes a Hybrid Reasoning mode, combining the speed of the Ling series with the logic of the Ring series, and supports a context window of up to 1 million tokens. Despite its size, only 5.1B parameters are active per token.

🌍 The emergence of such efficient MoE models narrows the gap between lightweight models and heavyweight systems, offering high knowledge density at low inference costs.

👤 The model can be tested for free via OpenRouter until August 3, 2026. It is OpenAI API compatible and is well-suited for programming and AI agents.

Source 1: https://www.aimadetools.com/blog/ling-3-0-flash-complete-guide/ Source 2: https://openrouter.ai/models/inclusionai/ling-3.0-flash