'Alice AI' by 'Yandex' no longer requires a human to choose which neural network will respond: the system itself assesses the complexity of the query, sends simple questions to fast models, and calculations, analysis of multiple documents, and step-by-step plans to 'Expert' mode, where the reasoning process is visible and multiple files can be uploaded. In an interview with Forbes, CEO of 'Alice and Smart Devices' Valery Stromov explained the architecture: the router operates on top of 3–5 models, and the harness maintains context and connects search, files, and services. Currently, all of this is available for free, but active use of 'Expert' is planned to be moved to the 'Alice Plus' subscription in the future.

image
image
image

What Happened

'Alice AI' has learned to automatically distribute queries between models based on their complexity: simple questions are processed by fast models, while calculations, analysis of multiple documents, and multi-step plans are sent to 'Expert' mode, where the user can see the reasoning process and attach multiple files at once. 'Expert' can be enabled manually, but it is not required: by default, routing works for everyone and automatically determines which level of model is needed for a specific query. In an interview with Forbes, Valery Stromov explained that the router is positioned on top of 3–5 models—'Yandex's' proprietary developments and open-source solutions for narrow scenarios—and the company plans to keep 80–90% of the task flow on its own models. A separate layer of the architecture is the harness, a wrapper that maintains the dialogue context and connects search, files, and services. Currently, 'Expert' mode is available to all users without restrictions, and, according to Stromov, it is best suited for purchases and financial tasks.

Context

The point of the update is not a new model, but a paradigm shift: instead of 'one model for everything,' orchestration is at work, where the system, not a human, assesses the complexity of the task. The cascade of 'a fast model by default plus a heavy reasoning model based on the router's decision' has long been known to the industry, and Perplexity, Cursor, and Manus are demonstrating the same path, so the novelty of 'Yandex's' approach lies not in the idea, but in bringing it to a mass product. With such an architecture, value shifts from the model itself to the harness—the code that stores the dialogue context and connects files, search, and services. The approach also has an economic dimension: according to Stromov, Google, with a single universal model, spends significantly more resources on inference, while routing allows maintaining quality at minimal cost. If the 'pyramid' takes hold, competition among mass assistants will shift from a race for a single best model to the quality of the router and harness.

Why This Matters for the Industry

For the industry, 'Alice AI' has become a live production example that the combination of 'router plus 3–5 models plus harness' works not in a demo, but in a mass product. According to Stromov, thanks to routing and the work of engineers optimizing inference, 'Yandex' achieved approximately 9 billion rubles in savings in the first half of 2025, making orchestration a verifiable way to reduce costs. For startups, the conclusion is twofold: the 'smart router' layer is rapidly being commoditized, so value shifts to the harness—data, integrations, and services around models. The complexity classifier, visible reasoning steps, and multi-file input are reproducible in proprietary agents, but there is no public API for this layer in the sources, so 'Alice AI' remains an architectural reference, not a platform. Orchestration is also moving into B2B: in September 2026, 'Alice AI for Business' will launch on approximately 80% of the same technology, which will be the first test of the approach's scalability for corporate scenarios.

Why This Matters for Users

Users no longer need to figure out which model is 'smarter': complex tasks—planning a trip with a budget, comparing multiple documents, calculating finances—can simply be thrown into the chat along with files, and the system will automatically enable heavy mode. The reasoning process is visible in the interface, making it easier to verify 'Expert' responses rather than taking them on faith. Currently, all of this works for free and without restrictions. A nuance to know in advance: active use of 'Expert' is planned to be moved to the 'Alice Plus' subscription, and workflows built on free access should be designed with future limits in mind.

What Is Still Unknown / Limitations

Key figures remain corporate statements: the approximately 9 billion rubles in savings on inference is a claim from a Forbes interview without a disclosed methodology; the source does not reveal the traffic structure, model sizes, hardware, or comparison baseline, so the figure aligns with the logic of routing but does not prove its advantage over a single universal model. There are no public benchmarks or routing accuracy metrics, and the main technical risk is a router error: if a complex query is sent to a fast model, quality will drop unnoticed by the user, and such misrouting is almost impossible to diagnose manually. The sources also do not reveal which specific model serves 'Expert' mode, how it is trained, and how it is evaluated. Finally, the criteria and limits for the future transfer of 'Expert' to the 'Alice Plus' subscription have not yet been disclosed.

Sources

Author

Look at AI, editorial team