Avito's recommendations team published a breakdown on September 24, 2026, in the AvitoTech blog on Habr about the restructuring of homepage recommendations: the avitofm collaborative model was migrated from rare batch updates to fine-tuning on 20-minute clickstream windows, which accelerated recommendation updates by 18 times, and the p99 path from event to recommendation no longer exceeds 25 minutes. To achieve this, chunk models fragmented by categories and regions were consolidated into one model per category, and an A/B experiment recorded +1% seller contacts.
What happened
The breakdown was written by Salavat Dinmukhametov, a senior ML engineer on the recommendations team. The old system relied on the batch collaborative model avitofm, which was retrained every six hours; fresh listings, which receive demand in the first hours after publication, entered recommendations with a delay and missed the peak of interest. The modernization took place in two steps. First, the team implemented incremental fine-tuning on top of existing batch embeddings: the first step yielded an offline gain of +6% Recall@K, and the final choice in favor of new architectures was verified by the team with an A/B test on live traffic. Then the near-time loop started working: the clickstream from Kafka is sliced into 20-minute windows, A100 and H100 GPU workers handle embedding fine-tuning, updated embeddings are dumped to Redis, and orchestration is managed by Avito's internal platform Aviflow, built on Kubeflow. In parallel, the model park was restructured: instead of 4,420 chunk models, one for each intersection of 52 categories and 85 regions, there are now 52 models, one per category.
Context
Chunking of the form "a model for each category and region intersection" has long looked like a reasonable compromise: local models learn on local data, and training and inference volumes remain predictable. The downside is a slow cycle: the smaller the segments, the less data each has and the less often it makes sense to retrain them, so embeddings inevitably lag behind feed events. Freshness in classifieds is not an infrastructure detail but a direct product property: the peak demand for a listing falls in the first hours of its life, and any gap between an event and a product entering the feed means a missed contact. The effect achieved by the team is built not on a new algorithm but on a structural trade-off: one model per category collects significantly more data, the park of artifacts for training, storage, and monitoring is sharply reduced, and the freed-up resources can be directed to frequent fine-tuning. The path itself is also telling: first a cheap incremental add-on on top of a working batch, and only after confirming the effect a full restructuring of the loop.
Why this matters for the industry
For engineers building recommendation systems under freshness pressure, this is a ready-made migration checklist, not an abstract idea: watermarks when reading windows, event deduplication by the key user, item, event, time, snapshots of affected embeddings in Redis, and automatic pickup of failed windows in Aviflow and Kubeflow form a mature streaming ML loop that can be copied in parts. The economics are also stated: fine-tuning all categories fits in approximately four GPU pods with a consumption of about 6 GB per category, and consolidating the model park reduced memory consumption by approximately three times. Teams that still keep a separate model for each segment split can compare their metrics with the published ones and calculate a pilot without major capital expenditures. If the recipe is repeated on other platforms, the combination of "one model per category plus frequent selective fine-tuning" will become a standard solution for marketplaces and classifieds, and chunking by regions risks becoming a recognized anti-pattern for large platforms.
Why this matters for users
For buyers, the difference is noticeable in a live example: a listing from fresh publications gets into the feed of those who need it faster and does not wait for the end of a multi-hour cycle until demand for it dries up. Engineer readers get a path of least resistance: the first step can be reproduced on their own data without restructuring the architecture, because incremental fine-tuning on top of working batch embeddings does not touch the rest of the pipeline, and Avito's transition started with exactly that. Next, it is worth comparing your own roadmap with the described loop: window fine-tuning, selective embedding dumps, automatic failure pickup, and honestly estimate how many pods and memory such a migration would require at your own volumes. Product teams get an argument for management: feed freshness translates into a measured business effect, not just a beautiful offline metric.
What is still unknown / limitations
The portability of the recipe to another platform is not shown in the material: merging regions increases requirements for event density per category, and on platforms with sparse interactions, merging chunks may not yield the described benefit. The figures cited are published by the team itself, there is no independent verification on external data, and the gap between offline and online metrics means that Recall@K cannot be used as a direct forecast of product effect. No open code or public benchmarks were attached to the breakdown, so until the recipe is repeated by other teams, its value lies in documented decisions and pitfalls, not in a guarantee of reproducibility.
Sources
Author
Look at AI, editorial team
