The Index LLM team at Chinese video hosting platform Bilibili (IndexTeam) released the open-source machine translation model family Index-Translate based on Qwen3.5 under the Apache-2.0 license. The models work with 150 languages, including Russian, and follow translation instructions: a specified glossary, preservation of markup, code, and placeholders. The 2B and 9B versions are already available for download on Hugging Face and can be run locally, while the larger MoE flagship is promised later.

image

What happened

The Index LLM team at Chinese video hosting platform Bilibili, known as IndexTeam, released the machine translation model family Index-Translate, fine-tuned based on Qwen3.5. The lineup consists of 2B and 9B versions, as well as the MoE flagship 35B-A3B, which has 30 billion parameters, of which about 3 billion are active per token; the entire family is distributed under the Apache-2.0 license. The models translate between 150 languages, including Russian, accept translation instructions — glossary, preservation of markup, code, and placeholders, style, and context — and additionally understand memes and slang from Chinese internet communities. Along with the main models, four derivatives are presented: Index-Echo S2TT translates speech directly into subtitles, Index-Echo S2ST does speech-to-speech with preservation of the speaker's voice for ZH and EN directions to six languages, Index-Homura-9B adjusts the translation to a specified number of syllables and became the best among 20 models on the Sandglass bench with a composite score of 0.8660, and Index-Nailong-9B translates an entire book in one pass. A technical report arXiv:2609.40181 was published with the release, and a collection was assembled on Hugging Face, where the 2B and 9B versions are already available for download, while the weights of the flagship 35B-A3B have not yet been published and the model is marked as preview.

Context

The release fits into the rapid growth of a separate niche of open machine translation, where Bilibili, Alibaba with the Hy-MT2 family, and Google with TranslateGemma are competing, with their own leaderboards and specialization for speech, dubbing, and long texts emerging. IndexTeam's bet is revealed in the technical report: on FLORES-200 with the COMET-22 metric, the flagship 35B-A3B scores 0.8794, the 9B version — 0.8789, DeepSeek-V4.1-Flash — 0.8762, GPT-5.6-Sol — 0.8650, while the open universal Hy-MT2-7B, Hy-MT2-30B-A3B, and TranslateGemma-12B remain behind. However, the determining factor is not pure translation quality, but the fulfillment of translation requirements: by the instTrans IFscore metric, the 35B-A3B version reaches 0.8336 compared to 0.7624 for GPT-5.6-Sol, meaning that precise adherence to the glossary, table structure, and placeholders provides a noticeable lead over universal frontier models. This profile indicates a deliberate specialization of the family on top of the base Qwen3.5, and the open license with a published report makes the recipe reproducible for any other team.

Why this matters for the industry

For developers and companies, the key becomes not the quality rating, but pipeline manageability, since the preservation of terms, table structure, and placeholders is critical for subtitle, localization, and data preparation pipelines, where universal LLMs break the format. Production tasks of the "markup, code, placeholders, glossary" class can already be deployed on a single video card with 24 GB of VRAM under Apache-2.0 instead of a paid API with unpredictable behavior on formatted data. Instruction-following translation is becoming a commodity: open weights on 8–24 GB of VRAM remove the price premium of the "just translate" operation and shift the margin to workflow layers — glossary memory, quality control, integration, and distribution, rather than to the model weights. A likely response wave from Hy-MT2 and TranslateGemma and independent metric restarts by the community will increase pressure on API translation prices, and the practical benefit lies in format robustness, not in a pure benchmark.

Why this matters for users

Practical interest begins with downloading: the IndexTeam/index-translate collection on Hugging Face contains the 2B and 9B versions under Apache-2.0, where the 2B runs locally on approximately 8 GB of video memory in bf16, and the 9B — on 24 GB. Russian is supported in both directions, and instructions allow specifying a glossary and preserving markup, code, and placeholders, which is immediately applicable to subtitles and technical documentation. The speech models of the family provide ready-made scenarios: Index-Echo S2TT assembles subtitles directly from speech, Index-Echo S2ST voices dubbing with preservation of the speaker's voice from Chinese and English to six languages, Index-Homura-9B adjusts phrase length for lip synchronization, and Index-Nailong-9B translates a book in its entirety in one pass. In the evening, you can set up local inference and build an MVP — a subtitle pipeline or documentation localization with a glossary, without passing data through external services.

What is still unknown / limitations

There are serious caveats to the cited metrics. The gap between the flagship and DeepSeek-V4.1-Flash on FLORES-200 of about 0.003 is at the level of possible statistical noise, and the superiority over GPT-5.6-Sol by 0.014 is not supported by data on statistical significance. The only non-shorted benchmark in the set, the WMT26 bench, paints the opposite picture: the 9B version scores 75.35 compared to 89.10 for GPT-5.6-Sol and 83.55 for DeepSeek-V4.1-Flash, with GPT-5.6-Sol acting as the judge, which creates a conflict of interest in the evaluations. In addition, the flagship 35B-A3B is still in preview status: its weights are promised later, so the main MoE result cannot be independently verified, and the conclusion about overtaking the frontier is based on the interpretation of the chosen metric, not on confirmed third-party tests.

Sources

Author

Look at AI, editorial team