Chinese banks, telcos, and restaurants are turning AI tokens — units of computation that a model spends on processing and generating text — into ordinary consumer bonuses. Rest of World describes how access to AI inference in China is beginning to be distributed through credit card cashback, mobile data packages, and free credits with lunch, following the mobile data model.

image
image

What happened

According to Rest of World, Moonshot AI in July, jointly with the Agricultural Bank of China and American Express, launched the Kimi credit card. China Merchants Bank gives new cardholders up to 1.8 billion MiniMax tokens, and Shanghai Pudong Development Bank provides subsidies of up to 3 billion tokens for Qwen models. China Telecom sells 10 million tokens for 9.9 yuan (about $1.40) per month, China Mobile in Shanghai — 400,000 tokens for 1 yuan, and China Unicom — 6–18 million tokens for 15–45 yuan. Beijing's Jingu Yuan dumpling restaurant gives away computation worth about 10 yuan for free after a meal, and the Haizhu district in Guangzhou launched a "token loan" — loans where the production and consumption of tokens are accepted as collateral.

Context

The background of the event is the sharp growth in AI computation consumption: in the discussion of the material, an estimate is cited that by mid-2026, about 500 trillion tokens per day will be consumed in China, compared to 100 billion at the beginning of 2024. The retail model became possible due to the cheapening of inference: Chinese open-weight models are claimed to be 60–90% cheaper than comparable models from OpenAI and Anthropic, which allows tokens to be sold at levels of 1–3 yuan per million. Analyst Pei Zhao (Hello China Tech) characterizes what is happening as a "supply-led experiment" — the market is driven by supply, not demand, and most users do not track their "token balance" in the process.

Why this matters for the industry

For the industry, this is a shift in the commercialization of inference: the token has for the first time become a consumer good, and the main front of competition becomes the mass user. Pricing power is effectively passing to distributors — banks and telcos, which are dumping access to AI computation and turning raw tokens into a commodity; for startups and developers, this is a signal to defend themselves not with a model, but with UX, integrations, and distribution channels. For engineering teams, this is a price anchor: retail levels of 1–3 yuan per million tokens strengthen the argument for self-hosting and cost-aware routing to cheap open-weight models in mass generative workloads.

Why this matters for users

For the reader, AI computation is becoming a utility like mobile internet: it can be bought in a package in a mobile plan, obtained for free with food, and used as collateral for a loan. While most people use AI through familiar apps and do not see a "token balance," it is precisely retail bonuses that are gradually embedding AI into everyday habits — from card cashback to a loan secured by tokens.

What is still unknown / limitations

The material has no technical result: not a single article, benchmark, or statement about new model capabilities. The claim that Chinese open-weight models are 60–90% cheaper than comparable models from OpenAI and Anthropic is presented without a methodology for calculating the price per token and without an assessment of their quality. The estimate of "about 500 trillion tokens per day by mid-2026" is given without indicating a source, and its order of magnitude needs to be verified before being repeated. The "token loan" in Haizhu is a financial construct: it only indirectly confirms the cheapening of inference, but says nothing about the quality of the models.

Sources

Author

Look at AI, editorial team