🤖 Qwen Releases E-Commerce Bench — a Store Simulator for AI Agents

An LLM agent with an initial capital of ¥100,000 runs an online store for a full year — 365 simulated days: negotiating with suppliers, setting prices, and managing inventory and cash flow. The environment includes 6,886 products and 576 suppliers, 152 of whom are fraudsters. GPT-5.6 Sol earned the most — ¥1,431,425, but ranked only 16th out of 18 in resilience to fraudsters; the best open model, Qwen3.8-Max-Preview, earned ¥416,252.

🌍 Profitability, negotiations, resilience to fraudsters, and solvency are independent axes: none of the 18 models leads on all seven. A deterministic negotiation core makes the results reproducible.

👤 Teams building agents for procurement, pricing, and inventory gain an open testbed with time and liquidity limits: models can be fairly compared before a live launch.

Source 1: https://qwen.ai/blog?id=e-commerce-bench Source 2: https://arxiv.org/abs/2608.30730