Chinese AI studio StepFun released Step 5 Preview on September 18, 2026 — a proprietary 600B-parameter reasoning model that accepts text and images as input and maintains up to 1M tokens in context. On the Artificial Analysis Intelligence Index v4.3 benchmark, it scored 44 points and ranked 25th out of 200 models, with a median of 25 points for reasoning models in a similar price class. Output token pricing is approximately 3.7x below the class median: $2.70 versus $10.00 per 1M. Limiting factors: the model is nearly twice as verbose as the median, generation speed is not yet disclosed, and the API is available from only one provider.

image

What Happened

Step 5 Preview launched on September 18, 2026 as a proprietary reasoning LLM: StepFun claims 600B parameters, image input support, text output, and a 1M token context window, allowing entire long documents to be loaded. In the independent Artificial Analysis Intelligence Index v4.3 evaluation, the model received 44 points and the 25th position out of 200, while the median for reasoning models in a similar price class is 25 points. Pricing is significantly more aggressive than class medians: $1.00 per 1M input tokens and $2.70 per 1M output tokens versus $1.88 and $10.00 respectively. The discount on cached tokens reaches 95%, the blended rate is $0.51 per 1M tokens, and the cost of one task per AA methodology is $0.71. According to AA, during evaluation the model generated approximately 160M output tokens, nearly twice the median, meaning it is substantially more verbose than a typical model in the class. Generation speed is not yet disclosed, and API access is provided by a single provider.

Context

The Artificial Analysis Intelligence Index is an aggregated evaluation that relates a model's intelligence level to the actual cost of performing a typical task, which is why summaries feature medians by price class rather than the entire market at once. The phrase "on the Pareto frontier," highlighted in the Hacker News discussion title, does not mean a record score, but rather a position on the "intelligence/price" curve: at comparable quality, the model is cheaper than competitors; at comparable price, it is more capable. This is the aggregator's interpretation, not a measurable property of the model. The methodology accounts for verbosity: output tokens are specifically billed, so a chatty model consumes part of the price advantage even with a low nominal rate. The appearance of Step 5 Preview fits into a broader market dynamic where competition in the reasoning class is shifting from raw benchmark scores to task cost, and cache economics with long context are becoming an independent selection factor because agentic pipelines repeatedly process the same prefixes.

Why This Matters for the Industry

For the industry, the release of Step 5 Preview is a continuation of intelligence commoditization: Chinese developers are consistently pressuring reasoning-class pricing, and another model with output token prices several times below the median intensifies pressure on Western labs, accelerating the race to lower costs. The comparison metric itself is changing: task cost takes precedence over per-token price, and model verbosity and caching efficiency can change the final calculation by orders of magnitude. The 95% cache discount is the technically most significant element of the pricing: for agentic pipelines with long, repetitive prefixes, it radically recalculates unit economics. For startups building document and agentic products, this is a potential direct reduction in COGS and an opportunity to turn profitable scenarios that previously didn't work economically. If speed data is confirmed and additional providers connect, within the next few months the model could take a place as a cheap tier in model routers for mass tasks, and within a couple of years, a context of around 1M tokens could displace part of the RAG wrapper and establish multi-provider routing with caching as a first-class optimization.

Why This Matters for Users

If you are choosing a model for working with long documents or agentic tasks, Step 5 Preview is a reasonable candidate for a comparative pilot. A logical first step is to connect the model through the available provider and run your own eval sets: 20–50 real scenarios, including long documents, agentic cycles, and image processing, then compare quality, cost, and latency with your current model. When recalculating costs, account for verbosity: the actual volume of output tokens is approximately twice the median, so the nominal rate underestimates the real bill. Caching repetitive prefixes is a separate savings reserve for pipelines with fixed system prompts. It is too early to put the model in production-critical paths due to its Preview status and incomplete operational data; defer conclusions for interactive scenarios until speed measurements appear, and track updates on the model's page at Artificial Analysis.

What Is Still Unknown / Limitations

The evidence base is currently narrow: the entire result relies on a single aggregated index without a breakdown by task type, so the model's capability profile — strengths and weaknesses in reasoning, code, and vision — cannot be derived from it. There is no StepFun technical report in the sources, so claims of 600B parameters and multimodal input are not independently verified, and the architecture (dense model or MoE) is unknown. Generation speed is not disclosed, which precludes conclusions about latency and suitability for interactive scenarios, and availability through a single API provider creates a single point of failure. There are no independent replications of the results yet, and the "Pareto frontier" phrasing is an interpretation by Artificial Analysis, not a measured property. The model has Preview status, so both pricing and access terms may change.

Sources

Author

Look at AI, editorial team