Cloudflare has released Auto Router for AI Gateway in public beta: instead of manually selecting a model, developers only need to specify the identifier 'cloudflare/auto' in their code, and the edge router will automatically select a model for the task, balancing expected quality and price. In Cloudflare's internal benchmark, this routing showed accuracy close to flagship models at significantly lower costs per successful task. While the beta continues, routing is free.


What happened
Auto Router operates at the network edge on top of Workers AI and GPU: a multi-functional classifier labels incoming requests across 14 task categories and assigns scores from 1 to 5 on four axes — complexity, ambiguity, cost of error, and context dependency. Next, a scoring matrix selects the executor using the formula 'expected quality minus adaptive price penalty,' while the candidate pool is filtered by credentials, spending limits, and model health during failures; the router also accounts for the cost of cache misses in long agentic sessions when making its selection. In an agentic scenario of 97 tasks, cloudflare/auto showed 86.6% accuracy (252 out of 291 successful attempts) at a cost of $0.0084 per successful task; for comparison, Claude Opus 5.5 reached 96.6% at $0.0210, and GPT-6 Sol — 84.2% at $0.0108. According to Cloudflare's self-report, internal use of the router yielded up to 30% savings compared to top models.
Context
AI Gateway is Cloudflare's gateway through which applications access different model providers via a single proxy, so the decision of which model to call previously remained entirely in the application code: engineers had to pre-classify tasks and set a cheap model in some places and a flagship in others. An error in either direction costs money: overpaying for a flagship on trivial requests or losing quality where it is truly needed. Long agentic sessions create a separate layer of complexity: swapping models resets the prompt cache, and context must be paid for again with each switch, so cache-awareness remains a non-trivial and rarely considered requirement for routing.
Why this matters for the industry
For the industry, the utility itself is less important than the shift in the decision point: AI Gateway transforms from a proxy into a control plane, and model selection becomes a pricing-quality policy of the gateway rather than an engineering task in each application's code. Such a control layer, through which all of a company's requests pass, gains leverage relative to model vendors, and competition shifts from individual models to the quality of routing, observability, and cache economics. For startups, this means a reduction in inference costs for products where a model was previously fixed as a constant. If the approach takes hold, hard-coding a specific model in code risks becoming an anti-pattern, teams will have to find differentiation above routing, and a task classifier at the network edge has every chance of becoming a standard infrastructure element — corresponding features in other gateways and routers look expected.
Why this matters for users
If traffic already goes through AI Gateway, testing does not require rebuilding the application: simply replace a specific model with 'cloudflare/auto' and compare accuracy, cost, and latency on your real tasks against manual selection — most practically on non-critical scenarios like drafts, classification, and summarization. The effect depends on the task structure: on simple requests, price weighs more and small models win, while on complex ones, quality becomes decisive, and here the router, according to vendor numbers, loses to flagships. Therefore, responsible production paths are still best left on a fixed model, and those who prioritize maximum quality without price compromises should wait for the planned cloudflare/auto-best profile.
What is still unknown / limitations
Key figures come from Cloudflare's internal benchmark, and the methodology, task list, and classifier calibration have not been published, so transferring results to other scenarios should be done carefully. The latency of the edge classifier and pricing after the beta ends have not yet been disclosed. The savings figure from internal use remains a self-report without a description of the sample, and the benchmark compares only two named models. The timing and conditions for the release of the planned cloudflare/auto-best profile have not been announced.
Sources
Author
Look at AI, editorial team
