OpenAI published a practical guide, "A model guide for the GPT-6 family," on its news page for startups building products on the GPT-6 family. The company describes not so much the capabilities of the models as the engineering around them: how to choose between GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna, how to manage reasoning depth and token costs, and how to ensure an agent is production-ready. The guide marks a notable shift: from demonstrations of what models can do to how to reliably build a product with them.

image
image

What happened

In the guide, OpenAI describes model selection as an "intelligence/price" trade-off: GPT-6 Astra is positioned for the most complex reasoning, GPT-6.1 Sol — for complex coding, research, computer use, and multi-agent workflows that are assembled in the Responses API (beta), and GPT-6 Luna — for high-volume, repetitive tasks with a clear goal. The reasoning effort parameter is configurable via the API from low to extra high/max and, according to the company's description, can be changed mid-dialogue without resetting the cache. Fast and Ultrafast modes are provided for quick responses, with the latter available on GPT-6 Astra. Computer use works on all three models and, according to the guide's wording, allows interaction with websites and desktop applications even where the target service has no API. As examples of teams applying these practices, OpenAI names Harvey, Cognition, Hex, and Invideo.

Context

The document itself is vendor engineering documentation, not a research result: it contains no methodologies, benchmarks, or independent comparisons. Its value lies elsewhere — OpenAI has formalized disparate production practices into a coherent pipeline: caching and context compaction, success rate and latency measurements, monitoring, and data control are described as connected steps in a single process, not as separate tips. This meets the demand of the moment: teams are moving agents from demos to working services, and vendors are increasingly focused on selling not the smartest model, but a reproducible way to bring it to a product.

Why this matters for the industry

For the industry, this is a platform play: OpenAI is taking on the layers of production engineering — model selection, reasoning effort configuration, context economics, and observability — offering teams a ready-made framework instead of DIY solutions. Three-tier routing of Astra, Sol, and Luna across workflow stages, dynamic reasoning effort, and cache as an architectural element reduce the unit economics of agentic products, but at the same time blur the advantage of thin wrappers over the API: a significant part of their value is now built into the platform. The guide also sets a communication bar for competitors, who will have to respond with comparable developer documentation; likely, in public model comparisons, weight will shift from intelligence benchmarks to economic-latency metrics — the cost of a successful agent step, latency, and cache efficiency.

Why this matters for users

For developer readers, the guide is a set of knobs that can be tried immediately: enable input token caching, which according to OpenAI data is up to 95% cheaper than standard, structure prompts for cache prefix hits, and use compaction for long agent sessions. Next, it makes sense to run your pipeline through reasoning effort levels, measure the success rate and latency of your current workflow, and calculate the actual savings on your own numbers. Computer use allows you to build a prototype integration for an application that has no API, and multi-agent scenarios can be tried in the Responses API while it is in beta status. For existing products, this is a direct recalculation of cost and margin; you can look to the Invideo case, where OpenAI records an approximately threefold increase in success rate on color correction tasks.

What is still unknown / limitations

All characteristics in the guide come from OpenAI itself: the document contains no methodologies, benchmarks, or independent measurements, so claims about model capabilities should currently be read as the manufacturer's word. The architectural claim about the independence of the cache from the reasoning effort level, which allows cheaply alternating reasoning depth within a dialogue, should be verified on your own traffic. The stated discount on cached tokens is an upper bound, not an average: actual savings depend on traffic patterns and context structure. No reliability measurements are provided for computer use, and multi-agent workflows remain in beta status, which limits production readiness.

Sources

Author

Look at AI, editorial team