OpenAI published a post titled “GPT-6 Astra: The next generation in intelligence for work” featuring internal use cases for the model released on September 3. The main demonstration — Astra, using computer use together with Codex, independently edited three hours of multicamera footage into a finished video, controlling editing applications without relying on a dedicated API.

image
image
image

What happened

OpenAI's developer and marketing teams completed a post-production task: Astra, using computer use, opened editing applications, found the necessary clips among three hours of multicamera footage, and assembled them into the video “First impressions of GPT-6 Astra from developers.” The video was uploaded on September 3 and gained over 550,000 views in 4 days. In the same post, the company provided two more use cases. Astra, using computer use, solves Financial Modeling World Cup tasks — a 2023 Excel championship — about 4 times faster than the human winner. An engineering team used the model to find a bottleneck in memory allocation in the Codex test environment: after switching allocators, turn latency dropped 25-fold, and peak memory increased by about 30%. The publication of the use cases accompanied the model's release, the start of which was reported by CNBC on September 3. Astra's public API pricing is $10/$50 per million tokens, and on Terminal-Bench 4.0, the model showed 57.9% compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1.

Context

Computer use is a model operating mode in which an agent controls existing software through the interface: it opens applications, finds the necessary objects on the screen, and performs actions without requiring developers to create special integrations. Professional video editing software (NLE) does not have dedicated AI integrations for Astra, so the demonstrated editing took place in existing applications. Financial Modeling World Cup is an Excel competition, and the company uses its 2023 results as a public benchmark for evaluating the model. For comparison, GPT-5.6 Sol — an earlier OpenAI model — and Claude Fable 5.1, a competitor's model, are cited on Terminal-Bench 4.0, so the stated figures can be read as a leap from the previous generation and from the nearest competitors.

Why this matters for the industry

For the industry, the demonstration shows that a general-purpose frontier model can perform professional post-production within existing NLEs without integrations, and competition is shifting from token price to the cost of a completed task: fewer steps and retries mean a cheaper finished result. The pricing and benchmark results published by OpenAI give companies a practical benchmark when choosing a model for agentic workloads. Specialized AI editing and generative startups are under direct pressure, as their value was built on a pipeline for one specific task: thin wrappers without their own moat lose arguments in front of clients and investors. For product teams, this is a new UX pattern of “task in, finished artifact out,” and orchestration of long computer tasks can already be built on top of Astra and Codex.

Why this matters for users

For the reader, the use case is a clear benchmark of the computer use level: the model does not just describe a video, but independently controls the application and completes the editing. The finished video can be watched on OpenAI's YouTube channel, and those with access to ChatGPT Work or Codex can try similar scenarios themselves and calculate the economics not by token price, but by the cost of a completed task. However, it is worth remembering that there is no dedicated API for editing: computer use in professional NLEs is not a ready-made integration, but a scenario that can fail in long chains of actions.

What is still unknown / limitations

OpenAI does not disclose the methodology of the demonstrations. The editing of a three-hour video is a one-off demo use case without a protocol, number of attempts, error data, or the cost of the tokens used, so it does not prove the reliability of long agentic tasks. For Terminal-Bench 4.0, the available materials do not include details of the agent configuration — pass@k, number of attempts, scaffolding, timeouts — or independent reproductions of the result. The statement about Financial Modeling World Cup mixes time and accuracy metrics and does not disclose launch conditions, and the allocator use case in Codex is a single example without error analysis. All the figures provided should be considered stated, not independently verified.

Sources

Author

Look at AI, editorial team