OpenAI has reconfigured the ChatGPT subscription quota consumption for the GPT-6 Astra model: heavy "long tail" operations — long agent sessions, 3D modeling, and similar — now consume up to 3-4 times less quota, according to the company, while the model's quality itself remains unchanged. The change was announced on September 6 on X by Codex and ChatGPT lead Tibo Sottiaux. The decision came against the backdrop of complaints: after the banked reset on September 5, Plus subscribers hit the limit in about 20 minutes of work. On the night of September 8, around 01:00 UTC, OpenAI carried out a second global usage reset for all paid subscriptions — nominal limits did not change, but the accounting of consumption changed.

What happened

On September 6, OpenAI's Codex and ChatGPT lead Tibo Sottiaux announced on X a change in GPT-6 Astra quota consumption: for heavy users — "long tail" operations — consumption from the subscription may become up to 3-4 times less, while model quality, according to his statement, remains unchanged. This is not a model update, but a recalculation of the cost of requests in subscription billing: nominal limits remained the same. On September 7 at 19:24 UTC, OpenAI announced a second global usage reset for all paid subscriptions; according to Sottiaux, it took effect around 6:00 PM Pacific Time, i.e., approximately 01:00 UTC on September 8. The reset affects ChatGPT Plus, Pro, and Business subscriptions.

Context

GPT-6 Astra was released on September 3 in a limited preview, and usage accounting became a problem literally in the first days. OpenAI compensated paid subscribers waiting for access with one banked reset for each day of waiting, and on September 5 carried out the first full banked reset. This did not solve the problem: according to user reports, Plus subscribers hit the limit in about 20 minutes of heavy work, while Pro burned through a significant part of their weekly quota in a day. The background here is economic: Astra is significantly more expensive in inference than GPT-5.x, but is sold within the same fixed subscriptions, so the laboratory has to rebalance not the limits, but the "price" of an individual request against the quota.

Why this matters for the industry

For the industry, this is a telling case of the economics of launching a frontier model: a model that is more expensive in inference enters the same fixed subscriptions, and the laboratory compensates for the difference by recalculating the weight of requests against the quota. The "price" of the model within the subscription turns out to be a flexible parameter that OpenAI reconfigures in live mode as real traffic reveals the cost per request. Two global usage reviews in four days — on September 5 and on the night of September 8 — show that launch-week quotas can be reviewed within 48 hours. For teams building products and agent workflows on consumer ChatGPT subscriptions without an API budget, this is a direct platform risk: the unit economics of heavy agent loads on Plus/Pro can change overnight, so measuring consumption in-house becomes a mandatory part of infrastructure.

Why this matters for users

If you use ChatGPT Plus, Pro, or Business, your Astra limits were reset a second time around 01:00 UTC on September 8 — you can start working from a clean slate. Long heavy sessions — agent tasks, 3D modeling, and similar — now consume significantly less quota, so extended agent runs have a chance to fit within a regular subscription. But you should measure the exact figure for your tasks yourself: in the days after the reset, track actual consumption on your task profile to understand the new norm. The experience of this launch week shows that accounting rules can change within 48 hours, so plans built on old consumption are better rebuilt after your own measurements.

What is still unknown / limitations

The figure "up to 3-4 times less" is not yet supported by methodology: it is not specified which operations exactly belong to the "long tail" class, on which workflows the reduction was measured, and by what method the measurement was carried out — no comparative data has been published. The statement "model quality does not change" is also provided without evidence: there are no A/B data, nor benchmarks before and after the change. An indirect technical signal — Astra's greater expense in inference compared to GPT-5.x — is consistent with a heavier frontier model, but there are no details about architecture and serving optimizations like caching, batching, and speculative decoding. Finally, it is unknown whether the new accounting will be maintained after the launch week ends or will be reviewed again as traffic grows.

Sources

Author

Look at AI, editorial team