On September 21, 2026, xAI, publishing under the SpaceXAI brand, released Grok 4.7 — its most capable model for coding and intellectual work. Price and speed remained at the level of the previous version, Grok 4.6, the context is 500,000 tokens, and the model is already available in Grok Build, all Cursor plans, and via API. The claimed gains are confirmed by vendor benchmarks, but independent measurements place the model noticeably below the leaders — Claude Fable 5.1 and GPT-6.

What happened
On September 21, 2026, xAI released a model with the identifier grok-4.7. The claimed update mechanism includes a larger base model, longer RL training on tasks requiring many hours, and improved output self-verification. Price and speed did not change compared to Grok 4.6: $2 per million input tokens and $6 per million output tokens, context — 500,000 tokens, knowledge cutoff — May 2026. Reasoning levels low, medium, and high (default) plus xhigh are available. The fast version, Grok 4.7 Fast, doubles output speed at double the price and is offered only in Cursor and Grok Build. According to official benchmarks, the model outperforms Grok 4.6 in all areas: CursorBench 4.0 — 46.3% vs. 40.4%, EEBench — 64.0% vs. 53.0%, Terminal-Bench 4.0 — 38.0% vs. 20.3%, Harvey Legal Agent — 19.6% vs. 15.8%, GDPval — Elo 1695 vs. 1605. Grok 4.7 is already the default model in Grok Build and is included in all Cursor plans.
Context
The release continues xAI's strategy of "near-frontier at the minimum price." Against $4/$20 per million tokens for GPT-5.6 Sol Max and $10/$50 for Fable 5.1 Max, the $2/$6 price remains the most aggressive offer at the same speed as Grok 4.6: according to the independent Artificial Analysis Intelligence Index v4.3.2, the model lags the leaders at 46 points vs. 53 with a five-fold lower output price — this combination is the company's bet. Competition in the agentic coding market is shifting from absolute records to the price-to-quality ratio. The claimed improvement mechanism is plausible and aligns with the industry trend toward RL for long horizons, but without a technical report, this is an interpretation, not a confirmed fact.
Why this matters for the industry
The main point for the industry is not records, but the shift in the price-to-quality boundary: the claimed gains in agentic coding at the same price and a 500,000-token context open up unit economics for an entire class of agentic products that did not work with more expensive models. If this ratio is confirmed on real workloads, large-scale agentic scenarios will shift to "near-frontier" models, and price pressure on OpenAI and Anthropic pricing will increase — cheap agentic plans may become the norm. The market, judging by xAI's strategy, is structuring into absolute frontier leaders and cheap mass-market models: teams will maintain routing by tasks, where the expensive model handles complex decisions and the cheap one handles volume, and their own evals will become the main selection tool. In the long term, if the combination of "RL on multi-hour tasks plus self-verification at the minimum price" becomes an industry standard, the capability lever will shift from pretraining size to post-training, structurally cheapening the frontier and accelerating the commoditization of agentic coding, while value shifts from model access to orchestration: memory, integrations, vertical workflows, and reasoning level management.
Why this matters for users
You can try the model today: for free in Grok Build at x.ai/build, in all Cursor plans, or via the Grok API with a key from console.x.ai, specifying the identifier grok-4.7. For existing Grok 4.6 users, this is a replacement without changing price and speed — just change the model id and run your evals. The most practical detail from the documentation: in the Responses API, set prompt_cache_key (or the x-grok-conv-id header in Chat Completions), otherwise the prompt cache often does not work and you pay the full price for input tokens. If you need to store inference in the US, the us.api.x.ai endpoint is available with a 10% surcharge.
What is still unknown / limitations
The claimed gains are confirmed only by vendor numbers: according to an independent run of Terminal-Bench 4.0 by Artificial Analysis, Grok 4.7 scores 26% vs. 60% for GPT-6 Astra, while xAI claims 38% — the 12-percentage-point discrepancy is explained by the dependence of agentic benchmarks on scaffolding, settings, and task selection, so the official picture cannot be transferred to other pipelines. There is no data on latency and the xhigh reasoning level. Claims about increased pretraining, longer RL training, and improved output self-verification remain without a technical report. Rumors about the next model, Grok 4.8 with 2.5T parameters trained via RL, are unconfirmed.
Sources
- Introducing Grok 4.7 | SpaceXAI (official xAI announcement)
- Grok 4.7 | SpaceXAI Docs (official API documentation)
- The Decoder: xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
Author
Look at AI, editorial team
