xAI released Grok 4.6 on August 12, and the headline number is a tie: the model matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index at 61 points, a composite of nine benchmarks. But the more telling numbers sit below that headline, in the agentic and coding evals where Grok 4.6 beats both GPT-5.6 Sol and Fable 5 Max on several long-horizon tasks. This is a release built around one idea: agents that can sustain work across many steps, not models that ace a single prompt.

The benchmark table tells a mixed story that rewards reading closely. On GDPVal-AA v2, a knowledge-work benchmark, Grok 4.6 scores 1753, ahead of GPT-5.6 Sol’s 1728 and Fable 5 Max’s 1741. On CursorBench v3.2, it hits 69.9%, edging GPT-5.6 Sol’s 67.2% but trailing Fable 5 Max’s 70.5%. DeepSWE v1.1 shows a wider gap: Grok 4.6 lands at 65.9%, well behind GPT-5.6 Sol’s 73% and Fable 5 Max’s 70%. The pattern inverts on FrontierCode v1.1 Extended, where Grok 4.6 posts 61.3% against GPT-5.6 Sol’s 60.6% and Fable 5 Max’s 63.6%. The model is competitive everywhere, best in some places, never dominant across the board.

What stands out is Terminal-Bench v3.0, a benchmark for command-line and terminal agent work. Grok 4.6 scores 26%, a sharp jump from Grok 4.5’s 15.7%, though both GPT-5.6 Sol at 34.6% and Fable 5 Max at 34.1% remain well ahead. The improvement from the prior xAI generation is the largest single-model gain in the table, and it signals where the training effort went. xAI says Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. The SFT trajectories were regenerated using Grok 4.5 across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work.

The agentic RL training is the substantive change. xAI lists domain-specific environments for kernel optimization, web development, and computer-aided design as part of the training mix. That is a deliberate bet that the next frontier in model capability is not raw reasoning on a single query but sustained, multi-step execution inside real tooling. The company’s own testing narrative supports this: xAI says the model is especially strong at turning a broad product idea into a working first version, researching unfamiliar domains, structuring an application, implementing core interactions, and refining through several rounds of feedback. On longer trajectories, xAI reports seeing more self-testing and verification, with the model checking its own work before moving on.

That self-verification behavior is the quietly important detail. Most frontier models can produce a plausible first pass on a coding task. The harder capability is knowing when the output is wrong and iterating before presenting it. xAI’s observation that Grok 4.6 checks its own work on longer trajectories suggests the RL training is producing something closer to a work loop than a single-shot generator. Independent verification is needed, and the company’s internal observations are not proof, but the benchmark pattern is consistent with the claim.

The safety framing deserves scrutiny. xAI says Grok 4.6’s safeguards have been improved and calibrated in line with the model’s capabilities, with its widest-ever suite of pre-deployment testing and extensive post-deployment and third-party testing. The company explicitly frames the safety stack as designed to maximize utility and security across legitimate use cases, including vulnerability patching and accelerating the engineering design cycle. That is a notable posture for a model that scores 26% on a terminal benchmark. A model that can operate a terminal effectively is a model that can do real damage if misused, and xAI’s framing of vulnerability patching as a legitimate use case is a reminder that offensive and defensive security work share the same underlying capability.

Pricing is where xAI is making its competitive move. Grok 4.6 starts at $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the price. That undercuts much of the frontier tier. The model is available today in Cursor and Grok Build, with 2x included usage in both for the first week, and through the API plus partners including OpenRouter, Vercel, and Cloudflare. The distribution strategy is aggressive: meet developers where they already work, make the first week effectively free, and let the benchmark scores speak.

The Artificial Analysis Intelligence Index tie at 61 is the marketing headline, but the competitive picture is more nuanced. Grok 4.6 matches GPT-5.6 Sol on the composite, trails on DeepSWE and Terminal-Bench, leads on GDPVal and FrontierCode. The honest read is that xAI has closed the gap to the frontier on aggregate intelligence while making specific, defensible gains in agentic coding and knowledge work. For a model released roughly three months after Grok 4.5, that is a fast cadence, and it reflects the longer supplemental training run plus the agentic RL focus rather than a fundamentally new architecture.

The deeper question is what this means for the agent economy. xAI is pricing Grok 4.6 at a level that makes long-running agent trajectories economically viable. A model that can work across a codebase for an hour at $6 per million output tokens changes the calculus for autonomous coding tools. The benchmark scores suggest the capability is real, and the pricing suggests xAI wants volume adoption, not margin maximization. That is a direct challenge to OpenAI and the other frontier labs that have priced their top models at a premium.

For AI builders, the practical takeaway is that the frontier has bifurcated. Single-turn intelligence is approaching a plateau where the top models cluster within a few points of each other on composite indices. The differentiator is now sustained, multi-step execution: how long a model can hold a task, how reliably it verifies its own work, and how well it operates inside real tools like terminals, code editors, and CAD environments. Grok 4.6 is the clearest signal yet that xAI is optimizing for that axis, and the benchmark gains on terminal and coding tasks, plus the self-verification behavior on long trajectories, suggest the training approach is working.

The competitive response will be telling. OpenAI’s GPT-5.6 Sol still leads on DeepSWE and Terminal-Bench, the benchmarks that most directly measure long-horizon software engineering and terminal operation. But the gap is narrowing, and xAI’s pricing pressure is real. The next few months will show whether the other labs respond with capability, price cuts, or both. The outstanding question is whether xAI can sustain this cadence and whether the self-verification behavior generalizes beyond the curated test environments into messier, real-world agent deployments. Grok 4.6 is available now, and the terminal benchmark says it can do more than chat.