Wednesday, Oct 7 | --:--
Back to home

Grok 4.6 Ties Sol on Artificial Analysis at $2/$6 — Without Owning DeepSWE

xAI released Grok 4.6 — a post-training upgrade of Grok 4.5 aimed at long-running agents, coding, and interactive visual work — that ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index, ships at $2/$6 per million tokens, and is live in Cursor, Grok Build, and the xAI API. DeepSWE and Terminal-Bench still trail Sol/Fable.

Times of AI Desk 5 min read San Francisco, CA View as Markdown
Cover illustration for Grok 4.6 Ties Sol on Artificial Analysis at $2/$6 — Without Owning DeepSWE

At $2/$6, matching Sol on the main independent composite is a direct shot at OpenAI and Anthropic list prices. Read 4.6 as a price–agent step that ties Sol on AA without owning DeepSWE or Terminal-Bench.

xAI released Grok 4.6, a longer supplemental training run on the Grok 4.5 foundation rather than a new base model. The lab positions it for long-running agents, coding, and more ambitious interactive/visual projects. On xAI’s published table it matches GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index (Grok 4.5 High was 56; Fable 5 Max is 62). Pricing stays $2 / $6 per million input/output tokens, with a fast variant at 2× that price. Access: Cursor, Grok Build (2× included usage for the first week), the xAI API, plus OpenRouter, Vercel, and Cloudflare.

What xAI announced

From the primary “Introducing Grok 4.6” post (x.ai/news/grok-4-6):

  • Training: Longer supplemental run than 4.5; Grok 4.5 regenerated SFT trajectories across reasoning efforts and agent harnesses; RL on knowledge work, general coding, and domain environments (kernel optimization, web development, CAD).
  • Behavior: Stronger first passes on visual/interactive apps; more self-testing before moving on in long trajectories.
  • Safety: Wider pre-deployment suite and third-party testing; no open-weights path.

Lab-cited evals (Grok 4.6 High vs named peers)

Eval Grok 4.6 High Grok 4.5 High GPT-5.6 Sol Max Fable 5 Max
AA Intelligence Index 61 56 61 62
GDPVal-AA v2 1753 1526 1728 1741
CursorBench v3.2 69.9% 66.7% 67.2% 70.5%
DeepSWE v1.1 65.9% 54% 73% 70%
FrontierCode v1.1 (Extended) 61.3% 56.6% 60.6% 63.6%
Terminal-Bench v3.0 26% 15.7% 34.6% 34.1%

xAI notes competitor figures are the best of self-reported or public leaderboard scores. DeepSWE and Terminal-Bench still trail Sol/Fable; the intelligence-index tie with Sol is the headline independent composite.

Independent rankings

Source Snapshot Standing
Artificial Analysis Intelligence Index Lab table + contemporaneous AA/VentureBeat writeups, Aug 12 Grok 4.6 = 61, tying GPT-5.6 Sol Max; Fable 5 Max 62 still one point ahead; +5 vs Grok 4.5 (56)
LMArena / other live boards Research window Aug 17 Not cited as ranked with a dated Elo in the primary; do not invent an Arena rank

Limits

  • No dated LMArena Elo in the primary.
  • DeepSWE and Terminal-Bench still trail Sol/Fable on the lab’s own table.
  • Fast variant at 2× price; usage pattern not yet public.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading