Models / xAI/ Grok 4.6

Grok 4.6

xAI · released Aug 12, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Closed
Input
$2.00
Output
$6.00
Cached
None held

List price · per 1M tokens · xAI at 500K context · machine-readable source ↗

Our take

Written Sep 17, 2026

Grok 4.6 is a hosted-only model from xAI, so using it means choosing a host rather than running it yourself. Its measured strengths sit in mathematics and reasoning, while its agentic task-outcome score falls below the board's neutral point.

Who should pick it

Reach for it on mathematical and reasoning-heavy work, where its measured scores are strongest, or for long-document analysis where the request capacity spares you from splitting files first. It also takes images and files alongside text. Skip it if you need to run the model yourself or need clear licence terms, or if you need an agent that reliably finishes multi-step tasks.

The case for it

  • Among the strongest on LiveBench Mathematics and Reasoning, at 92.57% and 90.51% — averages over competition-style and monthly-refreshed task sets.
  • The request capacity is large enough that long documents need not be split up first, though reliable recall across all of it is unverified in our data.
  • Text, images and files go into the same request, so a screenshot or a document does not have to be transcribed first.

The case against it

  • Arena Agent task outcome is negative, meaning it finished the session's task less often than the board's neutral point, while recovery, steerability and tool use are positive but small.
  • We list no download for it, so using it means choosing a host, and no licence terms are disclosed in our data.
  • Agentic coding is much weaker than its other coding measures: 57.02% on LiveBench Agentic Coding, measured inside an agent harness, against 76.78% on LiveBench Coding.
00

How good is it?

A closed text model from xAI for drafting prose, writing code and calling tools to carry out requests.

Good at
  • drafts, rewrites and editingArena Creative Writing · 33rd of 168
  • writing and completing codeArena Coding · 41st of 168
  • calling tools to carry out requestsArena Agent · Tool use · 9th of 55

EverydayGeneral questions and everyday reasoning

3.5 of 5

Arena Text (overall)47th of 168 · 1453

Arena Hard Prompts 42nd of 168Arena Maths 54th of 163LiveBench Reasoning 9th of 58LiveBench Mathematics 21st of 58LiveBench Data Analysis 33rd of 58

CodingWriting and fixing code on its own

4 of 5

Arena Coding41st of 168 · 1507

Arena Code (WebDev) 14th of 95LiveBench Coding 36th of 58

AgenticPlanning, calling tools, staying on task

2.5 of 5

Arena Agent23rd of 55 · 0.012

LiveBench Agentic Coding 19th of 58

WritingDrafting and rewriting prose

3.5 of 5

Arena Creative Writing33rd of 168 · 1445

LiveBench Language 15th of 58
How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one9th of 55
Steerabilitydoes what it was asked, and changes course when told16th of 55
Recoverygets back on track after a command fails14th of 55
Task outcomefinishes what the session set out to do31st of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 39th of 168LiveBench 15th of 58LiveBench Instruction Following 18th of 58Arena Agent · Tool use 9th of 55Arena Agent · Recovery 14th of 55Arena Agent · Steerability 16th of 55Arena Agent · Task outcome 31st of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
78.04source ↗
57.02source ↗
76.78source ↗
73.86source ↗
83.7source ↗
92.57source ↗
90.51source ↗
0.012source ↗
0.058source ↗
0.035source ↗
−0.033source ↗
0.004source ↗
1507source ↗
1445source ↗
1478source ↗
1451source ↗
1443source ↗
1453source ↗
1621source ↗
01

Where to rent it

Prices checked between 58 min and 1 hour ago — each listing carries its own date.

Cheapest published offer

xAI, direct and through OpenRouter

The lab is also the cheapest we hold. The strip above and this offer are the same one, so nothing on this page undercuts xAI on 500K of context.

per 1M tokens
$2.00 in / $6.00 out
Context served
500K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouterOpenRouter's own listing$2.00 / $6.00checked 1 hour ago500Knot measuredUnknownUnknownUnknown
xAIDirect and through OpenRouter$2.00 / $6.00checked 58 min ago directchecked 59 min ago through OpenRouter500K500K max reply direct450K max reply through OpenRouter53 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterYes30 daysDirectUnknownThrough OpenRouterConfirmed
xAIusThrough OpenRouter$2.20 / $6.60checked 59 min ago500K450K max reply50 tok/sNoYes30 daysConfirmed
Amazon Bedrockus-west-2Through OpenRouter$2.20 / $6.60checked 59 min ago500K450K max reply94 tok/sNoNoConfirmed
xAIpriority tierThrough OpenRouter$4.00 / $12.00checked 59 min ago500K450K max reply70 tok/sNoYes30 daysConfirmed

Across the 5 listings we hold: 4 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 1 does not say. 4 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenRouterOpenRouter's own listing✓✓✓
xAIDirect and through OpenRouter✓✓✓
xAIusThrough OpenRouter✓✓✓
Amazon Bedrockus-west-2Through OpenRouter✓✓✓
xAIpriorityThrough OpenRouter✓✓✓

Tool calling: 5 of 5 listings say yes. JSON output: 5 of 5 listings say yes. Strict schema: 5 of 5 listings say yes.

02

Models people weigh against Grok 4.6

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored 0.012 via xHigh on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.058 via xHigh on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.035 via xHigh on Arena Agent · Steerability
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.033 via xHigh on Arena Agent · Task outcome
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.004 via xHigh on Arena Agent · Tool use
What movedleaderboard
Sep 25, 2026BenchmarkScored 1507 via High on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1445 via High on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1478 via High on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1451 via High on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1443 via High on Arena Maths
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 5 listings does not say whether it trains on prompts, and 1 answers only through OpenRouter, not for its own listing.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and documents in, text out
Catalogue slug
x-ai-grok-4-6

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us