Grok 4.20
xAI · released Mar 31, 2026
- Type
- Proprietary
- Input
- $1.25
- Output
- $2.50
- Cached
- $0.20
List price · per 1M tokens · xAI at 2M context · source ↗
Our take
Written Aug 3, 2026Grok 4.20 is xAI's flagship hosted model that accepts text, images and files across a two-million-token request limit. It scores well on coding leaderboards and offers measured speed through xAI's own API, though provider choice is narrow and every tier sits at a premium price point.
Choose this for long-document or codebase analysis at two million tokens, where the request limit is among the largest we catalogue. Pick it when Arena Coding Elo is your relevant quality signal, or if you are already in the xAI ecosystem. Skip it if you need a cheaper tier, multi-vendor resilience, or strong mathematics performance — its own scores trail its coding peak by more than 50 points.
The case for it
- Two-million-token request limit, ten times the threshold often cited for long-context work.
- Highest measured coding score of its six Arena evaluations, at 1508.1 Elo.
- Arena scores barely moved between back-to-back evaluations — Text shifted 0.13 points and Coding 0.05 points.
- Native throughput of 156.5 tokens per second at the entry price tier through xAI direct.
The case against it
- Every tracked tier is expensive, with the cheapest input rate more than double that of several frontier alternatives and the top output rate reaching six per million tokens.
- Only four offers total, three from xAI itself and one from OpenRouter at identical pricing — no independent hosts or disclosed region options.
- Mathematics and instruction-following scores lag its coding peak by 56.0 and 59.2 Elo points respectively.
How good is it?
IntelligencePuzzles, maths, exam questions
Arena Text (overall)15th of 143 · 1474.4
CodingWriting and fixing code on its own
Arena Coding25th of 143 · 1508.1
Arena Coding is the only board that has scored it for this.
AgenticPlanning, calling tools, staying on task
Nobody we watch has scored Grok 4.20 for this. We would take the rating from Arena Agent (IPS).
WritingWe do not rate this
Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where Grok 4.20 placed and give it no mark out of five.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.
Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Or rent it from someone else
Why this differs from the header. The strip above quotes xAI's own list price. This is the cheapest live offer at the widest standard context we hold, whoever is serving it — a reseller undercutting a lab is ordinary commerce, not an error.
- per 1M tokens
- $1.25 in / $2.50 out
- Context served
- 2M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouter | $1.25 / $2.50 | 2M | not measured | Unknown | Unknown | Unknown |
| xAI | $1.25 / $2.50 | 2M | 198 tok/s | No | Yes30 days | Confirmed |
| xAIpriority tier | $2.50 / $5.00 | 2M | 137 tok/s | No | Yes30 days | Unknown |
| xAI | $2.00 / $6.00 | 2M2M out | not measured | Unknown | Unknown | Unknown |
Across the 4 listings we hold: 2 say they do not train on prompts, 0 say they do and 2 do not say. 1 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
- ✓
- Supported
- ✗
- Not supported
- Not published
- host gave no parameter list
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouter | ✓ | ✓ | ✓ |
| xAI | ✓ | ✓ | ✓ |
| xAIpriority | ✓ | ✓ | ✓ |
| xAI |
Tool calling: 3 of 4 listings say yes, 1 publishes no parameter list. JSON output: 3 of 4 listings say yes, 1 publishes no parameter list. Strict schema: 3 of 4 listings say yes, 1 publishes no parameter list.
Models people weigh against Grok 4.20
When we formed this view
Dates behind this page
Prices last checked 9d ago
What we do not know about this model yet
- 1 of 4 listings publish no parameter list, so what their API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 4 listings do not say whether they train on prompts.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.
Identifiers
- Modality record
- text+image+file->text
- Catalogue slug
- x-ai-grok-4-20