Models / xAI/ Grok 4.20

Grok 4.20

xAI · released Mar 31, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Proprietary
Input
$1.25
Output
$2.50
Cached
$0.20

List price · per 1M tokens · xAI at 2M context · source ↗

Our take

Written Aug 3, 2026

Grok 4.20 is xAI's flagship hosted model that accepts text, images and files across a two-million-token request limit. It scores well on coding leaderboards and offers measured speed through xAI's own API, though provider choice is narrow and every tier sits at a premium price point.

Who should pick it

Choose this for long-document or codebase analysis at two million tokens, where the request limit is among the largest we catalogue. Pick it when Arena Coding Elo is your relevant quality signal, or if you are already in the xAI ecosystem. Skip it if you need a cheaper tier, multi-vendor resilience, or strong mathematics performance — its own scores trail its coding peak by more than 50 points.

The case for it

  • Two-million-token request limit, ten times the threshold often cited for long-context work.
  • Highest measured coding score of its six Arena evaluations, at 1508.1 Elo.
  • Arena scores barely moved between back-to-back evaluations — Text shifted 0.13 points and Coding 0.05 points.
  • Native throughput of 156.5 tokens per second at the entry price tier through xAI direct.

The case against it

  • Every tracked tier is expensive, with the cheapest input rate more than double that of several frontier alternatives and the top output rate reaching six per million tokens.
  • Only four offers total, three from xAI itself and one from OpenRouter at identical pricing — no independent hosts or disclosed region options.
  • Mathematics and instruction-following scores lag its coding peak by 56.0 and 59.2 Elo points respectively.
00

How good is it?

IntelligencePuzzles, maths, exam questions

4 of 5

Arena Text (overall)15th of 143 · 1474.4

Arena Hard Prompts 23rd of 143Arena Maths 32nd of 139

CodingWriting and fixing code on its own

4 of 5

Arena Coding25th of 143 · 1508.1

Arena Coding is the only board that has scored it for this.

AgenticPlanning, calling tools, staying on task

not measured

Nobody we watch has scored Grok 4.20 for this. We would take the rating from Arena Agent (IPS).

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where Grok 4.20 placed and give it no mark out of five.

Arena Creative Writing 12th of 143 · 1461.1
Also scored, on boards we give no mark for
Arena Instruction Following 30th of 143

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
1508.1independentsource ↗
1486.2independentsource ↗
1452independentsource ↗
1474.4independentsource ↗
01

Or rent it from someone else

Cheapest published offer

Why this differs from the header. The strip above quotes xAI's own list price. This is the cheapest live offer at the widest standard context we hold, whoever is serving it — a reseller undercutting a lab is ordinary commerce, not an error.

per 1M tokens
$1.25 in / $2.50 out
Context served
2M
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouter$1.25 / $2.502Mnot measuredUnknownUnknownUnknown
xAI$1.25 / $2.502M198 tok/sNoYes30 daysConfirmed
xAIpriority tier$2.50 / $5.002M137 tok/sNoYes30 daysUnknown
xAI$2.00 / $6.002M2M outnot measuredUnknownUnknownUnknown

Across the 4 listings we hold: 2 say they do not train on prompts, 0 say they do and 2 do not say. 1 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
OpenRouter
xAI
xAIpriority
xAI

Tool calling: 3 of 4 listings say yes, 1 publishes no parameter list. JSON output: 3 of 4 listings say yes, 1 publishes no parameter list. Strict schema: 3 of 4 listings say yes, 1 publishes no parameter list.

02

Models people weigh against Grok 4.20

03

When we formed this view

Dates behind this page

Aug 2, 2026BenchmarkScored 1508.1 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1461.1 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1486.2 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1448.9 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1452 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1474.4 on Arena Text (overall)leaderboard
Jul 27, 2026Benchmark updategrok-4.20-beta1 enters LMArena at 1474 Elo26598 votes
Jul 26, 2026ListedListed on LLMapfirst indexed by our pipeline
Mar 31, 2026AnnouncedGrok 4.20 announced by xAI

Prices last checked 9d ago

What we do not know about this model yet

  • 1 of 4 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 2 of 4 listings do not say whether they train on prompts.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.

Identifiers

Modality record
text+image+file->text
Catalogue slug
x-ai-grok-4-20

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us