Models / xAI/ Grok 4.20 Multi-Agent

Grok 4.20 Multi-Agent

xAI · released Mar 31, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Closed
Input
$1.25
Output
$2.50
Cached
None held

List price · per 1M tokens · xAI at 1M context · machine-readable source ↗

Our take

Written Sep 4, 2026

Grok 4.20 Multi-Agent is xAI's flagship text model with a two-million-token request limit and built-in multi-agent tooling. It is a premium pick for coding and long-document work, though its creative and instruction-following scores trail its stronger suits.

Who should pick it

Choose this for coding assistance where measured preference peaks, or for analysing documents up to two million tokens in a single request. Use the standard tier for throughput; the premium tier costs double for slower speed. Skip it if you need open weights, if creative writing or precise instruction-following are central, or if you want measured quality data beyond Arena scores.

The case for it

  • Two-million-token request limit — ten times the 200,000 common among frontier models we list.
  • Arena-measured coding preference leads its own skill profile by 37.7 points over its overall text score.
  • Standard tier pricing is consistent across OpenRouter and xAI's own offer.

The case against it

  • Creative writing and instruction following lag its stronger suits by more than 20 points each.
  • Premium tier charges double for 19.5% slower throughput — a poor speed-for-price trade.
  • No MMLU, GPQA or SWE-bench scores in our data; quality picture is Arena-only.
00

How good is it?

A closed text model from xAI for everyday questions, drafting and coding, with no weak spots flagged on the board.

Good at
  • answering everyday questionsArena Text (overall) · 29th of 168
  • drafting and editing proseArena Creative Writing · 31st of 168
  • writing and completing codeArena Coding · 39th of 168

EverydayGeneral questions and everyday reasoning

4 of 5

Arena Text (overall)29th of 168 · 1470

Arena Hard Prompts 37th of 168Arena Maths 42nd of 163

CodingWriting and fixing code on its own

4 of 5

Arena Coding39th of 168 · 1508

Arena Coding is the only board that has scored it for this.

AgenticPlanning, calling tools, staying on task

not measured

Not yet scored on Arena Agent.

WritingDrafting and rewriting prose

3.5 of 5

Arena Creative Writing31st of 168 · 1447

Arena Creative Writing is the only board that has scored it for this.

Other boards it appears on
Arena Instruction Following 49th of 168

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.

Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
1508source ↗
1447source ↗
1483source ↗
1455source ↗
1470source ↗
01

Where to rent it

Prices checked between 1 hour and 35 days ago — each listing carries its own date.

Cheapest published offer

xAI, through OpenRouter

Why this differs from the header. The strip above quotes xAI's own list price; this is xAI through OpenRouter, which lists it at a lower price.

per 1M tokens
$1.25 in / $2.50 out
Context served
2M
Throughput
~349 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouterOpenRouter's own listing$1.25 / $2.50checked 1 hour ago2Mnot measuredUnknownUnknownUnknown
xAIDirect$1.25 / $2.50checked 35 days ago1M1M max replynot measuredUnknownUnknownUnknown
xAIThrough OpenRouter$1.25 / $2.50checked 1 hour ago2M1.8M max reply349 tok/sNoYes30 daysConfirmed
xAIpriority tierThrough OpenRouter$2.50 / $5.00checked 1 hour ago2M1.8M max reply333 tok/sNoYes30 daysConfirmed

Across the 4 listings we hold: 2 say they do not train on prompts, 0 say they do and 2 do not say. 2 appear in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenRouterOpenRouter's own listing✗✓✓
xAIDirect
xAIThrough OpenRouter✗✓✓
xAIpriorityThrough OpenRouter✗✓✓

Tool calling: 0 of 4 listings say yes, 3 say no, 1 publishes no parameter list. JSON output: 3 of 4 listings say yes, 1 publishes no parameter list. Strict schema: 3 of 4 listings say yes, 1 publishes no parameter list.

02

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored 1508 on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1447 on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1483 on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1444 on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1455 on Arena Maths
What movedleaderboard
Sep 25, 2026BenchmarkScored 1470 on Arena Text (overall)
What movedleaderboard
Aug 17, 2026Price changexAI cut Grok 4.20 Multi-Agent output pricing by 58% · machine-readable source ↗
What movedinput −38% ($2.00 → $1.25 per 1M tokens), output −58% ($6.00 → $2.50 per 1M tokens)
Jul 26, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Mar 31, 2026AnnouncedGrok 4.20 Multi-Agent announced by xAI

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • 1 of 4 listings publishes no parameter list, so what its API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 2 of 4 listings do not say whether they train on prompts.
  • We hold no batch or off-peak rate for any of its listings.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and documents in, text out
Catalogue slug
x-ai-grok-4-20-multi-agent

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us