Tools
Compare models
Add up to four models via ?models=slug-a,slug-b — or pick from any model page.
| Spec | Anthropic |
|---|---|
| Context | 1M |
| License | |
| Cheapest hosted (out/1M) | $25.00 · Anthropic |
| Released | Feb 4, 2026 |
| LiveBench | 74.52 |
| LiveBench Agentic Coding | 48.99 |
| LiveBench Coding | 78.18 |
| LiveBench Data Analysis | 69.89 |
| LiveBench Instruction Following | 63.31 |
| LiveBench Language | 83.27 |
| LiveBench Mathematics | 89.32 |
| LiveBench Reasoning | 88.67 |
| Arena Agent | 9th of 55 |
| Arena Agent · Recoverygets back on track after a command fails | 4th of 55 |
| Arena Agent · Steerabilitydoes what it was asked, and changes course when told | 11th of 55 |
| Arena Agent · Task outcomefinishes what the session set out to do | 22nd of 55 |
| Arena Agent · Tool usereaches for the right one, and does not invent one | 2nd of 55 |
| Arena Coding | 1548 |
| Arena Creative Writing | 1478 |
| Arena Hard Prompts | 1527 |
| Arena Instruction Following | 1500 |
| Arena Maths | 1507 |
| Arena Text (overall) | 1498 |
| Arena Code (WebDev) | 1537 |
| SWE-bench Verified | 75.6 |