Tools
Compare models
Add up to four models via ?models=slug-a,slug-b — or pick from any model page.
| Spec | OpenAI |
|---|---|
| Context | 1.1M |
| License | |
| Cheapest hosted (out/1M) | $30.00 · OpenAI |
| Released | Apr 24, 2026 |
| LiveBench | 80.19 |
| LiveBench Agentic Coding | 53.99 |
| LiveBench Coding | 82.15 |
| LiveBench Data Analysis | 81.58 |
| LiveBench Instruction Following | 70.73 |
| LiveBench Language | 87.36 |
| LiveBench Mathematics | 95.86 |
| LiveBench Reasoning | 89.65 |
| Arena Agent | 20th of 55 |
| Arena Agent · Recoverygets back on track after a command fails | 17th of 55 |
| Arena Agent · Steerabilitydoes what it was asked, and changes course when told | 19th of 55 |
| Arena Agent · Task outcomefinishes what the session set out to do | 35th of 55 |
| Arena Agent · Tool usereaches for the right one, and does not invent one | 2nd of 55 |
| Arena Coding | 1513 |
| Arena Creative Writing | 1456 |
| Arena Hard Prompts | 1499 |
| Arena Instruction Following | 1475 |
| Arena Maths | 1500 |
| Arena Text (overall) | 1477 |
| Arena Code (WebDev) | 1512 |
| SWE-bench Verified | 74.4 |