Models / Meta/ Muse Spark 1.1

Muse Spark 1.1

Meta · released Jul 16, 2026

Input: text, images, audio, video and documents. Output: text.InputOutput
Type
Closed
Input
$1.25
Output
$4.25
Cached
$0.15

List price · per 1M tokens · Meta at 1M context · source ↗

Our take

Written Sep 2, 2026

Muse Spark 1.1 is Meta's hosted-only multimodal model that can handle up to one million tokens in a single request. It scores strongly on mathematics and reasoning benchmarks, and carries identical pricing across its two providers.

Who should pick it

Pick this for long-document or multimodal analysis at million-token scale, or for mathematics and reasoning workloads where its benchmark scores lead. Use it for coding and web development given its strong arena scores, or if you are already in Meta's ecosystem with direct API access. Skip it if you need open weights, strong autonomous agent behaviour, or above-average steerability in live sessions.

The case for it

  • Exceptional mathematics performance on refreshed benchmark: 87.14% on LiveBench.
  • Strong reasoning on monthly-refreshed tasks: 87.73% on LiveBench.
  • One-million-token request limit for document and multimodal analysis.
  • Competitive coding scores in human preference arena: 1530.86 and 1539.69.

The case against it

  • Weak agentic coding in autonomous harness: 18.62 points below its standard coding score.
  • Below-average steerability in live agent sessions, with a negative inverse-propensity score.
  • Proprietary weights with no permissive licence available.
00

How good is it?

A general chat model for everyday questions, drafting and coding, though it can be slow to change course when you give new instructions.

Good at
  • getting answers to everyday questionsArena Text (overall) · 10th of 168
  • drafts, rewrites and editingArena Creative Writing · 14th of 168
  • writing and completing codeArena Coding · 10th of 168
Less good at
  • changing course when you give new instructionsArena Agent · Steerability · 42nd of 55

EverydayGeneral questions and everyday reasoning

4.5 of 5

Arena Text (overall)10th of 168 · 1491

Arena Hard Prompts 11th of 168Arena Maths 18th of 163LiveBench Reasoning 23rd of 58LiveBench Data Analysis 38th of 58LiveBench Mathematics 40th of 58

CodingWriting and fixing code on its own

4.5 of 5

Arena Coding10th of 168 · 1532

Arena Code (WebDev) 27th of 95LiveBench Coding 34th of 58

AgenticPlanning, calling tools, staying on task

2 of 5

Arena Agent39th of 55 · −0.039

LiveBench Agentic Coding 15th of 58

WritingDrafting and rewriting prose

4 of 5

Arena Creative Writing14th of 168 · 1465

LiveBench Language 45th of 58
How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one25th of 55
Steerabilitydoes what it was asked, and changes course when told42nd of 55
Recoverygets back on track after a command fails33rd of 55
Task outcomefinishes what the session set out to do30th of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 23rd of 168LiveBench Instruction Following 28th of 58LiveBench 30th of 58Arena Agent · Tool use 25th of 55Arena Agent · Task outcome 30th of 55Arena Agent · Recovery 33rd of 55Arena Agent · Steerability 42nd of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
75.3source ↗
58.54source ↗
77.16source ↗
72.55source ↗
69.64source ↗
74.34source ↗
87.14source ↗
87.73source ↗
−0.039source ↗
−0.018source ↗
−0.053source ↗
−0.019source ↗
0.003source ↗
1532source ↗
1465source ↗
1511source ↗
1489source ↗
1491source ↗
1541source ↗
01

Where to rent it

Prices checked 1 hour ago — each listing carries its own date.

Cheapest published offer

Meta, through OpenRouter

The lab is also the cheapest we hold. The strip above and this offer are the same one, so nothing on this page undercuts Meta on 1M of context.

per 1M tokens
$1.25 in / $4.25 out
Context served
1M
Throughput
~301 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouterOpenRouter's own listing$1.25 / $4.25checked 1 hour ago1Mnot measuredUnknownUnknownUnknown
MetaThrough OpenRouter$1.25 / $4.25checked 1 hour ago1M944K max reply301 tok/sNoYes30 daysUnknown

Across the 2 listings we hold: 1 says it does not train on prompts, 0 say they do and 1 does not say. 0 appear in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenRouterOpenRouter's own listing✓✓✓
MetaThrough OpenRouter✓✓✓

Tool calling: 2 of 2 listings say yes. JSON output: 2 of 2 listings say yes. Strict schema: 2 of 2 listings say yes.

02

Models people weigh against Muse Spark 1.1

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored −0.039 on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.018 on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.053 on Arena Agent · Steerability
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.019 on Arena Agent · Task outcome
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.003 on Arena Agent · Tool use
What movedleaderboard
Sep 25, 2026BenchmarkScored 1532 on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1465 on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1511 on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1473 on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1489 on Arena Maths
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 2 listings does not say whether it trains on prompts.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images, audio, video and documents in, text out
Catalogue slug
meta-muse-spark-1-1

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us