Models / Moonshot AI/ Kimi K2.5

Kimi K2.5

Moonshot AI · released Jan 1, 2026 · moonshotai/Kimi-K2.5

Input: text and images. Output: text.InputOutput
Type
Open weightsCustom licence
Params
1.1T
Context
262K

32B active per word · about 197K words of context · download allowed, licence restricts use

Our take

Written Aug 2, 2026

Kimi K2.5 is a large mixture-of-experts model from Moonshot AI that handles text and images with up to 262,144 tokens in a single request. It is built for long-document work and coding, with measured coding scores well above its general chat rating.

Who should pick it

Choose this for long-context document analysis and coding tasks where its 262,144-token limit and strong measured coding score matter. Use it for image-plus-text workflows that need frontier-scale parameter capacity, or when you want to shop across ten providers for the best rate. Skip it if you need a permissive licence, if creative writing is your main workload, or if you need output pricing under two dollars per million tokens.

The case for it

  • Strong measured coding performance: its coding score is 73.3 points above its general text rating on the Arena leaderboard.
  • Extremely large request limit for the active parameter count — 262,144 tokens with 32 billion active per forward pass.
  • Ten current offers with a 2.4× price spread, giving room to optimise for budget.
  • Fastest confirmed option runs at 44 tokens per second, 2.2× the slowest tracked host.

The case against it

  • Custom restricted licence, not a permissive standard like Apache or MIT, which limits commercial flexibility.
  • Creative writing lags other capability areas by 115.3 points on the measured leaderboard.
  • No provider achieves above 44 tokens per second, and two hosts have unverified throughput in our data.
00

How good is it?

IntelligencePuzzles, maths, exam questions

3 of 5

Arena Text (overall)53rd of 143 · 1431.6

Arena Hard Prompts 46th of 143Arena Maths 40th of 139

CodingWriting and fixing code on its own

4 of 5

Arena Coding31st of 143 · 1504.7

Arena Code (WebDev) 39th of 74

AgenticPlanning, calling tools, staying on task

Scored, not ratedSWE-bench Verified · 10th of 39 · 70.8via mini-SWE-agent

Kimi K2.5 is not on Arena Agent (IPS), which is where the rating would come from, so there is no rating here. It is on SWE-bench Verified, in 10th of 39 with 70.8.

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where Kimi K2.5 placed and give it no mark out of five.

Arena Creative Writing 62nd of 143 · 1389.4
Also scored, on boards we give no mark for
Arena Instruction Following 42nd of 143

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
1504.7independentsource ↗
1461.3independentsource ↗
1440.7independentsource ↗
1431.6independentsource ↗
1404.8independentsource ↗
70.8via mini-SWE-agentindependentsource ↗
01

Can you run it yourself?

Fits in memory
weights load entirely on the card
Spills to system RAM
some weights offload; much slower
Too large
will not load even with offload
est
size is calculated; the verdict could change by 10%
A card many people ownToo large

GeForce RTX 4090 · 24 GB

Weights at Q4_K_M667.5 / 24 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step upToo large

Apple M1 Pro (16-core GPU) · 32 GB

Weights at Q4_K_M667.5 / 32 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step downToo large

Radeon RX 7900 XT · 20 GB

Weights at Q4_K_M667.5 / 20 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

Q4_K_M
recommended
667.5 GBest
Too large
Q5_K_M
783.1 GBest
Too large
Q8_0
1169.8 GBest
Too large

All 3 sizes here are calculated, not measured. We hold no measured file for this model, so each size comes from the parameter count and every verdict above inherits that uncertainty.

Check against your own machine → · All 71 devices, with every size →

02

Or rent it from someone else

Cheapest published offer

Cheapest of 14 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per 1M tokens
$0.57 in / $2.85 out
Context served
262K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
Chutesint4$0.44 / $2.00262K20 tok/sNoYesunknown periodUnknown
DigitalOcean Gradient$0.38 / $2.02262K10 tok/sNoNoConfirmed
DeepInfrafp4$0.45 / $2.25262K26 tok/sNoNoConfirmed
DeepInfrafp4$0.45 / $2.25262Knot measuredUnknownUnknownUnknown
SiliconFlowint4$0.45 / $2.25262K55 tok/sNoNoConfirmed
AtlasCloudint4$0.49 / $2.50262K30 tok/sNoYesunknown periodUnknown
StreamLakefp8$0.54 / $2.70256K29 tok/sNoYesunknown periodUnknown
OpenRouter$0.57 / $2.85262Knot measuredUnknownUnknownUnknown
Novita AI$0.57 / $2.85262K26 tok/sNoNoConfirmed
Phala$0.60 / $3.00262K36 tok/sNoNoConfirmed
Novita AI$0.60 / $3.00262Knot measuredUnknownUnknownUnknown
Amazon Bedrockus-east-2$0.60 / $3.00262K65 tok/sNoNoConfirmed
Moonshot AIint4$0.60 / $3.00262K30 tok/sNoNoConfirmed
Venice AI$0.56 / $3.50256K26 tok/sNoNoConfirmed

Across the 14 listings we hold: 11 say they do not train on prompts, 0 say they do and 3 do not say. 8 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
Chutesint4
DigitalOcean Gradient
DeepInfrafp4
DeepInfrafp4
SiliconFlowint4
AtlasCloudint4
StreamLakefp8
OpenRouter
Novita AI
Phala
Novita AI
Amazon Bedrockus-east-2
Moonshot AIint4
Venice AI

Tool calling: 12 of 14 listings say yes, 2 publish no parameter list. JSON output: 11 of 14 listings say yes, 1 says no, 2 publish no parameter list. Strict schema: 10 of 14 listings say yes, 2 say no, 2 publish no parameter list.

03

Models people weigh against Kimi K2.5

04

When we formed this view

Dates behind this page

Aug 2, 2026BenchmarkScored 1504.7 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1389.4 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1461.3 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1434.4 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1440.7 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1431.6 on Arena Text (overall)leaderboard
Aug 2, 2026BenchmarkScored 1404.8 on Arena Code (WebDev)leaderboard
Jul 28, 2026Price changenovita repriced moonshotai/kimi-k2.5input $0.57 → $0.6, output $2.85 → $3 per 1M tokens
Jul 27, 2026Benchmark updatekimi-k2.5-instant enters LMArena at 1431 Elo8151 votes
Jul 27, 2026Price changenovita repriced moonshotai/kimi-k2.5input $0.57 → $0.6, output $2.85 → $3 per 1M tokens

Prices last checked 14h ago

What we do not know about this model yet

  • We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
  • 2 of 14 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 14 listings do not say whether they train on prompts.
05

Licence and identifiers

What the licence allowsCustom licence, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.

Licence

Custom licence

restricted_openCustom licence — review the terms

This model ships custom license terms that don't map to a known template. We haven't parsed them, so commercial use, redistribution and derivatives are unverified — review the original terms before shipping.

Identifiers

Architecture
Mixture of experts
Modality record
text+image->text
Catalogue slug
moonshotai-kimi-k2-5

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us