- Type
- Open weightsMIT License
- Params
- 358B
- Context
- 205K
active per word not recorded by us · about 154K words of context
Our take
Written Sep 2, 2026GLM 4.7 is a large downloadable text model from Z.ai with a permissive MIT licence and a 204,800-token request limit. Its measured coding skill sits well above its general chat level, and it is available from nine providers with competitive entry pricing.
Pick this for coding-heavy workloads where its Arena Coding score is the relevant signal, or for open-weights deployment at large scale with unrestricted commercial use. Use it for long-context tasks up to 204,800 tokens, or when you want provider choice at competitive rates. Skip it if creative writing quality matters most, or if you need image, audio or video input.
The case for it
- Coding is its standout skill: 43.6 points above its general chat score on the Arena leaderboard.
- Truly permissive MIT licence allows commercial use, modification and redistribution.
- Thirteen offers from nine providers, with entry pricing well under some alternatives.
- Up to 107 tokens per second from a major cloud host.
The case against it
- Creative writing is its weakest measured skill, 38.3 points below its overall chat score.
- Dense 358.3 billion parameters with no disclosed efficiency architecture.
- Throughput varies nearly ninefold by provider, so speed depends heavily on where you run it.
How good is it?
EverydayGeneral questions and everyday reasoning
Arena Text (overall)59th of 168 · 1441
CodingWriting and fixing code on its own
Arena Coding65th of 168 · 1484
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent.
WritingDrafting and rewriting prose
Arena Creative Writing63rd of 168 · 1404
Arena Creative Writing is the only board that has scored it for this.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.
Every published score for this model7 scoresEvery figure we hold, from 7 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Comfortable fit
Apple M3 Ultra (80-core GPU) · 512 GB
Room to spare. 149.3 GB spare means a 10% error in the size would not change the answer.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 57 min and 7 days ago — each listing carries its own date.
Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.
The only listing at 205K of context — the other 7 in the table below are not like-for-like. 4 cheaper rows there are outside that comparison: a different context length or a different quantisation.
- per 1M tokens
- $0.60 in / $2.20 out
- Context served
- 205K
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| DeepInfrafp4Direct and through OpenRouter | $0.40 / $1.75checked 58 min ago directchecked 13 hours ago through OpenRouter | 203K131K max reply through OpenRouter | 33 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| AtlasCloudfp8Through OpenRouter | $0.52 / $1.85checked 7 days ago | 203K182K max reply | not measured | No | Yesunknown period | Unknown |
| Venice AIfp4Through OpenRouter | $0.40 / $1.93checked 13 hours ago | 198K16K max reply | 26 tok/s | No | No | Confirmed |
| Novita AIfp8Direct and through OpenRouter | $0.60 / $2.20directchecked 58 min ago$0.54 / $1.98through OpenRouterchecked 57 min ago | 205K131K max reply through OpenRouter | 26 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Z.AIfp4Through OpenRouter | $0.60 / $2.20checked 57 min ago | 203K131K max reply | 26 tok/s | No | No | Confirmed |
| Google Vertex AIThrough OpenRouter | $0.60 / $2.20checked 57 min ago | 200K128K max reply | 35 tok/s | No | No | Confirmed |
| OpenRouterOpenRouter's own listing | $0.60 / $2.20checked 60 min ago | 205K | not measured | Unknown | Unknown | Unknown |
| Mancer 2fp4Through OpenRouter | $0.70 / $2.50checked 13 hours ago | 131K118K max reply | 21 tok/s | No | No | Confirmed |
Across the 8 listings we hold: 7 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 6 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| DeepInfrafp4Direct and through OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Venice AIfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✓ | ✗ |
| Z.AIfp4Through OpenRouter | ✓ | ✓ | ✗ |
| Google Vertex AIThrough OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Mancer 2fp4Through OpenRouter | ✗ | ✗ | ✗ |
Tool calling: 7 of 8 listings say yes, 1 says no. JSON output: 7 of 8 listings say yes, 1 says no. Strict schema: 5 of 8 listings say yes, 3 say no.
Models people weigh against GLM 4.7
When we formed this view
Recent changes
What moved
input +8% ($0.65 → $0.70 per 1M tokens)What moved
input +8% ($0.60 → $0.65 per 1M tokens)What moved
input −14% ($0.70 → $0.60 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- We do not hold the active parameter count for it, so how much of it runs on any one token is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 8 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- zai-org/GLM-4.7
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- z-ai-glm-4-7