Hy4 preview
Tencent · released Aug 27, 2026 · tencent/Hy4-preview
- Type
- Open weightsApache License 2.0
- Params
- 780B
- Context
- 1M
active per word not recorded by us · about 786K words of context
Our take
Written Sep 29, 2026Hy4 preview is a downloadable text model from Tencent, released in August 2026, whose measured strength is agent work: it places 10th of 55 on Arena Agent · Task outcome and 10th of 55 on Arena Agent · Recovery as of 25 Sep 2026. Its licence allows commercial use, changes and redistribution (Apache License 2.0), and it is also competitive on web-app building.
Pick it for agent sessions where the job is to finish a task and get back on track after a failed command, or for building web apps, where it places 10th of 95 on Arena Code (WebDev) as of 25 Sep 2026. The licence allows commercial use, changes and redistribution, so it will not need a legal review before you build on it. Skip it if you need reliable tool selection, or if you need a published parameter count to size your hardware.
The case for it
- 10th of 55 on Arena Agent · Task outcome and 10th of 55 on Arena Agent · Recovery as of 25 Sep 2026, so it is a reasonable first trial for sessions that have to reach a goal and recover from a failed command.
- 10th of 95 on Arena Code (WebDev) as of 25 Sep 2026, a board scored by human votes on web-app building tasks, which makes it worth trying on a small app before you commit.
- The request capacity takes long documents without splitting them up first, though reliable recall across all of it is unverified in our data.
- The licence allows commercial use, changes and redistribution (Apache License 2.0).
The case against it
- 36th of 55 on Arena Agent · Tool use as of 25 Sep 2026, a board measuring whether it calls the right tool and does not invent one, so check its tool calls before trusting a long chain.
- 26th of 55 on Arena Agent · Steerability as of 25 Sep 2026, so expect to restate a change of course rather than assume it lands first time.
- No parameter count is published, so you cannot estimate memory requirements from the model card.
How good is it?
An open text model built for multi-step agent work and getting back on track when a step fails.
- carrying out multi-step tasks for youArena Agent · 11th of 55
- getting back on track after a step failsArena Agent · Recovery · 10th of 55
EverydayGeneral questions and everyday reasoning
Not yet scored on Arena Text (overall).
CodingWriting and fixing code on its own
Not yet scored on Arena Coding. It is on Arena Code (WebDev), in 10th of 95 with 1631.
AgenticPlanning, calling tools, staying on task
Arena Agent11th of 55 · 0.044
Arena Agent is the only board that has scored it for this.
WritingDrafting and rewriting prose
Not yet scored on Arena Creative Writing.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 2 hours and 20 hours ago — each listing carries its own date.
- per 1M tokens
- $0.75 in / $2.25 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouterOpenRouter's own listing | $0.75 / $2.25checked 2 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| DeepInfrafp8Direct and through OpenRouter | $0.83 / $2.50checked 2 hours ago | 1M131K max reply through OpenRouter | 17 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Novita AIfp8Direct and through OpenRouter | $0.83 / $2.50checked 2 hours ago directchecked 20 hours ago through OpenRouter | 1M64K max reply through OpenRouter | 31 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| SiliconFlowfp8Through OpenRouter | $0.83 / $2.50checked 2 hours ago | 1M262K max reply | 38 tok/s | No | No | Confirmed |
| Tencentfp8Through OpenRouter | $0.83 / $2.50checked 8 hours ago | 1M64K max reply | 44 tok/s | No | No | Confirmed |
Across the 5 listings we hold: 4 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 4 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✗ | ✓ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Tencentfp8Through OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 5 of 5 listings say yes. JSON output: 4 of 5 listings say yes, 1 says no. Strict schema: 5 of 5 listings say yes.
Models people weigh against Hy4 preview
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- We do not hold the active parameter count for it, so how much of it runs on any one token is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 5 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsApache License 2.0, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Apache License 2.0
Fully permissive: commercial use, redistribution, and derivatives allowed. Requires attribution and a copy of the license. Includes an express patent grant.
Identifiers
- Hugging Face
- tencent/Hy4-preview
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- tencent-hy4-preview