Models / Qwen/ Qwen3.8 Flash

Qwen3.8 Flash

Qwen · released Aug 26, 2026

Input: text, images and video. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — we have no record of published weights

Our take

Written Sep 27, 2026

Qwen3.8 Flash is a hosted-only model with a lopsided profile: it follows instructions very well and builds web apps well, but it is a weak agent. We list no download for it, so using it means choosing a host.

Who should pick it

Use it for constrained rewriting and format-following work, where it places 4th of 51 on LiveBench Instruction Following as of 25 Jun 2026, or for building web apps, where it places 9th of 95 on Arena Code (WebDev) as of 25 Sep 2026. Skip it if you need an agent that reliably calls the right tool, if you want to run the model yourself, or if you need licence terms you can build a product on.

The case for it

  • Among the strongest here at following instructions: 77.11% on LiveBench Instruction Following, 4th of 51 as of 25 Jun 2026, a board of constrained-rewriting tasks rather than open-ended judgement.
  • Strong on set-piece maths and reasoning: 87.38% on LiveBench Reasoning and 85.82% on LiveBench Mathematics as of 25 Jun 2026, competition-style questions rather than reasoning about your own codebase.
  • Good at building web apps from a prompt: 9th of 95 on Arena Code (WebDev) as of 25 Sep 2026, a board of human votes that records which result people preferred, not whether the app was correct.
  • Pictures and clips go in with the question, and the request capacity takes a long report or a stack of documents beside it, though reliable recall across all of it is unverified in our data.

The case against it

  • Weak at agentic tool use: 45th of 55 on Arena Agent · Tool use as of 25 Sep 2026, a board scoring whether the model calls the right tool and does not invent one, so it is a poor fit for tool-calling pipelines.
  • Mid-to-lower field on general coding and language: 38th of 51 on LiveBench Coding and 37th of 51 on LiveBench Language as of 25 Jun 2026, against its 4th of 51 on Instruction Following.
  • We list no download for it, so using it means choosing a host, and no licence is supplied, so nothing here tells you what you may do with its output.
00

How good is it?

A text model for chat and everyday questions, though it trails most models at calling tools to carry out requests.

Less good at
  • calling tools to carry out requestsArena Agent · Tool use · 45th of 55

EverydayGeneral questions and everyday reasoning

Scored, not ratedLiveBench Reasoning · 24th of 58 · 87.38

Not yet scored on Arena Text (overall). It is on LiveBench Reasoning, in 24th of 58 with 87.38.

LiveBench Data Analysis 32nd of 58LiveBench Mathematics 44th of 58

CodingWriting and fixing code on its own

Scored, not ratedArena Code (WebDev) · 9th of 95 · 1636

Not yet scored on Arena Coding. It is on Arena Code (WebDev), in 9th of 95 with 1636.

LiveBench Coding 45th of 58

AgenticPlanning, calling tools, staying on task

2.5 of 5

Arena Agent29th of 55 · −0.005

LiveBench Agentic Coding 11th of 58

WritingDrafting and rewriting prose

Scored, not ratedLiveBench Language · 43rd of 58 · 74.64

Not yet scored on Arena Creative Writing. It is on LiveBench Language, in 43rd of 58 with 74.64.

How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one45th of 55
Steerabilitydoes what it was asked, and changes course when told37th of 55
Recoverygets back on track after a command fails35th of 55
Task outcomefinishes what the session set out to do12th of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
LiveBench Instruction Following 4th of 58LiveBench 24th of 58Arena Agent · Task outcome 12th of 55Arena Agent · Recovery 35th of 55Arena Agent · Steerability 37th of 55Arena Agent · Tool use 45th of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model14 scoresEvery figure we hold, from 14 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
76.19source ↗
61.62source ↗
72.55source ↗
74.24source ↗
74.64source ↗
85.82source ↗
87.38source ↗
−0.005source ↗
−0.026source ↗
−0.037source ↗
0.065source ↗
−0.009source ↗
1636source ↗
01

Where to rent it

Prices checked 2 hours ago — each listing carries its own date.

Cheapest published offer

DeepInfra, direct

Cheapest of 4 live listings.

per 1M tokens
$0.11 in / $0.38 out
Context served
1M
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
DeepInfraDirect$0.11 / $0.38checked 2 hours ago1Mnot measuredUnknownUnknownUnknown
OpenRouterOpenRouter's own listing$0.15 / $0.47checked 2 hours ago1Mnot measuredUnknownUnknownUnknown
Novita AIDirect$0.15 / $0.47checked 2 hours ago1Mnot measuredUnknownUnknownUnknown
Alibaba CloudThrough OpenRouter$0.15 / $0.47checked 2 hours ago1M131K max reply45 tok/sNoYesunknown periodUnknown

Across the 4 listings we hold: 1 says it does not train on prompts, 0 say they do and 3 do not say. 0 appear in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
DeepInfraDirect
OpenRouterOpenRouter's own listing✓✓✓
Novita AIDirect
Alibaba CloudThrough OpenRouter✓✓✓

Tool calling: 2 of 4 listings say yes, 2 publish no parameter list. JSON output: 2 of 4 listings say yes, 2 publish no parameter list. Strict schema: 2 of 4 listings say yes, 2 publish no parameter list.

02

Models people weigh against Qwen3.8 Flash

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored −0.005 on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.026 on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.037 on Arena Agent · Steerability
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.065 on Arena Agent · Task outcome
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.009 on Arena Agent · Tool use
What movedleaderboard
Sep 25, 2026BenchmarkScored 1636 on Arena Code (WebDev)
What movedleaderboard
Aug 27, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Aug 26, 2026ReleaseQwen3.8 Flash listed
Aug 26, 2026AnnouncedQwen3.8 Flash announced by Qwen
Jun 25, 2026BenchmarkScored 76.19 on LiveBench
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • 2 of 4 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 4 listings do not say whether they train on prompts.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and video in, text out
Catalogue slug
qwen-qwen3-8-flash

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us