Feed

News

Scanned daily from vendor feeds, Hugging Face and provider APIs. New releases land here within a day or two of announcement.

You're reading the changes — get them weekly.
Every Thursday: releases, price moves and new hardware, straight from this feed.

Weekly digest of model releases and price moves. No noise.

We do not yet track model retirements or licence changes on this feed — no event of either kind has ever been recorded. That is a gap in our coverage, not a sign none have happened.

1 October 2026
PRICEHost Mancer 2 cut gpt-oss-120b input pricing by 10%input: ↓ 10%input −10% $0.050 → $0.045 · output −8% $0.300 → $0.275prices per 1M tokensIf you buy: existing users of this model at this host can now run larger prompt-heavy jobs for the same spend, so re-check your usage limits and quota before the next top-up.storyRELEASEPareto 26.10 Preview listedIf you buy: A new option to test on this host, so check whether its context window and rate limits fit your workload before you point production traffic at it.modelPRICEGLM 5.3 Flash repriced across 2 hosts: InferenceNet input down 50%, Wafer output up 50%Wafer: output ↑ 50% $0.50 → $0.75 InferenceNet: input ↓ 50% $0.100 → $0.050 · output ↑ 33% $0.45 → $0.60 prices per 1M tokensIf you buy: On Wafer, this model now costs more to run, so check whether your workload leans on output before your next top-up.storyPRICEHost Relace raised GLM 5.3 input and cache-read pricing by 12%input and cache read: ↑ 12%input +12% $0.130 → $0.145 · cache read +12% $0.130 → $0.145prices per 1M tokensIf you buy: Budget for a higher input and cache-read cost on your next top-up here, and re-check whether your workload's cache-heavy pattern still fits this host.storyPRICEGLM 5.2 repriced across 5 hosts: Wafer input down 71%, Inceptron input up 85%Inceptron: input ↑ 85% $0.75 → $1.39 Wafer: input ↓ 71% $1.40 → $0.41 Reka: output ↓ 68% $4.40 → $1.40 Morph: input ↑ 20% $0.387 → $0.464 Relace: input and cache read ↓ 3% (input $0.135 → $0.131) prices per 1M tokensIf you buy: Inceptron now costs more to run for the same model, so re-check your budget before the next top-up and weigh the other hosts listed beside this item.storyPRICEHost Ionstream cut Qwen3.8 27B input pricing by 55%input: ↓ 55%input −55% $0.198 → $0.089 · output −2% $2.55 → $2.50prices per 1M tokensIf you buy: Input-heavy work on this model at Ionstream now costs far less to run, so long-prompt jobs become worth revisiting; re-check your output spend, which barely moved.storyPRICEKimi K3 repriced across 6 hosts: Relace input down 68%, Sail Research input up 781%Sail Research: input ↑ 781% $0.318 → $2.800 · cache read ↓ 25% $0.40 → $0.30 InferenceNet: input ↑ 526% $0.19 → $1.19 Morph: output ↓ 28% $13.55 → $9.81 · input ↑ 86% $1.17 → $2.18 Relace: input ↓ 68% $2.10 → $0.67 · cache read ↑ 219% $0.21 → $0.67 Wafer: input ↓ 4% $1.442 → $1.390 · output ↑ 56% $9.00 → $14.00 Together: all rates ↓ 10% (output $15.00 → $13.50) Wafer (US region): cache read ↑ 7% $0.28 → $0.30 prices per 1M tokensstoryPRICEGemma 4 31B rose across 2 hosts, by up to 67% at DekaLLM (input)DekaLLM: input ↑ 67% $0.060 → $0.100 DeepInfra through OpenRouter (fp8): input ↑ 15% $0.13 → $0.15 DeepInfra's own listing (fp8): input ↑ 15% $0.13 → $0.15 prices per 1M tokensIf you buy: Budget for a higher input cost on this host before your next top-up, and re-check whether your workload's prompt-heavy mix still fits the spend you planned.storyPRICEDeepSeek V4.1 Flash repriced across 8 hosts: DigitalOcean down 40% on all rates, Wafer output up 36%DigitalOcean: all rates ↓ 40% (output $1.20 → $0.72) Wafer: input ↓ 33% $0.075 → $0.050 · output ↑ 36% $0.44 → $0.60 Io Net: input ↑ 33% $0.090 → $0.120 Reka: output ↓ 33% $0.84 → $0.56 · cache read ↑ 76% $0.0080 → $0.0141 Morph: input ↓ 32% $0.078 → $0.053 Relace: input and cache read ↑ 32% (input $0.0200 → $0.0264) AtlasCloud: all rates ↑ 24% (output $0.456 → $0.564) OpenInference: input ↓ 22% $0.0198 → $0.0155 prices per 1M tokensIf you buy: On Wafer, output costs more while input and cache reads cost less, so re-check your budget if you generate long answers, and note prompt-heavy jobs now balance differently.storyPRICEDeepSeek V4 Pro repriced across 3 hosts: Reka input and cache read down 55%, Relace input and cache read up 36%Reka: input and cache read ↓ 55% (input $1.30 → $0.58) Relace: input and cache read ↑ 36% (input $0.123 → $0.167) Wafer: input and cache read ↑ 15% (input $0.48 → $0.55) prices per 1M tokensIf you buy: Relace now costs more to run for input-heavy and cache-reliant workloads, so re-check your usage mix and whether this host still fits before your next top-up.storyPRICEDeepSeek V4 Flash repriced across 3 hosts: StreamLake output down 79%, Wafer output up 100%Wafer: output ↑ 100% $0.35 → $0.70 · cache read ↓ 6% $0.053 → $0.050 StreamLake: output ↓ 79% $1.32 → $0.28 · cache read ↑ 100% $0.014 → $0.028 Relace: input and cache read ↓ 50% (input $0.0090 → $0.0045) prices per 1M tokensIf you buy: On Wafer, running this model now costs more per token, so check your usage mix and cache settings before your next top-up, and confirm the current rate on the host page.story
30 September 2026
PRICEHost Phala cut Qwen3.6 27B cache-read pricing by 80%cache read: ↓ 80%cache read −80% $0.150 → $0.030prices per 1M tokens$0.150 was · cache read per 1M tokens · Phala$0.030 nowstoryPRICEHost Inceptron raised Kimi K2.7 Code input and output pricing by 2%input and output: ↑ 2%input +2% $0.66 → $0.67 · output +2% $3.30 → $3.35prices per 1M tokensstoryPRICEDeepSeek V4 Flash repriced across 2 hosts: Relace input and cache read down 50%, output up 300%Relace: input and cache read ↓ 50% (input $0.018 → $0.009) · output ↑ 300% $0.32 → $1.28 Inceptron: input ↓ 11% $0.056 → $0.050 prices per 1M tokensIf you buy: Relace now suits chatty, cache-heavy workloads far better than long generations, so re-check your output volume before your next top-up.storyPRICEGLM 5.3 Flash repriced across 5 hosts: OpenInference output down 18%, InferenceNet output up 61%InferenceNet: output ↑ 61% $0.28 → $0.45 Inceptron: input ↑ 50% $0.150 → $0.225 Relace: input and cache read ↑ 40% (input $0.025 → $0.035) OpenInference: output ↓ 18% $0.300 → $0.247 Morph: input and output ↓ 10% (output $0.70 → $0.63) prices per 1M tokensIf you buy: Output-heavy workloads on this model here now cost more per call, so re-check your spend cap and whether shorter replies or another host fit your use.storyPRICEGLM 5.3 cut across 4 hosts, by up to 60% at Phala (input)Phala: input ↓ 60% $1.40 → $0.56 Relace: input and cache read ↓ 13% (input $0.15 → $0.13) AkashML: input and output ↓ 10% (output $3.96 → $3.56) Makora: output ↓ 6% $4.20 → $3.93 prices per 1M tokensIf you buy: Renting this model at this host now costs less to run, so it suits steadier, larger workloads; re-check your output budget and cache terms before your next top-up.storyPRICEGLM 5.2 cut across 3 hosts, by up to 72% at Decart (mxfp4) (input)Decart (mxfp4): input ↓ 72% $1.40 → $0.39 Relace: input and cache read ↓ 33% (input $0.200 → $0.135) Reka: cache read ↓ 23% $0.26 → $0.20 prices per 1M tokensIf you buy: Reka's cache-read cut makes long-context work with heavy prompt reuse easier to justify here, so re-check your cached-token share before your next top-up.storyPRICEHost Relace cut MiMo-V2.6-Flash input and cache-read pricing by 50%input and cache read: ↓ 50%input −50% $0.080 → $0.040 · cache read −50% $0.080 → $0.040prices per 1M tokensIf you buy: This cut makes the model easier to run on long prompts and repeated context, so re-check your cache-hit assumptions and whether your workload now fits a lighter budget.storyPRICEQwen3.8 27B repriced across 5 hosts: DekaLLM input down 39%, Ionstream input up 37%DekaLLM: input ↓ 39% $0.080 → $0.049 · output ↑ 20% $2.50 → $3.00 Ionstream: input ↑ 37% $0.145 → $0.198 AkashML: input and output ↓ 10% (output $1.98 → $1.78) Wafer: output ↓ 1% $4.40 → $4.35 Reka: output ↓ 1% $4.40 → $4.35 prices per 1M tokensIf you buy: Ionstream now costs more to prompt, so if your workload is input-heavy, re-check your budget before your next top-up; output-heavy use is less affected.storyPRICEHost Phala cut Qwen3.5-27B cache-read pricing by 80%cache read: ↓ 80%cache read −80% $0.150 → $0.030prices per 1M tokensIf you buy: If you lean on long prompts with reused context, this host now suits that pattern far better, so re-check your cache-read line before your next top-up.storyPRICEHost AkashML raised gpt-oss-120b input and cache-read pricing by 12%input and cache read: ↑ 12%input +12% $0.033 → $0.037 · cache read +12% $0.033 → $0.037prices per 1M tokensIf you buy: Re-check your expected input and cache-read spend on this model at this host before your next top-up, since both now cost more per token.storyPRICEKimi K3 repriced across 4 hosts: Sail Research input down 61%, Fireworks (US region) up 36% on all ratesSail Research: input ↓ 61% $0.809 → $0.318 InferenceNet: input ↓ 53% $0.40 → $0.19 · output ↑ 22% $9.00 → $11.00 Fireworks (US region): all rates ↑ 36% (output $16.50 → $22.50) Morph: input ↓ 5% $1.23 → $1.17 · output ↑ 27% $10.70 → $13.55 prices per 1M tokensIf you buy: Fireworks buyers now pay more on every rate for this model, so re-check your budget and whether a lighter model covers the same work before your next top-up.storyPRICEDeepSeek V4.1 Flash repriced across 6 hosts: Io Net input down 64%, Ionstream input up 97%Ionstream: input ↑ 97% $0.145 → $0.285 InferenceNet: input ↓ 42% $0.069 → $0.040 · output ↑ 67% $0.45 → $0.75 Io Net: input ↓ 64% $0.250 → $0.090 OpenInference: input ↓ 34% $0.0300 → $0.0198 Wafer: input ↓ 25% $0.100 → $0.075 Morph: input ↓ 3% $0.0805 → $0.0780 · cache read ↑ 270% $0.0027 → $0.0100 prices per 1M tokensIf you buy: Ionstream now costs more for the same model, so if you run long prompts there, re-check your spend before the next top-up and consider whether your workload still fits.storyPRICEDeepSeek V4 Pro repriced across 3 hosts: Relace input and cache read down 18%, Ionstream input up 104%Ionstream: output ↓ 6% $1.96 → $1.85 · input ↑ 104% $0.608 → $1.238 Relace: input and cache read ↓ 18% (input $0.150 → $0.123) Wafer: output ↓ 17% $4.20 → $3.49 · cache read ↑ 24% $0.38 → $0.47 prices per 1M tokensIf you buy: Ionstream now charges far more for input on this model, so if your workload is prompt-heavy, re-check your cost per call before your next top-up.story
29 September 2026
RELEASEGPT-6.1 Sol listedIf you buy: A new OpenAI model is now listed on this host, so check whether it fits your task and budget before you top up, and confirm the rate and limits shown on the page.modelRELEASEGPT-6.1 Sol Pro listedIf you buy: A new OpenAI model is now listed on OpenRouter, so check its context window, tool support and rate limits before you point existing integrations at it.modelPRICEGLM 5.3 Flash repriced across 4 hosts: OpenInference input down 60%, Inceptron input up 25%OpenInference: input ↓ 60% $0.050 → $0.020 Relace: input ↓ 38% $0.040 → $0.025 · cache read ↑ 67% $0.015 → $0.025 Morph: cache read ↑ 31% $0.0306 → $0.0400 Inceptron: input ↑ 25% $0.12 → $0.15 prices per 1M tokensIf you buy: If you run this model on Morph, re-check your cache-read spend before your next top-up, since prompt reuse now costs more there.storyPRICEGLM 5.3 repriced across 4 hosts: Makora input down 19%, Morph input up 163%Morph: input ↑ 163% $0.452 → $1.190 Inceptron: input ↑ 93% $0.311 → $0.600 Relace: output ↑ 21% $3.30 → $4.00 Makora: input ↓ 19% $1.05 → $0.85 prices per 1M tokensIf you buy: Renting this model at this host now costs more to run, so check your usage pattern and cache settings before your next top-up, and confirm the current rate on the host's page.storyPRICEGLM 5.2 repriced across 2 hosts: Relace input down 43%, Inceptron input up 110%Inceptron: input ↑ 110% $0.3565 → $0.7500 Relace: input ↓ 43% $0.35 → $0.20 · output ↑ 14% $3.50 → $4.00 prices per 1M tokensIf you buy: If you run long prompts through Inceptron, your input spend now climbs sharply, so re-check your budget and consider trimming prompt size before your next top-up.storyPRICEMiMo-V2.6-Flash repriced across 2 hosts: DeepInfra's own listing input down 43%, Relace cache read up 100%DeepInfra's own listing: input ↓ 43% $0.140 → $0.080 Relace: input ↓ 33% $0.120 → $0.080 · cache read ↑ 100% $0.040 → $0.080 DeepInfra through OpenRouter: cache read ↑ 4% $0.0027 → $0.0028 prices per 1M tokensstoryPRICEQwen3.8 27B repriced across 4 hosts: Wafer input and cache read down 65%, Ionstream input up 63%Wafer: input and cache read ↓ 65% (input $0.072 → $0.025) Reka: input and cache read ↓ 65% (input $0.072 → $0.025) Ionstream: input ↑ 63% $0.089 → $0.145 Phala: input ↓ 17% $0.24 → $0.20 prices per 1M tokensIf you buy: If you run this model on Ionstream, your input spend just rose, so re-check your usage mix and cache settings before the next top-up.storyPRICENemotron 3.5 Lightning cut across 2 DeepInfra listings, by up to 25% at DeepInfra through OpenRouter (input and cache read)DeepInfra through OpenRouter: input and cache read ↓ 25% (input $0.080 → $0.060) DeepInfra's own listing: input ↓ 25% $0.080 → $0.060 prices per 1M tokensIf you buy: this cut makes the model easier to justify for high-volume, cache-heavy work, so re-check your cache-read share and whether your routing still fits how you use it.storyPRICEKimi K3 repriced across 4 hosts: InferenceNet input down 60%, Sail Research output up 62%Sail Research: input ↓ 10% $0.90 → $0.81 · output ↑ 62% $7.86 → $12.75 InferenceNet: input ↓ 60% $1.00 → $0.40 · cache read ↑ 33% $0.30 → $0.40 Wafer: input ↑ 44% $1.00 → $1.44 Morph: output ↓ 4% $11.20 → $10.70 · input ↑ 40% $0.88 → $1.23 prices per 1M tokensIf you buy: If you lean on long outputs at Sail Research, re-check your budget before the next top-up, while prompt-heavy work now fits a different host.storyPRICEHost Inceptron cut Kimi K2.6 input pricing by 3%input: ↓ 3%input −3% $0.453 → $0.438 · cache read −2% $0.123 → $0.121prices per 1M tokensstoryPRICEDeepSeek V4.1 Flash repriced across 7 hosts: OpenInference input down 70%, InferenceNet input up 97%InferenceNet: input ↑ 97% $0.035 → $0.069 OpenInference: input ↓ 70% $0.100 → $0.030 Morph: output ↑ 64% $0.3101 → $0.5100 Relace: input ↓ 60% $0.050 → $0.020 · cache read ↑ 100% $0.010 → $0.020 Ionstream: input ↓ 48% $0.280 → $0.145 Wafer: output ↓ 37% $0.70 → $0.44 · input ↑ 1% $0.099 → $0.100 Phala: all rates ↓ 13% (output $1.38 → $1.20) prices per 1M tokensstoryPRICEDeepSeek V4 Pro repriced across 5 hosts: Relace input down 57%, Ionstream input up 169%Ionstream: input ↑ 169% $0.226 → $0.608 Relace: input ↓ 57% $0.35 → $0.15 · cache read ↑ 50% $0.10 → $0.15 Io Net: input ↓ 29% $0.99 → $0.70 · output ↑ 11% $3.15 → $3.50 GMICloud: input ↓ 24% $1.74 → $1.32 · output ↑ 14% $3.48 → $3.96 Wafer: output ↓ 17% $4.20 → $3.49 prices per 1M tokensIf you buy: Ionstream now costs more to run for input-heavy work, so check your usage mix and cache settings before your next top-up there.storyPRICEDeepSeek V4 Flash repriced across 5 hosts: AtlasCloud output down 65%, Inceptron output up 71%Inceptron: input ↓ 14% $0.065 → $0.056 · output ↑ 71% $0.38 → $0.65 AtlasCloud: output ↓ 65% $0.48 → $0.17 · cache read ↑ 64% $0.01008 → $0.01652 Venice: input and output ↑ 27% (output $0.275 → $0.350) Wafer: input ↑ 19% $0.059 → $0.070 Relace: input ↓ 14% $0.021 → $0.018 · cache read ↑ 12% $0.016 → $0.018 prices per 1M tokensIf you buy: On Inceptron, output costs more while input fell slightly, so re-check your input-to-output mix before your next top-up and whether this host still suits your workload.story
28 September 2026
PRICEGrok 4.7 rose across 2 xAI listings, by up to 25% at xAI through OpenRouter (all rates)xAI through OpenRouter: all rates ↑ 25% (output $4.80 → $6.00) xAI (priority tier): all rates ↑ 25% (output $9.60 → $12.00) prices per 1M tokensstoryPRICEHost AkashML raised gpt-oss-120b pricing by 10% on all ratesall rates: ↑ 10%input +10% $0.030 → $0.033 · output +10% $0.170 → $0.187 · cache read +10% $0.030 → $0.033prices per 1M tokensstoryPRICEHost AtlasCloud cut DeepSeek V3.2 Exp output pricing by 51%output: ↓ 51%input −50% $0.270 → $0.134 · output −51% $0.41 → $0.20 · cache read −75% $0.270 → $0.067prices per 1M tokensIf you buy: Output and cache reads now cost far less, so long generations and heavy context reuse are easier to justify; re-check your cache-hit assumptions and budget.storyPRICEHost AtlasCloud cut DeepSeek V3.2 input and cache-read pricing by 48%input and cache read: ↓ 48%input −48% $0.260 → $0.134 · output −47% $0.38 → $0.20 · cache read −48% $0.130 → $0.067prices per 1M tokensstoryPRICEHost AtlasCloud raised DeepSeek V3.1 Terminus output pricing by 5%output: ↑ 5%output +5% $0.95 → $1.00 · cache read +4% $0.130 → $0.135prices per 1M tokensIf you buy: budget for a slightly higher output cost on your next top-up here, and re-check your cache-read assumptions if your workload leans on cached prompts.storyPRICEHost AtlasCloud raised DeepSeek V3.1 output pricing by 5%output: ↑ 5%output +5% $0.95 → $1.00 · cache read +4% $0.130 → $0.135prices per 1M tokensIf you buy: Output on this model at this host now costs a little more, so check your usage mix and cache reads before your next top-up if you run long generations here.storyRELEASEClaude Sonnet 5.5 listedIf you buy: A new Sonnet is now listed on this host, so check whether your existing prompts and tool calls still behave as expected before you move any production traffic onto it.modelPRICEMiMo-V2.6-Flash repriced across 2 hosts: Relace output up 814%, DeepInfra through OpenRouter cache read down 4%Relace: output ↑ 814% $0.14 → $1.28 DeepInfra through OpenRouter: cache read ↓ 4% $0.0028 → $0.0027 prices per 1M tokensIf you buy: On Relace, output-heavy work now costs far more per call, so re-check your spend if you generate long replies there, while cache-heavy prompts on DeepInfra shift only slightly.storyPRICEGLM 5.3 Flash repriced across 4 hosts: AtlasCloud down 23% on all rates, Wafer input up 74%Wafer: input ↑ 74% $0.43 → $0.75 AtlasCloud: all rates ↓ 23% (output $0.500 → $0.385) Inceptron: input ↑ 9% $0.11 → $0.12 · cache read ↓ 43% $0.070 → $0.040 Morph: all rates ↓ 6% (output $0.568 → $0.536) prices per 1M tokensstoryPRICEGLM 5.3 repriced across 6 hosts: Relace input down 73%, output up 94%Relace: input ↓ 73% $0.55 → $0.15 · output ↑ 94% $1.70 → $3.30 Wafer: input ↑ 86% $0.98 → $1.82 Morph: input ↓ 62% $1.19 → $0.45 AtlasCloud: all rates ↓ 57% (output $4.40 → $1.89) Reka: output and cache read ↓ 56% (output $2.57 → $1.14) Inceptron: output ↑ 11% $2.51 → $2.79 prices per 1M tokensIf you buy: Relace now charges far less for input and far more for output, so long prompts with short answers suit it, while chatty work is worth re-checking before your next top-up.storyPRICEGLM 5.2 repriced across 4 hosts: Inceptron input down 27%, Wafer input up 30%Wafer: input ↑ 30% $1.40 → $1.82 Inceptron: input ↓ 27% $0.49 → $0.36 · output ↑ 25% $1.925 → $2.401 Cloudflare: input ↓ 16% $1.40 → $1.18 Relace: input ↓ 13% $0.40 → $0.35 prices per 1M tokensstoryPRICEQwen3.8 27B repriced across 5 hosts: Darkbloom input down 28%, Wafer output up 80%Wafer: input ↓ 13% $0.083 → $0.072 · output ↑ 80% $2.45 → $4.40 DekaLLM: cache read ↓ 54% $0.087 → $0.040 Darkbloom: input ↓ 28% $0.069 → $0.050 Reka: input ↓ 20% $0.090 → $0.072 Ionstream: input and cache read ↓ 15% (input $0.105 → $0.089) prices per 1M tokensIf you buy: On Wafer, output costs more while input and cache reads cost less, so this host suits prompt-heavy, short-answer work; re-check your output budget before your next top-up.storyPRICEKimi K3 cut across 2 hosts, by up to 65% at Morph (input)Morph: input ↓ 65% $2.50 → $0.88 Sail Research: input and output ↓ 13% (output $9.04 → $7.86) prices per 1M tokensstoryPRICEHost Inceptron raised Kimi K2.6 input pricing by 11%input: ↑ 11%input +11% $0.409 → $0.455 · output +3% $2.39 → $2.45 · cache read +39% $0.089 → $0.124prices per 1M tokensIf you buy: Budget for a higher input and cache-read cost on your next top-up here, and re-check whether your workload leans on cached prompts before you commit to a long run.storyPRICEDeepSeek V4.1 Flash repriced across 3 hosts: AtlasCloud down 62% on all rates, Wafer input up 136%Wafer: input ↑ 136% $0.042 → $0.099 AtlasCloud: all rates ↓ 62% (output $1.20 → $0.46) Morph: output and cache read ↓ 48% (output $0.60 → $0.31) prices per 1M tokensstoryPRICEDeepSeek V4 Pro repriced across 4 hosts: Io Net input down 11%, Relace input up 72%Relace: output ↓ 7% $3.78 → $3.50 · input ↑ 72% $0.204 → $0.350 Wafer: input ↑ 35% $0.293 → $0.397 Io Net: input ↓ 11% $1.24 → $1.10 Ionstream: input ↓ 11% $0.253 → $0.226 prices per 1M tokensIf you buy: Relace now charges more for fresh input but less for output and cached reads, so prompt-reuse and long-generation workloads gain while input-heavy jobs should re-check the bill.storyPRICEDeepSeek V4 Flash repriced across 5 hosts: Sail Research (US region) input down 16%, NextBit output up 252%NextBit: output ↑ 252% $0.300 → $1.056 · cache read ↓ 66% $0.035 → $0.012 Wafer: input ↑ 106% $0.0413 → $0.0850 AtlasCloud: output ↑ 70% $0.280 → $0.475 · cache read ↓ 64% $0.028 → $0.010 Inceptron: input ↑ 67% $0.039 → $0.065 · cache read ↓ 10% $0.030 → $0.027 Sail Research (US region): input ↓ 16% $0.0225 → $0.0190 Sail Research: input ↓ 12% $0.0215 → $0.0190 prices per 1M tokensIf you buy: NextBit now suits cache-heavy, input-light workloads, so re-check your output and cache-read mix before your next top-up there.story
27 September 2026
PRICEHost Wafer raised Kimi K3 input pricing by 19%input: ↑ 19%input +19% $1.00 → $1.19prices per 1M tokensstoryPRICEHost Wafer repriced DeepSeek V4 Pro: input up 20%, cache read down 30%input: ↑ 20%cache read: ↓ 30%input +20% $0.245 → $0.293 · cache read −30% $0.245 → $0.172prices per 1M tokensIf you buy: If you lean on cached context, your bill at Wafer shifts toward cache reads, so re-check your caching setup and how much fresh input you send before your next top-up.storyPRICEHost Wafer cut DeepSeek V4 Flash input pricing by 30%input: ↓ 30%input −30% $0.0590 → $0.0413 · cache read −36% $0.058 → $0.037prices per 1M tokensIf you buy: input and cache reads now cost less on Wafer, so this model suits steady, prompt-heavy work; re-check your cache-hit rate before your next top-up.storyPRICEHost Relace raised GLM 5.3 Flash input pricing by 75%input: ↑ 75%input +75% $0.040 → $0.070 · cache read +33% $0.015 → $0.020prices per 1M tokensIf you buy: Budget-conscious workloads on this model at this host now deserve a second look at your usage pattern before the next top-up.storyPRICEGLM 5.3 repriced across 3 hosts: Reka input down 48%, Inceptron output up 144%Inceptron: output ↑ 144% $1.03 → $2.51 Reka: input ↓ 48% $0.683 → $0.356 Wafer: input ↓ 30% $1.40 → $0.98 Inceptron: all rates ↑ 15% (output $0.898 → $1.032) prices per 1M tokensIf you buy: On Inceptron, output-heavy work now costs far more per call, so re-check your usage mix before your next top-up and lean on cached context where you can.storyPRICEQwen3.8 27B repriced across 4 hosts: Wafer output down 44%, Wafer input up 3%Wafer: output ↓ 44% $4.40 → $2.45 DekaLLM: output ↓ 43% $4.40 → $2.50 Reka: output ↓ 43% $4.40 → $2.50 Darkbloom: input ↓ 16% $0.0825 → $0.0690 Wafer: input ↑ 3% $0.0803 → $0.0828 Reka: input and cache read ↓ 1% (input $0.080 → $0.079) prices per 1M tokensstoryPRICEHost Darkbloom cut Nemotron 3.5 Lightning input pricing by 40%input: ↓ 40%input −40% $0.065 → $0.039prices per 1M tokensIf you buy: input costs less here now, so workloads that were too chatty to run at this host are worth re-costing before your next top-up.storyPRICEDeepSeek V4.1 Flash repriced across 3 hosts: DekaLLM output down 33%, Relace output up 50%Relace: output ↑ 50% $0.40 → $0.60 DekaLLM: output ↓ 33% $0.60 → $0.40 Wafer: input ↓ 14% $0.049 → $0.042 prices per 1M tokensIf you buy: Relace now costs more to run for output-heavy work, so check whether your usage leans on output or cache reads before your next top-up there.story
26 September 2026
PRICEHost Inceptron cut Kimi K2.6 input pricing by 6%input: ↓ 6%input −6% $0.434 → $0.409 · cache read −15% $0.105 → $0.089prices per 1M tokensIf you buy: Long-context work that leans on cached prompts now costs less to run here, so re-check your cache-hit assumptions before your next top-up.storyPRICEHost Reka cut Qwen3.6 35B A3B input pricing by 33%input: ↓ 33%input −33% $0.15 → $0.10prices per 1M tokensIf you buy: Reka's input rate for this model has dropped, so prompt-heavy workloads now cost less to run there; re-check your output-token spend before committing a larger budget.storyPRICEHost Reka cut Gemma 4 31B output pricing by 12%output: ↓ 12%input −11% $0.090 → $0.080 · output −12% $0.34 → $0.30prices per 1M tokensstoryPRICEHost Reka cut Gemma 4 26B A4B input and output pricing by 33%input and output: ↓ 33%input −33% $0.090 → $0.060 · output −33% $0.30 → $0.20 · cache read −30% $0.050 → $0.035prices per 1M tokensstoryPRICEGLM 5.3 cut across 2 hosts, by up to 43% at Inceptron (output), with cache read down 44% thereInceptron: output ↓ 43% $2.50 → $1.43 Relace: output and cache read ↓ 23% (output $2.20 → $1.70) prices per 1M tokensIf you buy: Long-context and cache-heavy work now fits a smaller budget here, so re-check your usage mix and cache settings before your next top-up.storyPRICEHost Inceptron cut GLM 5.2 output pricing by 40%output: ↓ 40%input −24% $0.646 → $0.489 · output −40% $3.18 → $1.92 · cache read −35% $0.208 → $0.136prices per 1M tokensstoryPRICEQwen3.8 27B cut across 2 hosts, by up to 2% at Reka (input)Reka: input ↓ 2% $0.092 → $0.090 Wafer: input ↓ 2% $0.093 → $0.091 prices per 1M tokensIf you buy: a slightly lower input rate at Wafer makes this a good moment to re-check your usage mix and confirm the new rate applies to your next top-up.storyPRICEKimi K3 cut across 2 hosts, by up to 23% at InferenceNet (input)InferenceNet: input ↓ 23% $1.30 → $1.00 Sail Research: input and output ↓ 14% (output $10.53 → $9.04) prices per 1M tokensstoryPRICEGLM 5.3 Flash repriced across 3 hosts: Relace input down 43%, Wafer input up 523%Wafer: input ↑ 523% $0.069 → $0.430 Relace: input ↓ 43% $0.070 → $0.040 · output ↑ 79% $0.28 → $0.50 OpenInference: cache read ↑ 33% $0.015 → $0.020 prices per 1M tokensIf you buy: Renting this model on Wafer now costs far more per token, so re-check your usage and whether a different host fits your workload before your next top-up.storyPRICEDeepSeek V4.1 Flash cut across 2 hosts, by up to 47% at Sail Research (output)Sail Research: output ↓ 47% $0.75 → $0.40 Relace: input and cache read ↓ 44% (input $0.090 → $0.050) prices per 1M tokensIf you buy: Output-heavy work at Sail Research now costs less to run, so re-check your usage split and whether your current plan still matches how you actually call this model.storyPRICEDeepSeek V4 Flash cut across 3 hosts, by up to 53% at Inceptron (input)Inceptron: input ↓ 53% $0.083 → $0.039 Sail Research: output ↓ 45% $0.55 → $0.30 Relace: input ↓ 30% $0.030 → $0.021 Sail Research (US region): input and cache read ↓ 25% (input $0.0300 → $0.0225) prices per 1M tokensIf you buy: Existing users at this host should re-check their usage and caching setup, since the lower rates now suit steady, high-volume work that was previously harder to justify.story
25 September 2026
PRICEHost Inceptron cut Kimi K2.6 output pricing by 20%output: ↓ 20%input −4% $0.45 → $0.43 · output −20% $2.97 → $2.39 · cache read −17% $0.126 → $0.105prices per 1M tokensIf you buy: output-heavy work on this model at this host now costs less to run, so long generations and agent loops are easier to justify; re-check budgets set against the old output rate.storyPRICEGemma 4 26B A4B cut across 2 hosts, by up to 25% at NextBit (all rates)NextBit: all rates ↓ 25% (output $0.300 → $0.225) Makora: input ↓ 20% $0.100 → $0.080 prices per 1M tokensstoryRELEASEPerceptron Mk1.5 listedIf you buy: a new model is now available to try on this host, so check its context window and rate limits before you point any production traffic at it.modelRELEASEGLM 5.3 FlashX listedIf you buy: a new option to weigh when picking a model for quick work on this host, so check its context window and rate limits before you commit a workflow to it.modelRELEASEQwen3.8 Omni Flash listedIf you buy: A new option to test for multimodal work, so check its context window and tool support against your current setup before switching any routine traffic.modelRELEASEClaude Opus 5.5 listedIf you buy: A new Opus tier is now reachable through this host, so check whether your existing prompts and tool calls still behave as expected before you move any production traffic onto it.modelRELEASEGPT-6 Sol listedIf you buy: A new OpenAI model is now reachable through this host, so check whether your existing setup already routes to it before you top up.modelRELEASEGPT-6 Sol Pro listedIf you buy: A new OpenAI model is now available to route to, so check whether your existing setup already points at it before your next top-up.modelRELEASEGPT-6 Luna listedIf you buy: A new OpenAI model is now listed on OpenRouter, so check whether it fits your workload and confirm the rate and limits before you commit spend to it.modelRELEASEGPT-6 Luna Pro listedIf you buy: A new OpenAI model is now available to route to, so check whether your current setup already points at it before your next top-up.modelRELEASEGLM 5.3 Prime listedIf you buy: a new Z.ai model is now selectable at this host, so check its context window and tool support before pointing existing GLM workloads at it.modelPRICEHost Inceptron repriced GLM 5.2: input down 28%, cache read up 11%input: ↓ 28%cache read: ↑ 11%input −28% $0.90 → $0.65 · cache read +11% $0.19 → $0.21prices per 1M tokensIf you buy: Prompt-heavy work now costs less here, but re-check your cache-read share before topping up, since that side moved the other way.storyPRICEGLM 5.3 Flash cut across 3 hosts, by up to 50% at OpenInference (input)OpenInference: input ↓ 50% $0.100 → $0.050 Inceptron: input ↓ 27% $0.15 → $0.11 Wafer: input ↓ 22% $0.089 → $0.069 prices per 1M tokensIf you buy: Renting this model at Wafer now costs less to run, so heavier prompt and cache traffic fits a tighter budget; re-check your usage tier and cache settings before your next top-up.storyPRICEGLM 5.3 repriced across 3 hosts: Sail Research (US region) input down 36%, output up 3%Sail Research (US region): input ↓ 36% $1.21 → $0.77 · output ↑ 3% $3.87 → $4.00 Inceptron: input ↓ 24% $0.79 → $0.60 Io Net: input ↓ 1% $0.78 → $0.77 prices per 1M tokensIf you buy: At Io Net the input rate barely moved, so if your workload is output-heavy this repricing changes little; re-check your cache-read share before your next top-up.storyPRICEQwen3.8 27B cut across 3 hosts, by up to 58% at Ionstream (input)Ionstream: input ↓ 58% $0.250 → $0.105 Reka: input ↓ 2% $0.094 → $0.092 Wafer: input ↓ 2% $0.095 → $0.093 prices per 1M tokensIf you buy: If you run this model at Wafer, the input side now costs less, so re-check your budget for long prompts and confirm the output rate before you commit to a bigger workload.storyPRICEDeepSeek V4.1 Flash repriced across 2 hosts: DekaLLM output down 40%, input up 275%DekaLLM: output ↓ 40% $1.00 → $0.60 · input ↑ 275% $0.040 → $0.150 NextBit: input and output ↓ 30% (output $1.20 → $0.84) prices per 1M tokensIf you buy: On DekaLLM, this now suits output-heavy and cache-reliant work far more than prompt-heavy jobs, so re-check your input-to-output mix before your next top-up.storyPRICEDeepSeek V4 Pro repriced across 2 hosts: Wafer input down 36%, output up 21%Wafer: input ↓ 36% $0.39 → $0.25 · output ↑ 21% $2.90 → $3.50 Ionstream: input ↓ 35% $0.390 → $0.253 prices per 1M tokensIf you buy: On Ionstream the balance shifts toward prompt-heavy work, so re-check your output share before your next top-up.storyPRICEKimi K3 repriced across 4 hosts: Makora input and cache read down 20%, InferenceNet output up 6%Makora: input and cache read ↓ 20% (input $2.55 → $2.04) Sail Research: input ↓ 11% $1.35 → $1.20 InferenceNet: input ↓ 7% $1.40 → $1.30 · output ↑ 6% $10.75 → $11.40 Wafer: output ↑ 3% $14.50 → $15.00 prices per 1M tokensstoryPRICEDeepSeek V4 Flash repriced across 2 hosts: Sail Research input down 21%, Inceptron output up 17%Sail Research: input ↓ 21% $0.038 → $0.030 Sail Research (US region): input ↓ 21% $0.038 → $0.030 Inceptron: output ↑ 17% $0.41 → $0.48 · cache read ↓ 17% $0.060 → $0.050 prices per 1M tokensIf you buy: On Inceptron, output now costs more while cache reads cost less, so re-check your mix of fresh output against cached input before your next top-up.story
24 September 2026
PRICEHost Inceptron cut Kimi K2.7 Code input pricing by 7%input: ↓ 7%input −7% $0.71 → $0.66prices per 1M tokensIf you buy: a lower input rate on this coding model at this host, so it now suits longer agent runs and heavy prompt traffic; re-check your cached and output rates before scaling up.storyPRICEQwen3.8 27B repriced across 3 hosts, from a 76% rise at DekaLLM (output) to an 11% cut at Ionstream (input)DekaLLM: input ↓ 4% $0.100 → $0.096 · output ↑ 76% $2.50 → $4.40 Ionstream: input ↓ 11% $0.28 → $0.25 Wafer: input and cache read ↓ 4% (input $0.099 → $0.095) prices per 1M tokensIf you buy: Output-heavy work at DekaLLM now costs noticeably more per call, so re-check your spend if most of your tokens are generated rather than read.storyRELEASEEmber-1 listedmodelPRICEHost Relace cut GLM 5.3 Flash input pricing by 30%input: ↓ 30%input −30% $0.100 → $0.070 · output −22% $0.36 → $0.28prices per 1M tokensIf you buy: a cheaper floor for anyone running high-volume prompts through this host, so re-check your spend cap and whether your workload still fits the smaller output budget.storyPRICEHost Io Net cut GLM 5.3 cache-read pricing by 22%cache read: ↓ 22%cache read −22% $0.18 → $0.14prices per 1M tokensstoryPRICEHost Inceptron cut GLM 5.2 input pricing by 11%input: ↓ 11%input −11% $1.01 → $0.90 · output −2% $3.24 → $3.18 · cache read −3% $0.194 → $0.188prices per 1M tokensstoryPRICEKimi K3 repriced across 3 hosts, from an 8% rise at Relace (all rates) to a 20% cut at Wafer (input and cache read)Wafer: input and cache read ↓ 20% (input $2.49 → $1.99) Sail Research: input ↓ 10% $1.50 → $1.35 · output ↑ 7% $10.76 → $11.50 Relace: all rates ↑ 8% (output $9.75 → $10.50) prices per 1M tokensIf you buy: Relace now costs more to run for both input and output, so check whether your workload still fits the budget you set before your next top-up.storyPRICEHost Inceptron cut Kimi K2.6 input pricing by 9%input: ↓ 9%input −9% $0.4972 → $0.4546 · cache read −1% $0.128 → $0.127prices per 1M tokensstory
100 stories, back to 24 September