The August 2026 Pricing Storm: DeepSeek's 4x Hike, Gemini's Clock, and the End of Flat API Pricing
In two weeks, three labs changed the meaning of their own rate card. DeepSeek moved to peak/off-peak billing that quietly tripled V4 Pro's output price, Google halved Gemini 3.7 Flash with a hike attached, and Anthropic cancelled one outright. Here is the real scorecard.
The Month a Published Price Stopped Being a Single Number
August 2026 ended the assumption that an API price is one number you can compare. Within two weeks, three labs changed the meaning of their own rate card, in three different ways, and only one of those changes was a cut.
DeepSeek moved from flat pricing to peak/off-peak billing. Google launched a model at half price with a doubling already scheduled. Anthropic cancelled a planned increase. And the open-weight players (Alibaba, Z.ai, Meta) kept prices flat while failing to ship the weights that were supposed to undercut everyone.
DeepSeek: The Hike Dressed as a Discount
DeepSeek's V4-Pro had been in preview since April at a promotional $0.435 input / $0.87 output per million tokens. The announced May 31 reversion to $1.74/$3.48 never took effect, so the promotional price quietly became the list price. On August 13, V4-Pro-0813 reached general availability alongside DeepSeek Harness v0.1, an MIT-licensed open-source rival to Claude Code, with native OpenAI Responses API and Codex integration. source
Then, on August 16 at 16:00 UTC, DeepSeek abandoned flat pricing. Peak hours run 01:00-04:00 and 06:00-10:00 UTC (Beijing 09:00-12:00 and 14:00-18:00); everything else is off-peak at exactly half the peak rate.
V4 Pro (per 1M tokens):
| Category | Until Aug 16 (flat) | New off-peak | New peak | Peak vs old |
|---|---|---|---|---|
| Cache-hit input | $0.003625 | $0.022 | $0.044 | +1,114% |
| Cache-miss input | $0.435 | $0.66 | $1.32 | +203% |
| Output | $0.87 | $1.98 | $3.96 | +355% |
V4 Flash (per 1M tokens):
| Category | Until Aug 16 (flat) | New off-peak | New peak |
|---|---|---|---|
| Cache-hit input | $0.0028 | $0.007 | $0.014 |
| Cache-miss input | $0.14 | $0.22 | $0.44 |
| Output | $0.28 | $0.66 | $1.32 |
Read the framing carefully. Off-peak is 50% of the new peak, but it is not a discount from today's prices. The cheapest hour of DeepSeek's new schedule costs more than twice what the same model costs now, and peak output is roughly 4.6x the old rate. Reuters flagged the sharpest line: V4 Pro now prices at about 9x V4 Flash on input and 14x on output. source A simple 1M-in/1M-out V4-Pro workload that cost $1.305 yesterday becomes $2.64 off-peak or $5.28 at peak. source
Google: Half Price, With a Clock Attached
Gemini 3.7 Flash launched on August 13, three weeks after 3.6 Flash, at an introductory $0.75 input / $3.75 output, with context-cache reads at $0.075. Google's own footnote: "Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply." That is a doubling, already scheduled. source
The same rate page lists Gemini 3.6 Flash, released in late July at $1.50/$7.50, at the identical $0.75/$3.75, a 50% cut roughly three weeks after launch, expiring at the end of the year. Google is buying the low-cost workhorse market outright, and telling developers exactly when the meter changes. On benchmarks Google cites against 3.6 Flash: FrontierCode 1.1 up to 43.6% (from 34.4%), DeepSWE v1.1 65.3% (from 49.0%), and 3.7 Flash is framed as "our most intelligent workhorse model yet for coding and agents." source
Anthropic: The Rarest Move, Cancelling an Increase
Claude Sonnet 5 launched June 30 at $2/$10, labeled introductory through August 31 with a $3/$15 standard rate scheduled for September 1. Around August 10, Anthropic quietly reversed course. The official pricing page now reads: "The $2/$10 per million input/output token pricing for Claude Sonnet 5 ... is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur." source
The consequence: Sonnet 5 now costs less than its predecessor Sonnet 4.6 ($3/$15) on the same rate card, the first time a newer flagship undercut the model it replaced. One caveat: Claude 4.7+ models use a newer tokenizer that produces roughly 30% more tokens for the same text, so the per-token rate understates per-task cost. Anthropic also hard-retired Claude Opus 4.1 on August 5 (60-day notice given June 5), pointing developers to claude-opus-4-8. source
The Rest of August: Flat Prices, Unshipped Weights
Elsewhere, prices held while delivery wobbled:
- xAI Grok 4.6 (Aug 12): 500K context, reasoning up to "xhigh," $2 input / $0.50 cached / $6 output under 200K prompt tokens (above that, the whole request bills at $4/$1/$12). On AWS Bedrock (Aug 19) the x.ai post says $2/$6 while AWS documentation lists $2.20/$6.60 for in-region and US-geo profiles. source
- Alibaba Qwen3.8-Max (Aug 3): 2.4T MoE / 95B active, ~1M context, $2/$6 on the hosted API; open weights shipped Aug 12-13, but text-only, without the hosted 1M context, under a custom revenue-sharing license, not Apache 2.0. source
- Z.ai GLM-5.3 (Aug 14): same 743B MoE base as GLM-5.2, 1M context, priced identically at $1.40/$4.40, API-only at launch, with weights promised "in stages following rigorous safety evaluations," roughly two weeks out. As of August 20, still not released, no license published. source
- Meta Muse Spark 1.2 (Aug 5): $1.25/$4.25 with 1M context; open weights announced as "coming soon" with no date or license as of mid-August. source
The Real Scorecard
| Vendor | Move | Direction for users |
|---|---|---|
| DeepSeek | Flat to peak/off-peak; every line item up | Effective 2.3-4.6x on V4-Pro output |
| 3.7 Flash intro at half 3.6's launch price; both double Jan 1 2027 | Cheaper now, planned +100% in January | |
| Anthropic | Sonnet 5 increase cancelled; $2/$10 permanent | Frozen, cheaper than Sonnet 4.6 |
| xAI | Grok 4.6 flat at $2/$6 (Bedrock in-region $2.20/$6.60) | Stable |
| Alibaba / Z.ai / Meta | API flat; open weights late or unshipped | "Announced is not shipped" |
What Developers Should Do
This month makes per-token comparison tables obsolete. Three rules instead:
- Price by time of day. If you use DeepSeek, shift batch or cacheable workloads off-peak; the spread between peak and off-peak is now 2x by design.
- Watch the clock, not just the rate. Gemini 3.7 Flash at $0.75/$3.75 is a one-quarter deal that doubles in January. If you standardize on it, budget for the step change or plan a re-evaluation.
- Verify weights before betting on open-source cost. Three of the month's "open" releases (GLM-5.3, Muse Spark 1.2, and Qwen3.8-Max's text-only build) shipped late, incomplete, or under licenses that are not Apache 2.0.
The lesson of August 2026: a published price is a decision, not a fact. source
Sources: DeepSeek API pricing docs, Reuters, VentureBeat, Google AI blog, Anthropic platform pricing and deprecations pages, x.ai, AWS, Artificial Analysis, Capital & Compute, August 2026. Facts current as of August 20, 2026.
Continue exploring
More decisions worth reading
Follow the thread from this article to the next practical buying question.
Plan comparison
01GPT-5.6 Sol vs Claude Fable 5: The Definitive 2026 Comparison
Open guideBuying advice
02The AI Price War Just Got a Scoreboard: Luna Down 80%, Opus 5 at Half Price, and Token Spend Is Moving
Open guideBuying advice
03The AI Price Arbitrage Era: How Third-Party Providers Are Undercutting Official API Pricing by 92%
Open guideBuying advice
04