Hubs AI and Technology Post
Join TrustHub to participate — every member is ID-verified
Sign Up Free
0

A Week After Its Cheapest Model Yet, DeepSeek Says Prices Are Going Up

On Thursday, DeepSeek put a short notice on its developer platform: prices for its API are going up "in the near future," and the increase will be "significant." No new rate card. No date. No numbers at all, just the warning.

The timing is what makes it a story. Exactly one week earlier, on July 31, DeepSeek had officially released DeepSeek-V4-Flash-0731, the cheapest serious model on its own rate card. Demand for it surged, per the South China Morning Post's reporting. And then the company that built its whole reputation on being absurdly cheap told its customers to expect to pay more.

The model itself, quickly

V4-Flash-0731 is the official release of the Flash model DeepSeek previewed in April. Per the company's changelog and the Hugging Face model card, the architecture did not change at all: 284 billion total parameters with 13 billion active per token, a 1 million token context window. The only thing that changed was post-training. The weights are public on Hugging Face under an MIT license, free for anyone to download and run.

The headline numbers are agent benchmarks. On Terminal Bench 2.1, which measures long-horizon terminal work, DeepSeek's published table shows V4-Flash at 82.7, well above the 72.1 posted by V4-Pro, DeepSeek's own much larger flagship, which is still in preview. Worth being precise about two things here. First, those are DeepSeek's scores from DeepSeek's own testing setup, not an independent lab's. Second, on that same table, Anthropic's Opus 4.8 scored 85.0. So the small model beat its bigger sibling, not everything on the market.

Artificial Analysis, which does run its own independent evaluations, scored V4-Flash at 52 on its Intelligence Index, third out of 101 models in its class, and measured it at 123 output tokens per second, which it rates as notably fast.

What it actually costs

Here is the current rate card, per DeepSeek's official pricing, unchanged with the release:

Input (cache miss) Input (cache hit) Output
DeepSeek V4-Flash $0.14 / 1M tokens $0.0028 / 1M $0.28 / 1M
DeepSeek V4-Pro $0.435 / 1M $0.003625 / 1M $0.87 / 1M

For a sense of scale, per Anthropic's pricing page, Claude Fable 5 costs $10 per million input tokens and $50 per million output. DeepSeek's output tokens cost roughly one one-hundred-and-seventy-eighth of that.

Two practical details most of the coverage skipped. First, per-token price is not the whole bill. Artificial Analysis measured V4-Flash generating 210 million output tokens across its evaluation suite, against a median of 100 million for comparable models. The model is verbose, it thinks a lot before it answers, and output tokens are the expensive line on any invoice. A cheap model that talks twice as much is not half as cheap as it looks. Second, the cache-hit price, $0.0028 per million, is one-fiftieth of the normal input price. For anyone running agents with a stable system prompt, that discount is where the real savings sit.

I pay for a Claude Max subscription and I run AI agents on my own projects most days, so output-token pricing is not an abstract number to me. It is the bill. That is why a $0.28 output price got my attention in the first place, and why a "significant" increase to it matters.

The warning

Here is what DeepSeek actually said on Thursday, per the statement on its developer platform, reported by the South China Morning Post and others: overall pricing for its API services will be raised in the near future, a significant increase is expected, users should plan their usage accordingly, and the specific plan will come in a later official notice. Bloomberg first reported the planned increase late on August 5, per RuntimeWire.

That is everything. There are no new figures to analyze, because DeepSeek has not published any.

It is also the second pricing move in about a month. The South China Morning Post reported on June 30 that DeepSeek planned to double V4 API prices during two daily peak windows in Beijing time, which the company attributed to resource allocation and service stability. Per DeepSeek's own documentation, that peak-hour surcharge has been announced but is not active yet. Thursday's warning is broader: it describes an increase in overall pricing, not just busy-hour pricing.

Per the South China Morning Post, the warning follows surging global demand for V4-Flash. eWeek reports that DeepSeek's aggressive pricing previously pushed competitors including ByteDance and Tencent to lower their own rates, so a real increase gives those companies room to hold their prices or raise them too. Not everyone is sympathetic to the timing. AI developer Michael Guo, quoted by the South China Morning Post, pointed out that newer American models are now competitive with DeepSeek on both capability and price, and asked whether raising prices right now is asking for trouble.

One more piece of context from eWeek's reporting: even a large percentage increase from this low a base could still leave DeepSeek cheaper than most competing services. University of Hong Kong professor Zhenhui Jack Jiang told the South China Morning Post that DeepSeek's mixture-of-experts design genuinely reduces the computation needed to run its models, which is part of how the prices got this low in the first place.

What happens next, as far as anyone actually knows

As of today, the current prices are still in effect. DeepSeek has not published a new rate card, named an effective date, or said which models the increase covers. The company said the final pricing plan will arrive in a future official notice.

So the honest summary is this: the lab that set the floor under AI prices, the one that forced everyone else to justify what they charge, has told its customers the floor is moving. By how much, and when, nobody outside DeepSeek knows yet.

0 Comments

Log in to join this hub and comment.

No comments yet. Be the first to reply!