The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence...
For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.
- 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.
- 86.6 on Terminal Bench 2.1. Pro 0813 is better.
- 55.9 on NL2Repo. Pro 0813 is better.
- 27 on Agent's Last Exam. Pro 0813 is a little worse.
- 72.5 on Toolathon-Verified. Pro 0813 is better.
- 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.
- 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.
I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.
bel8 6 minutes ago [-]
So it's a Fable class LLM?
DSV4Pro vs Fable5
HLE w tools 60.0 vs 63.0
Terminal Bench 2.1 87.9 vs 88.0
Cybergym 83.3 vs 83.1
DeepSWE 62.7 vs 70.0
Toolathlon-Verified 74.1 vs 77.9
AutomationBench (Public) 31.8 vs 29.1
DSBench-FullStack 71.1 vs 77.2
DSBench-Hard 67.2 vs 68.3
eli 3 minutes ago [-]
Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5
aftbit 4 minutes ago [-]
Fabble lol
goldenarm 4 minutes ago [-]
Geometric mean of all these benchmarks :
* GPT-5.6 Sol: 65.5
* Fable 5 (w/ fallback): 64.5
* Opus 5: 64.0
* DS-V4-Pro 0813: 62.5
* Kimi-K3: 62.3
* DS-V4-Flash 0731: 55.8
* GLM-5.2: 47.3
Gecko4072 35 minutes ago [-]
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount
1 minutes ago [-]
nchmy 10 minutes ago [-]
i dont see any price increase there... what am i missing?
minraws 11 minutes ago [-]
isn't it the same old pricing? did they increase V4 Pro pricing already?
igravious 8 minutes ago [-]
yup :)
i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)
how are you doing it?
am using Kimi K3 via kimi-code
and GLM 5.2 via ZCode
happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex
book_mike 11 minutes ago [-]
What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
okamiueru 6 minutes ago [-]
How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
xynelius 3 minutes ago [-]
If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]:
For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.
Cost per request for V4 Pro: $0.000875 per request.
Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.
How does it stack against the updated Deepseek Flash version?
pixelesque 26 minutes ago [-]
I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev).
Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
surgical_fire 6 minutes ago [-]
I use a plan -> implement wotkflow for this reason.
pro plans, flash implements. I am super happy with how flash behaves like that.
swiftcoder 23 minutes ago [-]
yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow
k__ 41 minutes ago [-]
Around 5 percentage points better. (E.g., 87% instead of 82%)
Gecko4072 36 minutes ago [-]
So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.
networked 23 minutes ago [-]
I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis pages (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. It critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.
saaga 30 minutes ago [-]
Yea that's what I was thinking.
Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.
I was running a session over a couple days and it didnt cross a dollar lol.
npn 13 minutes ago [-]
I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.
k__ 34 minutes ago [-]
I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.
Wasn't worth it.
sparkling 32 minutes ago [-]
deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.
saaga 29 minutes ago [-]
I feel the same too. I like the speed.
I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.
k__ 30 minutes ago [-]
I wouldn't exactly call it snappy, but faster than Pro, yes.
ericd 26 minutes ago [-]
Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.
JacobAsmuth 16 minutes ago [-]
Well sure but you're running on tens of thousands of dollars of hardware.
I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
ianm218 4 minutes ago [-]
I suspect if you follow dev groups in developing countries people are much more focused on token/ price efficiency.
For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale
sinuhe69 12 minutes ago [-]
Well, one reason is that we always have to work with the quirks of each model. So, a know model is often preferred over a new/unknown one because we have to be vigilant again. (Negative) surprises are mentally exhausting in the long run.
IMO, you can work much better when you know the model.
spacebanana7 14 minutes ago [-]
In an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.
BlackRabbit1 3 minutes ago [-]
There are EU/US providers offering Deepseek/Qwen/Kimi/etc.-as-a-Service. With zero ties of their infrastructure to China.
Fully compatible with the well known Antrophic API.
You only have to replace the URL and your key.
HawtAds 17 minutes ago [-]
Hacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.
BlackRabbit1 16 minutes ago [-]
A lot of it/infrastructure departments aren't aware that you can use Asian models hosted within the US or even EU.
Rendered at 17:23:20 GMT+0000 (Coordinated Universal Time) with Vercel.
For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.
- 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.
- 86.6 on Terminal Bench 2.1. Pro 0813 is better.
- 55.9 on NL2Repo. Pro 0813 is better.
- 27 on Agent's Last Exam. Pro 0813 is a little worse.
- 72.5 on Toolathon-Verified. Pro 0813 is better.
- 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.
- 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.
I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.
* GPT-5.6 Sol: 65.5
* Fable 5 (w/ fallback): 64.5
* Opus 5: 64.0
* DS-V4-Pro 0813: 62.5
* Kimi-K3: 62.3
* DS-V4-Flash 0731: 55.8
* GLM-5.2: 47.3
The prices on OpenRouter still look the same.
edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount
i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)
how are you doing it?
am using Kimi K3 via kimi-code
and GLM 5.2 via ZCode
happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.
Cost per request for V4 Pro: $0.000875 per request.
Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.
[1] https://opencode.ai/docs/go/#usage-limits
Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
pro plans, flash implements. I am super happy with how flash behaves like that.
I was running a session over a couple days and it didnt cross a dollar lol.
Wasn't worth it.
For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale
Fully compatible with the well known Antrophic API.
You only have to replace the URL and your key.