DeepSeek V4-Flash: a retrain beats the flagship at a third of the price
AnalysisA smaller model just outscored its own flagship, and DeepSeek pulled it off without touching the architecture. The V4-Flash build it shipped on July 31, a 13-billion-active mixture-of-experts model (a design that fires only part of itself per request to save compute), was re-post-trained rather than rebuilt. DeepSeek's own figures put it ahead of the larger V4-Pro preview on all nine agent and coding tests it published, including 82.7 on Terminal-Bench. Output runs about $0.28 per million tokens, near a third of Pro, and it now speaks OpenAI's Responses API, so Codex can call it directly. The scores are vendor-reported, but the lesson for anyone paying an inference bill is that post-training has become the cheapest place to buy capability.