New Deepseek Flash model matches OpenAI’s GPT-5.6 Luna at roughly 60 percent lower cost

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane July 31, 2026 1 min read
New Deepseek Flash model matches OpenAI’s GPT-5.6 Luna at roughly 60 percent lower cost

Deepseek has released V4 Flash “0731,” a budget model that now scores within one point of OpenAI‘s GPT-5.6 Luna while costing roughly 60 percent less per task. This performance gain comes alongside an 80 percent price cut from OpenAI, yet Deepseek maintains an 98 percent cache discount compared to the industry-standard 90 percent. The new version uses 12 percent fewer tokens than its April predecessor and improves across every tested category, with the largest gains appearing in agentic tasks. On the GDPval benchmark for complex office work, the model climbed from 1,189 to 1,559 Elo points while hallucinating less often. The architecture remains unchanged at 284 billion total parameters with 13 billion active, and the weights are available under an MIT license on Hugging Face.

The significance lies in the narrowing gap between premium and budget options for developers. Lower inference costs mean smaller teams can access capabilities previously reserved for larger budgets without sacrificing performance.

* The model retains a one-million-token context window
* Active parameters remain at 13 billion despite the upgrade
* Weights are distributed freely under an MIT license

Scroll to Top