DeepSeek published a formal release of DeepSeek-V4-Flash-0731, describing it as the official release of DeepSeek-V4-Flash and saying it supersedes the earlier preview while strengthening agentic capabilities.
The company says the new checkpoint keeps the same model structure as DeepSeek-V4-Flash-DSpark, meaning it uses a speculative decoding module in the same framework as the earlier Flash-DSpark model family.
DeepSeek’s model card also says V4-Flash-0731 outperforms the previous DeepSeek-V4-Pro preview in listed benchmarks while remaining broadly competitive with top proprietary options.
The release adds OpenAI-format support guidance via a dedicated encoding workflow and exposes three reasoning effort levels (low, high, max) for controlling deliberation before answers.
For serving, DeepSeek’s docs show DSpark-enabled vLLM launch flags and SGLang settings, and they recommend 384K maximum output length for high and max reasoning in agentic workflows.
The technical report linked from the card is titled Toward Highly Efficient Million-Token Context Intelligence, matching the same series emphasis on long-context performance and model-card guidance for this Flash branch.

