DeepSeek Launches V4.1 Flash With Faster Inference and Lower AI Costs

DeepSeek has launched V4.1 Flash, a faster multimodal AI model built for efficiency.

The new model adds to China’s increasingly competitive AI market, where developers are focusing not only on benchmark scores but also on the cost of running models at scale. Reuters also reported that DeepSeek is preparing for a potential listing on Shanghai’s STAR Market.
A Large Model With Lower Active Compute
DeepSeek V4.1 Flash uses a Mixture-of-Experts architecture with a 552-billion-parameter backbone. Its Causal Encoder-Decoder design activates around 8 billion parameters while processing input and 16 billion during output generation, rather than using the full model for every token.
The model supports both text and images. DeepSeek says the design is intended to improve inference speed and throughput while keeping computing requirements comparatively low.
Lower Cache Requirements Could Cut Costs
A key efficiency gain comes from the model’s KV cache, which stores information needed while generating responses. DeepSeek says V4.1 Flash requires about one-quarter of the HBM memory and one-eighth of the SSD storage used by the previous generation.
Those reductions could matter for AI agents and other applications that repeatedly process long prompts, where memory and storage costs can become a major part of deployment expenses.
DeepSeek says the lower resource requirements allow more requests to be served at reduced cost. The company also continues to offer cheaper off-peak API pricing for workloads that can run outside periods of high demand.
Earlier Flash Models Are Being Retired
V4.1 Flash is available through DeepSeek’s API under the model name deepseek-flash. The previous V4 Flash and V4 Flash Vision Experimental models have been retired, with their API names temporarily redirected to the new model.
DeepSeek also plans to phase out V4 Pro. From 14 September 2026 at 04:00 UTC, requests to V4 Pro are set to route to V4.1 Flash and use V4.1 Flash pricing until V4.1 Pro becomes available.
The company has released V4.1 Flash to the open-source community and says it will work with developers on broader inference and deployment support.
Efficiency Is Becoming a Bigger AI Battleground
For developers, the most important part of V4.1 Flash may be its economics. As AI models are built into agents and always-on services, lower memory use and cheaper inference could become as important as raw benchmark performance.
That gives DeepSeek another way to compete with larger Chinese and US AI developers while the industry looks for stronger models that do not require ever-growing infrastructure budgets.


