top of page

DeepSeek V4.1 Flash AI Model Outperforms Competitors, Boosts Efficiency

Writer: tech360.tv
tech360.tv
5 minutes ago
3 min read

DeepSeek has released its V4.1 Flash model. It claims to outperform its previous flagship while cutting inference costs and boosting speeds. This move furthers China's artificial intelligence market developments. The new model reportedly performs strongly against other notable AI offerings, according to the South China Morning Post (SCMP).


Smartphone screen showing DeepSeek new chat page with whale logo and the text Hi, I'm DeepSeek. How can I help you today?
Credit: UNSPLASH

The Chinese AI developer stated V4.1 Flash employs a new Causal Encoder Decoder architecture. Built on a substantial 552 billion parameter framework, the model relies on a Mixture of Experts (MoE) design. Unlike traditional AI systems that route every query through their entire structure, an MoE system directs tasks only to specific subnetworks. DeepSeek reported significantly reduced computing power per request by activating just 8 billion parameters for input processing and 16 billion for generating responses.


And DeepSeek presented V4.1 Flash as the smallest model in its new series. It offers native multimodal visual understanding. The company stated the model surpassed V4 Pro on benchmarks assessing coding, cybersecurity, and autonomous agent tasks. This launch occurs as Big Tech firms and pure play AI developers contend in a fierce battle to commercialise artificial intelligence across China.


Rising hardware costs and foreign chip export restrictions limit compute capabilities for Chinese players. These developers are thus accelerating efforts to offer efficient models. Such models deliver high end reasoning at fraction of a cent operational costs. On Terminal Bench 2.1, a test for real world compute tasks, V4.1 Flash achieved a score of 90.6. This exceeded OpenAI's GPT-5.6 Sol at 88.8, Moonshot AI's Kimi K3 at 88.3, and DeepSeek's own V4 Pro at 87.9.


On Cybergym, a benchmark for artificial intelligence models on cybersecurity tasks, V4.1 Flash scored 88.1. This placed it above V4 Pro at 83.3, Kimi K3 at 80, and OpenAI's GPT-5.6 Sol at 84.5. V4.1 Flash narrowly edged out Anthropic's Claude Opus 5 on the DeepSWE v1.1 software engineering benchmark. The American model, however, retained leads on other tests.


So, a main advancement of the new architecture involved memory optimisation. DeepSeek stated it reduced the model's key value (KV) cache. This temporary memory storage assists an AI in recalling earlier parts of a conversation. The company reported shrinking this memory footprint to 890 bytes per token. This represents a reduction from 3,514 bytes in the previous Flash version.


This reduction in memory footprint significantly lowered inference costs for AI agents. DeepSeek also adjusted its application programming interface (API) rates. These were set as low as 0.02 yuan (0.003 US cents) per million cached input tokens during off peak hours.


Cui Tianyi, head of DeepSeek's Harness team, addressed the change on social media platform X. He stated that it "would not be appropriate to provide DeepSeek users with the originally underperforming V4 Pro model at a higher price," given V4.1 Flash's superior performance across all metrics. DeepSeek informed users it would temporarily route all V4 Pro requests to the cheaper Flash model until its more powerful V4.1 Pro became available.


And this release shows a developing trend among Chinese developers. These organisations are now competing with "Flash" variants to attract clients. Last month, Z.ai unveiled its GLM-5.3-Flash, previously code named Ox Alpha. This model selectively activates 18 billion of its 320 billion parameters.


Alibaba Group Holding, the South China Morning Post owner, launched Qwen3.8-Flash-Next. This variant activates just 6 billion of its 125 billion parameters. Such introductions illustrate an ongoing focus on efficient and cost effective AI solutions within the competitive Chinese market.


  • DeepSeek's V4.1 Flash model claims superior performance and reduced operational costs.

  • The model utilises a Causal Encoder Decoder architecture and a Mixture of Experts design for efficiency.

  • It reportedly surpasses competitor models from OpenAI, Moonshot AI, and Anthropic on specific benchmarks.

  • Memory optimisation reduced the key value cache, significantly lowering inference costs.

  • DeepSeek has lowered its API rates and will redirect older model requests to V4.1 Flash.


Source: SCMP

Technology increasingly permeates every facet of our lives, making informed decision making an essential pursuit. We bridge this gap by combining the precision of AI with the irreplaceable discernment of human expertise. Our team produces rigorous product reviews that offer unique insights, honest critiques, and trustworthy recommendations. We also leverage AI to synthesise complex news from reliable sources into clear, actionable updates, ensuring that every story is carefully fact checked by our editorial staff before publication. Accuracy remains our priority. Should you identify any discrepancies, please contact us at editorial@tech360.tv. Your feedback is a vital part of our process in maintaining the high standards our readers deserve.

Tech360tv is Singapore's Tech News and Gadget Reviews platform. Join us for our in depth PC reviews, Smartphone reviews, Audio reviews, Camera reviews and other gadget reviews.

  • YouTube
  • Facebook
  • TikTok
  • Instagram
  • Twitter
  • LinkedIn

© 2021 tech360.tv. All rights reserved.

bottom of page