DeepSeek Launches V4.1-Flash: A Faster, Multimodal Mixture-of-Experts Model

by priyanka.patel tech editor
DeepSeek V4.1 Flash AI model release featured image

Chinese artificial intelligence startup DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, positioning the new mixture-of-experts model around faster inference, lower deployment costs, and native multimodal vision across its web app, mobile app, and first-party API.

Chinese artificial intelligence startup DeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in a new architecture family built with native multimodal vision. The release went live across the company’s web app, mobile app, and first-party API, with open MIT-licensed weights published on Hugging Face as deepseek-ai/DeepSeek-V4.1-Flash. According to the company’s API changelog, developers should call the model as deepseek-flash. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp continue to function as temporary aliases, but the previous Flash and Flash Vision Exp models have been retired, with those requests now served by V4.1 Flash at Flash pricing.

Architecture and Cache Efficiency Improvements

Rather than a minor tune-up, DeepSeek describes V4.1 Flash as built for a higher capability ceiling, faster inference, higher throughput, and room to scale into larger models in the same family. Independent write-ups summarizing the Hugging Face model card state that the instruct model uses a large mixture-of-experts backbone with sparse activation. It packs roughly 552 billion backbone parameters, activating about 8 billion parameters per token during prefill and 16 billion during decode, alongside a 1 million-token context window and up to 384K maximum output tokens.

The architecture also features a new causal encoder-decoder design. Engineering notes highlight cache-efficiency features such as compressed sparse attention and FP4 KV caching. Specifically, the startup said V4.1-Flash requires one-fourth the high-bandwidth memory and one-eighth the SSD storage for its key-value cache compared with the previous generation. This reduction aims to lower expenses for AI agents, where cache-hit charges can account for a substantial share of operating costs.

Vendor Benchmarks and Performance Claims

DeepSeek published instruct scores alongside the launch, reporting strong results across several standard technical evaluations. The vendor-reported figures include GPQA Diamond at 90.9, a Codeforces rating of 3471, MathArena Apex at 65.6, Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2, and HLE with tools at 63.9. Pure-text HLE without tools is listed at 36.8, with a 39.1 score noted on a text-only subset in the changelog.

DeepSeek Launches V4.1-Flash: A Faster, Multimodal Mixture-of-Experts Model
Photo: cxodigitalpulse.com

The company stated that extensive testing shows V4.1 Flash outperforms DeepSeek V4 Pro on performance, cost, speed, and total processing time. These capability gains stem from new pre-training methods and larger-scale reinforcement learning post-training. However, independent evaluation sites had not yet published full third-party measurements for V4.1 Flash in early coverage, leaving vendor documentation as the primary source for these benchmarks.

API Pricing and the Pro Model Retirement Plan

With the release, DeepSeek’s Models & Pricing page lists peak and off-peak rates for deepseek-flash. Off-peak prices per 1 million tokens stand at $0.003 for cached input, $0.15 for uncached input, and $0.60 for output. Peak rates are double those figures, running during peak hours of 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, with all other hours classed as off-peak.

Wide landscape preview
Photo: enterpriseai.economictimes.indiatimes.com

By comparison, the still-listed V4 Pro card remains far higher, charging $1.98 off-peak and $3.96 peak per million output tokens. Because V4.1 Flash undercuts Pro significantly, the upcoming routing change marks a substantial bill shift for developers. After 12:00 Beijing Time on September 14, 2026, and until a future V4.1 Pro ships, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at V4.1 Flash prices.

Developer Ecosystem and IPO Preparations

On the API side, V4.1 Flash supports JSON output, tool calls, the Responses API, Anthropic-compatible endpoints, and native vision in the base model. Concurrency limits on the pricing page show a higher Flash ceiling than Pro, registering at 2,500 compared to 500. The company is also working with the open-source community on inference support and deployment options, including large-scale deployments involving 2,000 GPUs and storage clusters.

DeepSeek V4.1 Flash Is INSANE – Is THIS the Best Open Model Yet?

This technical rollout coincides with corporate financial maneuvers. Reuters reported that DeepSeek has tapped CITIC Securities to prepare for an initial public offering on Shanghai’s technology-focused STAR Market. While the startup aims to begin the process this year, the timing, size, and valuation remain undecided. The company raised about $7.4 billion in June at a post-money valuation of more than $50 billion, while a funding round underway could value it at about 500 billion yuan, equivalent to $75 billion, to support computing infrastructure, model development, and talent retention.

DeepSeek V4 Flash Is INSANE – The Best Small Model Yet!

You may also like