Skip to content
Cogio
Sovereignty· 9 min read· by Alexandre Sauvageau

DeepSeek-V4.1-Flash: an open model built to run lean, analyzed for business

DeepSeek-V4.1-Flash under the MIT licence: 552 billion parameters, 8 to 16 billion active and a far smaller context memory. What it is worth to an SME.

A server’s drive bays lit in blue, status lights on
Photo: panumas nikhomkhai, Pexels

What came out

On September 10, 2026, DeepSeek introduced DeepSeek-V4.1-Flash as “smarter, faster, more efficient.” The weights went up on Hugging Face the same day under the MIT licence, alongside a technical paper. In the vendor’s API, the model is simply called “deepseek-flash” and it absorbs the older models: requests to V4-Flash are routed to it, and requests to V4-Pro have been since September 14.

That replacement says everything about the model’s ambition. According to DeepSeek, this “Flash” beats the company’s former top-tier model on agent tests. If the measurement holds, the vendor’s cheapest model is also its best.

An architecture designed for restraint

DeepSeek-V4.1-Flash has 552 billion parameters spread across 384 experts, 6 of which are called on per token. Its novelty is a “causal encoder-decoder” architecture: 20 layers read the input, 20 more write the answer. As a result, only 8 billion parameters work while reading and 16 billion while writing.

The main gain is in the KV cache, the memory the model keeps of the conversation in progress. DeepSeek says it takes a quarter of the GPU memory and an eighth of the SSD storage the previous generation needed. According to The Register, that means serving four to eight times as many users in the same memory footprint. For a business, that is the right metric: how many employees can one server handle at the same time?

The model accepts text and images, with a context window of about one million tokens and answers of up to 384,000 tokens in the API. Its reasoning effort can be tuned finely, on a scale of 1 to 100, so it can answer simple questions quickly and think longer about hard problems.

8B

active parameters to read the input, 16 billion to write the answer

¼

KV cache GPU memory compared with the previous generation, according to DeepSeek

235

tokens per second measured by Artificial Analysis on the API

DeepSeek-V4.1-Flash by the numbers (DeepSeek announcement, Hugging Face card and API documentation)
FeatureValue
Parameters552 billion in total, 8 billion active reading, 16 writing
InputsText and images
Context1 million tokens, 384,000 output (API)
LicenceMIT
Official weights510 GB
API price, off-peakUS$0.15 input, US$0.60 output, per million tokens
API price, peakUS$0.30 input, US$1.20 output, per million tokens

Performance, from press release to independent measurement

The model card lists a score of 74.2 on DeepSWE 1.1, which measures the ability to resolve real programming tasks. As always, that figure comes from the vendor.

Artificial Analysis gives an independent picture: 39 on the Intelligence Index, seventh out of 115 comparable models, at a speed of 235.1 tokens per second. That is slightly below GLM-5.3-Flash on the index (42), but more than five times faster. For an assistant a whole team uses all day, speed often matters more than a few leaderboard points.

As with the other recent Chinese models, we found no published evaluation of its French. A test bench on your own documents remains essential.

Hosting it yourself, in practice

The official weights weigh 510 GB. A large part of that volume is a memory table DeepSeek calls “Engram”: about 189 GiB in FP8 that the model looks up row by row rather than keeping entirely in memory.

That is what makes a modest setup possible. Developer Salvatore Sanfilippo (antirez) publishes a 2-bit version whose main weights come to 152 GiB. According to his documentation, it runs on a single 128 GB Mac by reading the Engram table straight from a fast SSD. A 4-bit version needs a 512 GB machine to keep everything in memory. In production, with several employees at once, a server with several professional GPUs remains the reference setup.

  • Official weights: 510 GB
  • Engram table: about 189 GiB, read from disk
  • 2-bit version (antirez): 152 GiB of main weights, a 128 GB Mac with a fast SSD
  • 4-bit version (antirez): 294 GiB of main weights, a 512 GB machine to keep everything in memory

Open weights or API: two different answers

The open weights of DeepSeek-V4.1-Flash, run on your server, send nothing to DeepSeek. Under the MIT licence, the model is a file you control entirely, with no subscription and no dependence on the vendor.

The API is tempting on price, especially off-peak. But DeepSeek processes requests on its own infrastructure, in China. Sending personal information there requires, under Law 25, a prior privacy impact assessment. For contracts, quotes or employee files, local hosting settles the question at the source.

What we recommend

DeepSeek-V4.1-Flash is the most interesting open model of the fall for a business that wants to serve many users from one server: its reduced context memory and its speed make it a strong candidate for an assistant shared by a whole team.

It is not, however, where a small business should start. A model of around 27 billion parameters on a 24 GB GPU covers most document needs for a fraction of the hardware cost, as our Qwen3.8 analysis explains. If you are torn between DeepSeek-V4.1-Flash and GLM-5.3-Flash, speed favours the first and the index score the second. Only a test bench on your own documents really decides it.

Frequently asked questions

Can DeepSeek-V4.1-Flash run on a single computer?

Yes, with trade-offs. According to the documentation of the 2-bit version published by antirez, a 128 GB Mac runs it by reading part of the model from a fast SSD. That is enough for testing or individual use. To serve a whole team, a server with several professional GPUs is a better fit.

What does “8 billion active parameters reading, 16 writing” mean?

The model has 552 billion parameters but only calls on a small share of them at each step. Its new architecture separates reading the input (8 billion active) from writing the answer (16 billion active). That reduces the compute needed and speeds up generation.

Does the MIT licence allow commercial use?

Yes, with no size or revenue restriction. The MIT licence allows use, modification and commercial redistribution, as long as the copyright notice is kept.

Does using DeepSeek send our data to China?

Only if you go through DeepSeek’s API, which processes requests on its infrastructure. Open weights hosted on your server connect to no service. For personal information, local hosting avoids the transfer outside Quebec and the prior assessment that Law 25 would require.

Should we choose DeepSeek-V4.1-Flash or GLM-5.3-Flash?

According to Artificial Analysis, GLM-5.3-Flash scores slightly higher (42 against 39), but DeepSeek-V4.1-Flash generates more than five times faster on the API (235 against 43.8 tokens per second). For a shared assistant, DeepSeek’s speed and smaller context memory weigh heavily. The final choice is made on your own documents.

Sources and references

This article is a plain-language summary, accurate as of the date shown. It is not legal advice: for your own situation, consult a legal adviser or contact the Commission d’accès à l’information.