GLM-5.3-Flash: what Zhipu’s open model changes for local AI in business
Zhipu’s GLM-5.3-Flash under the MIT licence: 320 billion parameters, multimodal and very cheap. The server it needs and what it is worth to an SME.

What came out, and how
GLM-5.3-Flash had an unusual launch. For weeks, an anonymous model called “Ox Alpha” circulated on the OpenRouter marketplace and the OpenCode agent platform. According to the South China Morning Post, it processed 62 trillion tokens there before Zhipu AI claimed it on Wednesday, August 26, 2026, publishing the weights on Hugging Face under the MIT licence.
The other headline is industrial: the trial ran entirely on a cluster of 100,000 Chinese-made chips, the same paper reports. Zhipu’s shares jumped more than 12% on the Hong Kong exchange the next day. For a business here, the question is practical rather than geopolitical: can a model of this calibre, free and open to modification, replace a subscription or run on your own server?
The spec sheet in plain terms
GLM-5.3-Flash is a mixture-of-experts model: 320 billion parameters in total, of which only 18 billion work on each token it produces. The published configuration lists 288 experts, 8 of them called on at each step. That is why its API price is so low: you pay for the compute of 18 billion parameters, not 320.
Zhipu also reworked the architecture. For the first time in the GLM series, three layers out of four use linear attention and the fourth uses sparse attention, a design that sharply cuts the cost of long contexts. The configuration declares a window of about one million tokens (1,048,576). The model was pre-trained on a 30-trillion-token multimodal corpus, according to its Hugging Face card.
18B
active parameters per token, out of 320 billion in total
MIT
licence of the weights, with no commercial restriction
42
score on the Artificial Analysis Intelligence Index, fourth out of 115 comparable models
| Feature | Value |
|---|---|
| Parameters | 320 billion in total, 18 billion active |
| Inputs | Text, images, video, scanned documents |
| Declared context | 1,048,576 tokens |
| Licence | MIT |
| Official weights (FP8) | 328 GB |
| 4-bit quantization (Unsloth) | 200 GB |
| Z.ai API price | US$0.15 input, US$0.50 output, per million tokens |
Performance, from press release to independent measurement
Zhipu says GLM-5.3-Flash outperforms GLM-5.2 at a tenth of the price and approaches the best commercial models on coding and agent tasks. Its card lists 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE, among others. These are the vendor’s figures, measured under its own conditions.
Artificial Analysis’s independent measurement adds nuance. The model scores 42 on the Intelligence Index, well above the median of 18 for comparable models, which puts it fourth out of 115. But it generates slowly, at 43.8 tokens per second against a median of 74, and it is wordy: 180 million tokens produced during the evaluation, against a median of 140 million. For an employee waiting on an answer, that slowness shows.
We found no published evaluation of its French at the time of writing. That is the first thing to test on your own documents before any rollout.
The server it takes to host it
This is where GLM-5.3-Flash parts ways with the models we usually recommend to SMEs. The official weights weigh 328 GB in FP8. The most widely used 4-bit quantization, published by Unsloth, comes down to 200 GB and the 2-bit version to 102 GB, at a noticeable cost in quality. That is a long way from the 17 to 19 GB a 27-billion-parameter model needs at 4 bits.
Two kinds of machines fit the bill. One is a unified-memory workstation with 256 GB or more, where the small number of active parameters keeps generation speed reasonable. The other is a server with several professional GPUs, or a hybrid engine such as KTransformers, listed on the model card, which keeps the experts in system memory and the rest on the graphics card. Either way, the hardware budget moves up an order of magnitude compared with a single-GPU server.
- Official FP8 weights: 328 GB
- Unsloth 4-bit quantization: 200 GB
- Unsloth 2-bit quantization: 102 GB, with degraded quality
- Vision encoder: about 1 GB more to read images
A Chinese model on your server: is it a risk?
Two uses need to be told apart. Open weights, downloaded and run on your server, send nothing to Zhipu: the model is a file and connects to no service. Under the MIT licence, you can inspect it, adapt it and run it without depending on the vendor. That is exactly the logic of local AI.
The Z.ai API is a different matter. Sending personal information to it means communicating it outside Quebec. Law 25 then requires a privacy impact assessment before the transfer, as our practical PIA guide explains. For sensitive business data, the API’s very low price does not make up for that compliance work.
Then there is alignment: models trained in China follow cautious guidelines on certain political subjects. For reading invoices, summarizing procedures or preparing quotes, the effect is nil in practice. A test bench on your own documents confirms it in a day.
What we recommend
For most Quebec SMEs, a model of around 27 billion parameters on a 24 GB GPU remains the best starting point: cheaper, faster and good enough for document work. Our Qwen3.8 analysis covers that tier in detail.
GLM-5.3-Flash becomes interesting in two cases. First, for a business running long, complex agents, where its level on agent tasks justifies a large-memory server. Second, for reading visual documents in volume, thanks to its native multimodality. In both cases, the choice is made on a test bench, with your documents and in French, not on a public leaderboard.
The hardware for such a project remains eligible for the C3I tax credit and the diagnostic that comes before it can be funded through ESSOR. Our comparison of local models places GLM-5.3-Flash among the other open options.
Frequently asked questions
Can GLM-5.3-Flash run on a 24 GB GPU?
No. Even Unsloth’s most aggressive 1-bit quantization weighs 93 GB, and the 4-bit version 200 GB. You need a unified-memory workstation with 256 GB or more, a server with several professional GPUs, or a hybrid engine that keeps the experts in system memory.
Does the MIT licence allow commercial use?
Yes, without restriction. The MIT licence allows use, modification and redistribution, including commercial, as long as the copyright notice is kept. No user or revenue threshold applies.
Does using GLM-5.3-Flash send our data to China?
Not if you host the weights yourself: the model runs on your server and talks to no service. The Z.ai API, however, processes your requests on the vendor’s infrastructure. For personal information, Law 25 then requires a privacy impact assessment before the transfer.
Is it better than Qwen3.8-27B?
On agent tasks, the published scores put it ahead, but they come from the vendors and are not measured under the same conditions. More importantly, the two models are not in the same hardware class: Qwen3.8-27B fits on a 24 GB GPU, GLM-5.3-Flash needs roughly ten times the memory. The right comparison is made on your documents and your budget.
Why is the model so slow according to Artificial Analysis?
The evaluation measures the Z.ai API, which produces 43.8 tokens per second against a median of 74 for comparable models. The model is also verbose, which lengthens its answers. On your own server, speed depends on the hardware and the quantization you choose.
Sources and references
- Official GLM-5.3-Flash Hugging Face card (architecture, benchmarks, MIT licence)
- Published GLM-5.3-Flash configuration (experts, layers, context)
- Z.ai, API pricing (GLM-5.3-Flash and GLM-5.3)
- Unsloth, GLM-5.3-Flash GGUF quantizations (file sizes)
- Artificial Analysis, independent evaluation of GLM-5.3-Flash
- South China Morning Post, Ox Alpha revealed as GLM-5.3-Flash on Chinese chips
- TechNode, Zhipu identifies Ox Alpha and releases the weights (August 27, 2026)
This article is a plain-language summary, accurate as of the date shown. It is not legal advice: for your own situation, consult a legal adviser or contact the Commission d’accès à l’information.
Read next
Sovereignty
Qwen3.8: what Alibaba’s new family changes for local AI in the enterprise
Read →
Sovereignty
Top 5 AI models to host on your own servers in 2026: an honest comparison
Read →
Sovereignty
DeepSeek-V4.1-Flash: an open model built to run lean, analyzed for business
Read →
A question about your own compliance?
The discovery call is free and takes half an hour. You leave with an honest read on your situation.
Let’s talk