Skip to content
Cogio
Sovereignty· 14 min read· by Alexandre Sauvageau

Top 5 AI models to host on your own servers in 2026: an honest comparison

The five best open-weight AI models for an SME as of July 2026: Qwen3.6, Gemma 4, Qwen3.5, Mistral Small 4 and GPT-OSS. VRAM, licences, French.

Graphics cards laid out side by side
Photo: Nana Dua, Pexels

Why host the model yourself, and how we chose

An AI model hosted on your own server is an assistant that reads your quotes, your procedures and your contracts without a single byte leaving the company. No per-token subscription, no terms of service that change under you, no CLOUD Act: the model is yours, literally, and Law 25 compliance gets simpler by the same measure.

The open-weight model market has moved a great deal in a year: Google put Gemma under the Apache 2.0 licence, Alibaba released two generations of Qwen, Mistral replaced its small dense model with a large mixture-of-experts one, and the open giants (DeepSeek, GLM, Kimi, Llama 5) grew until they no longer fit on an SME server. This comparison keeps five models on four criteria: the quality of their French, the video memory required at 4-bit quantization, the licence, and their fitness for RAG, the approach where the model answers from your documents rather than from its memory.

The video memory figures come from each model’s official Ollama or Hugging Face page, checked in July 2026. Models move fast: verify the current parameters before buying any hardware.

1. Qwen3.6-27B: the best quality-to-memory ratio on a single GPU

Released by Alibaba in April 2026, Qwen3.6-27B is the model we recommend by default for an SME with a single 24 GB GPU. Its Ollama page reports roughly 17 GB at Q4 quantization: it runs comfortably on a high-end consumer card, with room left for context.

Its strengths: a native context window of 262,000 tokens (enough to swallow a full tender with its appendices), leading reasoning for its size (its Hugging Face page shows 86.2 on MMLU-Pro and 87.8 on GPQA Diamond), an Apache 2.0 licence with no commercial restriction, and excellent multilingual behaviour, French included.

Its limits: its step-by-step thinking mode is verbose, which lengthens answers when it is not reined in, and the model is recent, so field experience is thin. Nothing blocking for an internal document assistant.

2. Gemma 4: careful French that fits in 8 GB

Google’s Gemma 4 family, released at the end of March 2026, marks a turning point: it is the first generation of Gemma published under Apache 2.0, without the restrictive terms of use that complicated commercial use of earlier versions. Three sizes interest us: 12B (7.6 GB at Q4 according to Ollama), 26B (18 GB) and 31B (20 GB).

The 12B is our pick for tight budgets and unified-memory mini-PCs: its French is remarkably natural for the size, with more than 140 languages supported and a 256,000-token context window. The 31B holds its own against the best in its class: Google ranks it third among open models on the LMArena comparison arena, a claim from its official announcement that we have not been able to verify independently.

Its limits: no long reasoning mode comparable to Qwen’s, and narrower language coverage than Qwen3.5’s 201 languages. For French-language RAG over business documents, those limits weigh little.

3. Qwen3.5-35B-A3B: the throughput champion for high-volume RAG

The 35B-A3B in the Qwen3.5 family (February 2026, Apache 2.0) is a mixture-of-experts model: 35 billion parameters in total, but only 3 billion activated per token. The result: the speed of a small model with the knowledge of a mid-sized one, for roughly 20 GB of video memory at 4-bit quantization according to the sizing guides.

This is the throughput choice: when the assistant has to handle hundreds of queries a day, summarize batches of documents or serve several departments in parallel, those 3 billion active parameters show up on the electricity bill as much as on response times. The family claims 201 languages and dialects, unheard of in open source, and the same 262,000-token context window as its bigger sibling.

Its limits: at equal memory, a dense model such as Qwen3.6-27B or Gemma 4 26B reasons more deeply about ambiguous cases. Mixture-of-experts wins on volume, dense wins on nuance.

4. Mistral Small 4: the best French, if you have two GPUs

Mistral Small 4, released on March 16, 2026, is “Small” in name only: 119 billion mixture-of-experts parameters (6 billion active), merging reasoning, vision and tool use. It has, as far as we can tell, the best native French among open-weight models: the publisher is Parisian and you can hear it in every turn of phrase.

The price of that quality is paid in memory: roughly 58 GB at IQ4_XS quantization according to the GGUF files published by the community, which means two 48 GB professional GPUs to be comfortable. Apache 2.0 licence, 256,000-token context window, remarkably concise output. For an accounting practice, an insurer or a law firm where every comma counts, the investment can be justified.

Its limits: at the time of checking, no official Ollama page was online, so you have to go through vLLM or llama.cpp, which assumes a technical team that is comfortable with both. And the hardware required doubles the server budget compared with the first three choices.

5. GPT-OSS: OpenAI’s reasoning, but French that lags

GPT-OSS, OpenAI’s pair of open models released in August 2025 (20B and 120B, Apache 2.0), remains a reference point: excellent reasoning for the size, reliable function calling, a very mature ecosystem with millions of downloads on Ollama. The 20B runs in 14 GB of video memory according to its Ollama page, the 120B needs 65 GB.

So why fifth? French. OpenAI’s official technical report places gpt-oss-20b between 73.2% and 80.2% on the multilingual MMMLU benchmark in French depending on reasoning effort, a clear notch below Qwen, Gemma and Mistral. On Quebec business documents that gap shows: anglicisms, calqued constructions, approximate terminology. The context window also tops out at 128,000 tokens, half that of its competitors.

The right use: workflows where reasoning and tooling matter more than language, such as an agent querying databases or orchestrating API calls, with short answers or answers in English. No successor had been released as of July 2026, despite the rumours.

Which model for which situation

No ranking replaces the opening question: what hardware, what language, what volume? Here is the grid we use at scoping.

  • For French-language RAG, the embedding model counts as much as the generation model: Qwen3-Embedding and BGE-M3 dominate the public comparisons in 2026.
  • The open giants (GLM-5.2, DeepSeek V4, Kimi K2.6, Llama 5) are outside the SME envelope: from 600 billion to more than 1,000 billion parameters, and Llama’s licence remains restrictive.
  • Honourable mention: NVIDIA’s Nemotron 3 Nano, very fast, but its house licence is less simple than Apache 2.0 and its French is not documented.
Recommendations by video memory tier (4-bit quantization, figures from official pages, July 2026)
Your hardwareFirst choiceAlternativeThe takeaway
8 to 16 GB (consumer GPU, mini-PC)Gemma 4 12B (7.6 GB)GPT-OSS 20B (14 GB) if reasoning matters mostThe 12B is enough for an SME document assistant
24 to 32 GB (one high-end GPU)Qwen3.6-27B (17 GB)Gemma 4 26B or 31B; Qwen3.5-35B-A3B for throughputThe sweet spot: quality, French and room for context
48 to 64 GB (two GPUs, dedicated server)Mistral Small 4 (58 GB)GPT-OSS 120B (65 GB) if French is not criticalThe tier for flawless French and heavy volumes

What this changes for a Quebec SME

First, compliance: a local model transforms your Law 25 analysis. Personal information leaves neither your company nor Quebec, which simplifies the privacy impact assessment and removes the question of transfers out of the province. Your documents are never used to train a third party’s model, by construction.

Then the bill: a server capable of running Qwen3.6-27B can be found under $10,000, and a dual-GPU server for Mistral Small 4 runs around $25,000 to $35,000 depending on the configuration. That hardware can qualify for the C3I tax credit (15% to 25% depending on the region, on the portion above the $5,000 threshold), and the implementation project itself can be funded: our overviews of Quebec and federal support set out the levers.

For transparency: Cogio deploys these models at client sites, that is our trade, and this ranking reflects our field experience as much as the spec sheets. The figures quoted were checked in July 2026 on the publishers’ official pages; they will change. For an independent view, public rankings such as LMArena or the Artificial Analysis Intelligence Index are useful counterweights.

17 GB

video memory for Qwen3.6-27B at Q4 according to its Ollama page: a single 24 GB GPU is enough

262,000

tokens of native context in Qwen3.6 and Qwen3.5: an entire tender in one go

Apache 2.0

the licence of all five models: free commercial use, with no royalty

Frequently asked questions

What exactly is an “open-weight” model?

A model whose parameters (its weights) can be downloaded and run on your own hardware, without going through a provider’s API. Open does not mean royalty-free: the licence decides. Apache 2.0, the licence of the five models kept here, allows commercial use with no royalty; Llama’s licence still imposes conditions.

What hardware do you actually need to start?

A single GPU with 24 GB of video memory covers the sweet spot: Qwen3.6-27B runs there at 4-bit quantization with room for context. For a first pilot, a 16 GB GPU and Gemma 4 12B are enough. Dual GPUs only become necessary for Mistral Small 4 or GPT-OSS 120B.

Does 4-bit quantization degrade quality?

Marginally, for most business uses. Quantization compresses the model’s weights to reduce the memory required; at Q4, the loss of quality is generally imperceptible in document RAG. For sharp reasoning tasks, finer quantization (Q8) can be justified if memory allows.

Do we need to retrain the model on our data?

Almost never. The RAG approach (the model consults your indexed documents at the moment it answers) covers the large majority of needs, updates continuously and leaves the model untouched. Fine-tuning is reserved for cases where a very specific style or vocabulary has to be internalized, and it complicates compliance.

Why is Llama not in the ranking?

Two reasons. Meta no longer offers a recent version in the SME size class: Llama 5 has 600 billion parameters, out of reach for an ordinary business server. And its licence is still not a recognized open-source licence, unlike the Apache 2.0 of the five models kept here.

Ollama, vLLM, llama.cpp: which one should we use to serve the model?

Ollama to get started and for individual workstations: installed in minutes, with a built-in catalogue. vLLM for multi-user production: better throughput, fine-grained memory management. llama.cpp when the model only exists as a community GGUF, as was the case for Mistral Small 4 at the time of writing.

Sources and references

This article is a plain-language summary, accurate as of the date shown. It is not legal advice: for your own situation, consult a legal adviser or contact the Commission d’accès à l’information.