RAG, fine-tuning, agents, automation: a decision-maker’s AI glossary
RAG, fine-tuning, agents, tokens, local inference: every term explained in business language, with when it is the right choice and when it is not.

Why a decision-maker’s glossary, and not one more dictionary
You do not need to know how to program to decide on an AI project. You need to understand a dozen words, because those words determine choices that cost money or make it: hosting on your own premises or renting an API, indexing your documents or retraining a model, buying a chatbot or building an agent. When the vocabulary is fuzzy, it is rarely the buyer who benefits.
So this glossary is a decision tool: each term is explained in business language, with the moment it is the right choice and the moment it is not. Many projects that go off the rails start from a misunderstanding about vocabulary, as we show in our analysis of why AI projects fail in SMEs.
Language model, open weights, parameters: what you are actually buying
A language model is software trained on vast bodies of text to produce the most plausible continuation of a sentence. It is a probability engine, not a database: it writes, summarizes, rephrases and reasons, but it “knows” nothing about your business until you show it your documents.
An open-weight model can be downloaded and run on your own server, like the ones in our comparison of models to host yourself. The parameters, the billions of numerical values that make up the model, are its “brain size”: the more there are, the more nuanced the model, and the more memory it demands. Quantization compresses those parameters so a model fits on reasonable hardware, with a marginal loss of quality in most business uses.
- An open-weight model: the right choice when your documents are sensitive or you want predictable cost; an unnecessary detour if your use is limited to occasional public text.
- Many parameters: the right choice for nuance and reasoning; a waste for sorting email or extracting simple fields.
- Quantization: the right default for respecting a hardware budget; to be avoided in its most aggressive form when every comma in a legal text matters.
RAG, fine-tuning, training: three words vendors confuse, sometimes deliberately
RAG (retrieval-augmented generation) has the model consult your documents at the moment it answers: it finds the relevant passages, then writes from them. Your files stay files, the model does not change. It is the default approach for querying a company’s knowledge, because it updates continuously and it can cite its sources.
Fine-tuning changes the model itself from chosen examples. It excels at internalizing a style, a format or a trade terminology; it is poor at learning facts, which go stale and are hard to correct. Full training means building a model from scratch: that is the business of large laboratories, not of an SME, and it is almost never the right answer to your need.
- RAG: the right choice for answering from your procedures, contracts and history; not enough on its own when the problem is the tone or the format of the output.
- Fine-tuning: the right choice for an in-house style or a specialized vocabulary repeated at scale; the wrong choice for “learning” information that changes every month.
- Full training: rule it out from the start; if a vendor proposes it, ask them why RAG is not enough, and write the answer down.
Agent, assistant, chatbot, RPA: who does what while you sleep
A chatbot converses: it answers questions in a chat window, usually within a narrow scope. An assistant helps a person with their task: it writes, summarizes, prepares, but does not act on its own. An agent chains actions together: it queries your systems, cross-references data, prepares a complete deliverable, then submits it for approval. That is the sense we give the word at Cogio: AI prepares, people decide.
Classic automation (RPA, robotic process automation) is nothing like intelligent, and that is its strength: it replays identical steps, fast and without tiring. When a process is perfectly repetitive, RPA is more reliable and cheaper than an agent. The agent becomes relevant when you have to understand a document, judge a particular case or combine several sources; its cost depends above all on how many systems have to be connected.
- Chatbot: the right choice for the first line of repetitive questions; frustrating as soon as you ask it to act.
- Assistant: the right choice for equipping a person with their writing and their research; it replaces no process.
- Agent: the right choice when a deliverable requires consulting several systems; overkill for a task a simple rule already handles.
- RPA: the right choice for repetitive, identical actions; brittle as soon as the documents vary or the interface changes.
Tokens, context window, hallucination: the limits nobody sells you
Models do not read words, they read tokens: fragments of text, often shorter than a word. API billing is counted per token, and the context window, meaning how much text the model can consider at once, is measured in tokens too. A large window lets it swallow a full tender with its appendices; a small one forces you to slice it up, at the risk of losing the thread.
Hallucination is the most discussed flaw: the model asserts something false with complete confidence. It happens mostly when you push it to answer beyond what you have shown it. You do not eliminate it, you contain it: RAG that cites its sources, well-defined scopes, and human sign-off on anything that commits the company.
Local inference or cloud API: where is your data allowed to go?
Inference is the moment the model does the work: it reads your request and produces its answer. It can happen on your own server (local inference) or at a provider’s, through a cloud API. The fundamental difference is not raw quality: it is the journey your data takes.
Locally, your documents never leave the company: the Law 25 compliance analysis gets simpler and the US CLOUD Act drops out of the equation. Through an API, you reach the most powerful models without buying hardware, but you have to examine the contract, the data residency and what is done with it. Our comparison of local AI and ChatGPT for business helps you decide according to how sensitive your documents are.
- Local inference: the right choice for sensitive documents, steady volumes and predictable costs; an unjustified investment for sporadic use on public content.
- Cloud API: the right choice for prototyping fast and reaching state-of-the-art models; the wrong default for personal information, until the analysis has been done.
Embeddings and semantic search: finding by meaning, not by exact word
Embeddings turn each passage of text into a sequence of numbers that captures its meaning. Two passages about the same thing get neighbouring numbers, even with no words in common. That is the engine behind semantic search: type “conveyor failure” and find the record titled “line 3 shutdown.”
It is also the hidden half of a RAG project: the quality of the answers depends as much on the model that searches as on the model that writes. The right choice when your employees lose time digging through the shared server; of limited interest if your entire document set fits in one filing cabinet.
What do you want? The mapping grid
Vocabulary only matters if it leads to a decision. Here is the grid we use at scoping: start from what you want to get, and the rest follows.
RAG
the default approach for querying company knowledge, with no model retraining
Agent
a word to reserve for systems that act: AI prepares, people decide
Local
inference on your own premises: your documents never leave the company
| You want to… | The right approach | What it involves |
|---|---|---|
| Query your procedures, contracts and quotes in plain language | RAG on an existing model | Prepare and index the documents; no retraining of the model |
| A very precise in-house tone or vocabulary in every text produced | Fine-tuning, alongside RAG | Build a set of quality examples; redo the exercise when the model changes |
| Answer clients’ repetitive questions on your website | A chatbot connected to your content (RAG) | Bound the topics covered; plan the handover to a human |
| Copy data from one system to another, always the same way | Classic automation (RPA), with no AI | Cheaper and more predictable; brittle if the interface changes |
| Complete deliverables (a quote, a summary, a file) prepared for approval | A custom AI agent | Connect your systems; define who approves what: AI prepares, people decide |
| Keep sensitive data in house, as Law 25 requires | Local inference on an open-weight model | A suitable server; a model chosen to fit your hardware |
| Find a document by its meaning, not by the exact word | Embeddings and semantic search | Index the document set; maintain the index with every addition |
Where to start
Pick a concrete irritant (quotes that take too long, knowledge sleeping in inboxes, repetitive questions) and restate it in the words of this glossary: do you want to search, write, converse or act? That one-line sentence beats a long requirements document, and it protects you from solutions sold before the problem.
At Cogio, that is how every scoping starts: the business need first, the vocabulary second, the technology last. We design custom agents hosted at the client’s site, a form of enterprise brain that prepares the work while your people decide. If you want to sanity-check how you have framed a need, write to us: a conversation is often enough to restate a problem in AI terms, and it commits you to nothing.
Frequently asked questions
Do we have to retrain a model on our data for it to be useful?
Almost never. RAG, which has the model consult your documents at the moment it answers, covers the large majority of an SME’s needs: it updates continuously, cites its sources and leaves the model untouched. Fine-tuning is reserved for very specific styles and vocabularies, and full training is not a realistic option for a business.
Can an AI agent make decisions on our behalf?
Technically it can be configured that way; we advise against it. Our design rule is constant: the agent consults, cross-references and prepares, then a human approves anything that commits the company, whether that is a price, a contract or a message to a client. Responsibility cannot be delegated to software.
What causes hallucinations, and can they be eliminated?
A language model produces the most plausible continuation of a text, not the truest. When the question goes beyond what it was shown, it fills the gaps with something plausible. The phenomenon cannot be eliminated, but it can be contained well: RAG with source citations, a defined scope, instructions to decline when the information is missing, and human sign-off on the output.
Is quantization a cut-price version of the model?
It is compression, not amputation. It reduces the numerical precision of the parameters so the model fits in less memory, with a generally marginal loss of quality on a company’s document work. For very specialized tasks, less aggressive quantization can be justified if the hardware allows.
RPA or an AI agent: how do we decide for a given process?
Ask one question: does this process require understanding, or only repeating? If every run is identical (same fields, same screens, same rules), RPA is simpler, cheaper and more predictable. As soon as you have to read a variable document, judge a particular case or combine several sources, an AI agent becomes relevant.
Are open-weight models worse than the large commercial models?
On some cutting-edge tasks, the large commercial models keep an edge. For an SME’s everyday uses (querying documents, summarizing, writing, extracting), the best open models deliver more than enough quality, with one decisive advantage: they run on your premises, on your hardware, without your data travelling.
This article is a plain-language summary, accurate as of the date shown. It is not legal advice: for your own situation, consult a legal adviser or contact the Commission d’accès à l’information.
Read next
Sovereignty
Top 5 AI models to host on your own servers in 2026: an honest comparison
Read →
Sovereignty
Local AI or ChatGPT Enterprise: an honest comparison for a Quebec SME
Read →
Strategy
Is your data ready for AI? A 10-point self-assessment
Read →
A question about your own compliance?
The discovery call is free and takes half an hour. You leave with an honest read on your situation.
Let’s talk