Is your data ready for AI? A 10-point self-assessment
Ten concrete tests to run this week to find out whether your documents, access rights and formats are ready for AI, with the remedy for each symptom.

AI does not tidy your data: it inherits its condition
An AI assistant plugged into disorderly folders gives disorderly answers, with added confidence. It is one of the leading causes of disappointment, well ahead of the technology: our analysis of why AI projects fail in SMEs goes into it in detail. The good news: the state of your data can be checked with no consultant, no tool and no budget.
The effort is worth it. According to BDC research, roughly 30% of Canadian SMEs were using AI in 2025, and those that do are reported to be 24% more productive. The difference between the two groups is often decided upstream, in the preparation. Hence this self-assessment: ten points grouped into four workstreams, each with a concrete test to run this week and what a failure implies.
Workstream 1: can your documents be found?
Point 1, where the documents live. The test: ask two employees from different departments to find the latest version of the same document, a standard quote, a procedure, a contract. If they take two different routes, or if one of them answers “ask Sylvie,” you have your answer. A failure means the AI will index contradictory sources and serve the wrong version part of the time.
Point 2, duplicates and versions. The test: search for an important document by name and count the copies: shared server, inboxes, individual machines, USB sticks. Beyond one official source and one archive, every copy is a risk. A failure means a clean-up before indexing, without which the project starts on the wrong foot: the AI does not pick the right version, it reads them all.
Workstream 2: do your documents resemble each other?
Point 3, naming conventions and templates. The test: open ten recent quotes. If every project manager has their own style, their own sections and their own file names, automatic extraction will have to cope with ten formats instead of one. A failure is not blocking, but every variation adds preparation work: a common template adopted today pays off on the very first project.
Point 4, usable formats. The test: take your twenty most-consulted documents and sort them into two piles: native (Word, Excel, a PDF generated by software) or scans and paper. Native text extracts faithfully; a scan requires character recognition, whose errors propagate all the way into the answers. A failure means setting priorities: digitize what is still in use, let the rest sleep.
Point 5, the quality of the French in your documents. The test: reread three internal documents picked at random. In-house abbreviations, jargon never defined, telegraphic sentences: what your best employee decodes out of habit, the model guesses at, sometimes wrongly. A failure means fixing the templates first, the back catalogue second, and recording the abbreviations in an internal glossary the AI can consult.
Workstream 3: who is allowed to see what?
Point 6, access rights. The test: open the shared server with a recent employee’s account and search for “salaries,” “dismissal” or “performance review.” Whatever they can see, an AI assistant plugged into the same server will be able to find and summarize for anyone who asks. A failure means restricting rights by role before any indexing: AI does not invent leaks, it industrializes the ones that already existed.
Point 7, personal information to redact. The test: in your project and client files, look for social insurance numbers, dates of birth, health information. Law 25 governs their use, and an AI project constitutes a new use to be assessed. A failure means identifying, moving or redacting before indexing; our 12-step checklist gives the order of operations.
Workstream 4: who answers for the data?
Point 8, structured history. The test: ask how the company solved a problem that came up three years ago. If the answer lives in one person’s memory or in an inbox, your history does not exist as far as the AI is concerned. A failure means structuring resolved cases into a searchable format, starting with the knowledge of the employees who will retire in the next few years.
Point 9, a data owner per department. The test: ask who answers for the quality of production, sales and HR data. If the answer is “IT,” that is a failure: IT hosts the data, it does not know the trade behind it. A failure means naming an owner per domain, with a one-page written mandate: decide what is official, what gets archived, what gets corrected.
Point 10, governance. The test: look for the document that says who can create a shared folder, who archives, who deletes, and what your employees are allowed to paste into a public AI tool. If it does not exist, everyone improvises. A failure means writing those rules before the project, not after: our AI acceptable use policy template provides the frame.
Symptom, remedy, effort: the correction table
No company passes all ten tests first time, and that is not the point. The point is to know where you stand, then correct in the right order. Here is the table we use to prioritize.
| Symptom | Remedy | Effort |
|---|---|---|
| The same document exists in four versions across three locations | Designate a single source per document type, archive the rest | Moderate: a supervised clean-up, department by department |
| Nobody names files the same way | Adopt a naming convention and templates, rename as you go | Low: a one-page convention and some consistency |
| The key archives are scans, sometimes handwritten | Digitize what is still in use first, with controlled extraction | High: keep it within the pilot’s scope |
| Everyone has access to everything on the shared server | Restrict rights by role before indexing anything at all | Moderate: an access inventory, then a rework of the groups |
| Personal information is lying around in working files | Identify, move or redact before any indexing | Moderate to high: it is also a Law 25 requirement |
| The history lives in emails and in individual memories | Structure resolved cases into a searchable format | High: start with the knowledge closest to retirement |
| No one is accountable for data in each department | Name an owner per domain, with a written mandate | Low: a management decision, not a technical project |
| The French in the documents is loose or full of in-house abbreviations | Fix the templates first, document the abbreviations second | Moderate: aim at the documents the AI will read most |
A diagnostic that can be half funded
This preparation work is precisely what ESSOR stream 1B, Investissement Québec’s program, funds. It covers up to 50% of the eligible expenditures of a digital diagnostic and implementation plan, to a maximum of $20,000, for SMEs with 250 employees or fewer and revenue of at least $2.5M. Our overview of AI funding in Quebec sets out the conditions and how to build the application.
The support is never guaranteed: the program is discretionary and every figure remains to be confirmed against your eligibility. But a half-funded diagnostic changes the arithmetic. The self-assessment above tells you whether you are ready; the formal diagnostic tells you what to fix, in what order and at what cost. To find the applicable programs, our Funding page walks through them.
50%
share of the eligible expenditures of a digital diagnostic covered by ESSOR stream 1B, non-repayable support
$20,000
the stream 1B ceiling, for SMEs with 250 employees or fewer and revenue of at least $2.5M
Where to start
Block out an hour this week and run the ten tests, without cheating: the aim is not a good grade, it is an honest map. Then sort the failures into two piles: what blocks a pilot (access rights, personal information) and what merely slows it down (templates, scans, naming). The first pile gets settled before any project; the second can be settled by the project, scope by scope.
At Cogio, every engagement starts from this field truth: we would rather postpone a project than build an enterprise brain on soft foundations, and we say so when that is the case. If your results leave you puzzled, talk to us about it: comparing your assessment with what we see elsewhere takes a conversation, not a contract.
Frequently asked questions
Do we have to clean everything up before launching a first AI project?
No, and that is the opposite trap: waiting for perfect data means never starting. The right approach is to pick a pilot scope (one department, one document type) and prepare it thoroughly, rather than cleaning the whole company superficially. Only the security points (access rights, personal information) have to be settled everywhere before indexing anything.
Our archives are on paper or in scans: is that a deal breaker?
No, but it is a cost to plan for. Modern character recognition gives good results on clean documents, less good on copies of copies or handwritten notes. The practical rule: digitize first what operations still use, check extraction quality by sampling, and let the dead archives sleep.
How long does data preparation take?
It depends entirely on the scope and the starting condition, and that is exactly what the diagnostic exists to cost out. A pilot limited to one well-kept department prepares quickly; a history scattered across emails and scans takes far longer. Be wary of anyone who quotes a duration without having looked at your data.
Who should lead this diagnostic internally?
A pair: someone from management who decides (what is official, what gets archived) and someone on the ground who knows the documents day to day. IT supports, but does not lead alone: data quality is a question of the trade before it is a question of servers.
Can AI be used to clean the data itself?
Yes, in part: detecting duplicates, proposing classifications, spotting personal information or extracting text from scans. But AI proposes and people decide: choosing which version governs, or who has access, remains a business decision. Same rule as everywhere else: AI prepares, people decide.
Is the diagnostic genuinely fundable?
Stream 1B of Investissement Québec’s ESSOR program funds up to 50% of a digital diagnostic and implementation plan, to a maximum of $20,000, for SMEs with 250 employees or fewer and revenue of at least $2.5M. The support is discretionary and remains to be confirmed against your own application: check your eligibility before committing any expense.
Sources and references
This article is a plain-language summary, accurate as of the date shown. It is not legal advice: for your own situation, consult a legal adviser or contact the Commission d’accès à l’information.
Read next
Strategy
Why so many AI projects fail: the numbers, then the seven avoidable mistakes
Read →
Funding
AI funding in Quebec in 2026: who pays for what, and how to build an application
Read →
Strategy
RAG, fine-tuning, agents, automation: a decision-maker’s AI glossary
Read →
A question about your own compliance?
The discovery call is free and takes half an hour. You leave with an honest read on your situation.
Let’s talk