Your Copilot isn’t failing because of the model, it’s failing because of your data
Nearly every AI pilot that stalls does so for the same reason, and it has nothing to do with which model was chosen.
The conversation repeats itself: Copilot gets switched on, tested for a week, and the verdict is that it “gives generic answers” or “makes things up”. The usual reaction is to go looking for a different model. It is almost never the model.
An assistant can only be as good as what it can read
When an assistant answers a business question badly, there are three possible causes and only one belongs to the model:
- It never found the document. It exists, but it sits in a folder the assistant does not index, or in a scanned PDF with no text layer.
- It found three versions and picked the old one. The 2019 procedure, the 2023 one and a draft all live in the same place with nothing marking which one is in force.
- The model reasoned badly. It happens, but it is the minority.
The first two are document management problems dressed up as AI problems. Switching models fixes neither.
The ten-minute test
Before investing in a pilot, do this: take the five questions your team is asked most often and look them up yourself on the intranet, the way the assistant would. If you struggle to find the right answer, the assistant will struggle more.
If your people take twenty minutes to find a procedure, the problem is not solved by putting an AI on top. It is solved by fixing the procedure — and then the AI becomes useful.
What has to be in place before the pilot
- One source that governs. For each topic, one document marked as current. The rest archived out of the assistant's reach.
- Permissions that match reality. Copilot respects existing permissions: if half of SharePoint is open to the whole company, the assistant will make that obvious on day one. That is not an AI failure, it is a free audit.
- Actual text. A scanned PDF is an image. If your procedures are scans, no assistant can read them.
- Minimum metadata. Area, validity and owner. With those the assistant can break ties between similar documents.
Where to actually start
Do not start with “the whole company”. Pick one bounded domain where questions repeat and answers are written down: HR policies, safety procedures, a technical catalogue. Tidy that domain, measure accuracy with real questions, and only then widen the scope.
An assistant that gets things right in one area builds trust. One that is merely adequate everywhere destroys it, and winning it back costs far more than starting slowly would have.