There are two common misconceptions about data and AI, and both are expensive. The first holds that modern models are so good that the state of your data no longer matters. The second holds that you must put your entire data landscape in order before you can begin with AI at all.
The first misconception produces assistants that give wrong answers. The second produces programmes that are shut down after eighteen months without ever having shown a benefit. The truth sits uncomfortably in the middle.
What AI actually does with poor data
A language model does not check whether information is correct. It judges whether a formulation is plausible. If your storage holds three versions of the same price list, the assistant will use one of them: fluently written, convincingly delivered and possibly two years out of date. The user sees no difference.
That is more dangerous than an empty answer. An empty answer prompts a follow-up question. A wrong, well-written answer prompts a decision.
An assistant makes your data quality visible. It does not improve it.
The three problems that really count in practice
After many assessments, the same three points turn out to be decisive, and none of them is a classic data quality issue in the narrow sense.
- Permissions: where storage is open too widely, an assistant shows people content they are technically allowed to see but would never have found. Salary lists and draft dismissal letters are the usual suspects.
- Duplicate and outdated versions: where the same document exists in five versions, there is no reliable answer, only a random one.
- Missing context: a table without a description, a field without a definition, a metric without a stated calculation. People fill those gaps from experience. Systems do not.
The pragmatic route
Instead of a data programme, I recommend scoping the work around use cases. You pick one concrete use case, clarify the data situation for that case specifically, and put it right there. That is achievable in weeks rather than years, and it produces visible value that funds the next stage.
In practice that means asking: which sources does this one use case need? Who owns them? Are the permissions on those sources clean? Is there a valid version, and is it obvious which one it is? Only once those four questions are answered for the first case do you move to the second.
Across five or six use cases this creates exactly the order that a large data programme aims for, but paid for out of value already delivered and without anyone losing interest.
Where a platform pays off
Above a certain size, there is no way around a shared data foundation. Microsoft Fabric, for example, brings storage, preparation and analysis together in one place and makes origin and ownership traceable. That is a sensible step, but as a consequence of the first successful use cases, not as their precondition.
Sequence decides the outcome. Start with the platform and you will spend two years discussing architecture. Start with a use case and you have a result after twelve weeks, and you then know very precisely which platform you need.
Your first step
Take the use case your leadership team already wants most urgently. Ask the four questions above. If the answers are uncomfortable, you have learned more about your data in a week than any audit report would tell you, and you know where to start.