Skip to main content

AI-Ready Data Starts With Master Data

· 7 min read
Dan Peacock
Chief Hustler

"AI-ready data" is now everywhere, and most definitions of it agree: clean, governed, contextualised, accessible, traceable. All true, and all hard to argue with.

The trouble is that these describe the destination. They do not explain why so few organisations arrive.

The reason is narrower and more awkward than "data quality". Most organisations cannot agree what their core business objects actually are — and no AI system will resolve that for you. It inherits the disagreement and scales it.

The foundation everyone skips past

We have written before that AI depends on three things: good-quality data, data with business context, and data stored in a consistent structure.

Quality gets attention because it is measurable. Structure gets attention because it is technical. Business context is the one that quietly stalls programmes, because it is neither.

Business context is master data: the definitions of the things your business is actually about. Customer. Product. Location. Asset. Service. Everything else — every transaction, every event, every metric — is a statement about one of those things. Get them wrong and every number downstream inherits the error.

Why master data projects fail

Not for the reason people expect.

The tooling has been mature for two decades. Matching, survivorship, deduplication and stewardship workflows are solved problems. Master data programmes rarely fail on technology.

They fail on a single question: what is a customer?

Ask five parts of a business and you get five answers, each entirely defensible in its own context:

  • Sales — anyone in the CRM, including prospects who have never bought anything
  • Finance — an entity that has been invoiced, which excludes most of the CRM
  • Operations — a delivery site, so one company may be forty customers
  • Legal — the contracting party, so forty sites may be one customer
  • Support — whoever raised the ticket, which may be an individual person

Every one of these is correct locally. None of them is reconcilable with the others enterprise-wide. And because each is genuinely right within its own function, nobody concedes.

This is the part that derails the programme. The question "what should our master data be?" has no objectively correct answer, so it does not get settled on merit. It gets settled by seniority, by stamina, or by whoever is still in the room after eighteen months of workshops. The output is usually a definitions document that nobody implements, because by the time it is finished the organisation has moved on.

Then a new source system arrives, and the whole argument runs again.

Why this is fatal for AI specifically

A reporting error and a model error do not behave the same way.

When a report is wrong, a human who knows the business usually notices. The number looks off, someone queries it, and the definition mismatch surfaces. The error is contained by the fact that a person read it.

A model trained across five definitions of "customer" learns all five as though they were one. Its outputs then contradict known operational results, and the failure presents as a model accuracy problem — which sends teams to retrain, tune and add data, none of which touches the actual cause. We have covered this pattern in why AI initiatives stall due to inconsistent data definitions.

The compounding factor is scale. AI applies the flawed definition to thousands of decisions without anyone reading the intermediate step. Inconsistency that was survivable in a monthly report is not survivable in an automated one.

The CryspIQ® difference: do not have the argument

This is where the CryspIQ® approach diverges from every other route to AI-ready data.

Most methods hand you a modelling capability and leave the definitions to you. That sounds like flexibility. In practice it hands the organisation a blank page and an unwinnable debate.

CryspIQ® ships with the master elements already defined. The patented enterprise data model is pre-built, and the core business objects are part of it: entity, product, location, service and asset. They are not proposals to be workshopped. They are the structure.

So nobody convenes a committee to decide whether "customer" qualifies as a master data object, or which attributes it should carry. In CryspIQ® a customer is an Entity. The model already provides the Entity Business Key, the Entity Name and the Entity Type. The work is connecting your source systems to that, not negotiating what it ought to be.

That reframes the project completely:

Conventional MDMCryspIQ®
First taskAgree the definitionsConnect the sources
Governed byConsensusThe model
Completion criteriaNone obviousSources mapped
New source systemRe-open the debateMap it

The first of those is a debate, and debates have no completion criteria. The second is a task, and tasks finish.

Single instance is a constraint, not an agreement

This is the part that matters most, and it is easy to read past.

Conventional master data management produces a golden record by agreement, then relies on organisational discipline to preserve it. Discipline decays. A new system arrives with its own customer table, a project is under deadline pressure, an exception is granted "just for now", and within two years there are three customer masters again.

In CryspIQ® the single instance is enforced by the model itself. A customer is defined once in the enterprise data model — not because a committee agreed, but because the structure provides exactly one place for it. Every source system carrying customer data connects to that same Entity. There is nowhere else to put it.

That is the difference between a policy and a constraint. Policies accumulate exceptions. Constraints do not.

It is also why the definition does not drift. There is no second definition available to drift towards.

What this changes in practice

  • A new source system is a mapping exercise, not a governance event. You connect it to elements that already exist.
  • Definitions cannot fragment, so reconciliation between systems stops being ongoing work.
  • AI trains on one meaning of each business object, which removes an entire class of model failure before any model is built.
  • Effort moves from arguing to connecting — and connecting is measurable, delegable and finishable.

Where to start

Before evaluating any tooling, run a count.

Take the three business objects your organisation reports on most — customer, product and site, in most cases — and count how many separate definitions exist across your systems, reports and pipelines. Not how many should exist. How many do.

Most organisations expect two or three and find eight.

That number is the size of your AI-readiness problem. It is not a data quality problem and no amount of cleansing will address it, because each of those eight definitions is internally clean. It is a master data problem, and it is the first thing to fix — because everything you build on top of it inherits it.


Further reading