hybridresourcing Sign up for the waitlist

Kennisbank

How mature is your data management when it comes to AI

AI models are only as good as the data they get to see. That is not a slogan but a practical problem: a language model that works with scattered, outdated or undefined data produces answers that are just as scattered, outdated and undefined. Many organizations only notice this after a pilot has started, when it turns out that the data the model needs is not stored in one place anywhere, or that no one knows anymore which version of a file is the correct one. Data management is one of the fundamental dimensions: the question is not which tasks AI could take over, but whether underneath those tasks there is a layer of data solid enough to build on.

What this dimension measures

Data management is about where data comes from, who is responsible for it, what state it is in, and how easy it is to find and combine. This includes whether data is stored centrally or in a scattered way, whether an owner has been assigned per dataset, whether definitions of concepts such as 'customer' or 'active project' are consistent across departments, and whether there is visibility into the quality of the data used on a daily basis. An organization can score highly on IT infrastructure here and still rank low on data management: the systems are present, but no one has ever documented exactly what is in them or who is allowed to change anything about it.

The five levels, in brief

At the baseline level, data exists mainly in silos: each department keeps its own files, without a shared structure or ownership. At the foundation level, there is a beginning of organization: there are agreements about where certain data should be stored, but execution still depends on individual habits. At activation, data ownership has been assigned and there is a shared framework of definitions, so that different departments use the same term in the same way. At the insight level, data quality is measured and tracked, and there is visibility into where the weak points in the data chain are. The intelligence level means data management is embedded in daily working practice: quality, provenance and access are continuously monitored, not as a separate project but as a fixed part of how work is done.

What indicates where you stand

A practical indicator is the question of how long it takes to answer a simple question, such as how many customers purchased a particular product in the past year. If the answer comes from a single system within a few minutes, that points to a higher level. If three people first need to be called to find out which file is current, that points to baseline or foundation. Another indicator is what happens when someone with knowledge of a dataset leaves: does the information remain accessible, or does it disappear with that person. Also relevant is whether there is someone who bears responsibility for a dataset, or whether that responsibility effectively lies with no one.

The plotting round: why the spread itself is information

Within the maturity measurement, this dimension is not scored by one person but by several people separately. An IT manager may rate data management highly because the systems are technically in order, while an operational manager scores it low because in practice the data does not feel findable or reliable. That spread becomes visible in the plot and is itself a signal: a large distance between scores often means the organization is further along on paper than in daily practice, or that different departments are working with different versions of the truth. A narrow spread, even at a lower level, indicates a realistic and shared picture to work from.

What moving up a level requires

The step from baseline to foundation often does not require new technology, but agreements: who owns which dataset, and where which information is stored. The step from foundation to activation requires a shared framework of definitions, so that departments do not define what an 'active customer' or a 'completed project' is at cross purposes with one another. What a specific step requires in terms of cost and time depends on the size of the organization and the state of the underlying systems; that is worked out on the page about what it costs to move an organization up a level and the page about what it costs to move IT infrastructure up a level, because data and infrastructure in practice often advance together.

The connection with other dimensions

Data management does not stand apart from the rest of the measurement. Without clear agreements about who may use which data and for what purpose, questions arise that actually belong to privacy and security. Without people who understand what data quality means and how to pay attention to it in their work, the dimension people and skills falls behind. And without a view on who decides what may and may not be done with data, the connection with ethics is missing. The seven dimensions of the measurement are interconnected; data management is one of them, and usually not the last one that deserves attention.

From carrying capacity to tasks

This measurement shows whether the ground under AI applications is solid enough: whether data is findable, reliable and owned by someone. That is a different question from which part of the work can itself be taken over by AI. That question is answered by the work scan from FTE TO AI, which calculates per task which part is suitable for transfer to AI, based on the data structure present at that time. The two measurements are meant to be read one after the other: first the carrying capacity, then the tasks that can rest upon it.

The maturity measurement, including the data management dimension, is currently under construction. Anyone who wants to take the measurement as soon as it becomes available can sign up for the waiting list.

Robbyde assistent van de volwassenheidsmeting

Vraag maar wat er moet staan voordat AI in uw organisatie kan landen.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.