hybridresourcing Sign up for the waitlist

Kennisbank

Beyond the chat window: what the trial session doesn't measure

A chat window is a doormat, not a house

Anyone working with a language model for the first time sees something work. An answer comes back, the answer is right often enough, and the conclusion seems obvious: this can do more. That conclusion isn't wrong, but it is based on a single moment in a single window. A chat window shows whether a model can answer a question. It does not show whether the organization behind it can properly formulate the question, supply the right data, and let the answer land in a decision that actually matters.

That distinction is why pilots often start but rarely scale. The pilot proves that the technology can do something. It proves nothing about the seven dimensions that determine whether the organization can carry that capability as well: from organization and leadership to IT infrastructure, data management, processes, people, governance, and culture. A chat window tests one use case, at one moment, with one user. It says nothing about what happens when a hundred users ask the same question, with data that isn't always clean, in a process that isn't set up for an answer that varies.

What the maturity measurement does do

The maturity measurement from hybridresourcing does not look at a task or an answer, but at the layer underneath: the question of whether the organization as a whole is ready to work with AI, and not merely to experiment with it. This happens across five levels — baseline, foundation, activation, insight, intelligence — and across seven dimensions that together determine where an organization stands. Not every dimension carries equal weight at every moment. The foundational dimensions, such as organization and IT infrastructure, take precedence over the dimensions that depend on them. That is not a choice we make, but the order in which it works: without a functioning infrastructure, good data management has little grip, and without clear decision-making, good infrastructure has little effect.

An important part of the measurement is the plot round. Multiple people within the organization score separately, without seeing each other's answers. What comes out of that is not only a position across five levels, but also the spread between what different people believe to be true. A leadership team that believes the organization is at the activation level, while IT places itself at baseline, has a problem bigger than the average. That spread is often more informative than the score itself: it shows where the organization disagrees with itself about itself, and that is exactly where a pilot gets stuck without anyone being able to point out why.

What this method cannot do

The maturity measurement gives no answer to the question of which tasks AI can take over. It says nothing about how much time a role spends on a particular activity, and nothing about what happens if that activity disappears. It measures readiness, not takeover. An organization can score high on all seven dimensions and still not have identified a single task ready to be handed over. Conversely, an organization can be full of tasks that lend themselves excellently to automation, while the organization itself is not yet set up to let that happen responsibly.

The measurement also does not indicate which decisions may or may not be automated — that is a question that deserves its own attention, addressed among other things in the overview of decisions that should remain outside the scope of automation. And the score says nothing about what is missing in practice in terms of documentation: anyone wanting to know what belongs in a decision inventory will not find that in a level or a dimension, but in a separate process.

What the measurement does do is provide clarity about the starting position. A low score on one of the seven dimensions does not automatically mean something is fundamentally wrong; it means two routes are possible, and which of the two applies differs per organization — a difference explored further on the page describing the two scenarios for a low score. Anyone wanting to start with their own assessment can do so on the pages addressing the maturity of the organization itself and the maturity of the IT infrastructure as two of the seven dimensions on which the full measurement is built.

Why this boundary exists

This separation is deliberate. An organization that first wants to know whether AI delivers something for it asks a different question than an organization that wants to know whether it is prepared to deploy AI responsibly. Both questions are legitimate, but they call for different instruments. Anyone who lets both questions blend together gets an answer that fits neither question well: a score that sounds like a task analysis, or a task analysis that pretends to say something about organizational readiness.

The tool is under construction

The maturity measurement, with its five levels, seven dimensions, and plot round, is currently being built. There is not yet an environment where you can log in and score today. Anyone interested can sign up for the waiting list and will be kept informed as soon as the measurement becomes available.

The question that comes next

Once the organization's readiness is mapped, the other question remains unanswered: which part of the work itself can actually be taken over. That is not something this measurement addresses, nor is it meant to. That question is answered by the work scan from FTE TO AI, which calculates per task which part of the work qualifies to be taken over by AI — a calculation that only means something once the organization it lands in is actually ready to carry it.

Robbyde assistent van de volwassenheidsmeting

Vraag maar wat er moet staan voordat AI in uw organisatie kan landen.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.