The pattern is consistent enough now to describe. An organisation runs a successful AI pilot. Throughput increases. Review becomes a bottleneck. Experts become a bottleneck. Outputs become inconsistent across cases that seem similar. Prompts proliferate. A second review layer is added. Then a third. The programme slows.

The instinct is to look at the model. Change the prompt. Fine-tune. Switch providers. What organisations are increasingly discovering is that the model was rarely the problem. The problem is that the institutional judgment the model was supposed to apply was never formally specified in the first place.

What organisations have historically relied on

For decades, organisations have relied on expertise that was never formally specified. The credit analyst knows when a score should be overridden. The claims reviewer knows which anomalies matter. The risk manager knows when an exception is genuinely material. The lawyer knows which clause deserves escalation. The operations manager knows when a KPI breach should trigger intervention and when it should be tolerated.

These decisions are rarely arbitrary. They are the product of experience, incentives, history, and institutional memory. But they are seldom documented in a way that can be tested or reproduced. People learn them socially. Systems do not.

As long as throughput remained low, this arrangement worked. Human review filled the gaps. Experts compensated for ambiguity. Decisions were adjusted through experience and discussion. AI changes the economics. As throughput increases, experts become bottlenecks. Review becomes sampling. Tribal knowledge stops scaling.

AI exposes what organisations never had to specify

Most organisations assume that deploying AI means teaching the model. In practice, they discover something different. The model can only operate on what has been made explicit. Thresholds. Tolerances. Escalation conditions. Materiality judgments. Contextual assumptions. These things always existed. They simply lived inside experienced people rather than inside systems.

The problem is no longer whether the model is intelligent enough. The problem is whether the organisation understands its own judgment well enough to specify it.

Most organisations are more under-specified than they realise

When AI outputs become inconsistent, the instinct is to blame the model. Change the prompt. Add another review step. Fine-tune. Switch providers. Yet many failures are specification failures rather than capability failures.

Ask three experts where escalation should occur and the answers are often similar but not identical. Ask which exceptions matter. Ask what material means. Ask what constitutes acceptable evidence. Ask when context should override the rule. The organisation discovers that it has relied on alignment through experience rather than alignment through explicit design. Humans tolerate this ambiguity remarkably well. Systems expose it.

Where the work actually begins

Organisations expecting to spend their time on model selection often discover that the harder work begins elsewhere. Workshops become discussions about escalation criteria. Subject matter experts debate what constitutes acceptable evidence. Risk teams disagree on materiality thresholds. Operations teams realise that exceptions were managed through experience rather than policy.

These answers exist. They simply exist in fragments, distributed across people rather than captured in a form that systems can inherit.

AI does not create these judgments. It exposes the fact that they were never formally captured.

AI is turning judgment into infrastructure

Reliable AI systems are not built by eliminating judgment. They are built by making judgment explicit.

This work rarely begins with prompts. It begins with conversations. Experts are asked where escalation should occur, what evidence matters, which exceptions are tolerable, and what acceptable looks like when the rules conflict. Thresholds become policies. Escalation conditions become workflows. Materiality becomes specification. Semantic meaning becomes governed definitions. Ambiguity becomes a designed review process rather than an accidental one.

The role of experts changes. They stop acting primarily as reviewers of outputs and start acting as designers of the conditions under which outputs are judged.

That transition is harder than model selection. And it is increasingly where AI programmes succeed or fail.

The challenge is no longer extracting answers from models. It is extracting judgment from people.

The next bottleneck is not reasoning

Most organisations do not suffer from a shortage of intelligence. They suffer from a shortage of specification.

Because the most valuable asset in many enterprises was never the model. It was the institutional judgment that lived inside people.

AI is not replacing that judgment. It is forcing organisations to specify it.

The organisations that scale successfully will not necessarily be the ones with the best models.

They will be the ones that understand themselves well enough to explain what good looks like.