I spend a lot of my time helping organisations understand how to work effectively with AI. One of the frameworks I find most valuable is the Purpose - Execution - Judgement (PEJ) framework (Dixon, 2026). PEJ proposes that every productive human-AI activity comprises three distinct domains:
-
Purpose: the why, meaning the goals and direction that determine what you're trying to achieve.
-
Execution: the how, the doing, creating, and analysing.
-
Judgement: the which, determining which output actually matters and being accountable for the choice.
I find this framework particularly useful in practice because it gives organisations a clear map of where human value sits in an AI-assisted workflow. Right now, the relative importance of these domains is shifting faster than most people have noticed. While Execution is accelerating, the bottleneck is moving. For most of professional history, that bottleneck was simple: the time it takes to produce things.
Think about what it used to mean to develop a marketing campaign. A team might spend days researching the ideal customer, developing concepts, producing copy and creative, rounds of revision, and sign-off processes. The ideas might come quickly, but the execution took time. The same pattern held almost everywhere. A consultant preparing a client proposal would spend hours writing, structuring, and formatting. A finance team producing a quarterly analysis might take days on the data-gathering and modelling alone. A project manager building a new product: painstaking work, line by careful line. In each case, producing the thing was where the hours went.
Organisations built themselves around this. Teams were sized to match how much they could produce, timelines reflected how long making things actually took, and rising to seniority often meant being fast and good at the work itself. Execution was the scarce resource.
AI is changing this, and it is changing it fast. What used to take a day now takes an hour. What used to take an hour takes minutes. From market analyses to complex project plans, the time cost of production has collapsed, and I frequently see these gains with my clients. But something else is happening alongside it, and it is starting to matter more than the headline efficiency gains.
Let’s return to that marketing team. Previously they might produce a single campaign asset or piece of content in a given cycle. Now, with AI, they can produce variations at scale: multiple versions of copy, creative directions, or audience-specific adaptations in the same amount of time. The increase in output is real and meaningful, but so is the burden it creates, because someone still has to evaluate what is actually worth using.
An even clearer example can be seen in software development. A team working on a new feature might previously have scoped, designed, and built one or two features over a given period. With AI accelerating execution, they can now prototype and build many more features in the same timeframe. The constraint shifts immediately, and the question is no longer "can we build this?" but "is this fully tested and ready for release?" The volume of possible outputs increases, but the capacity to make high-quality decisions about them does not.This is where the new bottleneck sits: not in producing the work, but in deciding what is good, what is relevant, and what is worth acting on.
AI scales execution. It does not scale judgement. Consequently, judgement becomes the new constraint.
This is becoming increasingly visible as AI adoption picks up. A 2026 survey by UnlikelyAI of over a thousand senior decision-makers found that they were spending almost as much time checking AI-generated work as using AI in the first place: around two and a half hours a week on verification, against roughly two hours and forty minutes of use. These are not the numbers of people who have solved the problem of AI evaluation; they are the numbers of people absorbing its cost, one careful hour at a time.
These costs accumulate in ways that rarely show up in productivity measurements. Organisations count the efficiency gains but rarely account for the growing cost of evaluating what has been generated. The result is a gap between adoption and value that the headline numbers obscure. While adoption is up, value is not keeping pace: only 22% of organisations reported significant returns on their AI investments three years after the technology entered mainstream business use. The constraint has shifted, and most organisations have not caught up with where it has moved to.
This is where PEJ becomes directly relevant. In PEJ terms, Execution is the domain AI has transformed. The cost of the "doing" has dropped dramatically and will continue to drop. This exposes Judgement as the immediate downstream consequence of abundant Execution. More output requires more evaluation, and the bottleneck moves.
The professionals and organisations that understand this shift now, and start designing for it deliberately, will have a real advantage.
Effective evaluation of AI output is not a technical skill; it is a domain skill. The capability you need to judge whether an AI-generated concept is right for a specific audience is not knowledge of how the AI works. It is deep knowledge of your field. A machine can generate a hundred versions of a strategy document, but it cannot know whether any of them is right for this client, this moment, or this context. It cannot catch the recommendation that would work on paper but fail given what you know about how the organisation actually operates. Those things need someone who has done the work long enough to develop an internalised sense of what "good" looks like that no brief fully captures.
Expertise does not transfer automatically. Knowing your field well does not mean you will evaluate AI output well by default. That judgement needs to be actively applied through the right habits of scrutiny and, at the organisational level, through the right structures.
The era of effortless production is arriving, but the era of effortless judgement never will.
My next article will explore The Scalable Judgement Model. If judgement is the new bottleneck, how do you design for it at scale?
