Photo by Tima Miroshnichenko on PexelsLeadership and AI
Why Your Judgement Still Matters
The most important work begins where the obvious answer ends.
AI can draft, sort and recommend at speed. This story is about the harder leadership work that remains: deciding whether an answer fits the moment, the person and the consequence.
The most important work begins where the obvious answer ends.
AI can draft, sort, summarise, compare and recommend at extraordinary speed. It can turn a blank page into a plausible answer before a person has finished describing the problem.
That is useful. It is not the same as judgement.
Judgement is what happens when the facts are incomplete, the context is uneven, the trade-offs are real and somebody still has to decide what should happen next. It is the ability to notice what the system cannot see, weigh what matters here, and accept responsibility for the consequence.
As AI becomes easier to use, that distinction becomes more important, not less. The question is no longer simply, "Can a machine produce an answer?" It is, "Who decides whether this answer belongs in this moment, for this person, with these consequences?"
We are entering the age of plausible answers
The first great advantage of generative AI is also its most seductive quality: it makes work look finished.
A proposal has headings. A reply sounds polished. A strategy contains familiar language. A report arrives in a confident tone. The rough edges that once signalled uncertainty have disappeared.
But presentation is not proof. Fluency is not understanding. A response can be clear, useful and wrong at the same time.
That creates a new kind of risk for organisations. In the past, unfinished work often looked unfinished. Now weak work can arrive dressed as authority. The burden shifts from producing the first draft to recognising whether the draft is sound.
This is why judgement is becoming a practical business capability. Teams do not only need people who can prompt a system. They need people who can interrogate an answer, trace it back to evidence, understand the situation around it and know when the normal rule should not apply.
The machine has made the answer cheaper. It has not made the consequence cheaper.
Judgement is not a mysterious instinct
People sometimes talk about judgement as if it were a gift: something a seasoned leader simply has. That makes it sound vague and impossible to design into everyday work.
Good judgement is more concrete. It is context made accountable.
It brings four things together:
AI can support every one of these. It can find evidence, summarise context, model outcomes and record an audit trail. But it cannot take responsibility in the human sense. It cannot stand in front of a client, colleague or community and say: I understood what was at stake, and I chose this.
That final act of ownership changes the quality of a decision. It forces a person to look beyond whether an output is technically acceptable and ask whether it is appropriate.
Four elements of accountable judgement
Good judgement becomes practical when four things are visible: the evidence behind the answer, the context around this moment, the consequence of being wrong and the person prepared to own the decision.
| Element | Question It Answers |
|---|---|
| Evidence | What do we actually know, and how reliable is it? |
| Context | What is different about this person, moment or situation? |
| Consequence | What happens if this is late, wrong, unfair or misunderstood? |
| Responsibility | Who is prepared to own the decision and explain it? |
Photo by cottonbro studio on PexelsGood judgement is context made accountable. It is not a vague feeling; it is evidence, consequence and responsibility brought into the same room.
Paolo Di Terlizzi, Founder, Refresh
The judgement gap
Automation performs best when the world behaves as expected. The input is known. The categories are stable. The acceptable output can be described. The cost of a mistake is limited and recovery is easy.
Judgement becomes valuable when one of those conditions breaks.
Imagine a system that drafts replies to customer enquiries. Most messages are routine. Opening hours, delivery dates and service details can be answered quickly from an approved source.
Then a message arrives from a long-standing client. The words are polite, but the relationship is strained. The client mentions a delay without explaining that it has already disrupted a launch. The correct answer is not merely the correct policy. It depends on history, tone, trust and the cost of getting this moment wrong.
The system can draft. A person must recognise the exception.
This is the judgement gap: the distance between an answer that fits the pattern and a decision that fits the situation.
Photo by Mikhail Nilov on PexelsWhere the system runs out of road
Routine automation follows the road it already knows. Judgement begins at the gap: the strained relationship, the hidden consequence or the exception that a system can describe but cannot responsibly carry.
Four moments when a human should move closer
Four conditions signal that a person should move closer to the final decision. As uncertainty, consequence, relationship and irreversibility rise, so should human attention.
Photo by RDNE Stock project on PexelsAmbiguity
The request is unclear, the evidence conflicts, or the goal itself is disputed. AI can offer interpretations, but somebody must choose which problem is actually being solved.
Exception
The situation falls outside the normal pattern. A valuable client, an unusual vulnerability, a one-off constraint or an emerging risk changes what a reasonable response looks like.
Consequence
The decision affects money, safety, reputation, opportunity or trust. Even a low probability of harm may deserve more scrutiny when the impact is difficult to reverse.
Relationship
The manner of the decision matters as much as the outcome. A technically correct message can still damage confidence if it ignores emotion, history or power.
A decision should have a shape
The phrase "human in the loop" is too blunt. It does not say what the human is there to do. A person who only clicks approve may add delay without adding judgement.
A stronger model gives different kinds of work different decision shapes.
This is not a maturity ladder where every task should eventually become automated. Some work should remain human-led because interpretation is the work.
The goal is not maximum automation. It is the right allocation of attention.
Four decision shapes, not one loop
“Human in the loop” is too blunt. Different work deserves different decision shapes, from routine audit to human-led exploration. This is not a maturity ladder; it is a way to put attention where it creates value.
| Decision Shape | Suitable Work | Ai Role | Human Role | Minimum Evidence |
|---|---|---|---|---|
| Automate and audit | Routine, reversible, well-defined | Complete the task | Sample and improve | Approved source, logs, exception rate |
| Draft and review | Context-sensitive, moderate impact | Prepare a recommendation | Check, edit and approve | Source links, assumptions, confidence |
| Analyse and decide | High consequence or contested | Compare options and surface trade-offs | Make and record the decision | Corroborated evidence, alternatives, rationale |
| Explore and lead | Novel, ambiguous or strategic | Expand the field of possibilities | Frame the problem and set direction | Explicit unknowns, experiments, review point |
The danger of delegated certainty
When people use AI under pressure, a subtle change can happen. The system's confidence becomes a substitute for their own examination.
The draft looks complete, so the reviewer scans rather than reads. The recommendation includes reasons, so the manager assumes the evidence has been checked. The answer resembles previous answers, so nobody asks whether this case is different.
Responsibility has not formally moved, but attention has.
There are warning signs:
- Nobody can say which source the answer depends on.
- Review means checking tone and spelling rather than checking reasoning.
- The same approval process is used for low-risk and high-risk decisions.
- Exceptions are treated as failures to comply with the system.
- Teams measure time saved but not corrections, complaints or missed context.
- People cannot explain what would make them reject the AI recommendation.
These are not failures of the model. They are failures of work design.
The ten-second judgement test
Before accepting an AI-assisted decision, pause for ten seconds. The ritual interrupts the dangerous habit of accepting a plausible answer before anyone has considered the situation.
What is the system assuming?
Every answer rests on a frame. Is the customer asking for information or signalling a loss of trust? Is the report describing performance or hiding a change in definitions?
What can it not see?
History, emotion, informal agreements, organisational politics and emerging events may never appear in the prompt.
What happens if it is wrong?
Some mistakes are easy to reverse. Others create a record, deny an opportunity, spend money or damage a relationship.
Who owns the next move?
If nobody is willing to explain the decision, the process has created output without ownership.
Build evidence into the experience
Good judgement should not depend on heroic people remembering everything. The workflow can make better decisions easier.
Show the source beside the recommendation. Make assumptions visible. Separate facts from inferences. Flag the conditions that pushed a task into human review. Record not only who approved the output, but what they changed and why.
Design the review around consequences. A routine content summary may need a quick sample. A client promise, pricing decision or public claim may need named evidence and a second pair of eyes.
Give reviewers permission to stop the process. If the interface makes rejection feel like failure, people will approve weak work to keep the queue moving.
And create feedback that teaches the system and the organisation. Which recommendations are repeatedly changed? Which exceptions appear most often? Which source is missing when the answer goes wrong? Judgement leaves clues. Capture them.
Judgement is a leadership practice
Leaders have always made decisions with incomplete information. AI changes the speed and volume at which those decisions arrive.
A manager may now review more proposals, messages, plans and recommendations in an hour than they once saw in a day. Without a clear decision architecture, the organisation can become fast at generating work and slow at understanding it.
Leadership therefore includes deciding where certainty is false, where speed is useful and where a pause protects value.
It also means modelling intellectual honesty. Saying "we do not know yet" can be a stronger act than approving a polished answer. Asking for the source is not resistance to innovation. Changing course when context shifts is not inconsistency. These are signs that judgement is alive.
The best leaders will not compete with AI on recall or output. They will create the conditions in which people and systems can make better decisions together.
Keep the responsibility visible
AI will continue to become faster, cheaper and more capable. More work will be drafted, classified and completed without a person touching every step.
That can be a good thing. People should not spend their lives moving information between boxes.
But removing effort is not the same as removing responsibility. The more invisible the machinery becomes, the more deliberately an organisation must show where judgement lives.
Name the decisions that matter. Define the evidence they require. Decide who owns them. Make exceptions visible. Review the consequences, not only the efficiency.
Then let automation do what it does well, without asking it to carry what only people can carry.
Efficiency is only useful when it accelerates the right decision. The future of good work is not human or machine. It is clear responsibility, intelligently supported.
Paolo Di Terlizzi, Founder, Refresh
Photo by Kindel Media on PexelsChoose one recurring decision this week
Map the evidence, context, consequence and owner. Decide which of the four decision shapes it deserves.

