CONTEXT Judgment
Judgement comes from repetition, not from information
An experienced managing director senses within minutes whether a project scope holds. She sees in a plan where it will tear. With AI projects this judgement is missing remarkably often. Research explains fairly precisely why that is. It also says what a reliable gut feeling grows out of in the first place.
What we notice in project conversations
In conversations with managing directors and division heads we meet a recurring pattern. On topics they have been responsible for over years they judge within seconds: that cannot be true in principle, that makes no sense, here I see a risk. With AI this judgement comes hesitantly or fails to appear at all. What is possible, what makes sense, where a risk arises and where none does: the answers stay open.
The reasons we are given stand side by side with equal weight:
- the field is new and moves too fast
- fear of approaching AI and resentment towards it
- missing filters for the flood of information
- the lack of time to get into it themselves
This is an observation from our own conversations. It leads to the question at issue here.
Gut feeling is recognition
In decision-making circles the matter is called judgement, or the ability to think critically: being able to assess something without diving in deep. Research is less romantic. From 1985 onwards Gary Klein observed fire chiefs, paramedics and power plant technicians at work. His finding: the experienced person compares no options at all. He recognises a pattern. With the pattern come the matching cues and the first plausible action, and he runs that through once in his head.
The most quoted definition comes from Herbert Simon (1992). Daniel Kahneman and Gary Klein signed it jointly: a situation supplies a cue, the cue gives the expert access to stored knowledge, and that yields the answer. “Intuition is nothing more and nothing less than recognition.” The two add that this demystifies intuition and that no magic is at work.
That holds the good news and the bad news at once. Gut feeling is experience in compressed form. Where the experience is missing, the pattern is missing too.
Two conditions, and both must be met
Kahneman and Klein stood for opposing schools for years. One researched fast judgement as a source of systematic distortion, the other as the achievement of skilled practitioners. In 2009 they brought their positions together in a joint paper. Its starting point: professional intuition is sometimes marvellous and sometimes flawed. The real question is therefore when each of the two applies.
Joint papers by avowed opponents are rare, because they demand concessions from both sides. That is what makes them solid, and what stands written there has survived scrutiny from 2 opposing directions. The paper sorts existing research and draws a shared conclusion from it.
Their answer names two conditions:
- The environment must supply usable cues. A stable link must exist between recognisable cues and what happens afterwards.
- There must have been an opportunity to learn these cues. That takes long practice of one's own and feedback that arrives quickly and is unambiguous.
Medicine and firefighting meet both. Picking individual stocks and long-term political forecasting meet almost none of it, and there the gut feeling is worthless according to this paper. Uncertainty on its own is no reason for exclusion: the authors cite poker and war as environments that supply stable cues and are still highly uncertain. Missing regularity is the problem.
How AI fares in this grid
Apply these test points to the judging of AI projects and the result is unfavourable on three counts.
- The cues are unstable. A controlled field experiment with 758 consultants found a jagged capability frontier. Inside that frontier AI raised performance markedly, with around 12 percent more tasks completed and about 25 percent less time. On a task deliberately placed outside it, performance fell below that of the control group. Where the frontier runs is impossible to see from outside.
- The environment keeps moving. In the USA alone around 50 notable models appeared in 2025, and a widely used programming benchmark rose from about 60 to nearly 100 percent within a year. What was a valid cue last year is useless today.
- The feedback arrives late and ambiguous. In a survey of 4,454 chief executives 56 percent said their AI investments had neither raised revenue nor cut costs within 12 months. Only 12 percent reported both.
That puts the judging of AI closer to stock picking than to firefighting. This classification is our conclusion from three sources and not a statement by Kahneman and Klein. In 2009 the question did not yet arise in this form.
The uncomfortable part
Unfortunately the gut feeling does not stay silent in such a situation. It speaks up as always, only unreliably. The two authors describe the case explicitly: experts with genuine skill in one field are regularly asked to judge in another. The boundary of one's own expertise is hard to pin down, for the experts themselves as much as for their observers.
On top of that comes the paper's most important sentence for practice: subjective confidence is an unreliable indicator of whether an intuitive judgement is sound. A sound gut feeling and a baseless one feel exactly the same. An experiment with around 500 participants from October 2025 points the same way: working with an AI chat, everyone overestimated their own performance. Those who considered themselves especially versed in AI overestimated themselves the most.
What the numbers from the executive floor give us
Here the evidence base gets thin, and that belongs on the record. Usable figures come almost exclusively from in-house studies by consultancies and associations. One of them surveyed 351 chief executives and 274 supervisory board members in May 2026. Result: 75 percent of the board members rate their own AI knowledge as equal to or better than that of their peers. 63 percent of the chief executives believe the boards overrate themselves.
Measured objectively the starting position is thin. Among the 500 largest US companies the share of directors with AI expertise rose from 1.5 to 2.7 percent between 2021 and 2025.
These figures do not prove the observation from our conversations. They show a gap between self-assessment and the view of others, nothing more. We found no academic survey on the judgement confidence of executives specifically in the case of AI, and that blank is itself a result.
Why the uncertainty is rarely spoken aloud
There is an explanation for this blank, and it comes from our conversations again. In leadership roles people barely talk about weaknesses of their own. In a position where confidence is expected, nobody likes to ask about something half understood. That holds even in a room of peers, because a sense of competition runs along there too. The behaviour is understandable. Admitted ignorance is expensive in this role.
Two things follow from this.
- Responding to it is impossible. What stays unspoken cannot be worked on. It remains an unreported number.
- Decisions get made anyway. Every day, about budgets, projects and partners. The uncertainty stays invisible, the decisions do not.
That also makes clear why nothing here can be measured. An unreported number by its very nature cannot be collected. Anyone who keeps their own ignorance to themselves shows up in no survey that asks exactly about it. Something like this becomes visible only in the view of others. The gap between self-assessment and the view of others further up could therefore be more than a curiosity. That survey measured AI knowledge, however, rather than how openly somebody speaks about uncertainty. As far as one can be read from the other, it is the only hint the figures give here.
What such judgement grows out of
The condition from the 2009 paper is inconvenient: plenty of repetition of one's own plus feedback that arrives quickly and is unambiguous. Reading fails to meet that, and listening does so just as little. Four findings fill this out.
- Runs of your own beat watching. An analysis of more than 6,500 procedures by 71 cardiac surgeons shows: people learn more from their own successes than from their own failures, and more from the failures of others than from the successes of others. Watching is good above all for seeing how something goes wrong.
- Practice environments have a measurable effect. A summary of more than 600 studies from medical training finds strong effects of simulation on knowledge and manual skills. On the outcome for the patient the effect is smaller, yet it remains.
- Self-assessment can be trained, even in poor environments. In a forecasting tournament running over several years short training and repeated forecasts with a subsequent resolution improved accuracy by around 10 percent. That worked of all things on long-term political forecasts, where intuition counts as worthless.
- For weak environments there are tools. In a premortem the team assumes the project has failed and collects the reasons. According to a study from 1989 this direction of view raises the hit rate by around 30 percent.
For executives and AI none of it is proven. We found no study showing that training or hands-on use improves judgement about AI projects. The step from surgery and a forecasting tournament to the executive floor is, however, a plausible analogy. In the consultant experiment a short introduction also failed to prevent the poor performance beyond the capability frontier. Short training therefore does not replace experience.
What we make of this
This is exactly where our AI Enabled Executive programme comes in. It makes up for what elsewhere builds up on the side over years: runs of your own on an AI working environment, on a secured server of your own where nothing can break. Repetition and feedback within the run time of the programme instead of spread across years. The same condition applies here as everywhere else, and Kahneman and Klein wrote it down.
As of 2026-09. The two conditions come from a review and consensus paper by Kahneman and Klein (American Psychologist, 2009) and not from a measurement of our own. The classification of AI as an environment with weak cues is our conclusion from it, supported by the consultant experiment of Dell'Acqua and colleagues (2023) and the AI Index 2026. The figures from the executive floor (BCG, PwC, The Conference Board) are in-house studies by players with an interest of their own and without a random sample. The absence of the gut feeling in the case of AI is an observation from our project conversations and not a survey. Carrying the findings on practice and simulation over to executives and AI is an analogy that lacks solid evidence.