The Perfection Trap
Last week, someone shared a screenshot of ChatGPT getting a factual detail wrong, and the caption underneath it — "And people want to replace humans with thi...
Last week, someone shared a screenshot of ChatGPT getting a factual detail wrong, and the caption underneath it — “And people want to replace humans with this” — carried the weary condescension of a person who had been waiting for exactly this moment, as though a single error in a probabilistic system constituted proof of something they’d always believed. It got thousands of likes. The same week, a solicitor in Leeds was struck off for submitting fabricated case law to a court, having invented citations and fabricated judgements and presented them as real, and nobody screenshot that either. A brief news item, then silence.
One hallucination becomes evidence that AI is fundamentally unreliable. One human catastrophe becomes a footnote worth roughly forty-eight hours of collective attention before the feed scrolls past it. The gap between those two reactions isn’t rational. It’s psychological, driven by four distinct mechanisms that all push in the same direction: judge the machine harder.
The Identity Threat
Social Identity Theory tells us that humans categorise the world into in-groups and out-groups automatically and often unconsciously, a survival mechanism that served us well on the savannah and serves us poorly in almost every modern context where nuance actually matters. We do it with nationality, profession, age, football teams, and, it turns out, with species.
When a human colleague makes a mistake, they’re in-group. They’re “us.” We extend them the benefit of the doubt because their failure reflects on a group we belong to, so we rationalise: they were under pressure, the brief was unclear, the deadline was unreasonable, the system let them down. We find context because finding context protects our own identity, and we do this so automatically that we don’t even recognise it as a choice.
When an AI makes the same mistake, it’s out-group. It’s “them.” We don’t look for context — we look for evidence that confirms what we already suspected, which is that this thing doesn’t belong here. The mistake isn’t a moment to be understood. It’s a verdict to be accepted.
This is why the solicitor’s fraud barely registers but ChatGPT’s hallucination goes viral. The solicitor is one of us, operating within a system we recognise, doing work we could imagine ourselves doing, and therefore deserving of the contextualising instinct that protects in-group members from the full weight of their failures. The AI is not. And the standards we apply to “us” and “them” have never been equal, in any domain, at any point in human history, and there’s no reason to think we’ll start now unless we recognise the pattern for what it is.
Prevention Mode for Machines, Promotion Mode for Humans
Regulatory Focus Theory distinguishes between two motivational orientations that shape how we evaluate everything around us, and the distinction matters here more than almost anywhere else in the conversation about AI. Promotion focus is about seeking gains, pursuing opportunities, being tolerant of risk and occasional failure because the upside justifies the variance. Prevention focus is about avoiding losses, preventing errors, being intolerant of mistakes because any single failure threatens the integrity of the whole system.
When most people interact with AI, they default to prevention focus. They’re scanning for what goes wrong, auditing every output, checking every claim — not because they’re rigorous, but because they don’t trust the system, and prevention focus is what distrust looks like in practice. When the same people evaluate human work, they default to promotion focus instead, looking for what’s right, assessing the overall quality, contextualising errors rather than cataloguing them, because trust activates a fundamentally different evaluative stance.
Same error. Two completely different frameworks. The AI gets audited like a suspect. The human gets assessed like a colleague.
This isn’t a conscious choice, and that’s precisely what makes it so difficult to override. Regulatory focus operates largely below awareness, which means the same quality of output will be judged as “not good enough” from a machine and “pretty solid, actually” from a person. The bar isn’t higher for AI because AI needs a higher bar. The bar is higher because distrust puts us in a psychological state where everything looks like a risk to be mitigated rather than an asset to be developed, and once you’re in that state, the evaluation is already distorted before a single piece of evidence has been weighed.
System 1 Sees the Error. System 2 Knows Better.
Dual Process Theory explains the mechanism that makes all of this so difficult to override, even when you know it’s happening, even when you’ve read articles like this one and recognise the pattern in real time, because knowing about a bias and being immune to it are very different things.
System 1 is fast, automatic, and emotional. It’s the system that sees a screenshot of an AI error and immediately feels vindicated, processing the error as a pattern without deliberation, without comparison to human baselines, without any of the contextualising work that would change the conclusion. System 1 doesn’t ask “compared to what?” It just files the evidence and moves on, confident in a judgement that feels like insight but is actually reflex.
System 2 is slow, deliberate, and analytical. It’s the system that would ask the important question: compared to the average human doing this task, is this AI better or worse? System 2 would look at the denominator. System 2 would check the solicitor’s record. System 2 would notice that we’re holding machines to a standard we’ve never applied to ourselves, and that the standard is not the product of rational analysis but of a cognitive shortcut that evolved for a world without language models.
But System 2 is expensive, and most people won’t spend that effort on a question they think they’ve already answered. The screenshot of the AI error is designed — not intentionally, but by its very nature — to activate System 1. It’s visual, it’s simple, it’s emotionally satisfying in a way that spreadsheets of comparative performance data will never be. The rational comparison, “AI performance vs average human performance across thousands of comparable tasks,” is abstract, statistical, and deeply unsatisfying to a brain that evolved to process threats, not datasets.
The Emotional Appraisal Nobody Makes
Appraisal Theory of Emotion says something that sounds simple but has profound implications for how we judge everything around us: we don’t respond to events directly. We respond to our interpretation of those events, and that interpretation is shaped by everything we believe about the thing that caused them.
When a human makes a serious error, we appraise it through a lens of shared vulnerability. “That could have been me.” The emotion might be sympathy, anxiety, or even relief that it wasn’t us this time, but the appraisal is fundamentally relational — it connects us to the person who made the mistake through a recognition of shared fallibility, and that connection softens the judgement in ways we rarely notice.
When an AI makes a serious error, the appraisal is categorically different. There’s no shared vulnerability, no “that could have been me” because it couldn’t. The appraisal becomes one of betrayal, or confirmation of suspicion, or righteous indignation at a system that promised competence and delivered failure. The emotion is colder, more distant, more final, because there’s no human bridge to carry it from one mind to another, and without that bridge, the error becomes evidence rather than experience.
This is why the reaction to AI failure is so disproportionate to the actual severity of the error. It’s not that the error is worse. It’s that the emotional appraisal transforms the same event into something that feels fundamentally different. A human error gets appraised as a moment of fallibility. An AI error gets appraised as evidence of systemic failure. Same event, different emotion, different consequence — and all of it happening below the level where you could intervene even if you wanted to.
The Operator Problem Nobody Talks About
All of this psychological machinery produces a convenient externalisation that anyone working with AI should recognise: when AI fails, it’s the AI’s fault. When a human fails while using AI, it’s still the AI’s fault. But in my experience, most AI “failures” are actually operator failures that get attributed to the tool because the psychological systems described above make that attribution feel correct, and feeling correct is dangerously close to being correct in most people’s minds.
When someone gets a rubbish output from ChatGPT, the most common cause by far is a rubbish prompt. Vague instructions, no context, no constraints, no examples — the equivalent of briefing a new hire with “do the thing” and then being surprised when they do the wrong thing. I’ve watched people type three-word prompts into a language model, get generic output, and declare that AI can’t write. The same people would never dream of briefing a human colleague that way. They’d provide context, examples, desired tone, audience, purpose. They’d review the first draft and offer refinements. They’d invest in the relationship because they understand intuitively that good output requires good input.
With AI, they expect magic from a standing start. And when they don’t get it, the psychological machinery kicks in: prevention focus locks onto the failure, System 1 processes it as a pattern, the out-group appraisal confirms the bias, and the conclusion writes itself before System 2 has even been consulted. AI isn’t ready. But the truth is that the operator wasn’t ready, and the four psychological systems described above ensure that distinction is never made.
What This Actually Means
The honest position requires holding two things in your head at once, and the discomfort you might feel reading them together is itself evidence of the bias at work: AI is imperfect, and AI is better than the average human at an increasing number of tasks. Both are true. The question is not “is AI perfect?” because nothing is, and demanding perfection from one category of tool while accepting mediocrity from another isn’t quality control — it’s the Identity Threat, prevention focus, System 1, and a cold emotional appraisal all working together to defend a verdict you reached before the evidence was in.
The people who get the most from AI are the ones who treat it the way they’d treat a talented junior colleague: clear briefs, iterative feedback, realistic expectations, and an understanding that the first draft is a starting point rather than a final product. They’re the ones operating in promotion focus, engaging System 2, appraising machine errors through the same lens they’d use for human errors, and recognising that the tool’s limitations are often their own limitations wearing a different face.
The people who get the least from AI are the ones who expect it to read their minds and then blame it when it can’t, their evaluation shaped by identity, motivation, cognition, and emotion all converging on the same conclusion: judge it harder, because it’s not one of us.
We demand perfection from machines because four separate psychological systems all push us in that direction simultaneously. We tolerate mediocrity from humans because those same systems push us the other way. Until we recognise that pattern for what it is — a bias rather than a judgement — we’ll keep holding AI to a standard we’ve never applied to ourselves, and we’ll keep calling it quality control when it’s really just the way our brains are wired.
The STAR Framework
If you enjoyed this essay, you'll find the full argument — and the framework behind it — in the book.