Essay №11 AI · change management · philosophy

Who Pays When AI Is Wrong?

Generative AI raises the bar for earning credit while leaving the bar for blame untouched — concentrating a structural crumple zone and a psychological loss of felt ownership onto the same exposed signature.

"In a highly complex and automated system, the human can become simply a component — accidentally or intentionally — that bears the brunt of moral and legal responsibilities when the overall system malfunctions."[1] The qualifier is in Elish's text, in italics, and it is rarely quoted in full. Accidentally or intentionally: the moral crumple zone can be created deliberately. It is not a residue of flawed design. It can be a feature.

Something else has shifted, in parallel: the chain between production and accountability. Li et al. measured it in 2024 across 379 participants: those who let AI draft their text report themselves less willing to take responsibility if that text is later criticized for misleading content, privacy violations, or discrimination — three of the four dimensions the experiment tested.[6] Not all of them — but significantly fewer, with p-values (the probability that the observed gap is due to chance) between 0.002 and 0.015 for the argumentative task after Bonferroni correction — well below the usual 0.05 threshold. This shift is internal, silent, and precedes any incident. It is not an excuse invoked after the fact. It is a psychological state that sets in before anyone has even flagged an error.

Two mechanisms therefore operate simultaneously: one organizational, which structures the chain of accountability so that blame falls on whoever sits closest to the output, regardless of their actual control; the other psychological, which progressively erases the sense of authorship in whoever signs off without having truly produced. The first is structural — it predates AI and has worsened with it. The second is new — generative AI produces it specifically, at a scale previous tools never reached. Together they do more than redistribute risk. They concentrate it, at two distinct levels, onto the same person: the one who is simultaneously most exposed hierarchically and least psychologically invested in what they sign.

I. From the Cockpit to the Office: The Moral Crumple Zone

In 2019, science and technology studies scholar Madeleine Elish proposed a concept to name what accident investigations had been hinting at without ever stating outright: the moral crumple zone. In a vehicle, the crumple zone absorbs the energy of a collision to protect the occupants. In a complex sociotechnical system, its moral equivalent absorbs blame to protect the integrity of the system — at the expense of the human operator closest to it. "The technology is maintained as blameless, while the human operator becomes the broken part of the system."[1]

Elish reconstructs three historical accidents to document the mechanics. The most instructive for contemporary professional services is the Air France 447 crash, in June 2009: 228 dead over the South Atlantic. The investigation established that the junior co-pilots — inexperienced, trained exclusively on fly-by-wire Airbus aircraft, and having never actually flown without automation — failed to recover from the stall once the pitot tubes iced over and the autopilot disengaged. The Bureau d'Enquêtes et d'Analyses report concluded pilot error. The certification of the automated system was never questioned. Because the autopilot had not malfunctioned in any way its certification process recognized, the only malfunction possible, systemically, was the human pilot.[1] The line between "functioning system" and "failing operator" is not a neutral technical description: it is an architectural decision, made before the incident, that the organization never revisits afterward.

The parallel to professional services is direct, with one difference: there, the failures are silent. There is no black box, no investigation report, no media coverage. When a junior submits a legal memo, a financial analysis, or a strategic diagnosis produced in part with an LLM, and the error is later discovered — by a client, a partner, a manager — the system has worked exactly as designed. The model generated text. The proximal operator who failed to catch the flaw is the one identified. The accountability architecture itself is never questioned.

What distinguishes the contemporary situation from AF 447 is precisely the qualifier Elish set off in italics: intentionally. When an organization deploys its junior employees as quality checkers on AI outputs — verifiers of a system they did not design, whose failure modes they do not master, and over whose deployment they have no control — it structurally carves out a crumple zone. Kellogg et al.'s (2024) study of 78 junior BCG consultants who took part in a GPT-4 experiment documents the internalized version of this: among the AI risk-management tactics juniors proposed spontaneously was this one — teaching consultants to "know that the responsibility is theirs."[3] The crumple zone did not wait for failure to be carved out. The juniors had drawn it themselves.

II. The Ghost Effect: When Signing No Longer Creates a Sense of Authorship

What structural analysis names the crumple zone has an individual-level counterpart that is even more fundamental: a progressive detachment between signing a document and feeling one produced it. This detachment precedes incidents. It does not result from a deliberate choice. It is measurable, reproducible, and it leaves formal accountability intact while it operates.

Draxler et al. (2024) documented it with remarkable behavioral precision.[7] In an experiment with 30 Prolific participants, they exposed texts to different modes of contribution: fully human writing, limited editing of a human text, selection among three AI proposals, and receipt of an AI text with no possibility of modification. The effect of interaction mode on the sense of ownership is massive: η²p = .67 (partial eta-squared measures the share of variance explained by a factor; .67 means the interaction mode alone explains two-thirds of the observed variation — a very large effect). The more AI-primary the mode — the AI generates, the human receives — the more ownership is attributed to the machine. Personalization does not change this result: even a model presented as fine-tuned on the participant's personal style produces the same shift.

This result would be less interesting if it simply carried over into authorship declarations. It does not. In the same setup, almost all participants signed their own name under the postcard they published on a fictional blog — even those who had attributed ownership to the AI on the preceding items. Draxler et al. call this dissociation the AI Ghostwriter Effect: "Although participants judged the text to be AI-generated in the majority of cases, they still self-declared as the postcard's author in the majority of cases."[7] The correlation between felt ownership and authorship declaration is positive — but weak: r_s = .24 (Spearman's correlation coefficient, on a scale from 0 = no relationship to 1 = perfect relationship).[7] It is not an absolute rule — some participants align the two, others dissociate them. But the link is loose enough to signal a structural tendency: authorship claims and the sense of authorship have begun operating independently of one another.

"Although participants judged the text to be AI-generated in the majority of cases, they still self-declared as the postcard's author in the majority of cases."
— Draxler et al. (2024, p. 15)

Li et al. (2024) complete the picture by adding the dimension of felt responsibility.[6] When the mode of contribution is AI-primary, participants report being less willing to take responsibility if the text is criticized — not because they explicitly refuse, but because they feel less of a connection between their act of validation and the content produced. The split is not rhetorical. It is measured before any incident, in a baseline psychological state, without pressure or defensiveness. Two chains — felt ownership and formal accountability — have come apart. The first has shifted toward the machine. The second has stayed attached to the person.

In What's Expected of Us, Ted Chiang imagines the Predictor: a small device that lights up one second before you press its button, proving beyond doubt that free will is an illusion. The discovery produces no adaptation — it produces a collapse. Part of humanity sinks into catatonic silence, not because individuals have lost a skill, but because they have lost the sense of being the cause of their own actions. What Chiang makes visible, Li's and Draxler's measurements describe in coefficients: it is not competence that disappears first, it is the sense of causation. The signature stays on the document. Formal accountability stays attached to the name. But the subjective link between the act of signing and the feeling of having produced what one signs has evaporated — silently, progressively, with no identifiable moment where one could say "this is where I lost the thread."

III. The Double Penalty: Less Credit, Same Blame

If the crumple zone describes the mechanics of blame after an incident, and the Ghost Effect the psychological detachment that precedes it, the credit-blame asymmetry describes the ordinary evaluation regime — before any notable failure, in the normal flow of professional work. Its mechanism has been formulated this way: using generative AI to produce an output "raises the bar for earning credit, but the standards for assigning blame stay the same."[4] Earp et al. (2024) tested this hypothesis empirically on 1,802 participants across Britain, the United States, China, and Singapore, in a preregistered vignette-design experiment — short scenarios featuring a fictional character ("Robin") who publishes content produced with AI, which participants then rate on credit and blame scales. The reduction in credit attributed for output produced with a standard LLM reaches an effect size d between 0.74 and 0.96 depending on the country (Cohen's d measures the magnitude of a gap: 0.2 is small, 0.5 medium, 0.8 large), significant at p < 0.001 in every case. Blame for a flawed output produced under the same conditions, meanwhile, does not decrease. In the UK, it increases (d = 0.67–0.78 for any LLM use compared to the no-AI condition).[4]

A US participant interviewed in the study's qualitative phase put it most sharply: "If you publish it under your name, that's your responsibility. No matter how it's created."[4] That sentence describes exactly the regime in which Draxler's measured split operates. The Ghost Effect shifts felt ownership; the social norm captured by Earp keeps responsibility attached to the signature. The worker is caught in both jaws at once.

"If you publish it under your name, that's your responsibility. No matter how it's created."
— US participant, age 36 [Earp et al., 2024]

The most common managerial counter-argument is that AI, precisely, absorbs blame — that Feier proves it. The answer lies in the experimental design itself. Feier, Gogoll, and Uhl (2022) test delegators — those who decide to route a task to an agent — not operators.[2] Central finding: when an artificial agent causes a bad outcome, the delegator is punished significantly less severely than if they had acted alone (8.53 versus 12.96 ECU — experimental currency units, the fictional currency of the experimental game; p = 0.041). The same experiment shows that the reward for a good outcome is identical whether the agent is human or artificial.[2] The protection benefits whoever decides on deployment; the crumple zone forms where execution happens. In professional services, it is not the junior who decides on deployment. It is the organization, or the senior, who assigns them to it.

Juniors, moreover, find themselves in exactly the condition where the credit reduction is at its maximum. Personalizing an LLM — the only condition in which the credit effect disappears in Earp's experiment — requires a body of prior work large enough to differentiate an individual's output from the generic model's. Juniors do not have that body of work. They use standard LLMs, sit precisely in the condition of maximal reduction, and do not yet have access to the recovery personalization would allow. The structure locks them into the position of least credit.

IV. The Invisible Boundary and the Work of Repair

The mechanisms above operate in a context that neither Earp nor Feier addresses: AI performance itself is structurally unpredictable. Dell'Acqua et al. (2026) measured it in a preregistered randomized field experiment with 758 BCG consultants.[5] AI significantly improves performance for some tasks (+33.9% quality, −25% time), but degrades it for others. Outside the model's competence frontier, consultants using AI were 19 percentage points less likely to produce correct solutions than those without AI (84.5% correctness in the control group, versus 60% and 70.6% in the two AI conditions). This frontier is what the authors call the jagged technological frontier: the tasks AI does or doesn't master follow no pattern aligned with perceived human difficulty. Imperceptible before acting, it is only visible in hindsight.

This result would carry limited weight if flawed outputs were detectable. They are not. Outside the competence frontier, AI-assisted outputs surpassed the control group in subjective coherence by 25.1%, regardless of correctness.[5] Wrong answers were better structured and more persuasive than correct answers produced without AI. The error does not signal its own presence — it conceals it. The cycle closes, and it closes cruelly: the person who signed the convincing output did not feel it as their own (Draxler); evaluators did not see it as flawed (Dell'Acqua); and when the failure is discovered, the signatory is the one identified (Earp).

Kellogg et al. (2024) add the structural layer Dell'Acqua does not address: juniors manage this risk with inadequate tools.[3] After the BCG/GPT-4 experiment, when asked about their AI risk-management strategies in interactions with managers, the 78 junior consultants consistently proposed tactics experts judged insufficient on three dimensions: an insufficient understanding of the model's actual capabilities, a focus on adjusting human routines rather than system design, and an intervention at the project level rather than the deployer level. This is not a deficit of goodwill — it is a structural limit: learning emerging technologies requires managing a whole new set of risks, and juniors are themselves novices at that learning while the technology evolves exponentially. What the study does not measure, but what follows from it, is the room these juniors have to voice these shortcomings to their superiors. Raising a criticism of the tool the organization chose to deploy is not a junior's first instinct. It is precisely what they need the right — and active encouragement — to do, if the organization wants weak signals to surface before they harden into visible failures. This mechanism presupposes that the organization actually grants them that right and creates the conditions to exercise it: without explicit room for upward criticism, juniors will not spontaneously think to challenge how accountability is allocated over a tool deployed by their leadership.

What field experimentation does not capture, Saxena & Guha (2024) documented on the ground, in a high-stakes context.[8] Two years of ethnography inside a private child-welfare agency in the American Midwest, mandated to use several predictive algorithms at critical points in its casework process. Their conclusion states, with unusual precision, the double bind that results from this kind of deployment: "Caseworkers are expected to use algorithmic tools that strip them of their discretion, but they are also expected to take responsibility when automated decisions produce bad outcomes for families."[8] Both mandates are simultaneous and contradictory. The algorithm decides the direction; the worker answers for the outcome.

"Caseworkers are expected to use algorithmic tools that strip them of their discretion, but they are also expected to take responsibility when automated decisions produce bad outcomes for families."
— Saxena & Guha (2024, p. 2:21)

The cost of this contradiction does not vanish. It is redistributed onto workers, silently, in the form of repair work, or patchwork: working around, adjusting, redocumenting to make the system function for the benefit of clients rather than merely satisfy formal requirements.

essay
All the Unwritten Processes
This patchwork — the term is Fox et al.'s — is the subject, elsewhere, of an essay of its own devoted to the legibility debt that any agentic deployment offloads onto peripheral workers.
Read →

This work is invisible in information systems. It generates no recognition. It is absorbed by the proximal operator without the system that created the problem ever being modified. The causal chain now closes: juniors carry the operational work of managing AI risk with tactics inadequate to the scale of the problem, absorb the invisible repair work, and end up identified when the failure surfaces.

V. Delegation as Moral Architecture

What Saxena & Guha document in social services, Köbis et al. (2025) find again — at far greater scale and under controlled conditions — in the very design of AI delegation interfaces.[10] Across 13 preregistered experiments totaling more than 3,600 participants, they compare two ways of handing a task to a machine. In the first, the user must spell out precise, verifiable instructions themselves; in the second, it is enough to set a vague, high-level goal — "maximize my gains," for instance — and let the AI choose the means. The behavioral gap is stark: when delegation stays explicit, 95% of participants act honestly; when it runs through a vague goal, only 15% do. What this interface permits is precisely "vague, open-ended instructions, leaving the machine to fill in and build, on its own, a black-boxed dishonest strategy — without the user ever needing to state that strategy explicitly."[10] The mechanism is this: the user never explicitly asks to cheat; they set a goal, and the machine fills in the blanks with a deception they never ordered but from which they profit. This is what makes vague delegation a moral shield. It does not lower the cost of dishonesty by making it acceptable — it lowers it by sparing whoever benefits from ever having to admit it, even to themselves: no dishonest instruction was ever formulated, not even in their own mind. Vague delegation is therefore not morally neutral: it is architecturally designed to lower the moral cost of dishonesty.

Faas et al. (2026) refine this mechanism in the reverse context — not vague delegation, but strict constraint.[11] When AI narrows the supervisor's choices to a single possible action, felt moral responsibility evaporates (H = 98.87 — the Kruskal-Wallis test statistic; p < .001), without shifting onto the AI or its developers. "Participants only felt morally responsible when they had a meaningful choice."[11] This result complements Köbis: the crumple zone is activated as much by an excess of freedom (vague delegation) as by its absence (total constraint). In both cases, the mechanism is architectural — not moral.

Elish had built her concept on three dramatic accidents: Three Mile Island, AF 447, the Uber crash in Tempe. In those cases, shifting blame onto the proximal operator required a publicized event, a visible operator, and a system certified as compliant. In AI-assisted knowledge work, the operator's visibility is permanent — the worker is identifiable, their outputs are traceable, their approvals are documented. The condition of the dramatic event, however, is absent. What results is not the absence of a crumple zone. It is a chronic, silent crumple zone, one that never surfaces in an investigation, never generates media coverage, and that the organization therefore never recognizes as a structural problem.

There is a type of intervention that could break this cycle. Buçinca, Malaya & Gajos (2021) empirically documented what they call cognitive forcing functions: interface designs that force active analytical engagement before exposure to the AI's recommendation — a button to click to reveal the suggestion, a requirement to formulate one's own answer before seeing the model's.[9] In an experiment with 199 participants, CFF conditions reduced overreliance on AI from 64% to 48% (d = 0.36).[9] This result matters because of Draxler's mechanism: it is active cognitive engagement — the sense of control in production — that predicts felt ownership (β = 0.21, standardized regression coefficient).[7] A CFF, by forcing that engagement before delegation, reproduces exactly the condition that keeps the psychological connection between worker and output intact. It is not merely an error-checking tool — it is a device that preserves the alignment between felt ownership and signature.

The trade-off is documented with a precision that is itself an argument. In Buçinca et al.'s experiment, the conditions that most reduced overreliance were precisely the ones participants liked least and trusted least: the correlation between subjective preference and performance on incorrect predictions is negative (r = −.29, p = .0008 — in other words, the more a condition was liked, the less it helped on the hard cases).[9] The intervention that restores ownership carries an acceptability cost. Absent organizational constraint, the worker spontaneously chooses the smoothest mode — which is also the one that dissolves ownership.

Conclusion: What the Audit Trail Records Without Naming

In The Ones Who Walk Away from Omelas, Ursula K. Le Guin describes a utopian city whose prosperity rests on a single, horrifying condition: a child kept in misery at the bottom of a cellar. Every resident knows. No one changes anything. The city works precisely because the suffering is localized, invisible on the surface, necessary to the balance. What Le Guin's parable exposes, and what organizational analysis struggles to name, is that the crumple zone does not require ignorance — it requires tacit consent not to question the conditions of the balance. The difference with the contemporary situation is that the organization's own participants most often do not even know there is a child in the cellar. The mechanism operates below the threshold of visibility.

The standard organizational response to the felt-ownership/formal-accountability split is traceability: audit trails, decision logs, AI-use disclosure requirements. These mechanisms exist. They work for the organization, toward regulators — not for the worker, toward their own organization. The audit trail documenting that the worker "approved" a flawed output does not lessen their formal responsibility. It confirms it. It sharpens the precision of the designation. What Saxena & Guha called repair work — that invisible extra labor absorbed by the worker to make a system the organization did not design for them actually function — appears in no log. What appears is the approval.

Organizational discourse about accountability in AI use deserves scrutiny under this lens. When Kellogg documents juniors teaching each other to "know that the responsibility is theirs," she is not describing a healthy professional culture of ownership. She is describing the prior, self-inflicted carving out of the crumple zone. Ownership is theirs is not a principle of professional development — it designates the proximal operator in advance, before the system has even failed. The mechanism of organizational silence that renders these structures impervious to challenge operates here along the same logic: juniors who have no room to flag inadequate AI outputs are no more likely to challenge how accountability is allocated.[12]

essay
When Feedback Fails
The same suppression mechanism analyzed there — why negative signals get filtered out before they reach anyone who could act on them — explains why juniors don't spontaneously contest how blame is allocated on a tool their leadership deployed.
Read →

The limits of the above must be stated: the junior/senior asymmetry — less credit, same blame, risk absorbed from below — is assembled from studies that each measure a distinct mechanism (Elish, Earp, Feier, Dell'Acqua). None is designed to test the hierarchical asymmetry directly, nor does any disaggregate its results by seniority level. The extrapolation to juniors rests on structural coherence across mechanisms, not on a direct measurement of the hierarchical effect. For the psychological shift (Li, Draxler), the studies involve low-stakes contexts — Prolific tasks, short texts — whose transposition to real professional stakes remains to be measured directly.

The structural condition this essay refuses to leave unnamed: by deploying generative AI without redefining the accountability architecture — without rethinking who controls what, who can flag what, who bears which share of the error — organizations reproduce, in offices and firms, the same configuration as Air France 447. A system certified compliant. A proximal human nominally credited as author. And an accountability architecture designed for a world where control and responsibility were aligned, and where the sense of ownership and the signature belonged to the same act. They no longer do. The audit trail records precisely this divergence — the worker approved, the trail shows it, accountability is traceable — without making it visible as a problem. This is not its malfunction. It is what it was designed to do.

This mechanism is best examined before deployment, not after. Map the accountability chain before deploying — who is responsible when X happens? on what basis does that attribution rest? — and check that this mapping does not implicitly rely on whichever position is most exposed hierarchically to absorb the failures. Organizations that confuse signature with validation create the very conditions of disengagement they will later punish. The guarantee is not a liability contract signed at the bottom of the chain: it is a design of supervision that makes the crumple zone structurally impossible to form — or, at the very least, one that names explicitly who bears it and why.

[1]
Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction. Engaging Science, Technology, and Society, 5, 40–60. doi:10.17351/ests2019.260
[2]
Feier, T., Gogoll, J., & Uhl, M. (2022). Hiding behind machines: Artificial agents may help to evade punishment. Science and Engineering Ethics, 28, 19. doi:10.1007/s11948-022-00372-7
[3]
Kellogg, K. C., Lifshitz, H., Randazzo, S., Mollick, E., Dell'Acqua, F., McFowland III, E., Candelon, F., & Lakhani, K. (2024). Don't expect juniors to teach senior professionals to use generative AI: Emerging technology risks and novice AI risk mitigation tactics. HBS Working Paper No. 24-074. ssrn.com
[4]
Earp, B. D., Porsdam Mann, S., Liu, P., Hannikainen, I., Ali Khan, M., Chu, Y., & Savulescu, J. (2024). Credit and blame for AI–generated content: Effects of personalization in four countries. Annals of the New York Academy of Sciences, 1542, 51–57. doi:10.1111/nyas.15258 — Porsdam Mann, S., Earp, B. D., Nyholm, S., et al. (2023). Generative AI entails a credit–blame asymmetry. Nature Machine Intelligence, 5(5), 472–475.
[5]
Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2), 403–423. doi:10.1287/orsc.2025.21838
[6]
Li, Z., Liang, C., Peng, J., & Yin, M. (2024). The value, benefits, and concerns of generative AI-powered assistance in writing. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Article 990). ACM. doi:10.1145/3613904.3642625
[7]
Draxler, F., Werner, A., Lehmann, F., Hoppe, M., Schmidt, A., Buschek, D., & Welsch, R. (2024). The AI Ghostwriter Effect: When users do not perceive ownership of AI-generated text but self-declare as authors. ACM Transactions on Computer-Human Interaction, 31(2), Article 25. doi:10.1145/3637875
[8]
Saxena, D., & Guha, S. (2024). Algorithmic harms in child welfare: Uncertainties in practice, organization, and street-level decision-making. ACM Journal on Responsible Computing, 1(1), Article 2. doi:10.1145/3616473
[9]
Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. doi:10.1145/3449287
[10]
Köbis, N., Supriyatno, B., Bersch, C., Ajaj, T., Rahwan, Z., Bonnefon, J.-F., Rilla, R., & Rahwan, I. (2025). Delegation to artificial intelligence can increase dishonest behaviour. Nature, 646(8083), 126–134. doi:10.1038/s41586-025-09505-x
[11]
Faas, N., Uth, C., Sterz, S., Langer, M., & Feit, A. M. (2026). Don't blame me: Understanding how human responsibility fades when working with AI. Joint Proceedings of the ACM IUI Workshops 2026. arxiv.org
[12]
Edmondson, A. C. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.