Vivian Voss

The Evaluation Penalty

philosophy ai shadow ai leadership

The professor of the old school dictated, and the secretary made the lecture. She typed it, corrected it, rescued its grammar without being thanked, and on a good day supplied the better half of its sentences. The hall then applauded a man delivering what was, in the strictest sense, a collaboration. Nobody called the lecture secretary-made, and nobody was supposed to. The arrangement was universal and perfectly respectable, provided no one said it out loud.

That silence did not retire with the secretary. It is the oldest working agreement of professional life: assistance is used by nearly everyone and acknowledged by nearly no one, because reputation is priced on the fiction of unassisted work. Social psychology has a name for the stable version of this, pluralistic ignorance: each person does the thing, believes the others do not, and so everyone keeps quiet together. What is new in 2026 is that the punishment for breaking the silence has been measured.

The Penalty, Now with Data

In May 2025 a Duke team published, in PNAS, four experiments with more than four thousand participants under a title that names this piece: there is a social evaluation penalty for using AI. Colleagues who disclose AI use are judged lazier and less competent for the same output. Hiring evaluators mark down candidates who admit to it. And the detail that gives the game away: the penalty largely disappears when the evaluator uses AI themselves. The judgement is not a principle. It is a posture, and it depends entirely on who is posing.

The scale of the quiet side is documented too. Microsoft and LinkedIn, in their 2024 Work Trend Index, measured three quarters of knowledge workers using generative AI at work, most of them bringing their own tools, and roughly half declining to admit using it for their most important tasks, fearing it makes them look replaceable. Ethan Mollick gave the species its name, secret cyborgs: people who have automated real parts of their jobs and tell nobody, in organisations whose elaborate rules taught them that honesty is the risky option. The professor at least got applause. The secret cyborg gets the applause only as long as the secretary stays hidden.

The silence, measured (Work Trend Index 2024, PNAS 2025) 75% of knowledge workers use generative AI at work 78% bring their own tools, unsanctioned and unseen 52% decline to admit it on their most important tasks THE TELL (PNAS, four experiments, 4,400+ participants) The penalty for disclosed AI use largely disappears when the evaluator uses AI themselves. The judgement tracks the judge, not the work.

The Bill Arrives as Data

Address this now to the leadership that believes the silence is working in its favour, because the first bill is already being paid. An employee who is punished for disclosure and handed a sanctioned tool that cannot do the job will not stop using AI. He will stop using yours. The work still flows to the model; it simply flows through the channels the company secures least. The security trade calls the phenomenon Shadow AI, and the measurements, from Cyberhaven Labs among others, are not subtle: about a third of what employees paste into AI tools is now sensitive company data, a share that has tripled in two years, and roughly three quarters of workplace AI use runs through private accounts rather than anything the company can see or audit.

The sensitive share of what employees paste into AI tools 2023 10.7% 2024 27.4% 2025 34.8% Tripled in two years; roughly three quarters of workplace AI use runs through private accounts. Cyberhaven Labs, AI Adoption and Risk Report. The least secured route carries the most sensitive cargo.

Call that what it is. The data loss was not suffered; it was organised. Punishing honesty did not protect the crown jewels, it selected the least secured route for them, and the policy that was written to demonstrate control is the instrument by which control was given away. A leadership that wants the flow back inside its walls starts by putting the subject openly on the table, and then has exactly one trade to offer: the best tools the data-protection rules allow, in exchange for amnesty for everyone who had been using better ones.

And in Europe the tool question is harder than the enterprise label suggests: the German Datenschutzkonferenz found in 2022 that GDPR-compliant operation of Microsoft 365 could not be demonstrated, and the EDPS, the EU's own data-protection supervisor, ordered the European Commission in 2024 to suspend the data flows of its Microsoft 365 use; compliance arrived only in 2025, after eighteen months of supervised rework. An EU company that treats the default stack as automatically lawful has already declined the first audit.

The Accusation as a Weapon

The second bill is subtler and nastier. In a climate where everyone uses the tool and nobody admits it, the sentence “this is AI-made” becomes a weapon that needs no evidence and no consistency. The accuser need not abstain himself; the accusation lands in the current window of opinion, dismantles valid work on arrival, and he knows it. The Duke data shows the penalty is contingent on the evaluator's posture rather than on any quality of the work; a verdict that changes with the judge's own habits is not a verdict, it is leverage.

These windows close. The same charge was once levelled at the calculator, and today a person boasting that he does long division by hand is not admired, merely late (the slide-rule generation said the same about him, and so on back to the abacus). Provenance accusations have a shelf life that quality judgements do not. Until this one expires, though, it will be used precisely where it hurts: against work too good to attack on its merits.

The museum of expired accusations THE ABACUS charge expired; nobody remembers it THE SLIDE RULE charge expired; now a collector's item THE CALCULATOR charge expired; long division by hand is late AI the same charge, window still open While the window is open, the charge is aimed precisely at work too good to attack on its merits. Provenance accusations have a shelf life. Quality judgements do not.

The Limit

The thesis has edges, and they should be said plainly. Some prohibitions are simply right: customer data, regulated records and genuine secrets do not belong in a third-party prompt window, full stop, and a company that forbids that is not posturing. And some AI-assisted work is genuinely bad; the accusation is sometimes true, slop exists, and defending the honest cyborg is no defence of it. This piece argues against exactly two things: the punishment of honesty, and provenance offered as a verdict. What matters in the end is the quality of the result, not the instrument that produced it; in able hands, the same thesis reads the same from a fountain pen, a typewriter, a laptop, or whatever comes next. Quality is testable. Test it.

The Point

Nearly everyone works assisted; nearly everyone always has. The organisation that punishes the admission does not restore the fiction of unassisted work, it merely buys secrecy at the price of its own visibility, and pays the balance in leaked data and in good work destroyed by cheap accusation. Judge the work and provision tools worth declaring; honesty then becomes the cheap option.

The penalty never stopped the tool; it only stopped the truth about it.