CONTEXT Governance
The AI Telephone Game: Why Prose Is the Real Problem
The AI telephone game. Everyone knows the children's game, called telephone in the United States and Chinese whispers in Britain: a sentence is whispered from ear to ear, and what stands at the end differs from what went in at the start. Working life has moved a good deal closer to that game than most of us realise. Anyone writing an email today often has a program phrase it. Anyone receiving one often has a program summarise it. 1 sentence becomes 3 paragraphs, and 3 paragraphs become 1 sentence again. In between sit 2 translations, and each one drops something: the clause carrying the condition, the figure from the attachment, the qualifier in front of the promise. Our answer to it stands at the end of this text and it is uncomfortable: the prose itself is the problem.
2 Calculations That Are Both Correct
Writing is where artificial intelligence shows its most tangible benefit in daily office work. In May 2026 the polling institute Gallup surveyed around 22,500 employed Americans: 52 percent use AI at work, and among them writing and editing ranks first at 51 percent. The more interesting question sits one level below that. What happens when a tool sits at both ends of the same message? Who gains the time, and where is it spent again?
The figures on benefit and the figures on effort contradict each other less than it first appears: They measure different ends of the same message. On the sending side employees consistently report a gain.
- Gallup, May 2026: 68 percent of those who use AI for writing and editing report a gain in productivity from it. The base is therefore the users of that one application rather than all AI users.
- Microsoft Work Trend Index 2026: 66 percent of the AI users surveyed say the technology gives them more time for high-value work.
- US Census Bureau, reported in August 2026: about a third of those who used AI in the previous week completed tasks 1 to 2 hours faster.
A word on how much weight such figures carry, before the counter-calculation. Almost everything published on this subject is self-reported: Somebody is asked and estimates their own time saving themselves. That applies to the whole field and is a weakness of no single study. The margin of error comes on top. A survey always catches a sample only, and with around 1,000 respondents a result swings by about 3 percentage points either way for that reason alone. A gap of 2 points between 2 surveys therefore says nothing yet. Such numbers are best read the way a shop supervisor reads her crew's timesheets: useful for the direction, imprecise in the second digit.
On the receiving side the sign flips. A survey run in the winter of 2025 into 2026 collected answers from 2,500 knowledge workers and IT decision makers across 10 countries, among them Germany, Austria and Switzerland. It reports 3 findings:
- 77 percent say AI-generated work needs more review time than work produced by people.
- 66 percent say that reviewing other people's AI text creates additional work for them.
- 43 percent have used AI-generated content even though they suspected it might be low quality or contain errors.
Those figures carry a caveat. The survey was commissioned by a provider of cloud communication and IT support, in other words a company that sells software for exactly this problem. And the respondents reported on themselves.
A better known study measures the same effect under the term workslop. It means AI text that looks tidy and hands the actual work onward: neatly formatted, fluently phrased, substantively unresolved. Whoever receives it does the thinking the sender spared themselves. Around 1,000 US employees answered in September 2025. 40 percent had received such text in the previous month, and per incident they reported just under 2 hours of rework. The researchers extrapolate that to 186 US dollars per employee per month. That sum rests on 2 estimates made by the respondents themselves: estimated salary times estimated time. The study was run by a provider of AI-assisted coaching together with a university lab. Whoever describes the problem here also sells the remedy, and that belongs beside every figure drawn from it.
A second wave from the same group in 2026 shows how carefully such numbers should be handled: 962 respondents, 38 percent affected, 3.4 hours of rework per month. Prevalence therefore failed to rise compared with 2025. Calling it a decline goes too far as well: the gap from 40 to 38 percent sits inside the margin of error. For 2025, moreover, either 40 or 41 percent was reported depending on the publication. The unit changed as well, since 2025 counted just under 2 hours per incident and 2026 counts 3.4 hours per month. The coverage built a growth story out of that. A valid comparison it is however not, as little as hours per job can be added to hours per month.
The size of the surface all this acts on comes from the one figure here that is a measurement rather than a survey. A special report of the Microsoft Work Trend Index from June 2025 evaluated actual usage of the company's own office software, counting along instead of asking: on average an employee receives 117 emails and 153 chat messages per working day. That holds for mailboxes with this one provider. A German survey published in January 2026 arrives at 53 work emails a day, and 14 percent of the employed respondents receive 100 or more. Every one of those messages is a possible point where a tool reads along or writes along. The real message in this body of data is pleasantly simple: The same people report a gain when sending and a loss when receiving. Looking at those 2 sides separately is what allows a fix in the right place.
Where the Distortion Begins
The process takes less than 5 minutes and has 3 steps.
- A sender has a single substantive sentence in mind, for instance that a delivery date slips by 2 weeks. He gives that sentence to an assistant, meaning a program that produces text on request. From it he has a polite email made.
- The email goes out: 3 paragraphs, with an opening, a justification and thanks.
- The recipient sees 3 paragraphs in a crowded inbox and has them summarised into one sentence.
It began with 1 sentence, and it ends with 1 sentence. In between sit 2 translations, and each one drops something. In the children's game that is the whole joke. In business correspondence it is a question of commitment: the clause carrying the condition, the figure from the attachment, the qualifier „subject to approval". The loss is the same one you get from a copy of a copy. Every single pass looks clean, and something is gone at the end all the same: the AI telephone game.
What the Research Shows, and What It Leaves Open
One 2025 paper carries the name of the game in its title. Researchers sent 150 documents each through 100 iterations and checked after every round how many statements still held true. The score declines steadily, and the losses add up from round to round.
The decisive part for our case is a side experiment. The researchers took the most conspicuous ingredient out and dropped the change of language. The text was therefore only rephrased, in the same language and with an explicit instruction to preserve the meaning in full. The decay remained all the same. That is exactly the case at issue here. One assistant phrases an email, and a second one summarises it. Both do nothing other than rephrase.
2 caveats belong in the same section. What was measured is 100 iterations rather than 2. The direction is therefore established, while the magnitude fails to transfer. Converting the collapse after 10 iterations into an exchange of emails means lying with correct numbers. And the obvious assumption, that 2 different programs in one chain do more damage than one, is precisely what this paper fails to support. The curve falls most steeply where the chain is most complicated, with 5 languages involved. The finding on several different models is mixed: with one language the combination amplified the distortion, with another it dampened the decay. For the constellation „my tool writes, your tool summarises" that remains an open question, and we state it here explicitly as open.
2 further studies support the direction.
- A 2025 study of transmission chains between models shows that small distortions, invisible in any single pass, accumulate across repetitions. Open-ended instructions drift markedly more than tightly framed tasks. That is the argument against „just write this up nicely".
- An analysis of more than 200,000 simulated conversations across 15 models compares 2 routes to the same information. Where it arrives piece by piece over several steps rather than all at once, results come out 39 percent worse on average. The cause is less a lack of capability than rising unreliability: The models commit early and rarely recover from a wrong turn. That was measured on conversations rather than email chains. What transfers is the finding that information delivered piecemeal lands worse.
A preprint from financial analysis names 2 mechanisms that explain the effect. Preprint means the work was published before peers reviewed it, which lowers its weight. The first mechanism is decontextualisation: A statement is separated from the qualifying clauses that accompany it in the original. The second is dependence on the tool, since different summarisers read the same text differently. The result reads fluently and plausibly and still changes the decision behind it.
The effect on people has been measured too. An experiment with 547 participants rated texts by their authorship. The question and the analysis were fixed publicly in advance, so that the result could be massaged afterwards by nobody. Participants judged AI-generated and AI-assisted texts as less trustworthy, less authentic and less suitable for taking up knowledge. One of the 2 settings examined was explicitly the work email. What this measures is perception rather than objective information content. For the sender it is still the more uncomfortable finding: The message arrives and is taken less seriously.
That leaves the finding that is missing most. A study measuring exactly this 2-step chain under control and quantifying the loss has yet to appear. A preprint from November 2025 describes the picture word for word, while explicitly labelling its results preliminary. The AI telephone game is therefore a plausible extrapolation from established parts rather than a measured quantity. That is how it stands in this article. The practical yield is immediate all the same: The loss arises at the translation rather than at the tool. Sending the one sentence you had in mind anyway saves both sides the 2 translations.
One related subject stays out of this text on purpose: the subject-line marker #NoKI, with which a sender demands on data protection grounds that a single message never reaches a tool at all. In our own programs a marker of that kind has filtered out every message carrying it since 22 September 2026. A separate article will cover it.
Prose Is the Problem
Taking single messages out solves one part of the problem, namely the part concerning confidentiality. The larger part remains standing, and it has nothing to do with secrecy. From here on what follows is our own view rather than a finding from a study.
The core problem is convenience. A nicely phrased email at the push of a button is too good to pass up, and that is precisely why everybody reaches for it. Someone who used to spend 10 minutes on a polite rejection now has it in 20 seconds. The price for it falls due at the other end of the inbox.
Our thesis in 1 sentence: prose is the problem. The padding a program produces is what makes the telephone game possible in the first place. It is like a stereo with the bass turned up too far: the words of the song are still playing, and hearing them is impossible. The actual message sits between the opening, the justification and the thanks, and that is where it drowns. Only that volume of text gives the recipient a reason to have it summarised again.
The counter-test is simple. If we actually sent hard and clear single statements only, we would not have this problem at all. „The delivery date slips by 2 weeks to 14 October, and approval is still pending." That is 2 statements in 1 line, and they survive any relay unharmed. There is nothing left to summarise there, because everything dispensable is already absent.
Back to the Short Message
Our proposal follows from that, and it is uncomfortable: back to short and pointed replies. A house convention is even conceivable, one that rules out having text written for you in the first place. The information is pushed across short and pointed, with no prose in front of it and no prose after it.
In daily practice that would mean 4 things.
- A reply runs 1 to 3 sentences and carries the decision, the figure and the condition.
- The salutation and the thanks appear once at the start of a correspondence rather than in every single email.
- Full prose is written for the customer and for the contract, while internal coordination goes without it.
- An email that is short already needs no summary. The second translation therefore falls away by itself.
The side effect is the real gain. Writing the one sentence yourself requires knowing beforehand what you want to say. That thinking is exactly what the fully phrased email hands onward today, and the recipient ends up doing it.
Honesty requires the objection. A convention that abolishes politeness has a hard time of it. The opening and the thanks are older than any office software and they serve a purpose: they show respect and keep a business relationship warm. An email of 6 words reads as brusque, and on a first contact it may well close a door. This tension resists resolution, and we leave it standing here on purpose.
Our position on it is clear all the same. Politeness rests on the person and on the reliability of a promise, and it rests on the volume of text in no way at all. A short sentence that is correct and triggers no rework for the recipient is the more polite message. Agreeing that trade within your own house pays at both ends of the inbox: less effort in writing and less checking in reading. And the message arrives the way it was meant.
On the sources. The usage figures come from Gallup (surveyed 6 to 20 May 2026, 22,573 employees, margin of error 0.9 percentage points; the 68 percent refer to respondents who use AI for writing and editing), the Microsoft Work Trend Index 2026 and an August 2026 report by the US Census Bureau, whose methodology we did not examine ourselves. The effort figures come from the „Pulse of Work in 2026" survey commissioned by a cloud communication provider and from the workslop study by BetterUp Labs and the Stanford Social Media Lab. Both commissioning parties sell remedies for the problem they describe. The key figures of the 2025 workslop wave differ between its own publications (1,004 or 1,150 respondents, 40 or 41 percent), which is why this article says „around 1,000" and „just under 2 hours". For the 2026 wave we found media coverage only and no primary publication; the growth story traces back to that report and its syndications. At sample sizes around 1,000 the gap from 40 to 38 percent sits inside the usual margin of error of about 3 percentage points, which is why this article says „failed to rise" and claims no decline. All prevalence and effort figures are self-reported. Measured rather than surveyed are the 117 emails and 153 chat messages per working day; they appear in the special report „Breaking down the infinite workday" of 17 June 2025 and come from Microsoft 365 usage telemetry. The 53 work emails a day and the 14 percent receiving 100 or more come from a telephone survey by Bitkom (1,002 people aged 16 and over, among them 532 employed internet users, fielded in calendar weeks 41 to 46 of 2025, published 13 January 2026); both values refer to those employed internet users and are therefore self-reported as well. For Germany there are figures on AI use by employees overall, yet none breaking out how many emails are written or summarised by AI. The chain effect rests on LLM as a Broken Telephone (ACL 2025), When LLMs Play the Telephone Game (ICLR 2025), LLMs Get Lost in Multi-Turn Conversation, a preprint on compressed financial analysis and the study by Sahebi, Formosa and Bankins on how AI-mediated text is perceived. In the first of those works the steepest factual decay occurs in chains involving 5 languages. Its finding on several models in one chain is mixed and therefore carries no claim here. These works measure 10 to 100 iterations, conversations or perceptions. The 2-step email chain described here is measured under control in none of them; a preprint from November 2025 names it and calls its own findings preliminary. The AI telephone game is therefore our extrapolation from established parts. The proposal to do without fully phrased prose is our own position. It comes from us and appears in none of the works named, and its effect has yet to be measured. The subject-line marker mentioned against the processing of single emails is likewise our own decision rather than a standard; it has been filtering in our programs since 22 September 2026 and gets an article of its own.