Skip to main content
JournalAIN° 011 / 2026

Spotting the tells is not the same as reading

People have swapped reading for tell-spotting, and the evidence says they can't do it anyway. On detection that fails, a government report nobody checked, and what actually separates thinking from its absence.

We have got very good at spotting AI writing and very bad at reading.

The tell-spotting is a genre now. Em dashes are evidence. "Delve" is evidence. So is a tidy tricolon, a parallel structure that lands a bit too cleanly, a paragraph that resolves without a loose thread anywhere in it. There are checklists. There are people who post the checklists, and people who arrive under someone's writing with a one-word verdict they reached in about four seconds. And the skill is real, in the narrow sense that these patterns exist and the people naming them are often right about where the text came from.

What nobody seems interested in is the question the skill quietly replaced.

The complaint is fair, which is why it needs answering properly

Start with the aesthetic objection, because it deserves the most agreement and the least argument. A lot of AI writing is flat. It has a house style: smooth, evenly weighted, agreeable, every sentence carrying the same amount of load, every section arriving at a conclusion nobody could object to. It reads like a person with no preferences. If you have spent the last two years watching that texture spread across your feed, the irritation is earned and I am not going to talk you out of it.

There is a separate objection about how these models were trained and who was paid for it, and it is a serious one. It is also a different argument from this one, and I am not going to pretend I have settled it in passing.

The aesthetic complaint, though, is doing something the people making it rarely admit. It is a claim about texture being used as a claim about worth. Flat prose is a fair reason to find something dull. It says nothing at all about whether anyone thought.

What reading was supposed to be

The stronger objection is the one about authenticity, and it is worth taking seriously rather than dismantling.

When you read someone, you are meant to be in contact with a person. Their judgment, the things they have seen, the argument they are prepared to defend if you push. That is the whole transaction, and it is a reasonable thing to feel protective of. Nobody signed up to read the averaged output of everything ever written. If that is what people are guarding against, they are guarding something worth having.

Here is where I part company with them. Tell-spotting does not protect it. Not even slightly.

It also does not work, which ought to end the argument on its own. Fiedler and Döpke put short excerpts in front of 63 lecturers at a German university of applied sciences and asked them to sort the human writing from the machine writing. The lecturers managed 57 per cent on the AI texts and 64 per cent on the human ones. Against a coin flip. The AI passages written to a professional standard were correctly identified by fewer than one in five. These were educated readers, given the task in advance, taking their time. Barely above guessing.

A checklist can tell you that a sentence has a certain shape. It cannot tell you whether anyone thought.

A checklist can tell you that a sentence has a certain shape. It cannot tell you whether anyone thought. Those are different questions, and the first one has quietly eaten the second because it is faster, it feels like discernment, and it can be performed in public at no cost. Somebody who has thought hard for a week and used a model to get the words onto the page will fail the test. Somebody who has thought about nothing at all and written it up longhand in their own uneven prose will pass. If your detector rewards the second and punishes the first, it is not defending the human part of writing. It is defending a costume.

And the reflex has an obvious blind spot, which is everything that gets waved through because it looks like it was checked.

The same refusal, pointed the other way

In July 2025, a report landed on the website of Australia's Department of Employment and Workplace Relations. It had been commissioned from Deloitte for AU$440,000, it reviewed the compliance system that automates penalties inside the country's welfare apparatus, and it ran to more than two hundred pages.

It contained a quote attributed to a federal court judge that the judge never said, and references to academic research that did not exist. The Associated Press reported that the errors surfaced only after Chris Rudge, a Sydney University researcher of health and welfare law, went through it and told journalists it was full of fabricated references. The Australian Financial Review ran the story in late August. Deloitte later published a revised version disclosing that a generative AI tool chain had been used, and repaid the final instalment of its contract.

Nothing caught it. No process, no reviewer, no sign-off. One person read it and followed the footnotes, which is the lowest bar imaginable and was apparently more than anyone else had cleared in two months.

The report was fluent. That is the entire mechanism. It read like a document somebody had checked, so nobody checked it, and the fluency did the work that verification was supposed to do. This is the same failure as the four-second verdict, running in the opposite direction. In both cases the surface decided, and in both cases nobody looked underneath.

What happened next is the more revealing half. Deloitte's position was that the corrections did not affect the substance, the findings or the recommendations, and the department agreed that the substance had been retained. Rudge's view, and he is a participant in this rather than a neutral observer, was that you cannot trust the recommendations when the foundation is a flawed and undisclosed methodology. Two ways of reading the same document. One of them treats the surface as the thing. The other one asks what is holding it up.

The prototype that looks like every other prototype

There is a version of this where the surface reading is entirely correct and still tells you nothing.

Someone without a technical background can now build a working demo of an idea in a weekend. It will look like every other thing built that way: same layout, same components, same palette, same rounded corners, a house style you can spot across a room. Anyone calling it derivative is right. They are also answering a question nobody asked, because a prototype exists to settle something else entirely.

The point is finding out whether anybody wants the thing. Five people with a working demo in front of them will tell you more in an afternoon than six months of describing it ever will, and if four of them shrug, that founder has learned something for the price of a weekend rather than a seed round. Then an engineer takes the idea, throws the code away, and builds the real one.

They are right about the surface and wrong about what it is worth. Those two things sit together comfortably, and the reflex has no way of holding both at once.

Where the barrier actually sits

Put two findings next to each other and something uncomfortable falls out.

Zhao, Chen and Cox at Sheffield surveyed 124 students with disabilities about what these tools actually do for them. Seventeen singled out something the authors call focalisation, and they were mostly the dyslexic and neurodivergent ones: help holding an argument in view, tracking whether they were still answering the question they had been asked. What they wanted was order, not eloquence.

Seventeen people is an indication and I am not going to dress it up as more than that. These tools are three years old. The research into who they actually help is younger still, and it concerns a group whose working needs the average employer has only lately started taking seriously. Nobody has the numbers yet. What we have is a direction, and the direction is worth following.

Order is exactly what the detectors punish. Structured, consistent, plainly worded prose scores as machine-made, because predictability is the thing being measured. So the people getting the most out of these tools are, by the design of the tools built to catch them, the people most likely to be accused of not having done the work. The same trait, read as help on one side and as guilt on the other. Nobody planned that, which somehow makes it worse.

The research is not uncritical and neither am I. The same body of work raises skill erosion and dependency, and the students themselves named inaccuracy and the cost of a subscription as real barriers. All three are fair.

I am on both sides of that pincer, so take what follows as a declaration of interest rather than as evidence.

I am dyslexic. The ideas were never the difficulty; the transcription always was. I have ADHD, so my thinking arrives out of order and assembling it into something a reader can follow is a tax I pay on everything I write. AI does not supply my judgment. It closes the distance between having the thought and getting it down, and that distance was only ever logistics.

The argument is mine. What to keep, what to cut, which objection to concede and which to press, whether the Deloitte example proves what I want it to prove or only looks like it does. That is the work. It is also invisible on the page, which is the whole problem, because the visible part is exactly the part a model can help with.

So the writing I do now carries the signals. Cleaner sentences than I could have produced unaided ten years ago, better structure, fewer of the seams that used to show. By the four-second test, that is disqualifying. I would only ask what the test actually measured.

Read it, then decide

The reflex is not going to survive contact with the next few years anyway. Every tell on every checklist is a moving target, the models are already learning around them, and the em dash detectives will be wrong more often each year until they are just wrong. Anyone building their judgment on surface signals is building it on something with a short shelf life.

What does not expire is the harder thing. Does this claim hold. Does the example prove what it is being asked to prove. Would the person who wrote this be able to defend it in a room, out loud, without the document in front of them. Those questions worked before any of this and they will work after, and they have exactly one requirement, which is that you read the thing.

Whether anybody thought is the only question in any of this that has ever mattered, and reading is still the only way anyone has found to answer it. Four seconds will not get you there. Nothing else will either.

Written by
Federico Corradi
Published
July 28, 2026
Reading time
10 min read
Topics
AI, Leadership
Edition
N° 011 / 2026
AILeadership