Take a jacket. A copywriter describes the morning it was made for, the walk to the station in a wet October, the particular relief of a hood that actually works. A retrieval system, asked whether the jacket is waterproof, finds no answer, because the copywriter never used the word. It moves on to a competitor whose page says 10,000mm hydrostatic head, taped seams, tested to EN 343.
That is the problem in one paragraph, and nobody in the business has solved it yet.
What the machine is actually reading
Retrieval systems do not read a page the way a person does. They split it into chunks and match those chunks against a question. The engineering teams behind these systems are unusually direct about what breaks: when a definition or a key detail sits several paragraphs away from the thing it explains, the chunk arrives without its context and becomes useless. A self-contained section retrieves. A carefully built argument does not, which rules out most of what good copy does: withholding the payoff, letting a claim accumulate across three paragraphs.
The questions have changed too. Search used to receive two words. It now receives a sentence: the best fabric for outdoor upholstery that will not fade in sun. Copy that contains no answer to a sentence like that is invisible to the system asking it, however well the sentence is written.
And why readers now distrust it
Here is where it turns awkward. The writing that retrieves cleanly is the writing readers have learned to recognise as machine-made.
Researchers who studied people identifying AI-generated text found that readers rely on two signals above all. Vocabulary accounted for 53.1 percent of correct identifications: repetitive phrasing and words chosen for range rather than precision. Sentence structure accounted for 35.9 percent, and the giveaways named were consistent three-item lists and a monotony of length. Separate linguistic work measured the gap directly. Human writing carries roughly twice the sentence-length variance of machine writing, uses contractions where machine writing uses almost none, and scores far higher on plain readability.
So the brand that writes for retrieval writes short uniform declaratives, front-loads the answer, avoids contractions and reaches for a list. Then a reader arrives, recognises the shape, and quietly discounts the whole page.
The writing that retrieves cleanly is the writing readers have learned to distrust.
The tension at the centre of every product page in 2026
Where the two readers actually agree
The overlap is larger than the argument suggests, and it sits in one place: specificity. A number, a test standard, a measured claim. Ten thousand millimetres is retrievable and it is also more persuasive than premium waterproof protection. Machines discount promotional language because it carries no information; readers discount it for the same reason, having seen it a thousand times.
Vagueness is the common enemy. Best-in-class, thoughtfully designed, unrivalled quality. None of it retrieves, and none of it sells.
Stop asking one text to do both jobs
The mistake most brands are making is treating this as a writing problem when it is an architecture problem. Structured data exists precisely so that machines can be answered without a human having to read the answer. A product schema carries dimensions, compatibility, materials and test standards in a form no shopper ever sees, which leaves the visible copy free to be written for the person.
Documentation teams reached this conclusion earlier than marketers, and their analogy is a good one: writing for machines resembles writing for screen readers. Nobody argues that alt text ruins photography. It sits alongside, doing a different job for a different reader. Product data works the same way, and the brands still trying to compress both jobs into one paragraph of body copy are the ones producing pages that serve neither.
What structured data cannot carry
The architecture answer is tidier than the situation deserves, because schema only covers the part of a product that fits in a field. Dimensions go in a field. Whether the thing is comfortable after four hours does not. Assistants reach for that second kind of information constantly, and they find it in reviews, in editorial coverage, in forum threads and in the prose on the page itself.
So the visible copy is still being read by both parties after all. The difference is what it is being read for. The machine is not looking for persuasion in that prose. It is looking for a claim it can attach to a question, which gives the writer a workable brief: every sentence that makes someone want the thing should imply a fact that can be checked somewhere on the page. That rules out the vague superlative, which was never doing any work anyway, and leaves the paragraph intact.
The problem is usually organisational
Most pages fail this test for a reason that has nothing to do with craft. Copy belongs to marketing. Product data belongs to operations, or to whoever owns the e-commerce platform, or to an agency that set up the feed once and moved on. The two groups rarely see each other work, and neither has been given the brief that the page has two audiences.
The teams getting this right have not found a new way to write. They have put the copywriter and the person who owns the feed in the same conversation, which is a management fix wearing the clothes of a creative one.
Sources: kapa.ai, documentation practice for retrieval-augmented systems · Research on human identification of AI-generated text, arXiv (2025) · Linguistic comparison of human and machine text, arXiv (2026) · Google, structured data and JSON-LD product markup guidance · Wordbank, on AI Overviews and localised copy · Locomotive Agency, on semantic chunking and query length.
