AI

Should You Write for the Model or for the Retriever?

Written by
Pravin Kumar
Published on
Oct 1, 2026

Should I write for the model or for the retriever?

For the retriever, because it goes first. A language model can only quote what something handed it, and the handing happens in a retrieval step that has never read your page as a whole. If the retriever does not surface your passage, the quality of your writing is irrelevant, because no model ever saw it.

Whether that model is Claude or anything else makes no difference to this. This ordering is the single most useful thing I know about getting cited, and it is the thing most AEO advice skips. People optimize the prose for an imagined intelligent reader, when the gate before that reader is a much blunter matching process that works on fragments.

So the practical question is not how to impress a model. It is how to write a page whose individual pieces survive being pulled out of it.

What is the retriever actually doing?

Finding candidate passages that look relevant to a query, usually by comparing numerical representations of meaning, sometimes combined with keyword matching. It operates on pieces of documents rather than whole ones, because whole documents are too large and too mixed in subject to match a specific question well.

That is the retrieval half of what is usually called RAG, or retrieval-augmented generation. The important consequence is that your page is not the unit of competition. A section of your page competes against a section of somebody else's. You can lose a citation while having the better page overall, because one of their fragments matched better than any of yours did.

I should be clear about the limits of what anyone knows here. The specific retrieval stacks behind ChatGPT, Perplexity, Gemini and Google's AI surfaces are not public, and anybody telling you exactly how they rank is guessing. What is public is how retrieval systems in general work, including the embedding models published by Anthropic, OpenAI and Voyage AI, and that is enough to write well.

Why does a chunk lose its meaning?

Because the sentence that explained it is in a different chunk. Anthropic, writing about its Contextual Retrieval approach, puts the problem directly: this kind of splitting works well for many applications, but it can lead to problems when individual chunks lack sufficient context. That is the failure in one sentence.

Think about what that means for a typical article. Your introduction establishes that you are discussing Webflow's CMS. Four sections later you write that the limit is generous but worth checking. Pulled out alone, that passage is about nothing. It mentions no product, no limit and no reason, so it matches no query.

The writing habit that causes this is the one good prose teachers encourage, which is not repeating yourself. Within a document read start to finish, carrying context forward is elegant. Within a document read in fragments, it is how your best material becomes unquotable.

How much does context inside a chunk matter?

Measurably, according to the people who tested it. In that same write-up, Anthropic reports that adding context to each chunk reduced its top-20-chunk retrieval failure rate by 35 percent, from 5.7 percent to 3.7 percent, in its own evaluation of the technique.

It reports larger gains when the approach is combined with other methods. Pairing contextual embeddings with a keyword-matching technique it calls Contextual BM25 reduced the failure rate by 49 percent, from 5.7 percent to 2.9 percent, and adding a reranking step brought the reduction to 67 percent, from 5.7 percent to 1.9 percent.

Two honest caveats. Those are results from Anthropic's own evaluation of its own technique, not a measurement of the open web, and they describe what a retrieval system can do to improve its handling of your content rather than what you can do from outside. What they establish for a writer is direction, which is that missing context inside a passage is a real and quantified cause of retrieval failure.

So how should you write a section?

As if it is the only part anyone will read. Name the subject in the section rather than relying on the heading above it or the paragraph before it. State the condition that makes the claim true inside the same passage. Give the number and what it refers to together, never split across a paragraph break.

In practice that means accepting a small amount of repetition that would look clumsy in an essay. If a section is about Webflow's CMS, the words appear in that section. If a claim only holds on a particular plan or under a particular condition, the condition sits beside the claim rather than three paragraphs earlier.

I have come to think of this as writing passages that can be kidnapped. Any paragraph should be able to be lifted out, shown to a stranger, and still make a true and complete statement. That test has changed my drafting more than any keyword tool, and it is how I write answer blocks meant to be cited.

Does this change your headings?

Yes, in one specific way. Headings should be phrased as the question a person would actually type, because the heading and the passage underneath it travel together and the question form gives the retriever an obvious match. A heading reading simply "Limits" matches almost nothing.

What headings cannot do is carry the context for the passage below them. It is tempting to treat the heading as the place where the subject is established, then write the body in shorthand. If heading and body are separated, or if the body alone is pulled, you have lost the subject again.

So I write headings as questions and then repeat the subject in the first sentence underneath. That feels redundant when you read the page top to bottom. It is the difference between a passage that stands alone and one that depends on its neighbor.

What about words like this, it and they?

They are the most common way a good passage becomes useless. A paragraph beginning "This means you should" is meaningless on its own, because the thing it refers to lives in the previous paragraph. Pronouns pointing across a paragraph boundary are the cheapest and most frequent self-inflicted retrieval wound.

The fix costs almost nothing. Replace the pronoun with the noun at the start of any paragraph. Not "this means", but "a stale sitemap date means". It reads slightly heavier and it makes each paragraph a complete statement, which is the trade I would make every time.

I would apply the same discipline to sentences beginning "as mentioned above" or "as noted earlier". Those phrases assume a reading order that fragment-based retrieval does not preserve, and they tell a reader who arrived mid-page that they have missed something, which is a reason to leave.

Where does this idea go too far?

When the page stops being readable by humans. I have seen pages written so defensively for retrieval that every paragraph restates the full premise, and they are exhausting. A page nobody finishes does not earn links, is not shared, and gives you nothing beyond the one fragment that matched.

The other excess is fragmenting everything into tiny disconnected blocks. Retrieval likes self-contained passages, and that does not mean it likes thin ones. A passage still has to contain an actual idea with enough substance to answer something, and chopping a developed argument into fragments leaves you with pieces that match a query and then disappoint.

My rule is that each section should be self-contained and each page should still read as one argument. Those are compatible with a bit of deliberate repetition, and incompatible with either extreme. Writing for machines at the cost of people is a bad trade, as I argued when weighing how AI search picks between two similar pages.

How do you test whether your sections stand alone?

Copy one paragraph into a blank document and read it cold. If you cannot tell what product, what condition and what claim it concerns, the passage fails. Do that for the three paragraphs on the page you most want cited, because those are the ones worth the effort.

A harder version is to paste the paragraph alone into an assistant and ask it what the text is about. If the answer is vague or wrong, the retriever would have had the same problem. This is not a measurement of any specific engine's behavior, and it is a reasonable proxy for whether the passage carries its own meaning.

Then look for the structural tells, which are quick to scan for. Paragraphs starting with a pronoun. Numbers without their unit or source nearby. Sections whose subject appears only in the heading. Those three checks catch most of it, and they overlap with how embeddings decide which of your pages gets cited.

What should you do next?

Take the page you most want cited and read each section in isolation. Fix the ones that cannot stand alone, mostly by replacing pronouns with nouns and moving conditions next to the claims they qualify. That edit takes an hour and changes the page more than rewriting it would.

Then change your drafting habit rather than only your published pages. Writing self-contained sections from the start is easier than retrofitting them, and after 350 articles I can say the habit becomes invisible quickly. You stop noticing the repetition and start noticing when a paragraph depends on its neighbor.

If you want someone to look at whether your best pages actually survive being read in pieces, reach out. It is usually a handful of edits rather than a rewrite, and it is the cheapest AEO work available.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.