How do embeddings decide what your page is actually about?
By turning your writing into numbers and comparing those numbers to the numbers made from a question. Nothing in that process looks for your keyword. It measures whether the meaning of your passage sits close to the meaning of the query, which is a different test from the one most people are still optimising for.
This is the part of modern retrieval that founders find hardest to accept, because it removes the comforting mechanical relationship between putting a phrase on a page and being found for it. The phrase still matters. It just stopped being the thing that decides.
I have written over 350 articles about answer engines, schema, and how evidence of expertise gets recognised, and understanding this one mechanism changes more about how you write than any tactic I could give you. So it is worth twenty minutes to understand properly rather than approximately.
What is an embedding, exactly?
A list of numbers that stands in for a piece of text. OpenAI's documentation defines it plainly: "an embedding is a vector (list) of floating point numbers." The text goes in, the vector comes out, and that vector is what gets compared against other vectors rather than the words themselves.
The comparison is geometric. According to the same documentation, "the distance between two vectors measures their relatedness," and "small distances suggest high relatedness and large distances suggest low relatedness." So relevance becomes a question of proximity in a space you cannot see, which is why the whole thing feels unintuitive.
There are several ways to measure that distance and the docs note a preference, stating "we recommend cosine similarity." The specific measure matters less to you than the consequence: two passages that say the same thing in different words end up near each other, and two passages that share vocabulary while meaning different things do not.
Why does this mean keyword presence is not enough?
Because a page can contain your target phrase and still sit far away from the query in meaning. If your page mentions a term once in a list and then spends two thousand words on something adjacent, the passage that gets compared is mostly about the adjacent thing. The keyword is present and the meaning is elsewhere.
This is the mechanism behind a frustration I hear constantly, which is that a page targets a phrase, uses it correctly, and never surfaces for it. The page was written to contain a term rather than to answer a question, and containment is not what is being measured.
It is also why keyword stuffing became useless rather than merely risky. Repeating a phrase does not move a passage closer to a question in meaning space. It only makes the writing worse, which reduces the chance that a human finds the page useful once they do arrive.
Why can a page rank for a phrase it never uses?
Because meaning is what is being compared, and a passage can be close to a question without repeating its words. If somebody asks how to stop their website jumping around while it loads, a passage about layout shift can be a near-perfect match without containing the word jumping anywhere.
This cuts both ways and the second direction is the useful one. You do not have to guess and reproduce every phrasing a reader might use. You have to write something whose meaning is unmistakable, and the matching will handle variations you never thought of. That is a much more pleasant way to write.
The practical implication is that your job shifts from coverage of phrasings to clarity of subject. A passage that states one thing precisely will match many questions about that thing. A passage hedged across several possible subjects matches all of them weakly, which in a ranked comparison means it matches none of them.
What does this mean for how you write a page?
Write passages that would survive being read alone. Since comparison happens at the level of a chunk of text rather than a whole document, every section needs to make sense without the paragraphs around it, including naming its subject rather than relying on a pronoun that refers back two screens.
Concretely, that means starting sections by saying what they are about rather than easing in. It means using the full name of the thing instead of it, and saying the product or concept again where a human editor might call it repetitive. Mild redundancy is a cost worth paying when passages travel separately from their context.
It also means putting the answer near the top of each section. A passage whose first sentences are throat-clearing is a passage whose meaning is diluted by filler, which literally moves it away from the question. I have written more about this in how pages get chunked and why one paragraph gets quoted.
Why does a page about five things lose to a page about one?
Because averaging a mixture produces something that is not close to any of its parts. A passage covering five related topics has a meaning that sits somewhere between all five, and that midpoint is further from each specific question than a focused passage would be.
This is the strongest practical argument for narrow pages, and it is a better argument than the usual one about competition. It is not only that a focused page faces fewer rivals. It is that a focused page is mechanically a closer match to the question it targets, because nothing in it is pulling the meaning sideways.
The same logic explains why long roundup posts underperform their word count. Ten shallow sections about ten topics create ten weak matches. One of those sections, expanded into its own page with real substance, would beat the entire roundup for its specific question.
Does this replace keywords entirely?
No, and treating it as a replacement is a mistake in the other direction. Words are still how meaning gets expressed, and using the vocabulary your reader uses is how your passage ends up near their question rather than near a differently worded version of the same idea.
What changes is the role. Keywords stop being targets to hit and become evidence of what a passage is about. Naming the specific products, standards, and terms involved makes the subject unambiguous, which is useful for exactly the same reason that vague writing is harmful.
It is also worth remembering that semantic matching is one layer among several. Google states that its generative features use retrieval to improve "quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages." Meaning gets you considered. Everything else about the page still decides whether you win.
How do you check whether a page reads as being about the right thing?
Read one section in isolation and write down what it is about in a sentence. If that sentence does not match the question you intended the section to answer, you have found the problem without any tools. This crude test catches most of it.
The second check is to have somebody who does not know the page read a single paragraph and tell you what it is about. Their answer is a reasonable approximation of how a passage will be interpreted, and the gap between what you meant and what they said is exactly the gap that hurts you in retrieval.
The third check is to look at what you actually get surfaced for. If you are being found for topics adjacent to your intent, your passages are drifting towards a neighbouring meaning. That is a diagnosable writing problem rather than a mystery, and it is often behind the complaint that answer engines cite competitors instead of you.
What should you do next?
Take one page that should be performing and is not, and read each section on its own. Ask what it is about, out loud, without looking at the heading. The sections where you hesitate are the sections doing nothing for you, and there are usually more of them than you expect.
Then fix the cheapest failure first, which is almost always pronouns and vagueness. Replace it and this with the actual name of the thing. Move the answer to the front of each section. Delete the hedging sentence that opens three of them. None of that is glamorous and all of it moves passages closer to the questions they should be matching. It also happens to be better writing for humans, which I have argued before in writing for answer engines versus writing for people.
If you have pages that read well to you and are not getting found, this is usually where the answer is, and it is quick to spot from the outside. Reach out if you want another pair of eyes on one.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.