AI

How Do You Build a Baseline for AI Visibility Before You Change Anything?

Written by
Pravin Kumar
Published on
Sep 28, 2026

How do you build a baseline for AI visibility before you change anything?

Write down a fixed set of questions your buyers would ask an AI assistant, run them today, and record three things for each: whether you were named, whether you were cited with a link, and whether the description of you was accurate. That record is your baseline. Nothing else counts until it exists.

Almost every conversation I have about AI visibility starts in the wrong place. Someone wants to know what to change on their site so ChatGPT or Perplexity will mention them. It is a fair question with no useful answer, because without a baseline you will not be able to tell whether anything you do afterwards worked.

I have written several hundred articles about how answer engines read and cite pages, and the single habit that separates the teams who make progress from the teams who churn is this one. They measure before they move.

Why does a baseline matter more here than it did in traditional search?

Because the output is not a ranked list. In classic search you could look up a position for a keyword and compare it next month. An AI answer is generated text that varies between sessions, models, and phrasings, so there is no scoreboard to check. You have to build the scoreboard yourself.

The second reason is that the failure modes are different. A page can rank well and still be described wrongly by an assistant, because the assistant is summarising rather than linking. Being absent and being misrepresented are both problems, and only one of them shows up in a ranking tool. A baseline that records accuracy catches the second one.

The third reason is timing. Changes to how a model describes you can lag your site edits by weeks, because retrieval and training operate on different clocks. Without a dated record of what the answer looked like before, you will attribute a change to whichever thing you did most recently, which is usually the wrong one.

What should you actually measure?

Measure four things per question. Whether your brand appears in the answer at all. Whether the answer links to a page you own. Whether the description of what you do is correct. And which competitors appear alongside you. Record them as plain values so you can count them later.

Presence and citation are different measurements and should never be collapsed into one. An assistant can describe your approach accurately without linking to you, which tells you your content is influencing the model but not earning attribution. The fix for one is not the fix for the other, so a single combined score hides the thing you need to act on.

Accuracy is the measurement most teams skip and the one that most often produces an urgent finding. I have run baselines where the assistant confidently described a client's pricing model as something they abandoned two years ago. No amount of new content fixes that until the stale source is found, and you only go looking because the baseline flagged it.

How do you build a prompt set that is worth repeating?

Build twenty to thirty questions grouped into three kinds: category questions where you would hope to appear, comparison questions naming you against competitors, and direct questions about your company. Write them the way a buyer would type them, not the way you would.

The discipline is that the set never changes once you start. The temptation after the first run is to reword the questions that returned bad answers, which quietly resets the experiment. Keep the wording frozen and add new questions to the end if you need them, so that the original twenty stay comparable across runs.

Include at least five questions you expect to lose. A prompt set where you always appear is a flattering document rather than a measurement instrument. The questions that consistently fail are where the actual work is, and they only stay visible if you refuse to edit them out.

What does Search Console give you and what does it not?

Search Console gives you Google's side of the picture, and Google keeps expanding what is reported. In September 2026 Google Search Central announced web multimodal Search performance reporting in Search Console. What it does not give you is anything about ChatGPT, Claude, or Perplexity, which are separate systems entirely.

That gap is the reason a manual prompt set still matters. No dashboard covers the assistants your buyers are actually using, and the vendors have little incentive to build one for you. Twenty questions run by hand once a month is unglamorous and it is currently more complete than any tool I have found.

The other half of the picture is referral traffic. Assistants that link out send visitors, and those visitors arrive with identifiable referrers in your analytics. Segmenting them is worth doing early, because it is the one signal that connects visibility to something a business cares about, a point I explored in pages AI cites but humans never visit.

How often should you re-run the baseline?

Monthly is right for most companies. Weekly produces noise you will over interpret, because answers vary between sessions even with no change at all. Quarterly is too slow to connect a result to a cause. Monthly gives you enough separation to see a trend and enough resolution to remember what you did.

Run it on a fixed date rather than when you feel like it. The value of the baseline comes from the comparison, and comparisons break when the interval wanders. I put it on the calendar the same way I put invoicing on the calendar, because both are boring tasks whose value is entirely in the consistency.

Run it from a clean session each time, without your account history attached. Personalisation will show you a friendlier picture than a stranger gets, and a friendlier picture is exactly what you do not want from a measurement. If the assistant already knows you, you are measuring your own history rather than your visibility.

What does a bad baseline look like?

A bad baseline has vague questions, a moving prompt set, subjective scoring, and no date. It feels like work and produces nothing comparable. The test is whether someone else could re-run it next month and get a result you would trust, and most first attempts fail that test.

The most common specific flaw is scoring on a feeling. If your record says the answer was good, you have written down an opinion. If it says you were named, not linked, and the description mentioned a product you no longer sell, you have written down facts, and facts survive a change of who is doing the measuring.

The second flaw is measuring only the questions where your product is the obvious answer. Those return good results and teach you nothing. The category questions, where a buyer has a problem and has not yet decided what kind of thing solves it, are where visibility is actually won and lost, which connects to the volume question I looked at in whether publishing more pages hurts your AI visibility.

What do you do with the baseline once you have it?

Sort the failures by type. Absence means there is no content of yours that answers the question well enough to be retrieved. Presence without citation means your content is being used but your page is not being credited. Inaccuracy means a wrong source is winning, and the job is to find and outrank it.

Each of those three needs a different response, which is why the sorting matters more than the total. Writing more articles is a reasonable response to absence and a poor response to inaccuracy. I have seen teams publish for two quarters against a problem that would have been solved by correcting one outdated page.

Pick one failure type and work on it alone until the next run. Changing several things at once against a monthly measurement gives you no attribution, and attribution is the entire reason you built the baseline. Slower and legible beats faster and unreadable.

What should you do next?

Spend one hour writing twenty questions a buyer would actually ask. Run them today from a clean session and record presence, citation, accuracy, and competitors for each. Put the file somewhere durable and put the next run on your calendar for a month from now.

Resist changing anything on your site this week. The value of a baseline comes from being taken before the intervention, and the urge to start fixing immediately is the most common way teams end up a quarter later with activity and no evidence.

If you want a second opinion on your prompt set or on what your first run is telling you, I am happy to look at it. Reach out with your questions and what came back.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.