BEAM

Seedlight BEAM: one place to run your whole eCommerce, with AI agents that know your business →

← AI visibility for ecommerce

Chapter 6 of 6

Measurement and maintenance

How to turn a one-off audit into a repeatable process: what to measure, how often, how to record results so they stay comparable, and in what order to fix what the measurement reveals. Plus an honest closing of the whole six-chapter path.

9 min read

Key points

  • Measurement answers several questions, not one: is the brand named, in what context, are the facts correct, does the model cite your sources, and where does it send the purchase.
  • Cadence: monthly while you are actively working on data and content, quarterly once catalogue and brand are stable. Anything less frequent breaks the link between a change in results and its cause.
  • Model answers are non-deterministic, so ask every question several times and record frequency. Draw conclusions from a trend across measurements, not from a single screenshot.
  • Fix in this order: wrong facts, then missing data, then content, then external signals. The order follows from what does the most damage and what is fastest to repair.

The external signals from the previous chapter share one inconvenient property: they work slowly and largely outside your control. The same is true of everything that came before. Product data, content and mentions only start working once models pick them up, and you will only learn that they did if you check.

The audit in chapter 2 was a snapshot of a single day. This chapter turns it into a process: what to measure, how often, how to record results so they are still comparable six months from now, and in what order to fix what the measurement shows.

What to measure: several questions, not one number

The temptation is to reduce AI visibility to one percentage on a slide. I would advise against it, because a single percentage never tells you what to fix. A workable measurement answers several questions, asked separately for each query in your control set.

Control questionWhat you checkWhat you record
Is the brand named at all?Whether the answer contains your brand or product nameYes or no, separately per model
In what context?Whether you are recommended, mentioned in passing, or cited as the weaker optionA short quote of the passage that concerns you
Are the facts correct?Price, availability, category, product attributes, markets served, company statusA list of errors and, where visible, the source the model took them from
Does the model cite your sources?Whether your pages appear among the links given, or only third-party sitesA list of cited domains
Where does the purchase go?Whether the product is attributed to your store, a marketplace or a competitorThe purchase destination named in the answer

Five columns in a spreadsheet are enough to start. What matters more than the tool is that the definitions stay identical at every subsequent measurement.

The naming is secondary, but worth knowing because it shows up in tool pitches. The industry uses terms such as Share of Model for a brand’s share of answers across a fixed set of queries. In our own service we work with three:

  • Brand Visibility: is the brand named at all.
  • Product Mention Rate: does the specific product appear when the question fits it.
  • Store Attribution: is the purchase credited to your store.

I deliberately give no reference values here, because meaningful benchmarks do not exist: results depend on category, language, the composition of the query set and the model. The only honest comparison is against your own previous measurement.

What has to stay fixed

The query set you built during the audit now becomes an asset, and its greatest value is that it does not change. Four things must stay constant at every measurement:

  • The wording of the questions: word for word, including phrasing you would happily improve.
  • The language and market: the same question in two languages is two separate measurements.
  • The list of models: the same one every time, with results recorded per model.
  • The scoring method: the same definitions, meaning the columns in the table above.

Add new questions as a separate group and compare them only from their own first measurement, rather than mixing them into the old pool and breaking comparability across the board.

The conditions you ask under

The conditions are part of the method too, though they are rarely written down. Ask in a session without chat history and without personalisation, ideally logged out. Otherwise you will measure your own earlier footprints instead of what a stranger sees.

Also note whether the model used web search for a given answer. An answer with search and one without are effectively two different measurements and should not be averaged together.

How often to measure

Match the cadence to how much is changing on your side:

  • Monthly: while you are actively working on data, content and external signals, because you want to see the effect of your own changes.
  • Quarterly: once catalogue and brand are stable and the measurement mainly guards against things quietly breaking.
  • Less often than quarterly: it stops being useful, because you can no longer connect a change in results to any specific cause.

More often than monthly is usually pointless, for a mundane reason: a change on your site has to be fetched, processed and settle into the sources a model draws on, and that takes time.

Outside the regular cadence, it is worth measuring in four situations:

  • After a platform migration: URLs, templates and the page source all change.
  • After a name or domain change: both versions live on the web side by side for a while.
  • After a major catalogue rebuild: different categories, different attributes, a different feed.
  • After a widely reported change at the model provider: a new model version or a new way of picking sources.

How to record results

The format has one purpose: to let you reconstruct, six months later, exactly what you saw. The minimum is seven fields in a single spreadsheet row:

  • Measurement date: the day you asked.
  • Model and version: the model name, plus the version where it is visible.
  • Market and language: separately, because everything else depends on them.
  • The exact question: pasted word for word, not summarised.
  • The full answer: copied in its entirety, not just the part about you.
  • The scoring: against the control questions in the table above.
  • The list of sources: the domains and links the model provided.

Two mistakes come up most often:

  • The score without the raw answer: a quarter later nobody can reconstruct why the context was judged neutral at the time, and arguing about it costs more than the measurement itself.
  • Screenshots instead of text: results you cannot search or count.

A plain spreadsheet with one answer per row holds up far longer than people expect, and it ports into any tool you adopt later.

Why a single measurement means nothing

Model answers are non-deterministic: the same question asked twice in a row can produce a different answer, a different set of recommended brands and a different set of sources. Four things feed into that:

  • How text is generated: randomness is built into the mechanism.
  • The model variant: the same query gets routed to different versions.
  • Web search: used in one run, skipped in the next.
  • Source updates in the background: the material the model draws on changes without your involvement.

According to industry sources, comparisons of answers to identical questions show the divergence extends beyond phrasing to who gets named and what gets cited.

The methodological conclusion is simple: ask every question several times and record frequency, not the bare fact that your brand appeared once. A single appearance is not a success and a single absence is not a failure. The signal only emerges from repeatability within one measurement and from the direction of change across the next ones.

Rule of thumb: one answer is an anecdote, not a measurement. Base conclusions on three consecutive measurements with the same query set, not on a screenshot somebody forwarded on a Friday afternoon.

What to do with the results: order of repairs

A measurement usually produces a long list of things to do, and this is exactly where teams stall. The order we use follows a simple calculation: first what does the most damage and is fastest to repair.

  • 1. Wrong facts. The model quotes an outdated price, claims you do not ship to a given country, confuses you with another company or describes a pre-rebrand offer. Fix this immediately, because that kind of error actively costs sales, and its source can usually be identified and corrected either on your side or on one specific external site.
  • 2. Missing data. The model has nothing to work with: attributes are absent, the feed is incomplete, structured data covers only part of the catalogue. This is the chapter 3 work, expensive once but long-lasting and applied across the whole catalogue at once.
  • 3. Content. The data exists, but nobody answered the question a customer actually asks. This is chapter 4, meaning the most editorial effort and an effect spread over time, though also the most durable.
  • 4. External signals. Reviews, rankings and mentions from chapter 5. The slowest layer, so start it in parallel with the rest, but do not expect it to move the next monthly measurement.

On your own or with a partner

Doing this yourself makes sense more often than tool vendors suggest. One market, one language, a few dozen control questions and a person on the team who will genuinely keep the cadence: that is a spreadsheet and roughly an hour a month. The biggest risk is not the absence of a tool but the fact that after the third month nobody runs it any more, at which point you lose your reference point entirely.

When a partner or a tool starts paying off

Usually in four situations:

  • Several markets and languages: the number of measurements multiplies.
  • A large catalogue: the query set has to cover many categories at once.
  • Several models to check: each one has to be queried and scored separately.
  • Measurement meant to produce repairs: in the data and on the site, rather than a chart for a meeting.

That is how we run it at Seedlight as our AI Visibility (GEO) service: a fixed query set, per-model measurement, Brand Visibility, Product Mention Rate and Store Attribution, plus the fixes in the data and content layer. The caveat belongs here, stated plainly: nobody guarantees a position, a mention or a citation, because those depend on variables no provider controls.

The whole path in one paragraph

We started with where models get their knowledge about your store and why your own site is only one source among many (chapter 1). Then came the audit: what AI says about you today and what your baseline is (chapter 2). On that foundation you put product data in order, feed and structured data as a single source of truth (chapter 3). Next you wrote content an answer can be lifted from, instead of content that merely reads well (chapter 4). Chapter 5 took the work beyond your own domain, into reviews, rankings and mentions. This chapter ties it into a cadence: measure, record, repair in a set order, measure again.

One thing has to be said plainly at the end, because without it the whole guide would be dishonest. You do not control what a model answers. Providers change models and sources without notice, the way citations are selected stays opaque, and nobody, ourselves included, can guarantee you a mention or a citation.

What you do control is five things:

  • The accuracy of facts about you everywhere they appear.
  • The completeness and accessibility of your product data.
  • The quality of your answers to real customer questions.
  • The consistency of your brand across the web.
  • The regularity of measurement and whether somebody in the company actually keeps it going.

That list is long enough to fill a year of work and concrete enough to start tomorrow. The rest, including the fact that some customers will decide without ever visiting your site, is a consequence we cover separately in our piece on zero-click commerce.

Questions

How often should I measure visibility in AI answers?

Monthly if you are working on product data, content and external signals in parallel, because you want to see the effect of those changes. Quarterly if catalogue and brand are stable. Less often than quarterly loses diagnostic value, because a change in results can no longer be tied to a specific cause.

Why do I get a different answer every time I ask the same question?

Because model answers are non-deterministic. The same prompt may be routed to a different model variant, may or may not trigger a web search, and text generation involves randomness by design. That is why a single result is not a measurement: ask each question several times, record how often your brand appears, and read the trend across several consecutive measurements.

Do I need a paid tool to measure AI visibility?

Not to begin with. For one market and a few dozen control questions, a spreadsheet and an hour a month are entirely sufficient, provided the definitions and the query set stay unchanged. A tool or a partner starts paying off with several languages and markets, a large catalogue, and when the measurement is meant to drive concrete repairs rather than produce a report.

All chapters in this guide

AI visibility for ecommerce

  1. 01Where AI learns about your storeAn AI assistant does not look at your store, it reads the facts it can extract from it without ambiguity. This chapter explains how AI decides what to recommend and the four groups of sources behind that picture: product data, machine-readable content, external signals, and what the model remembers versus what it fetches live.
  2. 02The audit: what AI says about you todayA repeatable afternoon procedure: which questions to ask assistants, how many assistants to use, how many times to repeat, and what to record in a spreadsheet so you end up with a baseline before you change anything. The headline number of a first AI visibility audit is not how often you are mentioned, it is how many facts about you are wrong.
  3. 03Product data: feed and schema as the foundationVisibility in AI answers starts with product data, not with the blog. How to set identifiers, a complete Offer and shopping feed hygiene, so an assistant has a price, an availability status and a product identity to work with.
  4. 04Content that lands in answersProduct data gives a model price and availability, but the sentences about who a product is for and how it differs from the alternative have to come from your content. How to write so those sentences can be lifted off the page and used as they are.
  5. 05External signals: reviews, rankings and mentionsA model does not build its recommendation from your site alone. Reviews, industry rankings, media mentions and your marketplace listings act as verification of what you say about yourself. What you can build there, how fast, and what money cannot buy.
  6. 06 · You are hereMeasurement and maintenance