The AI Visibility Audit: Where Your Brand Stands Across ChatGPT, AI Overviews, and Perplexity
Most published AI visibility audits produce a single number between zero and one hundred. That number is the least useful output the exercise can generate, because it merges two separate questions with two separate remediation paths. What AI systems currently say about your brand is one audit. Whether your site gives them anything usable to...
Last updated: 25 Aug 2026
CONTENTS
Most published AI visibility audits produce a single number between zero and one hundred. That number is the least useful output the exercise can generate, because it merges two separate questions with two separate remediation paths. What AI systems currently say about your brand is one audit. Whether your site gives them anything usable to say is a different one. A brand can score badly on the first while the second is in decent shape, and the work that follows is not the same in the two cases.
This is the version we run as a baseline: a fixed prompt set, three runs per platform, three scored axes, and four end states that each route somewhere different. It takes a day of focused work and no paid tooling. What it produces is a dated benchmark you can defend in a board meeting and beat in ninety days.
What this audit answers, and what it does not
The answer-side audit measures what generated responses currently say: whether your brand is named for the questions your buyers ask, whether the description is accurate, and which sources the answer drew from. It tells you where you stand.
The site-side audit measures whether AI systems can reach and read your pages: crawler access, rendering, product data completeness, structured data. It tells you why. Flatline publishes a free AI readiness analysis that covers the site-side view and returns a scored report with a ranked fix list, which pairs well with what follows here.
Running only the site-side audit gives you a list of technical work with no evidence about which item is binding. Running only the answer-side audit gives you a verdict with no explanation. The sequence that works is answer-side first, because it tells you which of the technical findings actually matters, then site-side to explain the gap.
One more boundary worth setting before you start. This audit produces a benchmark, not a forecast. Generated answers vary run to run, so the honest output is a range with a date attached rather than a precise score.

Build the instrument
Twenty to forty prompts, written by you, in the language your buyers use. Generated prompt sets are faster and worth less, because the instrument is the part that has to stay stable across quarters.
Draw them from four places. Your search query data, for the questions people already ask in words. Your sales and support conversations, for the constraints buyers actually state. Your category’s comparison language, since comparison questions are where recommendations get made. And your own brand name, to test what systems say about you when asked directly.
Cover four prompt types in roughly these proportions:
- Category recommendation (about half): “best [product type] for [use case]”, with the constraints your buyers really apply, including price bands and specific conditions.
- Comparison (about a quarter): your brand against a named competitor, and alternatives to a competitor.
- Attribute and suitability (about a fifth): “is [material or feature] good for [situation]”, the questions that decide a purchase rather than start one.
- Direct brand (the remainder): what is your company, what does it sell, who is it for.
Freeze the set and version it. Every future comparison depends on this file not changing.
Run protocol
Four rules, and the third is the one most audits skip.
- Pick three platforms and record them. Google AI Overviews and AI Mode, ChatGPT, and one of Perplexity, Gemini, or Copilot depending on where your buyers actually are. More platforms is more work without more clarity at baseline stage.
- Use a clean session. Logged out or in a fresh profile, with location set to your primary market, so you are measuring the system rather than your own history.
- Run each prompt three times per platform. This is the difference between a measurement and a single draw from a variable process. Record all three outcomes, not the best one.
- Date everything and capture the sources. For each response, save which brands were named, in what order, and which URLs the answer drew on. The source column is the most actionable data the audit produces.
The scoring sheet
Score each prompt on three axes. Keep them separate; do not add them into a composite.
| Axis | 0 | 1 | 2 | 3 |
| Presence | Not named in any run | Named in one of three runs | Named in two of three runs | Named in all three, and in the leading group |
| Accuracy | Absent, so not scorable | Described with a material error: wrong category, wrong price band, wrong market, discontinued product | Broadly right with gaps or outdated detail | Correct and current |
| Source | No source connected to you was cited | A third-party source about you was cited | One of your own pages was cited | (not used) |
Then produce four figures, each as an average across your prompt set, with the range across runs noted beside it.
- Presence average, which answers whether you are in the conversation.
- Accuracy average, scored only on prompts where you appeared, which answers whether being in the conversation is helping.
- Source profile, which shows whether recommendations rest on your own pages, on third parties, or on nothing connected to you.
- Competitive set, meaning the three to five brands that appear most often instead of you, and the sources those answers cited.
Four numbers instead of one. The composite score that most audits report hides exactly the distinction that determines what you should do next.

Four states, and how to recognise yours
| State | Signature | What it means |
| A. Absent | Presence average below 0.5. Your brand appears in almost no runs | Systems either cannot reach or cannot use what you publish, or nothing outside your site describes you in category terms |
| B. Unstable | Presence between 0.5 and 1.5, high variance across runs | You are a marginal candidate. Something about you is legible, and it is not consistent or strong enough to hold a place |
| C. Present but misdescribed | Presence above 1.5, accuracy below 2 | You are being recommended with wrong or stale information, which converts worse than absence in some categories |
| D. Established | Presence above 2, accuracy at or above 2 | The work shifts from earning presence to defending it and widening share against the competitive set |
State C deserves a note, because it surprises teams. A confident, wrong description travelling through generated answers is more expensive than not being named at all: a shopper acts on it, arrives with the wrong expectation, and the correction cost lands on your support team. It is also usually the fastest of the four to improve, because the underlying cause is often outdated public information rather than a structural gap.
Tie-breakers
Three cases sit between states, and each resolves the same way.
Strong on one platform, absent on another.
Score them separately rather than averaging. Divergence across platforms is a finding, not noise: it usually means one system is reading your site while another is relying on third-party sources you are missing from.
Present, but only third-party sources are cited.
Treat as state B regardless of presence average. Visibility resting entirely on other people’s pages is real and borrowed, and it moves when their content moves.
Accurate everywhere except price and availability.
This is not a content problem. Price and availability drift indicates that your structured data, your feed, or your page are disagreeing with each other, and it is a data reconciliation task rather than a visibility one.
What each state calls for
State A.
Start at access and legibility, in that order. Confirm that AI crawlers receive complete responses on product and category templates, since a partial response produces this exact result quietly. Then check whether your pages state your product facts in plain text rather than in tabs, images, and variant selectors. Content production is the last step here, not the first.
State B.
Completeness and corroboration. Fill product attributes on revenue-leading SKUs, since inconsistent matching against buyer constraints produces inconsistent presence. In parallel, map which third-party sources the answers cite and begin the slow work of appearing in them.
State C.
Correction first. Identify where the wrong information lives: your own outdated pages, stale third-party profiles, or directory entries nobody has updated since a rebrand. Make the current facts easy to find and consistent everywhere they appear. This is often weeks of work rather than quarters.
State D.
Defence and width. Expand the prompt set into adjacent use cases, monitor for competitors displacing you on specific constraint types, and keep the source profile healthy so your position does not rest on a single publisher’s roundup.
Re-run the full audit quarterly, against the frozen prompt set, and compare ranges rather than points. A presence average moving from 0.9 to 1.4 with an overlapping range is encouraging and not yet conclusive; the same move with separated ranges is a result.
What we deliberately left out
No benchmark percentages appear in this method, and that is a choice. The figures circulating for what constitutes a healthy mention rate come from vendor research with unpublished prompt sets, and a benchmark you cannot reproduce is not a benchmark. Your first dated run is your baseline. Beat that.
No composite score either, for the reason given at the top: the number that makes a slide look decisive is the number that removes the information you need to act.
If you would rather not build the instrument yourself, this is the baseline Flatline runs at the start of a GEO engagement. We are a Shopify Platinum Partner with AI consultancy and technical SEO practices, and the audit output is a dated benchmark, the four scored figures, your competitive set with cited sources, and a prioritised remediation plan tied to whichever state you land in. Request the audit and we will run the baseline and walk you through what it means for your catalog.
Frequently Asked Questions
How long does an AI visibility audit take?
A day of focused work for a twenty to forty prompt set across three platforms with three runs each, plus a few hours to score and write up. Building the prompt set well is the part worth slowing down for, since every future comparison depends on it staying stable.
Do I need a paid AI visibility tool to run this?
No. The method above uses the platforms directly and a spreadsheet. Paid tools become useful for cadence and scale once you have a baseline and want continuous monitoring rather than quarterly snapshots. They are a convenience layer over this method rather than a substitute for it.
Why score presence and accuracy separately?
Because they call for different work. Absence usually points to access, legibility, or missing corroboration. An inaccurate description points to outdated information in circulation. A combined score would put a brand with strong presence and wrong details in the same band as a brand nobody names, and send both the same remediation plan.
How often should the audit be repeated?
Quarterly, against the frozen prompt set, with the date recorded each time. Anything more frequent measures variance rather than progress, since generated answers differ across a single week. Re-run sooner if you change platforms, rebrand, or make a significant change to your catalog structure.
Key Takeaways
- Run two audits, not one. The answer-side audit tells you where you stand; the site-side audit tells you why. Answer-side comes first, because it identifies which technical findings are binding.
- Report four figures rather than a composite: presence, accuracy, source profile, and competitive set. The single score hides the distinction that determines what to do next.
- Three runs per prompt, dated, with the range recorded. A number from a single run of a variable system is not a measurement.
- Four states call for four different responses. Absence points at access and legibility, instability at completeness and corroboration, misdescription at correction, and an established position at defence and width.
The audit is worth running whether or not anyone acts on it immediately, because the dated baseline is the thing you cannot reconstruct later. Six months from now, the question will be whether the position improved, and that question is only answerable if someone wrote down where it started.
POPULAIR ARTICLES
GET IN TOUCH
To speak with us, call (+31) 613 326 179, send us an email, or reach out to us by chat or What’s App.