To run an AI visibility audit, you ask ChatGPT, Gemini, Perplexity and Claude the same set of buyer questions. Then you score each answer: does the engine recommend you, where do you appear, how positive is it, and are your facts right? AI engines answer differently almost every time, so you repeat the audit on a schedule and track the average score.
Step 1: Write a Fixed Set of Buyer Prompts
Start with the questions buyers ask when they're choosing in your category, plus a few that name you directly. Our guide to prompt design covers the four prompt types to use (broad, specific, comparison and brand-direct) and works through a 13-prompt example. Tag each prompt with its type, so you can see later if you only appear when buyers name you.
Once you've run the set, try not to edit it. You'll compare every later run against this first one, and a reworded prompt breaks that comparison. If you need new prompts, add them as a separate group and track them on their own.
Step 2: Run Every Prompt on Every Engine
Run each prompt in ChatGPT, Gemini, Perplexity and Claude. Use a fresh chat with memory turned off, so your past chats don't shape the answer. A ChatGPT visibility audit alone won't show what the other engines say, because each one picks its sources differently. If your buyers search on Google, add Google's AI Mode as its own engine, even if you already check AI Overviews. According to SISTRIX, the two cited different domains for the same prompt 83% of the time.
Log one row per prompt, engine and date. Paste the full answer into the row and list the sources the engine cited, marking whether your own site is one of them. Keep the full text so you can re-score old runs later or hand the sheet to an AI for scoring.
Step 3: Score Each Answer
Give each answer four scores you can fill in quickly and average later. We score each response out of 100: presence 40, position 30, sentiment 20 and context quality 10. Our explainer on AI search visibility covers that model. By hand, these four columns cover the same ground:
| Column | What to enter |
|---|---|
| Recommended | Yes if the engine recommends you, Mentioned if it names you without recommending you, No if you're absent |
| Position | Your place among the brands named: 1 if you're first, 2 if second. Leave blank if absent |
| Sentiment | 1 (negative) to 5 (strongly positive). Leave blank if absent |
| Accuracy | Yes or No. If no, note the wrong fact, such as an old price or a product you don't sell |
Use an AI to Score the Sheet
Once you have more than a few dozen rows, paste them into ChatGPT or Claude and let it do the scoring. Use a prompt like this:
Score each AI answer below for the brand [Brand].
Facts about [Brand]: [prices, product range, locations].
For each answer, return one row with these columns:
Recommended: yes, mentioned or no.
Position: [Brand]'s place among the brands named (1 = first), or blank.
Sentiment: 1 (negative) to 5 (strongly positive), or blank.
Accuracy: yes or no, listing any statement that contradicts the facts above.
Write "unclear" instead of guessing.Give it your real facts, because it can't judge accuracy without them. Use the same scoring prompt and the same engine every run, since the scorer's answers vary too. Then check a few rows by hand each time to confirm it scores the way you would.
Step 4: Repeat the AI Visibility Audit to Build a Baseline
Run the full audit several times before you trust a score, because the same prompt rarely gets the same answer twice. According to SparkToro, ChatGPT and Google's AI had under a 1 in 100 chance of naming the same brand list twice in 100 runs. The study ran 12 prompts 2,961 times across ChatGPT, Claude and Google's AI.
In the same study, the brands came in the same order only about once in 1,000 runs. Read your position score only as an average across runs.
SparkToro's advice is to ask a prompt at least 60 to 100 times and average the results. By hand, you won't get close. Run each prompt more than once per session where you can, and repeat the full audit on a fixed schedule, such as monthly. You need the numbers from Step 3 for this: you can average two sentiment scores, but not two written impressions. Your first few runs, averaged per engine and per prompt type, become your baseline.
Step 5: Find Out Why You Score Low
Sort your baseline by score and read the lowest-scoring answers in full. You'll usually see one of three patterns, and each points to a different first check.
- Absent on almost every prompt. AI crawlers may not be able to reach or index your pages. When we started on Three Squared Nine, Google hadn't indexed around 80% of the site's pages, and Bing had indexed only 3. Check your index status in Search Console and Bing Webmaster Tools, and check your raw HTML, before you assume a content gap.
- Mentioned but not recommended. Read the answers where a competitor got the recommendation, and note the reasons the engine gives for choosing them.
- Wrong facts or low sentiment. Read the sources the engine cited for those answers before you change your own pages or messaging. If it gets your pricing, product range or locations wrong, your brand data may be thin or contradicted elsewhere online.
Step 6: Track What Moves Your Score
Your baseline lets you check whether the work you or your agency do changes what AI says about you. Keep a dated log of every change, such as new pages, schema fixes, review campaigns or press coverage. Then compare each run's average against the baseline.
If the average moves after a change and holds over the next few runs, the change likely helped. A jump in a single run is more likely noise. Read the averages per engine and per prompt type, as well as overall. Otherwise you can miss a gain on one engine when another dips at the same time.
When a DIY AI Visibility Audit Isn't Enough
A spreadsheet audit works for a first baseline and a monthly trend. You'll find it hard to keep up once you need hundreds of prompts, competitors scored alongside you, and enough repeats to average out the noise. Our Brand Visibility Audit runs 200 or more structured prompts. On the Three Squared Nine project, we tracked core prompts daily or every 6 hours and averaged the results. Our guide to what a professional GEO audit covers walks through the five scored areas we use.
If you'd rather see where your brand stands before you build a sheet, get a free AI visibility assessment.
FAQs
How many prompts do I need for an AI visibility audit?
Enough to cover all four prompt types for each product line you want buyers to find you for. With only a handful of prompts, you learn about that handful and little about your category, so treat a set under 10 as directional.
What's a good AI visibility score?
There's no universal passing number, so compare yourself with competitors instead. Score the competitors named in the answers you've already collected, using the same four columns, so the benchmark costs no extra prompts.
Can I see AI visibility in Google Analytics instead?
Not directly. Analytics records visits, and a buyer who gets the answer from AI may never click through. According to Ahrefs' May 2026 study of 300,000 keywords, the number-one Google result loses 58% of its clicks once an AI Overview appears. Your analytics can show that drop, but not whether the AI named you in the answer.
The Bottom Line
Your first audit shows where you stand. Run the same prompts and score them the same way each time, and every later audit shows which of your changes moved you. Put your next hour or budget behind those.
