You can measure your brand's visibility in AI answers in about three hours, with a spreadsheet and no subscription. Run a fixed set of 20 prompts, log four things per prompt, then repeat monthly on the identical set. That is the whole method, and it produces a number you can defend in a meeting, which is more than several paid dashboards manage.
The reason to do it by hand at least once is unglamorous. Until you have run the prompts yourself, you have no idea what a tool's score is even counting, and you will nod along at a chart that means nothing.
Fair warning about the first number you get. It will be wrong, and the last section explains why that is fine.

What is an AI visibility audit?
An AI visibility audit measures how often an assistant names, recommends or cites your brand when somebody asks a buying question in your category. You pick the prompts a real buyer would type, run them under controlled conditions, and record the outcome as data rather than as a vibe.
The controlled part carries all the weight. Assistants personalise, they read your chat history, and they change their minds between runs. An audit that ignores all three is measuring your own account instead of your market, and it will flatter you, because your account has been reading about your company for a year.
Why bother when tools already exist?
Because a score you cannot reproduce is not evidence, and because this sector has an honesty problem that is worth understanding before you spend anything.
Somebody posted this to r/DigitalMarketing shortly before we started writing: "Our AI visibility reports feel made up. Am I the only one?" They were not. Ivan Palii, who builds SEO software and therefore has no reason to flatter the category, tested seven leading AI visibility tools around the same time and reported that not one of them showed a visibility trend at the level of an individual prompt. If a tool cannot tell you which prompt moved, its headline number is a mood ring.
There is a deeper problem underneath that one. A critical survey of generative engine optimization published on arXiv in July 2026 reviewed 45 studies and concluded that no evaluated technique consistently improves discoverability, downstream traffic or business outcomes across platforms. The one controlled experiment the field has, the Princeton GEO paper accepted to KDD 2024, found optimisation could "boost visibility by up to 40% in generative engine responses", with efficacy varying by domain.
Read those two together and the sensible order is measure first, buy later. Which is roughly the reverse of how most teams arrive. Buy dashboard. Panic at number. Ask what the number means. We have watched this sequence happen more than once and it never gets less funny from the outside.
Takeaway: you are not doing this by hand because manual is virtuous. You are doing it once so that you can tell whether a tool's number is worth £90 a month, which is a decision you cannot make from the outside.
How to run the audit: six steps
Step 1. Write 20 prompts across four intents
Twenty is enough to see a pattern and few enough that you actually finish. Split them four ways:
- Category prompts, where you are not mentioned: "best project management software for agencies"
- Comparison prompts, naming two rivals but not you: "Asana vs Monday for a 12-person team"
- Recommendation prompts with a constraint: "cheapest CRM that does email sequences under $50 a month"
- Objection prompts, the awkward ones: "is [your category] worth it for a solo founder"
Leave your brand name out of every single prompt. Asking "is Aiter any good" tells you how an assistant summarises our own website, which is a different and far less useful question than whether it would ever bring us up unprompted.
One exception, and it is worth the extra minute. Run "what is [your brand]?" once per assistant as a separate check, not as part of your scored set. You are not measuring visibility there, you are hunting for hallucinations in your entity data. Wrong founder, wrong pricing, wrong country, a feature you killed two years ago. We found a stale description of ourselves this way and it had been sitting there quietly misinforming people for months.
The builder below will draft the whole set for you.
Build your 20-prompt audit set
Everything runs in your browser. Nothing is stored or sent.
Category
Your brand is absent. This is the hardest set to win and the most valuable.
best your category for small teamstop your category tools in 2026what is the best your category softwarewhich your category do experts recommendyour category tools actually worth paying for
Comparison
Rivals named, you are not. Tests whether the model reaches for you unprompted.
Competitor A vs Competitor B, which is betterCompetitor A alternativesis Competitor A worth it compared to Competitor BCompetitor B or Competitor A for a team of fiveswitching from Competitor A to something cheaper
Recommendation
Constrained asks. These convert, because the buyer has already decided to buy.
cheapest your category that actually worksyour category for a five-person team under $100 a montheasiest your category to set up in a daybest your category when you have no budgetyour category that does not need a developer
Objection
The awkward ones. Skipped by almost everyone, and the answers are revealing.
is your category worth it for a solo founderdo I really need your categorywhy is your category so expensivecommon problems with your category toolscan I do your category manually instead
Run each one logged out, in a private window, and log whether you were named, recommended, or cited as a source. Three columns, not one.
Step 2. Run each prompt logged out, in a fresh session
Log out. Use a private window. Do not let the assistant see your history, because personalisation will quietly hand you a flattering result and you will believe it.
Run the same 20 prompts through each assistant you care about. Most teams pick three: ChatGPT, Google AI Mode and Perplexity. Add Claude if you sell to developers. Add Copilot if you sell into large companies where IT chose Microsoft for everybody.
Do not add all six. Every extra assistant multiplies the work by twenty prompts and adds noise to your average without changing a single decision you are going to make.
Step 3. Log four separate columns, not one score
Tools collapse these into a single number, and the detail they flatten is precisely the detail that tells you what to fix next week.
| Column | What it means | Why it earns its own column |
|---|---|---|
| Named | Your brand appears anywhere in the answer | Cheapest to win, weakest signal |
| Recommended | You are presented as a suggested option | The one that tracks pipeline |
| Cited | Your domain appears as a linked source | The one that drives actual traffic |
| Who else appeared | Every other brand in the answer | The most useful column on the sheet |
The first three separate three genuinely different diseases. A brand can be cited constantly and recommended never, which is a content problem. Or recommended without ever being cited, which means the model knows you but is not sending anybody your way. One number hides both.
The fourth column is the one almost everybody skips, and it is the one that pays. An SEO on r/SEO put it better than we would have: that last column tells you why competitors get picked, and the answer is usually comparison pages, directories and forum threads rather than anything on the competitor's own website. You are not just measuring yourself. You are getting a free map of which third-party pages the model trusts in your category, and those pages are almost always easier to get onto than they are to outrank.
Step 4. Score share of voice against three named rivals
For each prompt, record which competitors appeared. Then compute share of voice: your mentions divided by all brand mentions across the set.
Pick three rivals and hold that list steady. Swapping the comparison set between months is the easiest way in the world to fake progress, and if you are reporting to somebody else, they will eventually notice.
Step 5. Record the boring metadata
Date, assistant, model version if it is visible, country, and whether you were logged out.
Skip this and next month's comparison is worthless, because you will have no idea whether the number moved or the conditions did. We have skipped it. It was worthless.
Step 6. Repeat monthly on the identical prompt set
Same prompts, same order, same conditions. Change the set and you have started a new experiment rather than continued the old one.
Monthly is the right cadence for almost everybody. Weekly mostly measures noise, and quarterly is so slow that you will have forgotten which content change you were even testing.
What free tools should I use alongside the manual audit?
Four, all free, and between them they cover the half of the picture your prompt set cannot reach.
- Search Console's generative AI report. Impressions from AI Overviews and AI Mode, straight from Google. No clicks and no queries, so read what it shows and hides before you build a report on it.
- Bing Webmaster Tools. Verify the site even if Bing sends you nothing. Its AI Performance dashboard reports grounding queries, the phrases the assistant writes for itself when it goes looking for sources. It is the only free query-level data anybody publishes.
- GA4, segmented by referrer. Traffic arriving from chatgpt.com, perplexity.ai and friends. Small numbers today for most sites, but it is the only column in this list that connects to revenue.
- Your server logs or CDN dashboard. Which AI crawlers are fetching what, and whether anything is blocking them.
That fourth one deserves a caveat, because it is where a lot of GEO advice gets over-excited. Crawler access is real, and it is worth checking once. But it is rarely the thing that is wrong. One SEO who has audited around 17,000 small business sites this year summed it up on r/SEO: access is almost never the bottleneck, robots.txt is usually fine, GPTBot is usually not blocked. What kills those sites is that nothing on the page can be lifted out as an answer.
Takeaway: check crawler access once, fix it if it is broken, then stop looking at it. Treat bot logs as your debugging layer and the prompt audit as your actual KPI.
Do I need to be on Reddit to get cited?
Less than the advice suggests, and there is now a large dataset that complicates the standard answer.
Ahrefs analysed 1.4 million ChatGPT prompts and published the results in April 2026. ChatGPT cited roughly half of the URLs it retrieved. The interesting part is the split by source type: results pulled from its search index were cited 88.46% of the time, while Reddit came in at 1.93%, YouTube at 0.51% and academic sources at 0.40%.
So the model reads those places heavily and quotes them rarely. Being discussed on Reddit may well shape what a model believes about your category. It is a much weaker route to an actual citation than a page of your own that answers the question cleanly.
Worth saying plainly: Ahrefs sells an AI visibility product, so they have a horse in this race. We cite them because the sample is enormous and the method is published, not because they are neutral.
Two other findings from the same study are worth writing on a sticky note. Pages with natural-language URL slugs were cited 89.78% of the time against 81.11% for those without. And cited pages showed higher semantic similarity between their title and the query, 0.602 against 0.484 for pages that got retrieved and then passed over.
Which is a delightfully boring conclusion. Write a title that matches the question somebody actually asks. Use a readable URL. That is most of it.
How do you know a change is real?
Mostly you don't, and this is the part nobody in the tool business volunteers.
Assistant outputs are non-deterministic. Run the same prompt twice and you can get different brands in a different order, with no change anywhere on your side. So a move from 6 mentions out of 20 to 8 out of 20 looks like a 33% gain in a slide, and sits comfortably inside the range you would get by running the identical set twice on the same afternoon.
Two rules keep you honest.
- Run your baseline set twice on day one. Whatever gap opens between those two identical runs is your noise floor. Any monthly change smaller than that gap is not a result, however good it looks in the deck.
- Widen the set before you trust a small move. Twenty prompts gives you coarse resolution. If a two-point shift genuinely changes a commercial decision, you need 50 or 100 prompts, not a prettier dashboard.
There is a third option that costs nothing: run each prompt three times and average, rather than once. One run is a coin flip. Three runs is a measurement. It triples your afternoon, so we would only do it for the handful of prompts that actually matter commercially.
Honestly, most reported month-over-month movement in this field is noise wearing a suit.
What should I actually do about a bad score?
Fix your SEO first, and that is Google's advice rather than ours.
Google's guide to optimizing for generative AI features asks directly whether existing SEO practice still applies. Its answer: "In short, yes! The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems." It goes further, stating that "optimizing for generative AI search is optimizing for the search experience, and thus still SEO."
That same guide lists work you can skip: llms.txt files, content chunking, AI-specific rewrites, and chasing inauthentic mentions. That is Google telling you in writing that four popular GEO tactics are unnecessary. What it does recommend is ordinary and hard: a unique point of view, non-commodity content, meeting the technical requirements, decent page experience, less duplicate content.
Beyond that, your audit tells you which of three problems you have.
Cited but never recommended. Your pages describe features while the prompts ask about situations. That gap usually closes by writing about the job the buyer is doing rather than the thing you sell, which we covered in the piece on customer objections and jobs-to-be-done.
Recommended but never cited. The model knows you and is sending nobody. Usually the answer is that your own pages are not the cleanest available source for the claim, so it quotes somebody else describing you.
Neither, but competitors are everywhere. Go back to column four. Whatever third-party pages keep appearing are your target list, and getting onto a comparison page or a directory is normally faster than outranking one.
A word on FAQ blocks, since everyone asks. They help as structure and not as a trick, and Google has retired FAQ rich results for most sites, so there is no snippet prize any more. The thing that actually works is writing an answer that stands on its own without the paragraph above it, because a retrieval system splits your page into passages and scores each one separately. One SEO called the current wave of question-shaped H2s "the keyword stuffing of yesteryear", which we think is half right. The format is fine. Bolting it onto pages that answer nothing is what fails.
Doing all of this by hand stops being fun somewhere around prompt twelve. Running the research and drafting half automatically is roughly what we built Aiter to do, though the first audit is worth doing yourself, because that is where you learn what the numbers mean.
Anyway. Go and ask ChatGPT what the best tool in your category is, logged out, right now. Whatever comes back is your actual baseline, and it takes about forty seconds to find out.
Frequently asked questions
How long does an AI visibility audit take?
Around three hours for 20 prompts across three assistants, most of it copy and paste. Month two takes about an hour once the spreadsheet exists.
How many prompts do I need?
Twenty to see the shape, 50 or more to detect small changes with any confidence. Start at 20 and widen only when a real decision depends on the precision.
Should I run prompts logged in or logged out?
Logged out, in a private window. Logged-in results reflect your own history, which makes your brand appear far more often than it does for a stranger.
Can I automate it with the ChatGPT API?
You can, and the results will not match the consumer app. The web interfaces layer retrieval, tools and system prompts on top of the raw model. Automate for volume if you like, but do not present API output as what a buyer sees.
What is a good AI visibility score?
There is no benchmark worth quoting, because share of voice depends entirely on how crowded your category is. Your own trend line on a fixed prompt set is the only comparison that means anything.
Does being cited by AI actually bring traffic?
Sometimes, and less than the citation count suggests. Somebody who reads a summary mentioning you may never click through. Treat citations as evidence of inclusion, then check your referrer data in GA4 for the rest.
Is there a free AI visibility tool?
Search Console and Bing Webmaster Tools both report AI data for free, and between them they cover Google's and Microsoft's surfaces. For every other assistant, the free option is your own time and a spreadsheet.
Do I need to do this monthly?
Monthly suits most teams. Weekly mostly measures noise, and quarterly is too slow to connect an effect back to the content change that caused it.