Someone on the team opens ChatGPT, types the company name, and gets a friendly paragraph. The screenshot hits Slack. That is usually the last time anyone measures it.
The reply is not proof. The account you were logged into, earlier chats, custom instructions, and the exact wording of that one prompt all pulled the answer toward you. Measure AI visibility with a method you can repeat: a fixed set of questions, sessions that do not know who you are, more than one engine, and a scoring rule you wrote down before you looked at the results. A tool can run that loop. It cannot invent the questions. Read also: getting your company recommended by AI: supplier shortlist facts - the content gaps that show up as zeros on constraint questions.
In this article:
- Why is asking ChatGPT about your own brand worthless?
- What should you actually measure?
- Where do the questions come from?
- How do you run it, and how often?
- Spreadsheet, a SaaS seat, or self-host?
- How do you read the breakdown?
- How do you see AI traffic in analytics, and what do you report?
- Common questions
Why is asking ChatGPT about your own brand worthless?
Logged-in ChatGPT remembers you. Prior chats, custom instructions, and anything the product infers from your account all pull the answer toward brands you already talk about. You are the last person who should run the test, because the system is trying to be helpful to you, not to a buyer who has never heard your name.
One run is an anecdote anyway. Ask the same question twice in a fresh session and the shortlist often moves. That happens for three ordinary reasons, and none of them is a conspiracy against your brand.
The first is how the model writes. It does not look up a stored list of approved suppliers. It generates the answer one token at a time, and there is a bit of randomness in which word comes next. Two runs of the identical prompt can therefore name different companies, even when nothing on the web has changed.
The second is retrieval. For many questions the assistant also fetches pages while it answers, and the pages it finds are not a fixed library. A roundup that was crawled yesterday, a competitor's spec table that went live this morning, a forum thread that ranked a little higher today: any of those can enter the answer on the next run and drop out on the one after.
The third is wording. "Best marine bilge pumps for yachts" and "which 12 V submersible pump fits a 400 mm locker in salt water" look similar to a person and are different jobs for a model. They pull different sources and return different names. If you only ever ask the branded version, you will look visible and still be missing from the prompts that decide a purchase.
The method that follows is dull, which is why it works. Freeze the questions and write them down. When you change a phrasing, give it a version number so last month's score still means the same thing. Run the set in a clean state: logged out, or through an API, or through a tool that does not carry your history. Then repeat. Several runs per question, on more than one engine, before you treat a miss as a miss.
What should you actually measure?
You need four numbers. Each one hides something the others do not, so a single headline will always be incomplete.
Mention rate is the share of questions where the brand is named at all. That is the number you will put in a slide, because it answers "are we in the answer?" What it hides is where you appeared. A mention in sentence seven of a long reply is not the same as being the first name on a three-company shortlist.
Citation rate is the share of answers that link your domain. It often does not match mention rate. Models name companies they cannot point to, and they cite pages that never name you. Citation rate is the one that has a chance of becoming a visit, because a link is something a person can click. When citation rate stays at zero while mention rate looks fine, check whether fetchers can reach your pages at all - see can an AI actually read your website.
Position is where in the answer the brand appears. In a shortlist of three, first-named is the one the buyer remembers. If you count "mentioned somewhere" as a win, you will report a number you later cannot explain to anyone who actually reads the answers.
Share of voice is your mentions against all vendor mentions in the category. Boards like it because it looks like market share. It misleads quickly. The denominator is whoever the model happened to name that week, on that phrasing, not the real competitive set. Use it as a split you look at, not as the number you manage toward.
Write the scoring rule before the first run. Decide in advance what counts: a mention of a subsidiary, a mention without a link, a competitor page that names you, a "brands like X" list that includes you in passing. If you change the rule after you see the numbers, the baseline is already unusable, because you will not know whether the score moved or the exam did.
Here is one row, filled in, so the sheet is not abstract. The company is made up. The answers are invented to show the columns, not taken from a live run.
| Date | Engine | Q | Phrasing | Run | Mention | Citation | Position | What the answer actually said |
|---|---|---|---|---|---|---|---|---|
| 2026-08-26 | ChatGPT | 5 | v1 | 2 | No | No | none | Named two other pump makers. Cited a "best bilge pumps 2026" roundup. No Northwake. |
Keep the raw answer text next to the ticks. The wording is where you see why you missed: the model wanted a size you never published, it used last year's price, it confused you with a similar name, or it trusted a roundup that never asked you for facts.
Where do the questions come from?
Invented questions measure your assumptions. Real ones come from places buyers already talked.
Pull them from the sales inbox and lost-deal notes, from RFI documents, from support tickets, from recorded calls if you have consent, from Search Console queries that look like questions, and from the forums your buyers actually use. Keep the buyer's words, including the sloppy ones. Record the source and the date you added each question, so you can later see which sources produced the misses that mattered.
The set has to cover more than "do you know our name?" A buyer who already has your brand in mind is a different person from the one who asked an assistant to name three suppliers that fit a locker.
| Type | What it is for | Rough share |
|---|---|---|
| Branded | Can the model talk about you at all? | Small. Enough to know entity recognition works. |
| Category and "best of" | Do you appear when nobody named you? | Large. This is the shortlist. |
| Problem-aware | Do you survive a paragraph of constraints? | Large. This is how buyers actually ask. |
| Comparison | What happens when they already have two names? | Medium. |
| Cost | Can it quote a price, a range, or an honest "not published"? | Medium. |
| Segment | A different buyer, same product. | Some. |
| Geography and language | The same job in another market. | Split the set. Do not average them. |
A set that is mostly branded questions produces a comfortable number, and a useless one. You will look present and still be invisible on the prompts that decide the deal. Twelve to twenty well-sourced questions, with a few phrasings each, beat fifty that someone wrote in a workshop. What you want is a stable reading per type, not a round number that looks serious in a deck.
Below is a starter set for a made-up marine pump brand, Northwake. We do not sell pumps. The list is here so you can copy the shape, then replace every line with questions from your inbox.
- Is Northwake a good brand of marine bilge pumps?
- Who makes Northwake pumps, and where are they built?
- Best marine bilge pumps for yachts
- Which companies make 12 V submersible bilge pumps for salt water?
- I need a bilge pump that fits a 400 mm locker, runs on 12 V, and can sit fully submerged in salt water with grit
- Quiet bilge pump for a yacht cabin, under 15 A draw
- Northwake 400 versus Tidewell 400 for a 12 m yacht
- Which of those two has spare parts in Europe?
- How much does a marine bilge pump for a yacht cost?
- Does Northwake publish list prices for the 400 mm range?
- Bilge pump for a commercial fishing boat, not a leisure yacht
- Pump for a workboat that sits in harbour for months
- Marine bilge pump suppliers in Germany
- Who sells yacht bilge pumps with service in northern Europe?
Questions 1 and 2 will flatter you if the entity is known. Questions 5, 6, 11 and 13 are the ones that turn into pages, because they are the ones a buyer would actually type before they have a shortlist. If the answer says your IP rating or dimensions are missing, the fix is usually readable text on the page, not another branded FAQ - see text in images and SEO for why image-only specs fail both search and fetchers.
Freeze a core and leave it alone. When the market shifts, add questions with a new version number. If you silently replace line 5, you will think a content change moved the metric when you actually changed the exam.
For geography and language splits, treat each market as its own subset rather than averaging scores - the same pattern as a multilingual catalogue where one language can rank while another stays invisible.
How do you run it, and how often?
The pattern is simple, and it is the part that should survive a year from now: cover the places your buyers actually ask, and cover more than one kind of retrieval, because those surfaces disagree. A chat assistant, a citation-heavy answer engine, and whatever Google is doing with generated answers in search are different jobs. One of them can name you every time while another never does. That split is a normal first result. It is not a bug in the method.
Which products sit in those slots will move. As of writing, a worked example looks like ChatGPT, Perplexity, and Google AI Overviews or AI Mode, with Copilot, Claude or Gemini added only if sales conversations already mention them. Do not treat that list as the method. Ask your own team where deals start, put those surfaces in the sheet, and skip the ones nobody in your market uses just because a roundup named them.
You do not need every engine on day one. Two or three, of different kinds, already shows whether the miss is "we are unknown" or "we are unknown on this surface". Add another only if someone will actually read that extra split.
Some surfaces still need a human. A logged-in chat product, a sidebar in a browser, an answer that only appears after a follow-up: APIs and scrapers miss those. Budget a few manual runs for the questions that decide deals, and automate the rest.
Run a full baseline once, then a lighter subset on a schedule. The core constraint questions every two weeks or monthly is enough for most teams. Bring the long tail back when you have a reason, such as a new market or a page you just published. Daily scores on 200 prompts look busy. They rarely change what you write next week.
Record the answers, not just the ticks. A miss that says "no pump in this size is listed with an IP rating" is a content task: you never published the rating. A miss that names you but attributes last year's discontinued model is a correction task: the model found an old source. Those two score the same if you only store yes or no, and you will send the wrong person to fix the wrong thing.
Spreadsheet, a SaaS seat, or self-host?
Every paid AI visibility tool runs the same loop you just read: a question set, several engines, a score. Buying does not replace that loop. It scales it, so you can ask more questions, more often, without someone pasting into a chat all afternoon.
As of August 2026 the market splits into three ways to run it. The names below are examples of a category, not a ranking. The roster will move.
| Spreadsheet, clean sessions | Dedicated GEO platforms, or an SEO suite add-on | Self-host | |
|---|---|---|---|
| Coverage | A few engines, as many runs as you have hours | Broad, on a schedule | Whatever you connect |
| Who writes the questions | You | You, plus whatever default list the vendor ships | You |
| Raw answers | If you paste them | Varies. Some keep text, some keep a score | You store them |
| Competitor benches | Manual | Usually in the product | Only names you collect yourself |
| Cost | Time | Per seat, per prompt, or both | Engineering time plus API bills |
| Maintenance | You | The vendor | You |
Dedicated platforms in this category right now include Profound, Peec, and Otterly. SEO suites that bolted tracking onto an existing login include Semrush's AI toolkit and Ahrefs Brand Radar. We publish a self-hosted option, Ansvisor. It is MIT licensed, you hold the prompts and the answers, and the scoring formula sits in the repository. It will not invent competitor data you never collected, and it will not replace a human on surfaces that have no usable API. A spreadsheet plus a logged-out browser is still a legitimate first week.
Buying wins when you have no one to run scripts, you need competitor share of voice you cannot collect yourself, and a board wants a dashboard before the next meeting.
Self-hosting wins when the questions that decide deals have to stay yours, you need the raw answers six months from now, and per-seat pricing gets silly as the set grows. You also take on format changes, rate limits, and the gap between an API reply and what a user sees in the product. That last gap is real. A pipeline is not the same as sitting in the chat.
Most teams land on a hybrid: a bought tool for breadth, and a short owned set for the ten questions that actually close work. The owned set is the one you score by hand when the dashboard and the sales inbox disagree.
How do you read the breakdown?
The average is the least interesting number on the sheet.
A 40% mention rate can mean you appear on every branded question and on none of the constraint ones. It can also mean you appear in English and vanish in German, or that ChatGPT names you and Perplexity never does. Those three situations produce the same headline and three different backlogs, so the headline is not what you act on.
Split the results by engine, by language, and by question type before you brief anyone. What you will usually see is not a clean zero across the board. It is a decent overall rate with a hole in one cut. That hole is the work.
A miss becomes a page when the question is real, sales already cares about it, and nothing on a readable URL answers it. A list of misses is not a plan, because ten unrelated zeros will produce ten thin FAQs. Group them first. One spec table can close three constraint questions. Ten thin FAQs close none of them well. The pages that fix constraint misses are the same ones a CMS should already support - typed fields, tables, and crawlable HTML covered in 10 SEO features a modern CMS should have.
Do not treat a one-week move as proof that a new page worked. Answers jump around for the same reasons the first test lied: sampling, retrieval, phrasing. You need a frozen subset of questions, enough runs, and time before you claim a cause. Until you have that design, report presence and gaps. Do not report that the blog post "moved ChatGPT."
How do you see AI traffic in analytics, and what do you report?
Some assistants pass a referrer. In GA4 you can look for hostnames such as chatgpt.com and perplexity.ai and build a segment for them. Many chats pass nothing, so the visit lands as Direct, mixed in with people who typed the URL. First-touch and last-touch both lie in that situation, because the research happened in a conversation you do not own.
UTM tags on your own links still help for campaigns you control. They do not tag a citation inside ChatGPT. If the model names you and the buyer goes to your homepage by typing it, analytics will not know that ChatGPT was involved.
Google Search Console is a different window, and it only looks at Google. On 3 June 2026 Google launched Search generative AI performance reports. They show impressions when your URLs appeared in AI Overviews, AI Mode, and generative features in Discover. The first version is impressions: how often a page was shown inside those features. It does not show clicks, it does not show the queries that triggered the appearance, and it does not show ChatGPT. The reports were rolling out to a subset of sites, so if you do not see them yet, that is a rollout rather than a missing integration on your side.
Search Console therefore tells you whether Google's AI surfaces showed your pages. Your question set tells you whether ChatGPT named you. GA4 tells you about the slice of visits that still carry a referrer. None of the three is the whole picture. Put them next to each other. Do not add them into one "AI traffic" number and manage that. For the crawl and index side that feeds those Google surfaces, technical SEO is still the base layer.
What a board needs is smaller than a session chart. They need exposure: on the questions that decide deals, how often you are named, and on which engines. They need gaps: which question types or languages are still zero. They need corrections: where the model is wrong about you, and whether the source of that error is still live. Skip the sparkline of Direct traffic. It will bounce, and it will start an argument nobody can finish.
Common questions
Why does ChatGPT give me different answers to the same question?
Because it is not retrieving a stored fact sheet. Each time you ask, it generates the reply with a bit of randomness, so the next word is not fully determined, and two identical prompts can produce two different shortlists. On many questions it also fetches pages during the answer, and the pages it finds can change from one hour to the next. If you change even a few words of the prompt, you have given it a different job, so you should expect different names. If you are logged in, memory and earlier chats pull the answer toward you on top of all that. The way to measure through the noise is to freeze the phrasing, use a clean session, and treat several runs as the data. One reply is a story.
How much traffic actually comes from AI?
There is no single share that applies to every site. A slice of visits shows up as referrals from assistant hostnames. Another slice hides in Direct, because the chat never passed a referrer. A large share of the influence never becomes a session at all: the buyer got three names and went to the winner without touching your analytics. That is why the question set matters more than a traffic percentage. Measure how often you are named and cited. Treat GA4 as a partial view of visits, not as the size of the opportunity.
Can I see AI referrals in GA4?
Some of them. Build a referral segment for the assistant hostnames you actually see in the reports. Expect it to undercount, because many chats send the visitor with no referrer. Do not use that segment as the only number you report.
What is a good AI visibility score?
There is no universal good score, because the number is an average over your questions. A set loaded with branded prompts will look strong. A set loaded with constraint prompts will look worse and tell you more. Judge movement on the questions that decide deals, against your own baseline, not against a percentage you saw in someone else's screenshot.
Should I buy an AI visibility tool?
Buy one if you need volume, a competitor bench, and you will not run scripts. Keep the question set and the scoring rule yours. A seat that only runs the vendor's default prompts will report on a market that is not yours. A spreadsheet for a week is enough to learn whether you have a branded-only problem or a constraint problem. Decide after you have seen that, not before.
Run a baseline this week
Copy the scoring table from earlier in this article. Take twelve questions from the inbox, not from a brainstorm. Run them on two or three of the surfaces your buyers already use, a few times each, while logged out. Score mention, citation, and position, and read the zeros by type rather than as one average.
If you need volume and competitor benches, pick a seat from the category that fits how you already buy software. If you want the questions and the raw answers on machines you control, self-host. We maintain Ansvisor for that path.
If you already have a site and want to know whether it is in shape for models to find you and quote you, write to us. Send the URL. We will run an audit: whether the pages are readable, whether you show up on the questions buyers actually ask, and where the holes are. You get the sheet either way.