·11 min read
LLM visibility: how to find out whether ChatGPT recommends you
A growing part of purchase research and the buying process now happens inside ChatGPT, Gemini, Perplexity and other AI platforms. Official reporting already exists, but only in two places, and each of them shows something completely different. For the three assistants your customers actually open, there is no dashboard. That gap is exactly why you have to build the measurement yourself, and why it has to be a metric, not a screenshot.

Why one screenshot is not measurement
Most teams run into this topic the same way. Someone in marketing asks an assistant “what is the best [your category] in [your market]”, the company is not in the answer, and a slightly panicked meeting follows. The concern is justified. The method is not, and for a purely technical reason.
Assistant answers are not deterministic. The same question asked twice can return a different list of brands. Answers change with the wording, the conversation history, the user’s location and language, whether the assistant decided to search the web for that particular answer, and even which model version happens to be running that week. One answer is one random sample. You would not judge your visibility in Google by one person’s SERP screenshot either. Here that applies twice over.
The consequence is clear. Visibility in AI assistants only makes sense as a share from repeated, structured sampling. Everything below is about how to build that sampling.
Build your query set from data you already have
The queries you test are not invented in a workshop. They come from two sources.
- Search Console data. The non-branded queries that bring traffic to your category and product pages are the same information needs people now hand over to assistants. Take the commercial ones, “best X for Y”, comparisons, “X vs Y”, “is X worth it”. If you have the bulk export to BigQuery switched on, you get the data you need with one SQL query instead of an afternoon of exports.
- Rewriting into questions. People do not type keywords into an assistant, they ask questions. For each seed query, write how a real customer would ask it in a chat, with a budget, a constraint, a specific situation. “Best running shoes” becomes “I run 30 km a week on asphalt and I overpronate. What shoes should I consider under €150?”
A usable starting matrix has 20 to 50 queries, each assigned to a category and an intent, each in two or three wordings. Run each wording at least three times per assistant in one cycle. That is my own working threshold, not a figure from Google or a vendor, and it exists because one run of one wording is the same screenshot again, just in a spreadsheet. The matrix is small enough to run at a reasonable cost and structured enough that the results add up to something even a marketing director will read.

What to record for each answer
Run the matrix through each assistant’s API on a fixed schedule. A sensible default interval is once a week, with one-off runs before and after major content changes as the exception. For each answer, evaluate:
- Mention. Does your brand appear at all? The core metric is the share of answers that mention the brand within a given query group.
- Position in the list. If the answer is a list, where in it are you?
- Citation. When the assistant searched and lists sources, is your domain among them, and which URL?
- Tone. Is the mention a recommendation, a neutral listing, or a caveat (“X is popular, but users report…”)?
- Competitors in the answer. Which brands appear instead of you or next to you? Over time, this is the most strategically interesting column in the whole dataset.
Store the full text of the answers too, not just the scores. When a number moves, the first question is what exactly was written there. Running the query again will give you a different answer from the one that moved the number.
From my own practice. For me this runs as a scheduled pipeline. I keep the query matrix in a spreadsheet, call the APIs on a fixed schedule and write the evaluated results into a table together with the full text of each answer. The decision that matters is boring. Same queries, same schedule, every run. The moment someone “improves” the questions in the middle of a series, the trend breaks and the history stops being comparable.
What data the official tools offer
Official reporting exists in two places, and it is worth knowing exactly what each of them gives you. The difference is bigger than you would expect.
Bing Webmaster Tools gives the most. The AI Performance report is in public beta and shows the number of citations of your pages, the average number of cited pages, so-called grounding queries (the wording the answer is built on), citations per URL and the trend over time. The citations come from Microsoft Copilot, AI summaries in Bing and selected partner integrations.
Search Console gives much less. Google launched the reports for generative AI features on June 3, 2026, and since August 31, 2026 they have been available worldwide. There are two, one for Search and one for Discover. They show impressions, pages, countries, devices and the trend over time. Not clicks. Not queries. And watch out for one trap. This data is also included in the overall performance report. If you add it on top, you count it twice.
ChatGPT, standalone Gemini and Perplexity do not offer a comparable official visibility report for site owners.
So the order of data availability is clear, and it matters quite a lot for planning the work. Bing gives you grounding queries and citation counts for free, Google only impressions without queries, and for the three assistants your customers probably care about most, you have to build the measurement yourself. That is the reverse of their real market share, and it is worth factoring into the budget.
How to read the results honestly
To make this data useful, and not just something for a presentation, two principles help.
First, react to trends, not to individual runs. When the brand drops out of one weekly run, that is noise. A mention share falling for a whole month while a competitor’s rises is a signal. The samples here are small compared with search data, so treat week-to-week movements with caution.
Second, be honest about what can be known at all. Aggregate usage data exists. In February 2026, OpenAI published its Signals hub with an analysis of how people actually use ChatGPT, carried out with regard to user privacy. Alongside it, it publishes Enterprise Signals, a measure of enterprise adoption. The page was last updated on August 12, 2026. Before you quote from it, read the scope. It gives you context about the landscape, not a keyword tool. There is still no per-query volume data for your category, and no vendor dashboard changes that.
What you do have is a usable substitute, meaning your own repeated sampling plus one hard number from your own analytics, visits coming from assistants that link out. The truth lies between these two numbers. Anyone who promises you more precision is selling you an illusion.
What actually moves these numbers
To be honest, influencing AI assistant answers is partly the same work as search visibility, and the mechanisms are less documented than anyone selling “AEO packages” suggests. This much can be said with certainty:
- When assistants search the web, they cite content they can load. Answers built on live search draw on pages that OAI-SearchBot can access and crawl. Check which bot you are actually blocking. GPTBot and OAI-SearchBot are separate settings, and whether ChatGPT can load and cite you is decided by OAI-SearchBot.
- By blocking GPTBot, you signal that your content should not be used for training, and you do not affect anything else with it. A site without OAI-SearchBot will not appear in ChatGPT answers, but it can still appear as a navigational link. On Google’s side, Google-Extended, according to the documentation, does not affect inclusion in Google Search or ranking, it only controls training and grounding for Gemini. Blocking it therefore does not take you out of AI Overviews or AI Mode. And if you bury the answer under 800 words of introduction, you are hard to cite either way.
- A consistent description of your brand adds up. Assistants assemble answers from many sources that talk about you consistently, from your website, review platforms, comparison articles and industry overviews. A brand that is described the same way everywhere is easier to find and name than a brand every site describes differently, and that is a content problem before it is an AI problem.
- Mentions on other sites have value even without links. Being present in comparison content and in “best products” roundups that assistants draw on is visibility work even where those mentions carry no link.
What cannot be said with certainty is any formula like “do these five things and they will mention you”. Measurement exists precisely because the mechanism is opaque. You change something, track the shares and evaluate whether the results changed for the queries you track.
What to avoid
- Do not present a single screenshot to your director, however it turns out. One good answer is just as unrepresentative as one bad one.
- Do not treat mention share like rankings. There is no first place to hold. There is a probability that the assistant mentions or cites you, for each platform and each query type.
- Do not add impressions from the generative Search Console report to the overall performance report. They are already counted there, and adding them up gives you a number that does not exist anywhere.
- Do not mass-produce “content for AI”. Thin content written to be quotable fails for the same reason thin content always fails, and it also weakens the rest of the site that does the real work.
- Do not block AI crawlers by default and then commission an AI visibility project. Check robots.txt and the bot rules on your CDN first. I have seen this contradiction more than once.
- Do not skip the baseline measurement. The first measurement does not mean much on its own. Its value is that it is the point everything else is compared against.
- Do not buy a dashboard that measures someone else’s queries. A vendor “AI visibility” tool often tests different wording from what your customers use. The numbers look good and tell you nothing. Your query set should come from your own Search Console data.
- Do not overlook CDN rules. Blanket bot blocking “against scrapers” also blocks OAI-SearchBot and PerplexityBot. The site then does not appear in answers, regardless of its content. Before you measure, check robots.txt and the CDN rules for specific bots.
- Do not measure in a different language from the one your customers ask in. Measuring in English says nothing about your visibility in a non-English market. Queries belong in your customers’ language and location.
What to measure quarter by quarter
- Mention share per query group and assistant. The main trend.
- Citation share. How often your domain appears in answers with sources, and which URLs earn the citations. These URLs tell you what your most findable content looks like. Create more of it.
- Overlap with competitors. Who appears in answers where you do not, and whether that group is stable or changing.
- Traffic from assistants in your analytics. Small numbers on most sites today, but the only column in this whole dataset that relates directly to visits and revenue.
This work has no finish line you reach. It has a starting point you measure from, and you either build it now, or a year from now you will be explaining why you do not know whether anything got worse.

Want to know how AI assistants see your brand?
The AI SEO and brand citation measurement service starts with an audit that builds the query matrix, runs a baseline measurement across assistants and tells you what to change. The recommendations come from your data, not from a screenshot.