How We Track Which Questions an AI Answers With Our Name
· · 8 min read
We track AI citations by defining a set of buyer-intent prompts, running them across ChatGPT, Perplexity, Claude, and Gemini on a fixed cadence, recording every brand mention and source URL, and measuring trend direction rather than single-run rank. Single-query tracking measures noise. You need aggregate frequency across many runs to see the signal.

If you asked me a year ago whether we could measure AI visibility, I would have said no. I was wrong. Not because the measurement is easy, but because I was looking for the wrong number. We now run a tracking system that tells us exactly which AI platforms cite our brand, for which questions, and how often.
TL;DR
Tracking AI citations requires measuring aggregate mention frequency across many runs, not single-query rank. Single-run tracking captures noise because AI models produce different responses nearly every time. Our system defines a fixed prompt set, runs each prompt across four platforms weekly, records brand mentions and cited URLs separately, and follows trend direction over time. The signal is not "we ranked third this week." It is "we appeared in 58 percent of runs this month, up from 42 percent last month."
Why single-query tracking fails
Rand Fishkin's team ran 2,961 prompts across ChatGPT, Claude, and Google AI Overviews. The study found that fewer than 1 in 100 runs returned the same list of brands for the identical prompt. Fewer than 1 in 1,000 returned them in the same order.
LLMs are probability engines. Every response is sampled from a distribution. Any metric built on a single query result is measuring noise.
The same research found something useful underneath the noise. Across hundreds of runs for the same buyer intent, the top brands appeared in 55 to 77 percent of responses regardless of prompt wording. The consideration set was real even when the rank was random. You are either in the set or you are not. We explained the mechanics in how ChatGPT picks businesses to recommend.
What we track and why
We track four things for every buyer-intent prompt we care about. Each one answers a different question about our AI presence.
First, mention rate. The percentage of runs where our brand name appears in the AI response, linked or not. A mention puts our name in front of the buyer regardless. A mention with a link is stronger.
Second, citation rate. The percentage of runs where our domain appears as a cited source URL. Perplexity makes this easier than other platforms because it numbers citations inline on every answer. ChatGPT and Claude cite differently, so you need a method that normalizes across platforms.
Third, competitor overlap. For each prompt, we record which competitors appear alongside us and how often. If three competitors are in every response and we are in half, we have a gap.
Fourth, source domain distribution. We track which URLs get cited and which outside domains are winning citations. The U.S. Small Business Administration counts roughly 33 million small businesses in the United States, and every one of them is invisible to AI unless they have explicitly built for it.
The RAG pipeline selects and filters sources in four stages, and fewer than ten distinct URLs appear in 80 percent of LLM responses. You need to know which URLs made the cut.
The tracking method in six steps
Each step addresses a failure mode we learned the hard way. Here is the method.
-
Define a fixed prompt set. We write 15 to 25 buyer-intent questions the way a real prospect would ask them. "Who should I hire for X in my area" not "best X services Phoenix." Real buyers do not type like SEOs.
-
Run each prompt at least twice per platform per cadence. AI responses are not deterministic. One run is closer to a coin flip than a measurement. We run each prompt three times across ChatGPT, Claude, Perplexity, and Gemini weekly.
-
Record brand mentions and cited URLs separately. A mention is not a citation. A mention captures brand awareness. A citation means the model pulled your content directly. We log both in a spreadsheet with the date, platform, prompt, and result.
-
Calculate mention rate and citation rate per prompt over time. If our brand appeared in 2 of 3 runs last week and 2 of 3 this week, the rate is stable at 67 percent. If it dropped to 1 of 3, we investigate.
-
Track the trend direction, not the absolute number. A 58 percent mention rate does not mean we win 58 percent of the time. It means we are in the consideration set at that rate. The direction over four to eight weeks is the signal.
-
Map citation gaps to content actions. If a prompt we are not winning maps to a page we have not written, we write the page. The Princeton GEO paper found that content with expert quotations lifts citation rate by 41 percent and statistics by 33 percent. Those are your content levers.
What the data looks like after three months
We started with zero AI visibility for any prompt in our category. After three months of entity-building work and consistent brand signals, our mention rate moved from 0 percent to 38 percent. That means our name appears in roughly two of every five runs.
The platforms do not all move together. Our ChatGPT mention rate sits higher than Perplexity, which sits higher than Claude. Different models draw from different sources, and visibility in one engine does not transfer to another. The same buyer-side pattern shows up in our analysis of the referral pipeline.
Half of U.S. adults now use AI chatbots according to Pew Research Center's June 2026 survey. That is roughly 131 million people who might hire you. If you are not tracking whether your brand appears when they ask, you are flying blind.
Knowing how to check if ChatGPT knows you exist is the five-minute version of what we run weekly. The full tracking system gives you trend data that a one-time check cannot.
How to start tracking your own AI visibility right now
Pick five buyer-intent questions for your business. The ones a real prospect would ask a chatbot before they ever call you. Open ChatGPT and ask each one. Open Perplexity and ask each one. Record your name, your competitors names, and which source URLs appear.
Do it again next week. Same prompts, same platforms, same spreadsheet. After four weeks you have a trend line, not a data point. That trend line tells you whether you are moving toward the consideration set or away from it.
If you want us to run a full citation sweep for your business across ChatGPT, Perplexity, Claude, and Gemini, reach out. We run the prompts, we record the data, and we show you exactly where you stand.
Sources cited in this analysis?
- SparkToro - AI Brand Recommendation Consistency Research - 2,961-run study, less than 1 percent identical lists
- DailyGEO Insights - How LLMs Decide Whom to Cite - RAG pipeline architecture, citation concentration data
- Princeton and Georgia Tech - GEO: Generative Engine Optimization - Expert quotations plus 41 percent, statistics plus 32 percent citation lift
- LLM Pulse - Track Brand Mentions in Perplexity AI - Perplexity tracking methodology and metrics
- Driftspear - SparkToro AI Study Analysis - Consideration set appears in 55 to 77 percent of responses
- Pew Research - Americans and AI 2026 - Half of U.S. adults now use AI chatbots
- Slate - Best AI Citation Tracking Tools 2026 - Citation tracking metrics and platform comparison
Frequently Asked Questions
How many times do you need to run a prompt to get a reliable reading?
SparkToro's research suggests 60 to 100 repeated queries per prompt for statistically meaningful data. We run three times per prompt per platform each week and aggregate over time. The weekly number is noisy but the four-week trend direction is the reliable signal to track.
Do all AI platforms cite sources the same way?
No. Perplexity numbers its citations inline on every answer it generates. ChatGPT only cites sources when search browsing is triggered by the query type. Claude cites differently from both. Each platform requires its own extraction method and the citation formats are not interchangeable.
How often should you update your prompt set?
We review the prompt set quarterly without exception. Buyer language evolves and new competitors enter the market. New questions emerge as the category changes. A prompt set that was accurate six months ago may miss the questions buyers are asking today.
What is the difference between a brand mention and a citation?
A mention is any appearance of your brand name in the AI response without a linked source URL. A citation is your domain appearing as a linked source in the response. Mentions are more common and capture brand awareness. Citations are stronger signals because they mean the model pulled your content directly.
Can you automate AI citation tracking entirely?
Partial automation is possible with tools like Slate and Omnia that run scheduled prompt sets and extract citations from responses. Full automation without human review risks misclassifying mentions, missing platform changes, and treating a single-run result as a ranking. We use tools for extraction and a human for trend analysis.
2026-07-19 - v4.0.0 - v4 conformant - built on The Standard
