" /> Track AI Citations Free: Manual Method + Tools Compared (2026)
AEO Fundamentals

Track AI Citations Free: Manual Method + Tools Compared (2026)

Arielle Phoenix
Arielle Phoenix
Jun 26, 2026 · 11 min read

TL;DR

  • Citations are the new backlinks for the AI era. Ahrefs’ study of 75,000 brands found branded web mentions are the strongest correlate with AI visibility (0.656 for Google AI Overviews, 0.664 for ChatGPT, 0.709 for Google AI Mode). Link metrics like backlink count barely correlated at all.
  • Most AEO writing tells you to add schema and ship an llms.txt and never tells you to measure anything. That order is backwards. You cannot improve a citation rate you have never looked at.
  • A one-off check lies. LLM answers are non-deterministic. Ask ChatGPT the same question twice and you can get different sources, because, as Lily Ray puts it, responses are ‘non-deterministic and increasingly personalized’. A single screenshot proves nothing.
  • Only about 12% of URLs that ChatGPT, Gemini and Copilot cite rank in Google’s top 10 for that prompt (Ahrefs, 15,000 prompts). Perplexity is the outlier at 28.6%. Your rank tracker is mostly blind to where you actually get cited.
  • AI engines lean fresher. URLs cited by AI assistants average 1,064 days old versus 1,432 for organic results, about 25.7% fresher (Ahrefs, ~17M citations).
  • The manual method works and costs nothing: a spreadsheet, logged-out sessions, the same prompts run weekly. A tracker earns its keep when you move from “did it happen once” to “is the trend going up”.

Here is the opinion most AEO content is too polite to say out loud. The whole industry has the order backwards. Everyone publishes the same checklist (add schema, ship an llms.txt, optimise for fan-out) and almost nobody tells you to measure first. You cannot improve a citation rate you have never looked at. So before you touch a single piece of markup, you need a way to track AI citations, the act of checking whether ChatGPT, Perplexity, Gemini and Google AI Overviews actually name your site in their answers.

Why bother? Because citations are the new backlinks. Ahrefs studied 75,000 brands and found branded web mentions were the strongest correlate with AI visibility across every engine: 0.656 for Google AI Overviews, 0.664 for ChatGPT, 0.709 for Google AI Mode. Link metrics like backlink count and URL Rating barely moved the needle. If mentions are the currency now, measuring them is not optional. I dig into why they matter in how brand mentions and AI citations actually work. This post is about the other half: the act of tracking them.

And the scale is not small. ChatGPT alone handles roughly 2.5 billion prompts a day. Some slice of those is people asking the exact buyer questions your business answers. You either know whether you show up or you are guessing.

How do I check if ChatGPT or Perplexity is citing my website?

Use a fresh logged-out session, turn web search on, ask your real buyer query, and log whether your URL appears in the cited source pills. A single check tells you almost nothing, because the answer changes run to run. So you log it and repeat.

The free method, in order:

  1. Open a fresh, logged-out session (or a private window). Memory and account history skew results, and you want a clean read.
  2. Turn web search on. Without retrieval, the model answers from training data and cites nothing live.
  3. Ask the actual questions your buyers ask. Not “best CRM” but “best CRM for a 5-person plumbing business” if that is who you sell to.
  4. Look at the citation pills and footnoted sources, not just whether your brand name appears in the prose. Being mentioned by name and being cited as a source are different things.
  5. Log every run in a spreadsheet. This is the part everyone skips, and it is the whole point.

Here is the spreadsheet schema I actually use. Copy it:

AEO God Mode — Free WordPress Plugin Get your site cited by ChatGPT, Perplexity, and Google AI Overviews. Install in under 5 minutes.
Download Free
ColumnWhat goes in it
dateWhen you ran the check
engineChatGPT / Perplexity / Gemini / Google AIO
promptThe exact question, word for word
web_search_ony / n
citedy / n (were you a named source?)
cited_urlThe exact URL it pulled, if any
positionWhere in the answer (first source, buried, etc.)
competitors_citedWho else got named

Why log the exact cited URL and not just “yes”? Because AI engines do not always cite the page you would expect, and sometimes cite pages that are not even live. Ahrefs found AI assistants send users to 404 pages 2.87x more often than Google Search (0.43% versus 0.15% on average), and ChatGPT was the worst offender at 1.01% of its clicked URLs hitting a dead page, across roughly 8 million clicked and nearly 8 million cited URLs. If you only log “cited: yes”, you miss that the engine is pointing people at your old, deleted, or wrong page.

Why does the same prompt give different AI citations every time?

Because LLM answers are non-deterministic and personalised. Sampling, memory settings, location and live retrieval all shift which sources get named, so the same question can pull different sites two minutes apart. That is why you track a rate over many runs, never a single answer.

Lily Ray put the limit plainly. These responses are “non-deterministic and increasingly personalized”, accounting for a user’s conversation history, Memory settings, interests and geography. Any third-party tracker stays “fundamentally limited by the black box of the user experience”, and the prompt-volume numbers those tools sell you are “at best, highly directional and greatly sampled”.

Sit with that, because it is the contrarian beat the tool listicles will not give you. The vendors selling a precise “AI visibility score” out to one decimal place are quietly overselling. There is no stable single number to measure here, by the admission of the people closest to the data. What you can measure honestly is a trend: run the same prompt set every week, log every result, and watch the citation rate move over a month. One screenshot is a coin flip. Twenty logged runs is a signal.

Is tracking AI citations different from tracking my Google rankings?

Yes, and your rank tracker is mostly blind here. Being page one on Google does not guarantee you get cited by AI. Ahrefs ran 15,000 prompts and found only about 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google’s top 10 for that prompt. More than 80% of those citations do not rank anywhere in Google for the query.

The per-engine breakdown is where it gets useful, and where I think the single “12%” stat undersells the story:

EngineShare of citations from Google’s top 10
ChatGPT~8%
Gemini~8%
Copilot~9%
Perplexity28.6%
Google AI Overviews76%

Read the spread. ChatGPT, Gemini and Copilot barely touch Google’s top 10. Perplexity leans on it harder (28.6%) because it relies more on Google rankings than the others do. And Google’s own AI Overviews pull 76% from the top 10, which makes sense given they are Google. The takeaway: your rank tracker is a decent proxy for AI Overviews and a partial one for Perplexity, and close to useless for predicting ChatGPT, Gemini or Copilot. That is the gap a citation log fills.

AEO God Mode — Free WordPress Plugin Get your site cited by ChatGPT, Perplexity, and Google AI Overviews. Install in under 5 minutes.
Download Free

There is a second reason rankings mislead you. AI engines lean fresher than organic search. URLs cited by AI assistants average 1,064 days since publication versus 1,432 for organic Google results, about 25.7% fresher, across nearly 17 million citations. ChatGPT showed the strongest freshness pull. Your evergreen page can hold its Google rank for years and quietly fall out of AI answers as it ages. The rank tracker stays green while the citation rate bleeds. Only a citation log catches it.

What should I actually track over time to measure AI visibility?

Track citation rate, the exact pages that get pulled, your share of voice versus key competitors, and AI referral traffic. Trend lines matter more than one-off screenshots, because they show whether visibility is rising or falling week to week.

One framing that helps: citation tracking is the downstream outcome metric. It tells you whether you got cited. The upstream predictive metric is on-page, what I call your Citability Score and why it matters more than rankings. Citability predicts; citation tracking confirms. You want both, and you want to know which is which.

Can I see AI traffic in Google Analytics or Search Console?

Partly, and badly. Google does not separate AI Overview clicks at all. Ahrefs confirms Google blends AI Overview data into standard organic search, with no tools in Search Console or GA4 to identify AI Overview impressions, citations or clicks. An AIO click shows up as google / organic or with no referrer, no distinct label.

The chatbots are kinder. ChatGPT, Perplexity and Gemini referrals land in your reports as referral traffic you can filter by source. The concrete win the listicles skip: ChatGPT tags its outbound links with utm_source=chatgpt.com, so you can isolate ChatGPT referral sessions in GA4 with a clean filter. That is a free, practitioner-grade tactic. Set up a segment on that UTM and you have a running count of sessions ChatGPT sent you, no tool required.

What none of this catches is the bots themselves. A citation starts when an AI crawler fetches your page, long before any human clicks an answer. To see that, you read your server logs, which is its own job. I cover it in checking which AI bots are actually crawling your site. Treat the crawler log as the upstream signal: no fetch, no citation, ever.

When is a citation tracker worth paying for versus doing it manually?

Manual checks are fine for one site, under a dozen prompts, and a weekly trend. You want a tracker once you are handling multiple brands, more than 20 prompts, or you need competitor share of voice by engine. The manual method stops being enough the moment it eats an afternoon a week.

AEO God Mode — Free WordPress Plugin Get your site cited by ChatGPT, Perplexity, and Google AI Overviews. Install in under 5 minutes.
Download Free

Be honest with yourself about the break point. If you are checking 12 prompts on one site and only need a weekly trend, the spreadsheet wins. It costs nothing and it teaches you what your buyers actually ask. Do that first. If you move to multiple brands, more than 20 prompts, or need competitor share of voice by engine, the time cost starts to outweigh the free method. At that point, a tracker is paying for the extra history, faster comparisons, and less manual copying into the sheet. The difference shows up in the fields you already use: date becomes daily history instead of one-off checks, engine and prompt multiply across brands, cited and cited_url are still useful, but competitors_cited and position get hard to manage by hand once the volume grows.

That is the gap I built AEO God Mode to fill for WordPress sites, so you do not bolt a separate SaaS onto your stack. Its Citation Tracker logs whether ChatGPT, Perplexity and Gemini cite your site over time, with per-day history so you watch the trend not a screenshot. AI Referral Analytics catches the visits those engines send that GA buries in referral traffic. The AI Crawler Log shows which bots fetched which pages, the upstream signal before a citation ever happens. Three data streams in one place inside WordPress. To be straight about coverage: AEO God Mode actively addresses 16 of 23 factors (11 strong, 5 partial); the other 7 depend on content quality, off-site brand, or time. The manual method is real and free. Do that first. The plugin exists for when you are tired of doing it by hand every week.

And once you are measuring, fix the things that actually move citations. That means the Tier 1 work: can the crawler reach the page, do you rank for the prompt and its fan-out sub-queries, does the next paragraph answer the exact question asked. Schema and llms.txt are worth doing once and then forgetting, low effort, low impact. They are not the hero. Measurement is. The rest of the playbook lives in the full WordPress AEO checklist, and if you want the engine-by-engine detail, see how ChatGPT decides which websites to reference and how Perplexity selects its sources.

Frequently asked questions

Weekly is the sweet spot for most sites. Often enough to catch a trend forming, rare enough that you do not drown in noise. Because answers are non-deterministic, a single weekly run is still one sample. If a prompt really matters, run it two or three times in the same session and log each result, then read the rate, not any single answer.

Track the engines your buyers actually use, but do not assume they behave alike. Ahrefs’ 15,000-prompt study showed ChatGPT, Gemini and Copilot pull almost nothing from Google’s top 10 (around 8 to 9%), while Perplexity pulls 28.6% and Google AI Overviews pull 76%. If you only check one engine, you get a distorted read. ChatGPT plus Google AI Overviews is a sensible minimum because they sit at opposite ends of that spectrum.

Because Google rank and AI citation are mostly different games. Only about 12% of URLs cited by ChatGPT, Gemini and Copilot rank in Google’s top 10 for the prompt, and over 80% do not rank anywhere in Google for that query. ChatGPT also leans toward fresher pages, so an older page can hold its Google position while quietly dropping out of AI answers. Ranking is necessary for some engines, never sufficient for all of them.

Yes, and you should start there. A spreadsheet, logged-out browser sessions with web search on, and the same buyer questions run weekly will tell you most of what you need. For referral traffic, filter GA4 by source, and use the utm_source=chatgpt.com tag to isolate ChatGPT sessions. A paid tracker earns its place when the manual routine starts eating an afternoon a week.

A citation is when an engine names your URL as a source in its answer. A brand mention is when your name appears in the text, with or without a link, which still matters because branded mentions correlate strongly with AI visibility. AI referral traffic is the human clicks that land on your site from those engines. You can be mentioned without being cited, and cited without getting any clicks, so track all three separately.

A little, and not as much as the noise around it suggests. Schema is worth adding once because it helps machines parse your page, but the data does not support treating it as the thing that gets you cited. An llms.txt file is quick to ship and low impact. The factors that decide most citation outcomes are whether the crawler can reach the page, whether you rank for the prompt and its sub-queries, and whether the next paragraph answers the exact question asked. Measure first, then fix those.

Arielle Phoenix
Written by
Arielle Phoenix
AI SEO at AEO God Mode

Helping you get ahead of the curve.

AEO AI SEO Digital Marketing AI Automation
View all posts →
AI Search Optimized by AEO God Mode (opens in a new tab)