How AI Engines Decide Who to Cite
When ChatGPT, Perplexity, Gemini or Google's AI Overviews answer a question, they name a handful of brands and cite a handful of sources — and leave everyone else out. The selection isn't random, and it isn't a ranking you can buy your way into. It's the product of how these models are built and how they fetch information. Understand that, and “how do I get cited by AI?” stops being a mystery and becomes a to-do list.
This is an honest explainer, not a growth-hack list. No one outside OpenAI or Google can see the exact weights, and anyone promising a shortcut to guaranteed citations is selling something. But the factors below are well understood, consistent across engines, and — importantly — things you can actually influence.
The one principle underneath everything: AI mirrors the web
Every one of these engines is a reflection of the open web, just compressed and summarised in different ways. A model learns from the web during training; a search-grounded engine reads the web live at answer time. Either way, if a brand is well represented across the pages and sources these systems draw on, it tends to get named. If it's absent or invisible to crawlers, it doesn't. Almost every factor below is a specific version of this one idea.
Two modes, two different jobs
Before the factors, one distinction most write-ups skip — because it changes what you work on:
- Memory (training data): the model answers from what it absorbed during training, usually with no live lookup and no citations. Here you compete on how widely and credibly your brand appears across the web the model was trained on.
- Live search (grounded): the engine browses in real time, reads current pages, and cites specific URLs. Here you compete to be on the exact pages it pulls for that query.
“Get into the training data” and “get onto today's top pages” are different tasks. Our guide to generative engine optimization (GEO) goes deeper on the discipline; this post focuses on the deciding factors.
The 7 factors that decide who gets cited
1. How widely you appear across the web the model learned from
For memory-based answers, there's no live page to cite — the model is recalling patterns from training. Brands that are mentioned often, across many credible sources, are the ones it “remembers.” This is slow-moving and can't be gamed quickly: it's the accumulated result of genuine presence — reviews, articles, documentation, discussions — built up over time.
2. Whether you're on the pages the engine retrieves live
For search-grounded answers (Perplexity, Google AI Overviews, ChatGPT Search), the engine runs a query, pulls a small set of pages, and builds its answer from them. If your page isn't in that retrieved set, you can't be cited — full stop. Being genuinely useful and discoverable for the specific question is what gets you into the candidate pool.
3. Third-party mentions, not just your own pages
This is the factor brands underestimate most. AI engines lean heavily on independent sources — roundups, “best X” listicles, directories, documentation and active communities — because a third party saying you're good is more trustworthy than you saying it. A large share of citations point to pages the brand doesn't own. That means earning a place in the roundups and conversations the engines read is often more powerful than any change to your own site.
4. Whether crawlers can actually read your page
A surprising amount of content is invisible to AI crawlers. Many don't execute JavaScript, so if your key content only renders client-side, the crawler sees an empty shell. Blocking AI user-agents in robots.txt, slow pages, and login walls have the same effect. You can be perfect on every other factor and still be uncitable because the bot never sees your words. Clean, server-rendered HTML — and, increasingly, plain-text or markdown versions for LLMs — removes that barrier.
5. Clear, directly-answered content
Models favour content that answers the question plainly and is easy to extract: a direct answer near the top, clear headings, lists and tables, and structured data (schema) that spells out what a page is. This is the overlap with answer engine optimization — see what AEO is. Burying the answer under 1,500 words of preamble makes you harder to quote, so you get quoted less.
6. Topical authority and a consistent identity
Engines build an understanding of entities — who you are and what you're known for. A site that covers its topic thoroughly and describes itself consistently (same name, same positioning, clear “we are an X that does Y”) is easier to recognise and recommend than one with a thin or contradictory footprint. Depth on your actual subject beats scattered content.
7. Freshness and corroboration
For live answers, recency matters — engines often prefer current pages, and a claim repeated across several independent sources is treated as more reliable than one that appears once. Being corroborated (several trusted places saying the same thing about you) reinforces whether you get named, and keeping key pages current keeps you eligible.
What you can actually do about it
Notice that none of the seven is a trick. The honest levers follow directly from them:
- Publish genuinely useful content for the questions you want to win, with the answer stated clearly and early (factors 2, 5).
- Make your pages readable to bots — server-rendered HTML, no AI-agent blocks, fast loads, and ideally LLM-friendly plain-text versions (factor 4).
- Earn third-party mentions — get into the roundups, directories and communities the engines cite (factors 1, 3, 7).
- Build topical depth and a consistent identity rather than scattered one-off posts (factor 6).
- Add structured data so engines can parse what your pages are (factor 5).
What you can't do is force an engine to recommend you, pay for a citation slot, or predict a model update before it lands. Be skeptical of anyone who claims otherwise.
How to know if it's working
Because these factors move at different speeds, the only way to know what's helping is to measure the actual answers over time — which brands get named for your questions, which sources get cited, and when that changes. That's what an AI visibility tracker does, and a citation tracker shows you the exact sources shaping each answer so you know which factor to work on next.
Start with your baseline
Before optimising anything, find out where you stand today. Our free AI search visibility report checks your buyers' questions across ChatGPT, Perplexity and Google AI Overviews and shows whether you're named — and who's named instead.
Frequently asked questions
How do AI engines decide which brands to mention?
They reflect the web they're built on. For answers from training data, they recall brands that appear widely and credibly across many sources. For live, search-grounded answers, they pull a small set of current pages for the query and name the brands those pages support. Crawlability, clear content, third-party mentions, topical authority and corroboration all influence which brands make the cut.
Why isn't my brand cited by ChatGPT or Perplexity?
Usually because the sources the engine reads for your topic don't mention you, or because your own pages aren't readable to AI crawlers (for example, content that only renders with JavaScript). The fix is to earn mentions on the pages engines cite and make sure your own content is genuinely present and crawlable.
Can I pay to be cited by AI engines?
No. There's no paid citation slot in an AI answer, and no tool can guarantee a recommendation. Citations are earned through genuine presence across the web the engines draw on — be cautious of anyone promising to buy or guarantee them.
How long does it take to improve AI visibility?
It varies by factor. Fixing crawlability or adding a clear answer to a page can help the next time an engine reads it. Building the third-party mentions and topical authority that drive memory-based answers is slower — weeks to months — because it depends on the web around you changing, not just your own site.
Want to see where you stand in the AI era?
Launch the AI Scanner