Guide · updated 2026-09-13

How to get cited by AI engines.

AI engines cite sites that other sites already talk about, that publish original numbers, that answer the question in the first sentence, and that let the search crawlers in. Everything else is small.

This is what the evidence says moves citations in ChatGPT, Claude, Gemini, Grok and Perplexity, in order of how strong the evidence is, with what is different about each engine and a checklist you can work through in an afternoon.

01

What "cited" means, and why only one kind sends visitors

When an AI assistant answers "what is the best invoicing tool for freelancers", it does two things: it names products in the text, and it links to the pages it read. Linked means a page on your domain was one of those links. Named means your brand appeared in the text with no link. Linked is the one that sends a visitor; named is the one that tells you the engine knows you exist.

Across the sites we audit, the two come apart constantly. tally.so was named in 52% of answers and linked in 14%: the engines know it, but the click goes to a listicle about it. That gap is the whole game. The rest of this guide is about closing it.

02

What moves citations, ranked by strength of evidence

#SignalEvidenceStrength
1Other sites talking about you: YouTube, Reddit, listicles, review sitesAhrefs, 75,000 brands: YouTube mentions correlate with AI visibility at r = 0.74, web mentions 0.66, backlinks only 0.22. For SaaS, 82% of citations go to third-party pages, 18% to the brand's own site.Strongest, replicated
2Original statistics, quotations, cited sourcesThe Princeton GEO paper, a controlled experiment: adding statistics lifted visibility 34%, quotations 44%, citing sources 29%. Keyword stuffing cut it 8%. Half of ChatGPT-cited posts contain owned data.Causal
3Answer-first pages under question headings72% of ChatGPT-cited posts open with a 20-25-word direct answer. 44% of citations come from the first 30% of a page. Passages under a question heading are twice as likely to be quoted.Strong
4Letting the search crawlers inDomains that block GPTBot get 0.003 ChatGPT citations per Google ranking; domains that allow it get 0.417. Large publishers block and still get cited through syndication; a small site cannot.Large
5FreshnessPages AI engines cite are 250-458 days fresher than Google's top ten for the same query. Updating the visible date alone moves things about 13%.Medium
6Specifics: names, products, numbersCited passages are 20.6% proper nouns against 5-8% in typical web text. Vague copy is not quotable.Medium

The order matters. Most GEO advice starts at the bottom of this table because the bottom is what you can change on your own site this afternoon. The top is what actually decides it.

03

What does not move citations (much)

  • llms.txt. Ahrefs looked at crawler logs: 97% of llms.txt files received zero requests from AI crawlers; Slackbot fetched them more than PerplexityBot did. Publish one, it costs nothing, but it is not a lever.
  • Schema.org JSON-LD. A difference-in-differences test on 1,885 pages that added JSON-LD, against 4,000 controls, measured +2.2% ChatGPT citations and -4.6% in Google AI Overviews. Engines extract visible HTML.
  • Promotional tone. The strongest negative on-page signal in Ahrefs' data: promotional pages are 26% less likely to be cited. "Leading", "seamless" and "powerful" are the words of pages that do not get quoted.
  • Per-engine copies of the same page. "How to rank in Claude", "How to rank in Gemini" and so on, each repeating one playbook, is what Google's spam policy calls doorway pages. The 2026 spam updates removed page sets like that from results.
04

ChatGPT

How it searches: through Bing's index. If Bing cannot find your page, ChatGPT cannot cite it. Crawlers to allow: OAI-SearchBot is what search results depend on; GPTBot is the training crawler; ChatGPT-User fetches a page when a person asks about it. Blocking GPTBot alone does not remove you from search, but sites that block it are cited about 140 times less often in practice, so treat the three as a set.

What it cites: Wikipedia is 7.8% of all ChatGPT citations, then Reddit, Forbes and G2. Its mix moves abruptly: Reddit's share of its top sources fell from roughly 60% to 10% in a single month in 2025.

What we see in audits: ChatGPT's web search links news and listicle pages far more than product sites, for every domain we test. plausible.io was linked in 2% of ChatGPT answers and 80% of Grok answers for the same questions. A low ChatGPT number is an engine trait as much as a verdict on you; the fix is being in the listicles it reads.

05

Claude

How it searches: through Brave Search's index, not Google's or Bing's. Check that your pages appear on search.brave.com; many sites that rank on Google are thin there. Crawlers to allow: Claude-SearchBot builds the search index, Claude-User fetches on a person's request, ClaudeBot is training. Anthropic documents all three separately; robots rules for ClaudeBot do not cover the other two.

What it cites: no large published breakdown exists for Claude. In our runs it links product sites when a page answers the buyer's question directly in its opening lines, and it repeats the same handful of sources across runs more than ChatGPT does, which makes its results stable once you are in.

06

Gemini

How it searches: Google Search grounding. Ordinary Google indexing is the prerequisite; Google-Extended in robots.txt controls training use, not whether Gemini can cite you, and blocking Googlebot removes you from everything. What it cites: about three sources per answer, drawn from the top of the Google results: Wikipedia, Reddit and YouTube lead. Google's AI results have the highest share of brand domains of any engine, at roughly 60%, so a product page that ranks on Google has a real chance here.

What we see in audits: Gemini is the engine most likely to link a product site directly, and the one where classic SEO carries over most cleanly.

07

Grok

How it searches: xAI's live search over the web and over posts on X. It is the only engine where activity on X is a direct input. What it cites: no published breakdown. In our audits Grok links product sites more readily than ChatGPT and cites X posts about products alongside them; plausible.io was linked in 80% of Grok answers. It also costs the most per query to run, which is why the free snapshot skips it and the full audit includes it.

08

Perplexity

How it searches: its own index, built by PerplexityBot; Perplexity-User fetches on request. Allow both. What it cites: Reddit is 6.6% of all Perplexity citations and nearly half of its top-ten sources; YouTube, Gartner and, for B2B questions, LinkedIn and G2 follow. A Reddit thread in which people recommend you is worth more on Perplexity than any page you can write.

What we see in audits: Perplexity linked hyperping.com in 70% of answers where ChatGPT linked it in none. It links several sources per answer, so being one of them is achievable for a small site with a clear answer page.

09

The checklist

  1. Let the search crawlers in. Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and their user-fetch variants in robots.txt. Decide separately about the training crawlers.
  2. Be in Bing and Brave, not only Google. Submit your sitemap to Bing Webmaster Tools; check search.brave.com for your key pages.
  3. Serve the content without JavaScript. Fetch your page with curl; if the answer to the buyer's question is not in that HTML, the engines do not see it. In our audits this is the single most common failed check.
  4. Write the answer in the first sentence. For each question you want to win, a page whose first 25 words answer it, under a heading that is the question.
  5. Put numbers on the page. Your own: customers, uptime, prices, benchmark results. Numbers get quoted; adjectives do not.
  6. Publish one piece of original data a quarter. A survey, a benchmark, a count of something only you can count. This is the only lever with a controlled experiment behind it, and it is what earns the third-party pages that matter most.
  7. Get onto the pages the engines already cite. The listicles, comparison posts and review sites that show up for your category. The audit tells you which ones those are.
  8. Be present where people recommend things. A YouTube walkthrough with your name in the title; honest answers in the Reddit threads for your category. Never promotional.
  9. Keep pages visibly current. Refresh the data posts every one or two months and update the date on the page.
  10. Cut the marketing words. Delete "leading", "seamless", "powerful", "cutting-edge". Replace each with a fact.
10

Measure it before and after

One chat is one answer from one engine, and the answers change every time. The free snapshot asks four engines the five questions a buyer in your category asks, and shows who each one linked: you, a rival, or nobody. Run it, work the checklist, wait for the engines to re-crawl, and run it again.

Run a free snapshot

Sources

Sources