A data-backed breakdown of the factors that actually get your pages cited by AI answer engines like ChatGPT and Perplexity - and what matters far less than most checklists claim.

You already know how to rank a page on Google. Answer engines play a different game. ChatGPT, Perplexity, Gemini, Google AI Overviews and Claude don't hand out ten blue links - they read a shortlist of sources and synthesize one answer, citing a handful of them. Your job is no longer to be ranked. It's to be selected. That shift changes which factors matter. Some classic SEO signals still pull weight. Others - the ones agencies love to sell - turn out to have little independent effect on whether an AI names or links you. And a few factors that barely register in traditional SEO, like being discussed across Reddit and Wikipedia, sit near the top. This guide breaks down the AEO ranking factors the current evidence actually supports - factor by factor, each with a short "why it matters / what to do." It's built on public studies (AirOps, Kevin Indig's Growth Memo, Cyrus Shepard's Zyppy analysis, Ahrefs, Semrush), not vibes. You'll leave with a prioritized checklist, a repeatable way to audit your own pages, and a clear picture of what to stop wasting effort on.
In AEO, there is no ranked list for a user to scroll - there's a selection decision made by a model. An answer engine typically runs a pipeline: it expands your question into several sub-queries (a "fan-out"), retrieves a set of candidate pages for each, then decides which passages to pull into the synthesized answer and which URLs to cite. Ranking, in the classic sense, is only the first gate. Selection and citation are separate gates after it. The gap between "retrieved" and "cited" is large. In AirOps' analysis of 548,534 retrieved pages across 15,000 prompts, only about 15% of the pages ChatGPT retrieved actually showed up in the final answer - roughly 85% were pulled and then dropped. Getting retrieved is necessary but nowhere near sufficient. The factors below are really about surviving that second cut. So "AEO ranking factors" is a slight misnomer everyone uses anyway: there's no public ranking algorithm to reverse-engineer, only correlational evidence about which properties of a page or brand line up with being cited across large samples of real answers. Treat every factor here as a lever with evidence behind it, not a guaranteed input.
Other Articles
Why AI Confuses Your Brand - Entity Disambiguation Fixes
AI answer engines often merge or misattribute similarly named brands. Learn why it happens, how to diagnose it across ChatGPT, Perplexity, and Gemini, and the entity disambiguation fixes that make your brand unmistakable.
Wikipedia & Wikidata for AI: How to Earn (and Keep) a Presence
AI engines lean on Wikipedia and Wikidata to describe brands and entities. Here's how notability and sourcing rules really work, and the policy-compliant way to earn and keep a presence.

The single most consistent finding across studies is that tightly matching the specific question beats covering everything around it. AirOps found that retrieval rank was the strongest signal - a page in the top retrieval position was cited around 58% of the time versus roughly 14% at position 10, a fourfold difference - and that heading-to-query relevance was the strongest on-page factor. Pages whose headings closely mirrored the query were cited more often than pages that sprawled across many loosely related subtopics. Why it matters: engines fan a question into sub-questions and match passages to each. A page that answers one question cleanly is easier to slot into an answer than a megaguide where the relevant sentence is buried. What to do: map one primary question per page, put the query language in the H1 and H2s, answer it directly in the opening, and resist the urge to make every page "comprehensive."
AI models cite passages, not pages - so the answer has to be liftable as a self-contained chunk. Kevin Indig's study of 1.2 million AI answers found that roughly 44% of ChatGPT citations came from the first third of the content, a consistent "ski ramp" where citation likelihood is highest early and tapers off. If your clearest answer lives under a heap of preamble, it's less likely to be the passage that gets extracted. Why it matters: extractability is mechanical. Clear headings, short self-contained paragraphs, definitions, lists and tables give the model clean units to quote. Buried or context-dependent answers ("as we saw above…") don't travel. What to do:
Lead with the answer, then expand. Put a crisp, quotable statement in the first third.
Use descriptive H2/H3s phrased like questions or claims.
Write paragraphs that make sense out of context - no orphan pronouns pointing at earlier text.
Use lists, comparison tables and short definitions for facts you want quoted.
Don't block AI crawlers or over-restrict content previews; if the model can't fetch or preview the passage, none of the above matters. Cyrus Shepard's meta-analysis rated URL accessibility and preview controls among the highest-confidence citation factors.
Authority matters, but not the "domain authority score" version of it. This is where the evidence gets counterintuitive. AirOps found that classic domain-authority and backlink metrics had no positive correlation with being cited - if anything, slightly negative. Yet Ahrefs, studying tens of thousands of brands, found that off-site brand signals were the factors most strongly correlated with AI visibility. Both can be true: raw link authority is weak, but entity authority - being a recognizable, consistently described thing across the web - is strong. Why it matters: models build a sense of who's credible from how entities are described across their sources. Real-world experience, author expertise and clear organizational identity (the E-E-A-T cluster) feed that; a high Domain Rating on its own doesn't. What to do: strengthen entity signals, not just link counts. Keep author bios, an "About" page and organization details (name, description, sameAs links) accurate and consistent. Earn genuine expertise signals - original data, named authors with real credentials, first-hand testing. Treat E-E-A-T as reputation the model can corroborate, not a checklist you self-declare.
Answer engines lean toward current sources, and stale pages get quietly passed over. AirOps observed that pages more than about two years old were cited less often - consistent with the intuition that AI answers about anything time-sensitive prefer recent material. Why it matters: for anything evolving (tools, prices, "best X in 2026"), recency is a proxy for correctness. An engine minimizing the risk of citing something wrong favors the fresher source. What to do: refresh pages you already rank for rather than endlessly publishing new ones. Update figures, dates and examples in the content itself - and don't fake it; swapping the date without changing anything fools nobody and nothing.
A large share of AI citations goes to a small set of community and reference sites you don't own. Semrush's study of the most-cited domains in AI answers found user-generated and encyclopedic sources - Reddit, Wikipedia, YouTube and the like - occupying the top citation slots across engines, well ahead of typical SEO winners. If the AI's answer about your category leans on a Reddit thread or a Wikipedia section, your presence there shapes whether you're in the answer. Why it matters: engines treat these platforms as high-trust corroboration. A brand discussed favorably in the exact places the model already trusts gets pulled in indirectly, even when your own site isn't cited. What to do: show up authentically where your buyers already discuss the category - relevant subreddits, Q&A sites, review platforms, YouTube. Earn a well-sourced Wikipedia presence if you genuinely meet notability (don't manufacture it), and pursue mentions in credible third-party articles. This is closer to digital PR than on-page SEO.
Being mentioned - by name, consistently, across credible sources - appears to move AI visibility more than links do. Ahrefs' analysis of AI brand visibility found branded web mentions among the strongest correlates of appearing in AI answers, ahead of backlinks. As Ahrefs' Ryan Law framed it, the lever isn't a content-volume arms race - it's being discussed consistently across credible sources. There's a subtlety worth planning for: Kevin Indig's work on "ghost citations" found that AI frequently uses a source without naming the brand. Both the linked citation and the un-named mention matter, and neither shows up in your referral logs. Why it matters: consistent naming teaches the model what you are and what you're good at, building the entity association that gets you recommended even when no single page is cited. What to do: standardize how you're described (same name, category and positioning) everywhere you appear, seed accurate descriptions on the platforms that get cited, and track brand mentions in AI answers, not just clicks.
Several tactics that dominate AEO checklists have thin evidence behind them. Being skeptical here saves real budget.
Schema markup. In AirOps' data, structured data showed no independent effect on citations once domain factors were controlled. That makes schema a weak standalone citation lever, not a useless one: valid JSON-LD can clarify entities and content types and remains valuable for Search eligibility. Use the JSON-LD schema guide for AI citations for implementation, but don't expect schema alone to unlock AEO results.
Core Web Vitals / page speed. No significant effect on AI citation once domain factors are controlled. Good for users; not a citation lever.
Raw domain authority and backlink volume. Weakly or negatively correlated with citation in the studies above. Links still support ranking (which feeds retrieval), but chasing DR for its own sake is misdirected.
Sheer length and "ultimate guides." Multiple studies point the same way: focused pages outperform sprawling ones. Longer is not more citable.
Keyword density and old-school on-page tricks. Semantic match to the question matters; stuffing does not.
The through-line, echoed in Cyrus Shepard's ranking-factor synthesis: AEO and modern SEO are largely the same job - be retrievable, answer the exact question, and be a credible, widely-discussed entity. Most "AEO-specific" hacks are noise.
Run a free audit: see if ChatGPT, Gemini and Copilot recommend you, in about a minute.

A factor matters only when the team can turn it into a controlled change and a measurable outcome.
Lever | Controlled change | Leading check | Outcome |
Semantic match | Rewrite one section around the exact question and required entities. | The passage answers the prompt without surrounding context. | Citation rate for the mapped prompt cohort |
Extractability | Move the answer first, shorten the passage, and use the appropriate list or table. | The decisive answer can be copied intact. | Correct page and passage selected more often |
Entity confidence |
Change one primary lever per page cohort. If content, schema, internal links, and PR all change together, the result may improve but the team learns nothing reusable.
See your mentions across ChatGPT, Claude and Perplexity in real time, the moment buyers ask.
Rough priority order, highest-leverage first:
Priority | Factor | Core action |
1 | Retrievability / accessibility | Don't block AI crawlers; ensure pages are fetchable and previewable |
2 | Relevance & semantic match | One question per page; query language in headings; answer up top |
3 | Passage extractability | Lead with the answer; clean headings, short chunks, lists/tables |
4 |
Run this pass on the pages you most want cited. It takes an afternoon per cluster.
This is the loop Qwairy is built to close: it tracks your citations and brand mentions across ChatGPT, Claude, Perplexity, Gemini and AI Overviews, shows which sources the engines pull from, and surfaces which factors correlate with the answers you actually win - so the audit runs continuously instead of once a quarter. More method on the Qwairy blog.
Turn ranking signals into action: Follow the complete AEO guide, rewrite priority passages with the GEO content optimization framework, and measure outcomes with the AI citation playbook.
AEO isn't a new algorithm to game - it's the discipline of being selectable. Make pages retrievable, answer the exact question in an extractable passage, keep content current, and build a consistent, well-discussed entity across the places AI already trusts. Stop over-investing in schema, speed scores and raw domain authority as standalone citation levers; none substitutes for relevance, retrievability and extractable evidence. The brands that win AI answers are genuinely relevant and genuinely talked about - and now you can measure both.
In rough order: being retrievable and accessible to AI crawlers, matching the specific question (semantic relevance), passage extractability (a clean answer high on the page), off-site presence and mentions on trusted platforms, entity/brand consistency, and freshness. Classic domain authority and schema are weak standalone predictors of citation compared with relevance, retrieval access and passage quality.
They overlap heavily; the best current evidence suggests AEO and modern SEO are largely the same job - be findable, answer the question, be a credible entity. The difference is emphasis: AEO cares more about passage-level extractability, off-site brand mentions, and being selected into a synthesized answer rather than ranked in a list.
The evidence for a direct citation lift is weak. AirOps' controlled analysis found no independent effect of structured data on AI citations once domain factors were accounted for. Schema still helps machines resolve entities and content types and supports traditional Search features, so implement it correctly with the JSON-LD schema guide for AI citations - just don't treat it as a standalone AEO strategy.
That's the "ghost citation" pattern: answer engines frequently draw on a source without naming the brand or linking it, so the value never appears in your referral analytics. It's a strong reason to track brand mentions inside AI answers directly, not just clicks and citations.
Heavily. Studies of the most-cited domains consistently show community and encyclopedic sites like Reddit, Wikipedia and YouTube near the top of AI citation sources across engines. If the AI leans on those platforms to answer questions in your category, your accurate, favorable presence there can put you in the answer even when your own site isn't cited.
Because citations shift as models and their sources change - and barely overlap across engines - a one-time audit goes stale fast. Re-check your priority prompts at least monthly, or use continuous monitoring so you catch changes in who's being cited before they cost you visibility.
Track your mentions across ChatGPT, Claude, Perplexity and all major AI platforms. Join 1,500+ brands monitoring their AI presence in real-time.
Free trial • No credit card required • Complete platform access
No conflicting entity facts remain. |
Correct brand attribution and fewer conflations |
Freshness | Substantively update time-sensitive facts and document the review date. | Every changed claim has current evidence. | Recovery on freshness-sensitive prompts |
Off-site corroboration | Earn relevant mentions on sources already used for the topic. | New independent evidence is indexed and accurate. | Higher mention and citation share across engines |
Reddit, Wikipedia, YouTube, credible third-party coverage |
5 | Entity / brand consistency | Same name, category and description everywhere |
6 | Freshness | Refresh pages you already rank for; keep facts current |
7 | E-E-A-T signals | Named expert authors, first-hand data, clear org identity |
- | Schema, CWV, raw DR | Do for users/SEO; don't expect AEO gains |