Skip to content

AI Citations

What Makes AI Cite a Page?

What large-scale citation research suggests about content alignment, page format, authority, freshness and the sources different AI systems choose.

Emerson Kauffmann · · 22 min read

What does it mean for an AI system to cite a page?

A citation is a source an assistant attaches to an answer as the evidence for something it said. It is not a recommendation, it is not the same as having your brand named, and it is not a ranking position. Treating it as any of those things is the fastest way to misread the research in this article.

It helps to separate four outcomes that get collapsed into the word visibility. A page is retrieved when a system pulls it into the working set of material it considers. It is cited when it survives that set and appears as a source beside the answer. A company is mentioned when its name appears in the answer text. It is recommended when the answer puts it forward as the thing the reader should choose.

These come apart constantly. Semrush's 2026 analysis of nearly four thousand domain appearances found that 61.7% of them were citations without an accompanying brand mention, and a further 25.1% were mentions with no citation. Only 13.2% were both. So a page can be the evidence behind an answer that never names its author, and a company can be named in an answer that cites somebody else entirely.

This article is about the narrowest of the four: what characteristics of a page appear to accompany more citations. That is a genuinely useful question, because citation is the outcome most directly tied to what is on a page. It is also the outcome least directly tied to revenue, which is why the wider framing in our introduction to AEO and the measurement discipline in the AEO and SEO comparison both still matter.

What this research can and cannot tell us

It can tell us which characteristics of a page tend to accompany more citations across a large sample of pages that were actually cited. It cannot tell us how any assistant ranks sources, because none of them publish that, and none of the studies discussed here had access to it.

That distinction is worth holding onto, because most published AEO advice is either pattern-matched from search engine optimization or generalised from a handful of screenshots. Observational research on a large sample is a real improvement on both. It is still observational.

Almost every quantitative claim in this article comes from one study: Discovered Labs' 2026 analysis of roughly 2 million citation observations across four AI engines over a six-month window, joined to crawled and feature-engineered data on 10,000 of the cited pages. It is, as far as we know, the largest published analysis of citation predictors. It is also a single study of one population, and single studies are how fields end up confidently wrong.

  • Correlation is not mechanism. A feature that accompanies more citations may be causing them, may be a proxy for something else, or may be a consequence of the same editorial care that produced the citations.
  • Controls help, and they are not proof. The stronger findings below survive a multivariate model with domain fixed effects, content depth, page type and page age controlled for. That removes the most obvious confounders, not all of them.
  • The population is specific. The dataset is oriented toward business-to-business subject matter and four particular engines. A local trades business or a consumer retailer is not obviously the same population.
  • The systems change. These are measurements of products that were updated during the capture window and have been updated since. A coefficient is a photograph, not a law.
  • Only cited pages were crawled. Feature analysis on a sample of pages that already earned citations tells you what distinguishes them from each other. It says less about what separates them from the pages that were never cited at all.

The honest summary is that this evidence should change what you prioritise, and should not be quoted as if it were documentation. Where a finding below is strong, we say so. Where it is one number from one model, we say that too.

Content alignment appears to matter most

Of everything measured, how closely a page's language and concepts matched the way people actually prompt was by far the strongest page-level predictor of citation count. Its standardised effect was roughly three times the size of the next-strongest page feature.

The reported coefficient is +0.37, with a 95% confidence interval of +0.33 to +0.41. Within the study's model, a one standard deviation increase in alignment corresponded to roughly 30% more citations. The next three features on the list, in order, were FAQ sections at +0.07, a short summary or bottom-line-up-front block at +0.05, and an author bio at +0.02.

Empirical

Prompt alignment was the dominant page-level signal

Prompt-content alignment
+0.37
FAQ section
+0.07
TLDR or BLUF block
+0.05
Author bio
+0.02

Standardised effect on citation count

Standardised regression coefficients on citation count, from Discovered Labs’ analysis of roughly 2 million citation observations across 10,000 cited pages. Domain, content depth, page type and page age were controlled for, and the FAQ figure is the one reported for third-party pages. These are observed associations within one model on a business-to-business population, not a universal ranking formula.Source: Discovered Labs, 2026, What actually drives AI citations.

Alignment was also the most robust finding in the study rather than merely the largest. It remained significant after false discovery rate correction, was selected in every one of 200 stability-selection bootstraps, and survived a double machine learning treatment that strips out confounder variance, coming through at +0.35. Most single features in a study this size do not clear all of that.

That combination is what makes it worth acting on. A large coefficient that appears in one specification and vanishes in the next is a curiosity. A moderate coefficient that survives every attempt to break it is a finding.

What alignment actually means

Not keyword matching. A page is aligned when its language, concepts and factual coverage correspond to the questions people actually ask, at the level of detail they ask them.

The difference is easiest to see in a real buying question. Somebody evaluating compliance software asks how much SOC 2 tooling costs for a fifty-person company. An aligned page states a price or a defensible range, says what team size the range assumes, names what is in and out of scope, and explains what the implementation actually involves. A page headed with an invitation to transform your compliance journey covers none of that, and no amount of repeating the phrase SOC 2 changes it.

Alignment therefore starts with knowing the questions rather than with writing. The cheapest sources are the ones already in the business: recorded sales calls, the questions support answers repeatedly, the objections that come up before a decision, the comparisons prospects raise unprompted. A keyword tool will give you phrasings. It will not give you the fifty-person qualifier, which is the part that made the question specific.

Two failure modes are worth naming. The first is answering a more general version of the question, which is what happens when a page is written from a keyword rather than from a conversation. The second is answering the right question in language nobody uses, which is common in regulated and technical categories where the internal vocabulary and the customer vocabulary have drifted apart. Both look like coverage on a content plan and neither is aligned.

Where on the page the answer sits

The paragraph engines cited most often sat at a median depth of 0.36 down the page, where 0 is the top and 1 is the bottom. Roughly speaking, the best matching passage tended to be about a third of the way in.

Empirical

The strongest matching passage tended to appear relatively early

  1. 0.0Top of the page
  2. 0.36Median depth of the paragraph engines cited most often
  3. 1.0Bottom of the page

A median, not a cut-off. Half of the best-matching paragraphs in the sample sat deeper than this line.

From Discovered Labs’ crawl of 10,000 cited pages. The statistic describes where the most frequently cited paragraph sat on a page, expressed as a proportion of the page's length. It does not establish that engines stop reading at any particular depth, and page length varies enormously across the sample, so the same proportion means very different word counts.Source: Discovered Labs, 2026, What actually drives AI citations.

The tempting conclusion is that AI reads only the first third of a page. That is not what a median depth of 0.36 says. Half of the best-matching paragraphs sat deeper than that, and some sat a long way deeper. What the number does suggest is that pages which bury their strongest passage are competing against pages that do not.

There is also a simpler explanation available, and it is worth preferring: well-written pages tend to put their best material early, because that is how you keep a reader. The statistic may be measuring editorial competence as much as machine behaviour. The practical advice happens to be the same either way, which is a comfortable position to be in.

So the useful version of this is not a rule about the top third. It is a question to ask about any page you care about: if somebody quoted the single most useful paragraph on this page, where would it be, and is there a reason it is not closer to the beginning? Setup, background and qualification are often the reason, and they are usually movable.

Structural additions help, but much less than alignment

FAQ blocks, short summaries and author information all showed positive effects on citation count, and all of them were small. The largest, an FAQ section at +0.07, is under a fifth of the alignment effect and does not reach the conventional threshold for a small effect at 0.10.

This is the part of the study that most contradicts current practice, because the standard AEO checklist is largely made of these items. They are worth doing. They are not worth doing first, and they will not rescue a page that answers the wrong question.

The more interesting results in this part of the analysis were the ones that came back empty. After controlling for domain, content depth, page type and page age, the study found no significant independent effect on citation count from:

  • real-user Core Web Vitals, which collapsed into latent factors showing no significant effect once domain was controlled for
  • synthetic Lighthouse performance scores
  • schema markup, which the authors suggest contributes indirectly through content structure rather than as a standalone citation signal

It is important to read that correctly, because it is easy to turn into three bad conclusions. It does not mean speed does not matter, that structured data is useless, or that technical work can be skipped.

What it means is narrower. Among pages that were already being cited, these attributes did not independently predict how often. That is entirely compatible with them mattering earlier in the chain: a page that loads badly or renders only after scripts execute may not be reliably retrievable in the first place, and structured data removes ambiguity about what a page is describing. Neither of those effects would show up as a citation-count coefficient in a sample of pages that had already cleared the bar.

There is also a confounding explanation the study itself gestures at. Schema markup, decent performance and a clear author byline tend to co-occur with a competently run website. Once domain is controlled for, much of what those signals were standing in for has already been absorbed by the domain term. The signal has not disappeared. It has been reassigned.

The practical position, then: treat this layer as hygiene rather than strategy. Implement it consistently, keep it accurate, and stop expecting it to move a number on its own.

Page format changes the odds

What kind of page you publish mattered independently of how well it was written. In a controlled model that already accounted for domain, content depth, alignment and page age, pricing pages carried a coefficient of +0.39 while listicle reviews sat at -0.12.

Empirical

Page format still mattered after controlling for other factors

Pricing pages
+0.39
Listicle reviews
-0.12

Coefficient on citation count, against the reference category

Comparison and how-to formats sat near the reference category in the published analysis and are not plotted, because no coefficient was reported for them numerically.

Coefficients from Discovered Labs’ controlled multivariate model, which included domain fixed effects, content depth, alignment and page age. Because the effect persists after alignment is controlled for, format appears to carry information of its own rather than acting purely as a proxy for commercial intent. Only the two formats with published numerical coefficients are shown.Source: Discovered Labs, 2026, What actually drives AI citations.

The likely mechanism is structural rather than mysterious. A pricing page is shaped like the answer to a transactional question. A comparison page that names alternatives and puts their differences side by side is shaped like the answer to an evaluation question. A listicle of loosely ranked options is shaped like something written to occupy a search result, and it tends to contain less of substance per paragraph than either.

The wrong response to this is to convert your content plan into pricing pages. Format only helps when the format matches a question somebody is asking, and a pricing page for a service nobody prices, or a comparison page that compares nothing meaningfully, inherits none of the effect. The finding is about correspondence between the shape of a page and the shape of a question, which is really the alignment finding again, viewed from a different angle.

The right response is narrower and more useful: check whether the commercially important questions in your category have a page of the appropriate shape at all. Many businesses have twenty blog posts and no page that states what anything costs.

Why specific commercial information carries weight

Because it is the material an answer cannot be assembled without. A general-purpose model can produce a competent paragraph about what dental implants are or how SOC 2 audits work. It cannot produce what you charge, how long your waiting list is, or which of your tiers includes a named feature.

That makes a short list of facts disproportionately valuable, and they are usually the facts businesses are most reluctant to publish: prices or honest ranges, what is included and excluded, eligibility and prerequisites, realistic timescales, what the process involves step by step, what you do not do, and how your version differs from the two or three alternatives a buyer is actually weighing.

The reluctance is understandable and usually mispriced. Withholding a price does not keep you in the conversation, it removes the page from the set of pages that can answer the question. If a genuine range is impossible, the useful move is to say why and give the variables: a page explaining that implant costs run between two figures depending on bone grafting, sedation and the number of units is more usable than either a fixed number or silence.

Engines disagree about freshness

There is no single answer to how fresh content needs to be, because the four engines studied had measurably different appetites. The median age of cited content ran from 5.1 months on Claude to 8.0 months on ChatGPT.

Empirical

Different engines favour different content ages

Claude
5.1
Google AI
6.0
Gemini
7.8
ChatGPT
8.0

Median age of cited content, months

Median ages of cited content across Discovered Labs’ six-month capture window. These describe what each engine cited, not how often anything should be published or updated. A median also says nothing about the spread: an engine with a median of eight months still cited plenty of recent material.Source: Discovered Labs, 2026, What actually drives AI citations.

The same divergence shows up in the shape of the decay rather than only in its midpoint. Claude drew 60% of its citations from content less than six months old. ChatGPT drew 40% from the same band, meaning the majority of what it cited was older than half a year.

Empirical

Claude skewed more heavily toward recent content

Claude
60%
ChatGPT
40%

Share of the engine's citations on content under six months old

The two ends of the range in the same dataset: Claude showed the steepest decline in citations as content aged and ChatGPT the shallowest. Both figures describe an engine's aggregate appetite over the capture window, and neither predicts how long a specific page will keep its citations.Source: Discovered Labs, 2026, What actually drives AI citations.

The conclusion this most invites is a refresh cadence, and that is the conclusion to be most careful about. Nothing here shows that republishing a page with a new date earns citations. What it shows is that the pool of content these engines cited skewed toward the recent, more sharply on some than others.

The defensible reading is narrower than a cadence and more useful. Content whose facts move should not be allowed to become visibly stale, because a page with last year's prices, a discontinued service or a superseded regulation is wrong rather than merely old. Content whose facts do not move needs no schedule at all: an explanation of how a procedure works does not improve by being touched every quarter.

So the question to ask about each page is not when it was last updated but whether anything on it has become untrue. That distinction is what separates maintenance from the practice of rewriting timestamps, which costs the same and achieves nothing.

Authority sits above the page

The strongest predictor in the whole analysis was not a property of the page at all. A domain-level authority feature carried a mean absolute SHAP value of 0.38, against 0.06 for the strongest individual page-level feature other than alignment: roughly six times the influence, within that model.

Empirical

Domain-level authority dominated individual page signals in this model

AI-perceived domain authority
0.38
Strongest page-level feature beyond alignment
0.06

Mean absolute SHAP value

SHAP values report how much a feature moved this model’s predictions, which is not the same as how important a signal is in general. AI-perceived domain authority is a feature the study constructed, and it is not Moz’s Domain Authority, Ahrefs’ Domain Rating or anything Google publishes. Read it as evidence about where a citation model put its weight rather than as a causal weight in any engine.Source: Discovered Labs, 2026, What actually drives AI citations.

If the finding holds, it is the most commercially consequential one in the study, and the least convenient. It says the ceiling on what page-level work can achieve is set somewhere above the page, and that a challenger competing against a well-established domain is not going to close the gap with better formatting.

It also explains a pattern that otherwise looks arbitrary: a thinly written page on a widely referenced domain being cited in preference to a better page on an unknown one. That is not a judgement about the two pages. It is the domain term doing most of the work.

The caveat matters as much as the number. This is one model's attribution of its own predictions, using a feature its authors built. A different construction of authority would produce a different number, and possibly a different ranking. What survives the caveat is the direction: domain-level evidence appears to be upstream of page-level effort, which is consistent with how these systems are described and with what we see in client work.

What authority means in practice

Not a score, and nothing you can buy. In this context authority is an accumulation of independent evidence that a business exists, does what it says, and is regarded in a particular way by people who are not it.

In practice that means being referenced rather than only publishing: coverage in trade or local press, entries on professional registers and licensing bodies that actually govern your category, accurate listings on directories a human would consult, citations in other people's research, mentions from suppliers, partners and clients, and reviews with enough text to describe something specific. What these have in common is that a third party chose to say something about you.

The uncomfortable part is the timescale. Alignment work can change a page this week. Authority accrues over quarters, mostly as a side effect of operating visibly and doing things worth mentioning, and there is no version of it that runs faster because you spent more. The realistic strategy for a smaller business is not to win the authority contest but to compete where it is narrower: a specific service, a specific city, a specific question where the established domains have nothing precise to say.

A source can matter to one engine and barely register in another

Engines do not draw on the same web. In the same dataset, 97% of the LinkedIn citations observed came from Google AI alone. Claude and ChatGPT each drew under 0.4% of their third-party citations from LinkedIn, and Gemini produced none at all in that window.

That figure needs reading precisely, because it is easy to inflate. It is not LinkedIn's share of all citations, and it does not make LinkedIn important. It says that among the citations of LinkedIn that were observed, almost all of them came from one engine. The platform was close to irrelevant to the other three.

The same pattern shows up in how much weight each engine gives to a company's own pages. In this dataset 39% of ChatGPT's unique cited URLs were brand-controlled, against 14% for Gemini, which is a large difference in how much of the answer a business can supply itself. We look at that split, and what to do about it, in the guide to ChatGPT visibility.

The practical consequence is that a source strategy built by watching one assistant will be miscalibrated for the others. It is also a reason to be sceptical of any list of the platforms AI cites: the list is engine-specific, and the only reliable version is the one you build by reading the sources actually shown for your own questions.

There is no universal AEO checklist

The evidence supports a rough order of priority, not a list of tasks that produces citations when completed. The order is more useful than the list, because most of the items only work when the ones above them are already in place.

Read from the top. Answering the right questions is the precondition for everything else. Specific, checkable content is what makes an answer possible to assemble. Using a page format that matches the shape of the question multiplies content that is already aligned. Domain-level authority sets the ceiling and moves slowest. Keeping time-sensitive information current protects what you have. Structural additions and technical quality are cumulative and cheap, and they are the last place to look for a result.

None of that reduces to a checkbox, which is precisely the problem with how AEO is currently sold. A page with impeccable schema, an FAQ block, a fast load time and a named author, answering a question nobody asks, on a domain nothing references, will not be cited. Every item on the standard checklist is present. The two things that mattered most are missing.

What a citation-ready page looks like

A page that can be cited is one where a specific, checkable claim can be lifted out and attributed without distortion. That is a higher bar than being well written, and a different one from being optimised.

Conceptual

The parts of a page that make it quotable

A descriptive H1 naming the exact question the page answers

  1. Direct answerThe question resolved in the opening paragraph, in the words a buyer would use
  2. Specific factsPrices or ranges, timings, eligibility, exclusions and anything a general model could not know
  3. EvidenceFirst-hand experience, original data or named sources supporting the claims made
  4. ComparisonHow this option differs from the alternatives a reader is actually weighing, named
  5. Frequently asked questionsOnly where real questions genuinely remain after the body of the page
  6. Entity detailsWho publishes this, where they operate, and details that agree with every other source
  7. Current informationA visible date and a real review cycle, on the pages whose facts actually move
Conceptual. The two emphasised blocks are the ones the research in this article speaks to most directly; the rest are supporting, and the order shown is a reading order rather than a required template. A page that needs none of the lower blocks is not a worse page.

The test that catches most problems is extraction. Take the single most useful paragraph on the page and read it on its own. If it still means what it meant in place, it can be cited. If it collapses because it depended on the previous three paragraphs to establish what it was referring to, it cannot, and a system looking for an answer will use somebody else's paragraph instead.

In practice failures are almost always one of the same few things. The subject is a pronoun. The number has no unit or currency. The claim is hedged into meaninglessness. The geography is implied by an address in the footer. The comparison names no alternative. All of these are cheap to fix and none of them require writing for machines, which is the part most AEO advice gets backwards: extractable prose is just prose that says what it means.

What the evidence does not justify

Several claims currently circulating are either unsupported by the research above or directly contradicted by it. They are worth stating plainly, because most of them are being sold.

  • That schema markup makes you appear in ChatGPT. Schema showed no significant independent effect on citation count in this analysis. It reduces ambiguity, which is a reason to implement it accurately, not a mechanism for inclusion.
  • That an FAQ block earns citations. It was the strongest of the structural signals and its effect was +0.07, below the conventional threshold for a small effect. It is a small positive, not a lever.
  • That every page needs constant rewriting. The freshness findings describe what engines cited, not a publishing schedule. Pages whose facts do not move gain nothing from being touched.
  • That length wins. Content depth was a control in these models rather than a headline finding, and adding words to a page that answers the wrong question changes nothing about which question it answers.
  • That putting the answer first guarantees inclusion. A median matching depth of 0.36 is an observation about where good passages sat, not a rule about what gets read.
  • That domain authority is a score you can buy. The authority feature here was constructed by the researchers from observed evidence across the web. It is not a purchasable metric and it is not any vendor's published score.
  • That a citation is a recommendation. The majority of citations in Semrush's dataset came with no brand mention at all. Being the evidence behind an answer is not the same as being the answer.

The common thread is scale. Most of these claims take a real but small effect, or a descriptive statistic, and promote it to a mechanism. The research does not support mechanisms. It supports priorities.

How to apply the evidence

The sequence below is our reading of the research rather than part of anybody's dataset. It is ordered by how much the evidence suggests each item matters, and by how long each one takes to move, which are not the same thing and both belong in a plan.

Answered Labs interpretation of the research, not part of the source datasets
PriorityWorkWhy it sits here
HighAlignment between your pages and the questions people actually askThe largest and most robust page-level effect measured
HighPage purpose and format matched to the shape of the questionHeld up independently after alignment and depth were controlled for
HighDomain-level authority through third-party evidenceOutweighed every individual page feature, and moves the slowest
MediumSpecificity: prices, ranges, exclusions, process, timingsSupplies the facts an answer cannot be assembled without
MediumCurrency of time-sensitive informationEngine appetites for recent content differ, and stale facts are simply wrong
SupportingFAQ blocks, summaries, author information, schema, technical qualitySmall or indirect effects individually, cheap and cumulative together

Two notes on using it. The order is a default, not a diagnosis: a business whose pages cannot be crawled has a technical problem sitting above everything in the table, and a business already aligned and well referenced should be spending its time further down. And nothing here is measurable without a baseline, which means a fixed set of questions, run repeatedly, recorded.

What the evidence changes, in the end, is what you stop doing. It does not describe a new discipline. It suggests that a considerable amount of current AEO activity is concentrated in the layer with the smallest measured effects, and that the two things which mattered most, saying something specific about the right question and being referenced by other people, are the two things no checklist can complete for you. The practical version of that, for a single assistant, is in how to improve your visibility in ChatGPT.

Sources

  1. Discovered Labs (2026), What actually drives AI citations: a statistical analysis of 2M AI citations across 10K pages, the source of the alignment, page depth, page format, freshness and domain authority figures above.
  2. Semrush (2026), Why 62% of AI citations don’t lead to brand mentions, the source of the citation and mention split above.
  3. Google Search Central, Understanding page experience in Google Search results, on why performance work continues to matter for reasons other than citation count.
  4. Google Search Central, Introduction to structured data markup, on what structured data does and does not claim to do.

Where we describe how a system behaves without a citation, we are describing what we have observed rather than documented behaviour, and it may change.

Written by

Emerson Kauffmann

Co-founder, Answered Labs

About Answered Labs

We are an answer engine optimization agency. We test how AI systems find and recommend businesses, publish what we learn, and apply it to client work. More about us.

See where you currently appear

We will run a prompt set for your category and area, and show you who is being recommended today.