AI Citations
What Makes AI Cite a Page?
What large-scale citation research suggests about content alignment, page format, authority, freshness and the sources different AI systems choose.
What does it mean for an AI system to cite a page?
A citation is a source an assistant attaches to an answer as the evidence for something it said. It is not a recommendation, it is not the same as having your brand named, and it is not a ranking position. Treating it as any of those things is the fastest way to misread the research in this article.
It helps to separate four outcomes that get collapsed into the word visibility. A page is retrieved when a system pulls it into the working set of material it considers. It is cited when it survives that set and appears as a source beside the answer. A company is mentioned when its name appears in the answer text. It is recommended when the answer puts it forward as the thing the reader should choose.
These come apart constantly. Semrush's 2026 analysis of nearly four thousand domain appearances found that 61.7% of them were citations without an accompanying brand mention, and a further 25.1% were mentions with no citation. Only 13.2% were both. So a page can be the evidence behind an answer that never names its author, and a company can be named in an answer that cites somebody else entirely.
This article is about the narrowest of the four: what characteristics of a page appear to accompany more citations. That is a genuinely useful question, because citation is the outcome most directly tied to what is on a page. It is also the outcome least directly tied to revenue, which is why the wider framing in our introduction to AEO and the measurement discipline in the AEO and SEO comparison both still matter.
What this research can and cannot tell us
It can tell us which characteristics of a page tend to accompany more citations across a large sample of pages that were actually cited. It cannot tell us how any assistant ranks sources, because none of them publish that, and none of the studies discussed here had access to it.
That distinction is worth holding onto, because most published AEO advice is either pattern-matched from search engine optimization or generalised from a handful of screenshots. Observational research on a large sample is a real improvement on both. It is still observational.
Almost every quantitative claim in this article comes from one study: Discovered Labs' 2026 analysis of roughly 2 million citation observations across four AI engines over a six-month window, joined to crawled and feature-engineered data on 10,000 of the cited pages. It is, as far as we know, the largest published analysis of citation predictors. It is also a single study of one population, and single studies are how fields end up confidently wrong.
- Correlation is not mechanism. A feature that accompanies more citations may be causing them, may be a proxy for something else, or may be a consequence of the same editorial care that produced the citations.
- Controls help, and they are not proof. The stronger findings below survive a multivariate model with domain fixed effects, content depth, page type and page age controlled for. That removes the most obvious confounders, not all of them.
- The population is specific. The dataset is oriented toward business-to-business subject matter and four particular engines. A local trades business or a consumer retailer is not obviously the same population.
- The systems change. These are measurements of products that were updated during the capture window and have been updated since. A coefficient is a photograph, not a law.
- Only cited pages were crawled. Feature analysis on a sample of pages that already earned citations tells you what distinguishes them from each other. It says less about what separates them from the pages that were never cited at all.
The honest summary is that this evidence should change what you prioritise, and should not be quoted as if it were documentation. Where a finding below is strong, we say so. Where it is one number from one model, we say that too.
Content alignment appears to matter most
Of everything measured, how closely a page's language and concepts matched the way people actually prompt was by far the strongest page-level predictor of citation count. Its standardised effect was roughly three times the size of the next-strongest page feature.
The reported coefficient is +0.37, with a 95% confidence interval of +0.33 to +0.41. Within the study's model, a one standard deviation increase in alignment corresponded to roughly 30% more citations. The next three features on the list, in order, were FAQ sections at +0.07, a short summary or bottom-line-up-front block at +0.05, and an author bio at +0.02.
Empirical
Prompt alignment was the dominant page-level signal
Standardised effect on citation count
Alignment was also the most robust finding in the study rather than merely the largest. It remained significant after false discovery rate correction, was selected in every one of 200 stability-selection bootstraps, and survived a double machine learning treatment that strips out confounder variance, coming through at +0.35. Most single features in a study this size do not clear all of that.
That combination is what makes it worth acting on. A large coefficient that appears in one specification and vanishes in the next is a curiosity. A moderate coefficient that survives every attempt to break it is a finding.
What alignment actually means
Not keyword matching. A page is aligned when its language, concepts and factual coverage correspond to the questions people actually ask, at the level of detail they ask them.
The difference is easiest to see in a real buying question. Somebody evaluating compliance software asks how much SOC 2 tooling costs for a fifty-person company. An aligned page states a price or a defensible range, says what team size the range assumes, names what is in and out of scope, and explains what the implementation actually involves. A page headed with an invitation to transform your compliance journey covers none of that, and no amount of repeating the phrase SOC 2 changes it.
Alignment therefore starts with knowing the questions rather than with writing. The cheapest sources are the ones already in the business: recorded sales calls, the questions support answers repeatedly, the objections that come up before a decision, the comparisons prospects raise unprompted. A keyword tool will give you phrasings. It will not give you the fifty-person qualifier, which is the part that made the question specific.
Two failure modes are worth naming. The first is answering a more general version of the question, which is what happens when a page is written from a keyword rather than from a conversation. The second is answering the right question in language nobody uses, which is common in regulated and technical categories where the internal vocabulary and the customer vocabulary have drifted apart. Both look like coverage on a content plan and neither is aligned.
Where on the page the answer sits
The paragraph engines cited most often sat at a median depth of 0.36 down the page, where 0 is the top and 1 is the bottom. Roughly speaking, the best matching passage tended to be about a third of the way in.
Empirical
The strongest matching passage tended to appear relatively early
- 0.0Top of the page
- 0.36Median depth of the paragraph engines cited most often
- 1.0Bottom of the page
A median, not a cut-off. Half of the best-matching paragraphs in the sample sat deeper than this line.
The tempting conclusion is that AI reads only the first third of a page. That is not what a median depth of 0.36 says. Half of the best-matching paragraphs sat deeper than that, and some sat a long way deeper. What the number does suggest is that pages which bury their strongest passage are competing against pages that do not.
There is also a simpler explanation available, and it is worth preferring: well-written pages tend to put their best material early, because that is how you keep a reader. The statistic may be measuring editorial competence as much as machine behaviour. The practical advice happens to be the same either way, which is a comfortable position to be in.
So the useful version of this is not a rule about the top third. It is a question to ask about any page you care about: if somebody quoted the single most useful paragraph on this page, where would it be, and is there a reason it is not closer to the beginning? Setup, background and qualification are often the reason, and they are usually movable.
Structural additions help, but much less than alignment
FAQ blocks, short summaries and author information all showed positive effects on citation count, and all of them were small. The largest, an FAQ section at +0.07, is under a fifth of the alignment effect and does not reach the conventional threshold for a small effect at 0.10.
This is the part of the study that most contradicts current practice, because the standard AEO checklist is largely made of these items. They are worth doing. They are not worth doing first, and they will not rescue a page that answers the wrong question.
The more interesting results in this part of the analysis were the ones that came back empty. After controlling for domain, content depth, page type and page age, the study found no significant independent effect on citation count from:
- real-user Core Web Vitals, which collapsed into latent factors showing no significant effect once domain was controlled for
- synthetic Lighthouse performance scores
- schema markup, which the authors suggest contributes indirectly through content structure rather than as a standalone citation signal
It is important to read that correctly, because it is easy to turn into three bad conclusions. It does not mean speed does not matter, that structured data is useless, or that technical work can be skipped.
What it means is narrower. Among pages that were already being cited, these attributes did not independently predict how often. That is entirely compatible with them mattering earlier in the chain: a page that loads badly or renders only after scripts execute may not be reliably retrievable in the first place, and structured data removes ambiguity about what a page is describing. Neither of those effects would show up as a citation-count coefficient in a sample of pages that had already cleared the bar.
There is also a confounding explanation the study itself gestures at. Schema markup, decent performance and a clear author byline tend to co-occur with a competently run website. Once domain is controlled for, much of what those signals were standing in for has already been absorbed by the domain term. The signal has not disappeared. It has been reassigned.
The practical position, then: treat this layer as hygiene rather than strategy. Implement it consistently, keep it accurate, and stop expecting it to move a number on its own.
Page format changes the odds
What kind of page you publish mattered independently of how well it was written. In a controlled model that already accounted for domain, content depth, alignment and page age, pricing pages carried a coefficient of +0.39 while listicle reviews sat at -0.12.
Empirical
Page format still mattered after controlling for other factors
- Pricing pages
- +0.39
- Listicle reviews
- -0.12
Coefficient on citation count, against the reference category
Comparison and how-to formats sat near the reference category in the published analysis and are not plotted, because no coefficient was reported for them numerically.
The likely mechanism is structural rather than mysterious. A pricing page is shaped like the answer to a transactional question. A comparison page that names alternatives and puts their differences side by side is shaped like the answer to an evaluation question. A listicle of loosely ranked options is shaped like something written to occupy a search result, and it tends to contain less of substance per paragraph than either.
The wrong response to this is to convert your content plan into pricing pages. Format only helps when the format matches a question somebody is asking, and a pricing page for a service nobody prices, or a comparison page that compares nothing meaningfully, inherits none of the effect. The finding is about correspondence between the shape of a page and the shape of a question, which is really the alignment finding again, viewed from a different angle.
The right response is narrower and more useful: check whether the commercially important questions in your category have a page of the appropriate shape at all. Many businesses have twenty blog posts and no page that states what anything costs.
Why specific commercial information carries weight
Because it is the material an answer cannot be assembled without. A general-purpose model can produce a competent paragraph about what dental implants are or how SOC 2 audits work. It cannot produce what you charge, how long your waiting list is, or which of your tiers includes a named feature.
That makes a short list of facts disproportionately valuable, and they are usually the facts businesses are most reluctant to publish: prices or honest ranges, what is included and excluded, eligibility and prerequisites, realistic timescales, what the process involves step by step, what you do not do, and how your version differs from the two or three alternatives a buyer is actually weighing.
The reluctance is understandable and usually mispriced. Withholding a price does not keep you in the conversation, it removes the page from the set of pages that can answer the question. If a genuine range is impossible, the useful move is to say why and give the variables: a page explaining that implant costs run between two figures depending on bone grafting, sedation and the number of units is more usable than either a fixed number or silence.
Engines disagree about freshness
There is no single answer to how fresh content needs to be, because the four engines studied had measurably different appetites. The median age of cited content ran from 5.1 months on Claude to 8.0 months on ChatGPT.
Empirical
Different engines favour different content ages
Median age of cited content, months
The same divergence shows up in the shape of the decay rather than only in its midpoint. Claude drew 60% of its citations from content less than six months old. ChatGPT drew 40% from the same band, meaning the majority of what it cited was older than half a year.
Empirical
Claude skewed more heavily toward recent content
Share of the engine's citations on content under six months old
The conclusion this most invites is a refresh cadence, and that is the conclusion to be most careful about. Nothing here shows that republishing a page with a new date earns citations. What it shows is that the pool of content these engines cited skewed toward the recent, more sharply on some than others.
The defensible reading is narrower than a cadence and more useful. Content whose facts move should not be allowed to become visibly stale, because a page with last year's prices, a discontinued service or a superseded regulation is wrong rather than merely old. Content whose facts do not move needs no schedule at all: an explanation of how a procedure works does not improve by being touched every quarter.
So the question to ask about each page is not when it was last updated but whether anything on it has become untrue. That distinction is what separates maintenance from the practice of rewriting timestamps, which costs the same and achieves nothing.
Authority sits above the page
The strongest predictor in the whole analysis was not a property of the page at all. A domain-level authority feature carried a mean absolute SHAP value of 0.38, against 0.06 for the strongest individual page-level feature other than alignment: roughly six times the influence, within that model.
Empirical
Domain-level authority dominated individual page signals in this model
Mean absolute SHAP value
If the finding holds, it is the most commercially consequential one in the study, and the least convenient. It says the ceiling on what page-level work can achieve is set somewhere above the page, and that a challenger competing against a well-established domain is not going to close the gap with better formatting.
It also explains a pattern that otherwise looks arbitrary: a thinly written page on a widely referenced domain being cited in preference to a better page on an unknown one. That is not a judgement about the two pages. It is the domain term doing most of the work.
The caveat matters as much as the number. This is one model's attribution of its own predictions, using a feature its authors built. A different construction of authority would produce a different number, and possibly a different ranking. What survives the caveat is the direction: domain-level evidence appears to be upstream of page-level effort, which is consistent with how these systems are described and with what we see in client work.
What authority means in practice
Not a score, and nothing you can buy. In this context authority is an accumulation of independent evidence that a business exists, does what it says, and is regarded in a particular way by people who are not it.
In practice that means being referenced rather than only publishing: coverage in trade or local press, entries on professional registers and licensing bodies that actually govern your category, accurate listings on directories a human would consult, citations in other people's research, mentions from suppliers, partners and clients, and reviews with enough text to describe something specific. What these have in common is that a third party chose to say something about you.
The uncomfortable part is the timescale. Alignment work can change a page this week. Authority accrues over quarters, mostly as a side effect of operating visibly and doing things worth mentioning, and there is no version of it that runs faster because you spent more. The realistic strategy for a smaller business is not to win the authority contest but to compete where it is narrower: a specific service, a specific city, a specific question where the established domains have nothing precise to say.
A source can matter to one engine and barely register in another
Engines do not draw on the same web. In the same dataset, 97% of the LinkedIn citations observed came from Google AI alone. Claude and ChatGPT each drew under 0.4% of their third-party citations from LinkedIn, and Gemini produced none at all in that window.
That figure needs reading precisely, because it is easy to inflate. It is not LinkedIn's share of all citations, and it does not make LinkedIn important. It says that among the citations of LinkedIn that were observed, almost all of them came from one engine. The platform was close to irrelevant to the other three.
The same pattern shows up in how much weight each engine gives to a company's own pages. In this dataset 39% of ChatGPT's unique cited URLs were brand-controlled, against 14% for Gemini, which is a large difference in how much of the answer a business can supply itself. We look at that split, and what to do about it, in the guide to ChatGPT visibility.
The practical consequence is that a source strategy built by watching one assistant will be miscalibrated for the others. It is also a reason to be sceptical of any list of the platforms AI cites: the list is engine-specific, and the only reliable version is the one you build by reading the sources actually shown for your own questions.
There is no universal AEO checklist
The evidence supports a rough order of priority, not a list of tasks that produces citations when completed. The order is more useful than the list, because most of the items only work when the ones above them are already in place.
Read from the top. Answering the right questions is the precondition for everything else. Specific, checkable content is what makes an answer possible to assemble. Using a page format that matches the shape of the question multiplies content that is already aligned. Domain-level authority sets the ceiling and moves slowest. Keeping time-sensitive information current protects what you have. Structural additions and technical quality are cumulative and cheap, and they are the last place to look for a result.
None of that reduces to a checkbox, which is precisely the problem with how AEO is currently sold. A page with impeccable schema, an FAQ block, a fast load time and a named author, answering a question nobody asks, on a domain nothing references, will not be cited. Every item on the standard checklist is present. The two things that mattered most are missing.
What a citation-ready page looks like
A page that can be cited is one where a specific, checkable claim can be lifted out and attributed without distortion. That is a higher bar than being well written, and a different one from being optimised.
Conceptual
The parts of a page that make it quotable
A descriptive H1 naming the exact question the page answers
- Direct answerThe question resolved in the opening paragraph, in the words a buyer would use
- Specific factsPrices or ranges, timings, eligibility, exclusions and anything a general model could not know
- EvidenceFirst-hand experience, original data or named sources supporting the claims made
- ComparisonHow this option differs from the alternatives a reader is actually weighing, named
- Frequently asked questionsOnly where real questions genuinely remain after the body of the page
- Entity detailsWho publishes this, where they operate, and details that agree with every other source
- Current informationA visible date and a real review cycle, on the pages whose facts actually move
The test that catches most problems is extraction. Take the single most useful paragraph on the page and read it on its own. If it still means what it meant in place, it can be cited. If it collapses because it depended on the previous three paragraphs to establish what it was referring to, it cannot, and a system looking for an answer will use somebody else's paragraph instead.
In practice failures are almost always one of the same few things. The subject is a pronoun. The number has no unit or currency. The claim is hedged into meaninglessness. The geography is implied by an address in the footer. The comparison names no alternative. All of these are cheap to fix and none of them require writing for machines, which is the part most AEO advice gets backwards: extractable prose is just prose that says what it means.
What the evidence does not justify
Several claims currently circulating are either unsupported by the research above or directly contradicted by it. They are worth stating plainly, because most of them are being sold.
- That schema markup makes you appear in ChatGPT. Schema showed no significant independent effect on citation count in this analysis. It reduces ambiguity, which is a reason to implement it accurately, not a mechanism for inclusion.
- That an FAQ block earns citations. It was the strongest of the structural signals and its effect was +0.07, below the conventional threshold for a small effect. It is a small positive, not a lever.
- That every page needs constant rewriting. The freshness findings describe what engines cited, not a publishing schedule. Pages whose facts do not move gain nothing from being touched.
- That length wins. Content depth was a control in these models rather than a headline finding, and adding words to a page that answers the wrong question changes nothing about which question it answers.
- That putting the answer first guarantees inclusion. A median matching depth of 0.36 is an observation about where good passages sat, not a rule about what gets read.
- That domain authority is a score you can buy. The authority feature here was constructed by the researchers from observed evidence across the web. It is not a purchasable metric and it is not any vendor's published score.
- That a citation is a recommendation. The majority of citations in Semrush's dataset came with no brand mention at all. Being the evidence behind an answer is not the same as being the answer.
The common thread is scale. Most of these claims take a real but small effect, or a descriptive statistic, and promote it to a mechanism. The research does not support mechanisms. It supports priorities.
How to apply the evidence
The sequence below is our reading of the research rather than part of anybody's dataset. It is ordered by how much the evidence suggests each item matters, and by how long each one takes to move, which are not the same thing and both belong in a plan.
| Priority | Work | Why it sits here |
|---|---|---|
| High | Alignment between your pages and the questions people actually ask | The largest and most robust page-level effect measured |
| High | Page purpose and format matched to the shape of the question | Held up independently after alignment and depth were controlled for |
| High | Domain-level authority through third-party evidence | Outweighed every individual page feature, and moves the slowest |
| Medium | Specificity: prices, ranges, exclusions, process, timings | Supplies the facts an answer cannot be assembled without |
| Medium | Currency of time-sensitive information | Engine appetites for recent content differ, and stale facts are simply wrong |
| Supporting | FAQ blocks, summaries, author information, schema, technical quality | Small or indirect effects individually, cheap and cumulative together |
Two notes on using it. The order is a default, not a diagnosis: a business whose pages cannot be crawled has a technical problem sitting above everything in the table, and a business already aligned and well referenced should be spending its time further down. And nothing here is measurable without a baseline, which means a fixed set of questions, run repeatedly, recorded.
What the evidence changes, in the end, is what you stop doing. It does not describe a new discipline. It suggests that a considerable amount of current AEO activity is concentrated in the layer with the smallest measured effects, and that the two things which mattered most, saying something specific about the right question and being referenced by other people, are the two things no checklist can complete for you. The practical version of that, for a single assistant, is in how to improve your visibility in ChatGPT.
Sources
- Discovered Labs (2026), What actually drives AI citations: a statistical analysis of 2M AI citations across 10K pages, the source of the alignment, page depth, page format, freshness and domain authority figures above.
- Semrush (2026), Why 62% of AI citations don’t lead to brand mentions, the source of the citation and mention split above.
- Google Search Central, Understanding page experience in Google Search results, on why performance work continues to matter for reasons other than citation count.
- Google Search Central, Introduction to structured data markup, on what structured data does and does not claim to do.
Where we describe how a system behaves without a citation, we are describing what we have observed rather than documented behaviour, and it may change.
Keep reading
Related
A working definition, how it differs from SEO, what a business can actually influence, and what nobody can promise.
Read→ ChatGPT How to Improve Your Visibility in ChatGPTWhat visibility in an assistant actually means, why the answer changes, and the work that makes being named more likely.
Read→ AEO AEO vs SEO: What's Actually Different?The differences are real but narrower than the marketing suggests, and the overlap is where most of the value sits.
Read→See where you currently appear
We will run a prompt set for your category and area, and show you who is being recommended today.