Start a Project
AEO & GEO

Why Isn’t AI Citing Your Website?

9 min read
Eva Okhotnik

Eva Okhotnik

Marketing Executive, Web25

Why Isn’t AI Citing Your Website?

I started digging into why some pages get quoted by AI tools and others don’t, expecting a fairly clean answer. It isn’t one. The honest short version: it depends on which AI tool you’re asking about, whether that tool even looks at the live web for your kind of question, and whether your page gives it something clean enough to lift. Generic content is part of the problem, but it’s not the whole story, and I want to walk through the parts that usually get left out.

AI tools don’t all work the same way

This is the part that trips people up first. “AI citing my website” isn’t one behavior, it’s several different systems doing different things.

Perplexity is built around live web retrieval for almost every answer, so it’s constantly pulling in and citing current pages. Google’s AI Overviews work off Google’s own index, similarly retrieval-based. ChatGPT is different depending on the mode: plain conversational answers often come from what the model already learned during training, with no live citation at all, while its search-enabled mode does fetch and cite pages. Gemini sits somewhere in between depending on the product surface you’re using.

So the first real question isn’t “why isn’t AI citing me,” it’s “does this particular tool even search the web for this kind of question.” If it doesn’t, no amount of rewriting your page changes anything, because nothing on the live web is being consulted in the first place. That’s also the core idea behind answer engine optimization: getting content structured for the tools that do retrieve and cite, since not every AI product works that way.

When a tool does retrieve from the web, a page still has to clear a few more hurdles before it shows up: it has to be crawlable and indexed to begin with, it has to be relevant to the exact question being asked, not just the general topic, and its answer has to be extractable, meaning a system can pull out a clean fact without wading through paragraphs of context to find it. Source credibility and how recently a page was updated matter too, though how much they matter varies by tool and by topic.

What the Princeton research actually measured

A study that gets cited constantly in this space is “GEO: Generative Engine Optimization” (Aggarwal et al., presented at KDD 2024, researchers from Princeton, IIT Delhi, Georgia Tech, and the Allen Institute for AI). It’s worth being precise about what it actually tested, because the summarized version floating around online overstates it.

The researchers built a controlled experimental setup: for a given query, they took the top five sources Google already returned, fed all five into a generative model as fixed context, and then measured how visible each source was in the model’s synthesized answer using a metric based on word count and position. They then tested nine ways of rewriting a source, adding statistics, adding quotations, citing outside sources, adjusting tone, and so on, to see which changes shifted that visibility score.

In that setup, adding relevant statistics, quotations, and credible citations improved source visibility. Some improvements reached approximately 40%, though results varied by topic and by which tactic was used. That’s a real, useful finding. It is not the same as saying any page that adds a statistic will see its real-world AI citation odds jump by 40%. The study never tested whether a source gets discovered, crawled, or selected into that initial group of five in the first place, and it didn’t test schema markup at all. It measured how a source performs once it’s already sitting in front of the model, which is a narrower and more specific claim than “gets you cited more.”

What a vague statement looks like next to a specific one

Here’s a made-up but realistic pair, meant to illustrate the difference rather than report on a real page.

A vague version: “Our maintenance plans keep your website running smoothly and protect it from problems.”

A specific version: “Our Care plan runs weekly plugin and theme updates on a staging copy of your site before anything touches production, checks for broken links after every update, and sends a monthly report showing exactly what changed.”

Both sentences are making a similar promise. Only the second one gives an AI system, or a human reader, anything concrete to repeat. It names a process, a frequency, and a safeguard. The first sentence could describe almost any maintenance service on the internet, which is exactly the problem: there’s nothing in it that distinguishes one source from another, so a model summarizing several competing sources has no reason to reach for that one specifically.

What this means in practice

Based on what the research does support, plus how retrieval-based AI tools generally work, a few things are worth doing regardless of which specific AI platform ends up mattering most for your business.

Write content that would hold up as a specific answer to a specific question, not a general statement about a general topic. Where you have real detail, real numbers, or a real process to describe, use it instead of a broader claim. And keep the basics in place, a page has to be indexable and genuinely relevant to a question before any of this matters.

One limitation worth naming honestly: none of this guarantees a citation. AI answer engines change their retrieval and ranking behavior often (the whole search landscape is moving fast, which we’ve written about separately in whether SEO itself is dying), the research on this is young, and a tactic that helps in one controlled study doesn’t automatically translate to every platform or every query. Treat this as informed judgment, not a formula.

Web25 works on this kind of content restructuring, alongside classic SEO and AEO/GEO work. If you want a second opinion on whether a specific page of yours holds up under this kind of scrutiny, that’s a conversation worth having with us.

A few things people ask about this


No. In the Princeton study’s controlled setup, adding statistics improved a source’s visibility once it was already among the sources being synthesized, by as much as 40% in some cases. That’s different from a guarantee about real-world citation frequency, which depends on discovery, crawling, and relevance first.


It may not be retrieving from the live web for that kind of question at all. Plain conversational ChatGPT answers often come from what the model learned during training, not a live search, so no page gets cited regardless of how it’s written. Its search-enabled mode behaves differently.


It’s good practice for making a page’s structure clear to any automated system, including search engines, but the Princeton study specifically didn’t test it, so I can’t point to that research as evidence either way for AI citation specifically.


It can be, on topics where freshness matters to the question being asked. It’s one signal among several, not a rule that applies evenly across every topic.


Written by Eva Okhotnik, Marketing Executive, Web25. Last updated September 2026.

More from the blog