AI search
How to get your content cited by ChatGPT and AI search
9 min readPublished
How do I get my content cited by ChatGPT, Perplexity, and AI Overviews?
Write answer-first, cite credible external sources, and put attributed statistics in the text. The peer-reviewed work on generative engine optimization (Aggarwal et al., KDD 2024) found citations, quotations, and statistics were the strongest levers, and that keyword stuffing did essentially nothing. Sounding more authoritative changed nothing measurable. Evidence is the lever, not rhetoric.
Key takeaways
- Answer engines quote fragments, not articles. A sentence that needs the paragraph above it will never be lifted.
- The biggest levers are citing real sources and putting attributed statistics in the text. Both are writing changes, not engineering ones.
- Sounding more authoritative tested as no change at all. Evidence moves the needle; confidence does not.
- Google says outright that no special schema is needed for its AI features. Ship schema for feature eligibility and entity clarity, not for a citation boost nobody has demonstrated.
- Being cited and ranking are now separate outcomes. Dashboards built on rank-to-clicks will tell you a page is dying while it is being read to thousands.
Generative engine optimization, or GEO, is the work of getting your content quoted by AI answer engines: ChatGPT, Perplexity, Google's AI Overviews, Gemini, Copilot. It shares foundations with SEO and chases a completely different event. SEO wants a ranked link that earns a click. GEO wants a sentence of yours to show up inside somebody else's answer, with your name attached.
Almost all of the work is editorial. You can do it this week without filing a single engineering ticket.
You can also do most of it wrong, because this field has more confident advice than evidence. I have repeated some of that advice myself, in an earlier draft of this article, which is a large part of why the corrections below are here.
Three levers, and one that does nothing
The levers are about writing, not infrastructure. The one peer-reviewed study anybody actually cites here is Aggarwal et al., published at KDD 2024, and it tested a pile of content modifications to see which ones made generative engines quote a page more. Three won: cite your sources, add quotations, add statistics. Keyword stuffing, the reflex we all spent fifteen years building, produced what the authors politely call "little to no improvement".
Fifteen years. Little to no improvement.
Do these first
- Cite credible external sources
- Add attributed statistics
- Answer-first under every heading
Worth the work
- Rewrite passages to stand alone
- Build genuine subject depth
- Track citations directly
Skip entirely
- Keyword density tuning
- Thin content at volume
Do it, but not for citations
- Article and FAQPage schema
- Crawlability and render checks
Write so a paragraph survives being ripped out of context
The core skill in GEO is writing passages that stand on their own. An engine does not quote your article. It quotes a fragment, stripped of the heading above it and the example below it, and hands that fragment to somebody who will never see the rest. Any sentence whose meaning depends on the missing material is invisible, no matter how well it reads in place.
Your best paragraph might be unquotable. That is a hard thing to accept about your own writing.
- Name the subject in the sentence instead of leaning on a pronoun that points three paragraphs back.
- Put the answer in the first sentence under each heading, then support it. Never build to a conclusion.
- One claim per paragraph, so a retriever does not have to guess which half you meant.
- Keep the number and its source in the same sentence. A statistic separated from its citation is one an engine is right to distrust.
- Cut openings like 'as we saw above'. They explicitly mark a passage as unliftable.
Sounding confident buys you nothing
Turning up the persuasion does not get you cited. The GEO authors tested exactly that. They rewrote content in a more authoritative, confident register, measured it, and found no significant change. The engines did not care.
You will nonetheless find it stated everywhere that promotional writing carries a 26% citation penalty. That number does not come from the paper. It traces back to vendor correlation studies, and I could not find a primary source for it anywhere. I had it in this article for a week.
So the honest version is duller than the myth and considerably more useful. Rhetoric is neutral. Evidence is the lever. Every hour spent making a paragraph sound more authoritative is an hour not spent giving it a number and a source, and only one of those two things measured.
Write like documentation anyway. Say what the thing does, state the limits, name the cases where it is the wrong choice. Not because superlatives are punished. Because a page full of them almost never has a fact in it, and facts are the part that gets quoted.
Admitting a boundary is also the cheapest credibility available. It costs one sentence.
Schema: useful, and not what you were told
Google says plainly that you need no special markup to appear in its AI features. The exact wording in its own documentation is that there is "no special schema.org structured data that you need to add". Read that twice, then reread the last five posts you saw promising a specific percentage lift from schema. Every one of them is selling a vendor correlation.
Ship it anyway, for the two things it genuinely does: it makes you eligible for specific search features, and it disambiguates your entity so an engine is confident which company it is talking about. Those are real. "Schema gets you cited by ChatGPT" is not established, and the honest move is to say so rather than repeat it.
The retirements confused people here. Google restricted FAQ rich results to government and health sites in August 2023, then retired them completely on 7 May 2026. HowTo went earlier: announced 8 August 2023, gone from desktop by 13 September. Plenty of teams took that as a cue to rip the markup out of their templates.
The display treatment died. The description of the page did not.
- Article schema with honest datePublished and dateModified values, so freshness is legible.
- FAQPage schema whose questions mirror the visible page exactly.
- HowTo schema on genuine procedures, matching the steps a reader can actually see.
- BreadcrumbList, so the page's place in the site is unambiguous.
Your traffic dashboard is now lying to you
AI answers broke the chain from rank to clicks, so a traffic decline no longer means what it used to. A page can be quoted in an answer read by thousands of people and register one visit. Sometimes none.
A team still watching rank-to-clicks will read that as failure, and because the number really did fall, and because somebody has to explain the number in a meeting on Thursday, the page will end up on a list of underperforming content, and then on a list of content to consolidate, and then it will be gone, and the answer engine that was quoting it every day to a few thousand people will quietly start quoting a competitor instead. Nobody will connect the two events. They are four months apart.
- Track citation directly: ask the engines your buyers' questions on a schedule and record which sources get named.
- Expect impressions and clicks to decouple. Stop treating the ratio as a quality signal.
- Weight branded search and direct traffic more heavily, because a citation often produces a name lookup rather than a click.
- Keep publishing where engines actually crawl and quote, which in practice means your own site plus the discussion platforms that get indexed.
Sheevook scores every draft on these levers before you publish, and tells you when the platforms you picked are ephemeral enough that a GEO score there means very little. The scoring is deterministic and runs offline, so it costs nothing and runs on every draft rather than on request.
What is generative engine optimization (GEO)?
GEO is the practice of structuring and writing content so AI answer engines such as ChatGPT, Perplexity, and Google's AI Overviews retrieve it and cite it in their answers. It overlaps with SEO but targets citation inside a generated answer rather than a ranked link that earns a click.
Is GEO different from SEO?
They share foundations and diverge sharply on emphasis. Both need crawlable, relevant, credible content. GEO puts far more weight on answer-first structure, self-contained passages, and cited statistics, and effectively no weight on keyword density, which the research found ineffective for generative engines.
Does FAQ schema still help now that FAQ rich results are retired?
Keep it, but be clear about why. Google restricted FAQ rich results to government and health sites in August 2023 and retired them entirely on 7 May 2026, so the display payoff is gone. The schema still describes your page's question-and-answer structure for any parser, and Google states no special schema is required for its AI features, so treat it as hygiene and entity clarity rather than a citation lever.
Does promotional writing get penalized by AI search?
There is no peer-reviewed evidence that it does. The GEO paper tested a more authoritative, persuasive register and found no significant change in visibility. The frequently cited 26% penalty for promotional tone comes from vendor correlation studies rather than controlled research. Write plainly because it forces you to include facts, not because superlatives are punished.
How long does GEO take to show results?
It depends on how often each engine recrawls and refreshes its index, which varies by engine and by site authority. The practical move is to treat citation itself as the metric and check it directly by asking the engines your buyers' questions on a regular schedule, rather than waiting on a traffic number that may never move the way it once did.
Can I just use AI to write GEO content at scale?
Not safely on its own. The levers that drive citation are specific sourced facts, real statistics, and genuine admitted boundaries, which are exactly the things a model will approximate if left unsupervised. Unreviewed AI content published at scale also carries real risk under search quality guidelines. Use AI to draft and structure, and supply the facts and sources yourself.
Sources
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization (KDD 2024, arXiv:2311.09735)
- Google Search Central, AI features and your website ("no special schema.org structured data that you need to add")
- Google Search Central, FAQPage structured data (retirement note, 7 May 2026)
- Google Search Central, changes to HowTo and FAQ rich results (8 August 2023)
- Sheevook's research brief on generative engine optimization, the practices behind this article
- Sheevook's 2026 review of the search landscape, covering AI answers and their effect on clicks