Schema and llms.txt assume the engine reads your site. Often, it does not.
The most common technical recommendations for AI visibility are schema markup, an llms.txt file, and structured on-page data. These do real work when an AI engine is retrieving and reading your pages directly. The problem is that in many categories, the engine never retrieves your site. It is reading the directories, review platforms, and community forums that came back in its search step. A fix aimed at your own website cannot reach pages the engine is already pulling from somewhere else.
How engines retrieve pages before writing an answer
When someone asks ChatGPT, Perplexity, Claude, or Gemini a buyer-intent question, the engine does not answer from memory. It runs a search query first, pulls a set of pages from the results, and generates its answer from that retrieved content. The set of pages it pulls is determined by what that search query returns, not by anything you have done on your own site.
For a query like "who is the best internet provider near me," the search results are dominated by aggregator directories and community discussions. Yelp, Nextdoor, local Facebook groups, Reddit threads. The brand's own website, with its schema and llms.txt, may not appear in those results at all. If the engine never retrieves the page, it cannot read what is on it.
This is a retrieval problem, not a parsing problem. Schema and llms.txt address parsing: they improve how an engine reads a page it has already decided to fetch. They do nothing to change whether the engine fetches the page in the first place.
What the citation sources actually show
When we extract the citation URLs from a prompt panel, the list shows exactly which pages each engine retrieved and cited. In an audit for a regional ISP across 184 AI responses, five external URLs accounted for most of the citations. A single Facebook group post was cited 23 times. The brand's own domain was not in the citation list.
Adding schema to that brand's website would improve how Googlebot and AI crawlers parse the pages if they visited. It would do nothing about the 23 citations to a Facebook post that the brand does not control, or the directory listings the engines kept returning to.
The citation source list is the useful diagnostic. When your own domain appears in it, on-site fixes are in play. When it does not appear, the leverage is elsewhere.
When on-site technical fixes do move the needle
On-site improvements are not useless. They work in categories and query types where the engine's search step actually surfaces individual brand websites. Common cases include:
- B2B software queries, where product landing pages and documentation rank in search and get retrieved
- Informational and comparison queries, where the engine is pulling from articles, reviews, and guides that may include your content
- Branded queries asking specifically about your company, where your own site is the obvious source the engine reaches for
In these cases, schema markup helps an engine extract your product name, description, pricing, and reviews cleanly from your page. An llms.txt file gives AI crawlers a map of what your site contains and which sections matter. Both are worth doing. They just do not help with the queries where your site is never in the retrieved set.
The per-engine wrinkle
Different engines retrieve from different sources for the same query, which means the same on-site fix can matter for one engine and not another. We have seen ChatGPT and Claude pull more from brand websites than Perplexity does for local service queries. Perplexity, in the same panel, pulled heavily from Facebook, which ChatGPT cited zero times across 46 prompts.
That means a schema improvement could lift your ChatGPT citation rate for certain query types while doing nothing for Perplexity, because Perplexity is reading a different set of sources for those queries. The per-engine split in a citation audit shows you which engines are already retrieving your site and therefore where on-site work can compound.
Matching the fix to where the engine looks
The practical prioritization follows directly from the citation data. Pull the top cited domains for your category across each engine. For each domain where you are absent, ask whether a presence there is achievable: a complete business profile, a review listing, a post in a forum the engine indexes.
For local and regional services, that list typically includes directories (Yelp, Angi, HomeAdvisor), community platforms (Facebook groups, Nextdoor), and local news sites. Getting a complete, accurate, regularly-reviewed presence on those pages addresses the actual retrieval problem. Schema on your own site is the second step, not the first.
For B2B categories, the list looks different: G2, Capterra, Trustpilot, industry publication roundups. The fix playbook is the same: be present on the pages the engine is already reading, then optimize what those pages say about you, then improve how your own site is parsed when the engine does reach it.
A note on Gemini
Gemini routes its citations through vertexaisearch.cloud.google.com rather than returning specific domain URLs. Every citation in a Gemini response points to Google's own intermediary, not to the source page. This means the per-source fix playbook cannot be applied to Gemini the same way. You cannot audit which pages Gemini pulled from for a given answer just by looking at its citations. Google Search ranking and presence in Google Business data are the nearest proxies, but the direct citation-source analysis that guides remediation for other engines does not work the same way for Gemini.
The practical sequence
Run your citation audit and pull the per-engine source lists. For each engine, identify which domains appear most often and check your presence there. Close the gaps on the high-citation platforms first. Then look at whether your own domain appears in any engine's citation list, and if it does, apply on-site improvements for those engines. Schema, llms.txt, and structured data belong in the plan, but at step two, after you know which engines are actually reading your site.
A CitedWell audit extracts the citation URL list for your category across all four engines, so you can see exactly which pages are being read and where your brand is absent.
Order an audit, $490