Shruwd
Shruwd
GuidesLLM SEOGEOAI searchllms.txtschema

LLM SEO: five levers with evidence, and two without

Five LLM SEO levers with evidence behind them, graded by strength, and the reason schema markup and llms.txt aren't among them.

Devin D.
Founder

11 min read

LLM SEO: five levers with evidence, and two without

LLM SEO is the work of getting your brand named and your pages cited in answers written by large language models, such as ChatGPT and Google's AI Overviews. Five levers have evidence behind them: letting AI crawlers in, putting your content in the raw HTML, being findable in search, getting mentioned on the pages that answers draw from, and writing passages a model can quote. Two popular tactics, schema markup and llms.txt, have little or none so far.

I searched "llm seo" on 18 September 2026 (US, desktop) and checked the first five guides Google returned, skipping forum threads and videos. All five recommend schema markup. Two say nearly all the pages ChatGPT cites have it, and neither gives a source.

The evidence behind the five isn't equally strong, so each one below says what it rests on.

Find out what the answers say about you.

One measurement, your own crawler data, and the three access findings that matter most. Free, no card, about two minutes.

Start free

LLM SEO tactics at a glance

TacticWhat the evidence isHow strong
Let AI crawlers inOpenAI's, Anthropic's and Google's own documentationDocumented. You can check your site in minutes
Put your content in the raw HTMLA crawler-traffic study across one large network, December 2024Strong, and ageing
Be findable in searchGoogle's documentation, plus citation studies from Ahrefs and Seer InteractiveStrong for Google, thinner for the others
Get mentioned where answers get their sourcesCorrelation studies from Ahrefs and SemrushCorrelation only
Write passages a model can quoteOne lab test with a known flawWeakest of the five
Schema markup, for AI citationsA matched study of 1,885 pagesTested, and no lift found
llms.txt, for AI searchGoogle says its search ignores it, and in one site's logs the AI crawlers checked didn't request itNo evidence AI search reads it

Cyrus Shepard's review of 54 experiments, patents and case studies scored 23 possible factors by hand and landed in a similar place. The top three were URL accessibility, search rank, and rank in the follow-up searches an AI system runs, which it calls fan-out rank. Structured data came 20th and llms.txt came last.

What is LLM SEO?

LLM SEO means making your site and your brand easy for a large language model to find, read and cite when it answers a question. You'll see the same work called GEO, AEO or LLMO. GEO vs SEO covers how the terms differ, and mostly they don't.

A model can name you from memory, meaning what it absorbed in training, or from pages it retrieves when the question is asked. You can't edit what a model already learned. Everything below works on retrieval, the part you can change.

The five levers

1. Let the AI crawlers in

A page an AI crawler can't fetch is unlikely to be cited. OpenAI says so in its crawler documentation: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links."

GPTBot is OpenAI's separate training crawler, so you can refuse one and allow the other. Anthropic splits its crawlers the same way: ClaudeBot collects training data and Claude-SearchBot serves search.

A block can sit in three places:

  • A robots.txt rule naming the bot.
  • Bot protection on your CDN or firewall. Cloudflare has a setting that blocks AI bots, and a block there applies whatever robots.txt says.
  • For Google's AI features, the Search generative AI control in Search Console. Include is the default, and a site has to be included to appear.

To see how your site treats OpenAI's search crawler, run this in PowerShell. It requests a page under the name OpenAI publishes for its crawler. The version number in it may change.

powershell
$url = "https://example.com/pricing"
$ua  = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot"
try   { (Invoke-WebRequest -Uri $url -UserAgent $ua -UseBasicParsing).StatusCode }
catch { [int]$_.Exception.Response.StatusCode }

200 means the page was served. 403 or 429 means something refused the request. The test has a limit: some firewalls also check the crawler's IP address, so they can treat your test and the real bot differently.

To skip the typing, Shruwd's free crawler access check does the robots.txt half for every AI crawler token, and counts the readable text a homepage carries before JavaScript runs. No account.

Your server logs show what the real crawler received. Shruwd reads those logs. It verifies each hit against the vendor's published IP ranges, and flags a page that refuses a verified crawler while a normal browser gets it.

2. Put your content in the raw HTML

Vercel and MERJ analysed crawler traffic across Vercel's network in December 2024. None of the major AI crawlers rendered JavaScript. OpenAI's and Anthropic's crawlers downloaded JavaScript files and didn't run them. Googlebot does render pages.

If your content appears only after JavaScript runs, those crawlers may receive an empty shell. To check, view a page's source and search for a sentence you can see on the page. An earlier post has a PowerShell snippet that counts the text a crawler receives. Server-side rendering, static generation or prerendering fixes it.

The data is from late 2024 and comes from one hosting network. I haven't found a newer study of the same kind.

3. Be findable in search, for the follow-up questions too

AI answers that cite sources are built by running searches and reading the results. Google's documentation says a page must be "indexed and eligible to be shown in Google Search with a snippet" to appear as a supporting link in AI Overviews or AI Mode.

Ranking for the main keyword matters less than it did, though. Ahrefs found that 76% of the pages cited in AI Overviews ranked in Google's top 10 for the same query in July 2025, and 38% in March 2026. Ahrefs changed its tracking between the two studies, so don't read the drop as exact. Its reading is query fan-out: the model splits a question into several searches and cites pages that rank for those.

For ChatGPT, look at Bing as well as Google. Seer Interactive compared ChatGPT search citations with search results in February 2025. 87% matched a page in Bing's top 20 for the same question, against 56% for Google. The sample was 100 queries, and Seer calls it directional.

In practice: keep your SEO basics in order, answer the follow-up questions on the same page as the main one, and check that Bing has indexed you.

4. Get mentioned where answers get their sources

Ahrefs studied 75,000 brands in May 2025. Branded web mentions had a correlation of 0.664 with visibility in AI Overviews, where 1 would be a perfect match. Backlinks had 0.218. Its December 2025 follow-up added ChatGPT and AI Mode and found web mentions between 0.66 and 0.71, YouTube mentions strongest at about 0.74, and link metrics "very weak".

A Semrush study of 1,000 domains adds a detail. Nofollow links tracked AI mentions about as closely as followed links did, at 0.509 against 0.504. Whatever is at work here, it doesn't look like the ranking value a link passes.

These are correlations. Famous brands collect mentions and AI visibility alike, so a mention campaign may not move you. Google's own guide to its AI search features adds a warning that seeking inauthentic "mentions" across the web "isn't as helpful as it might seem."

In practice: ask your buyers' questions in ChatGPT and Google, note which pages the answers cite, and work on getting included in those. When I checked Google for 11 buying keywords in my own category in September 2026, one search each, a single third-party article was cited 4 times across the 11 AI Overviews.

5. Write passages a model can quote

This lever has the weakest evidence of the five. It rests mostly on one lab test, the 2024 paper that coined "generative engine optimization". In the authors' own test engine, adding quotations, statistics and cited sources to a page raised how much of the answer drew on it. A language model wrote those additions, though, and the prompts in the authors' published code allowed it to invent them. I went through the paper's method and its limits in an earlier post.

Google pushes the other way. Its guide says "You don't need to write in a specific way just for generative AI search", and that there's "no requirement to break your content into tiny pieces".

So don't rebuild pages for a model. Do the part that helps a reader and costs little: answer the question in the first lines under its heading, and use real numbers, names and dates. Lost in the Middle, a 2024 study of how language models read long inputs, found they often do best when the relevant information sits at the beginning or the end. That's one more reason to put the answer first.

Does schema markup help LLM SEO?

Not in the best test I know of. Ahrefs tracked 1,885 pages that added schema markup and matched them against 4,000 control pages. Citations in AI Mode and ChatGPT rose by about 2%, which was statistically indistinguishable from zero. AI Overview citations fell 4.6% against the controls, and Ahrefs says it can't clearly attribute that to schema.

Two smaller pieces of evidence agree. searchVIU tested five AI systems on one page, and none of them read facts that existed only in the schema markup when they fetched the page. Google's guide says structured data "isn't required for generative AI search".

The guides that say cited pages nearly all have schema may be right. That would show little, because well-run sites tend to carry schema whether or not an AI cites them.

The evidence on the other side is second-hand. An attendee reported that Fabrice Canel of Microsoft Bing said at SMX Munich in March 2025 that schema markup helps Microsoft's LLMs understand content. I couldn't find slides or a transcript.

Two caveats keep the question open. Every page in the Ahrefs study was already being cited, so schema might still help a page get discovered. And searchVIU notes that schema could be used when a page is indexed, which its test couldn't see.

Keep schema for rich results in classic search, where it does work. Shruwd flags missing schema too. That check compares your page with the pages being cited, so the findings reference gives it the lowest of its three confidence labels: a lead worth testing.

Does llms.txt help LLM SEO?

I've found no evidence that it does. llms.txt is a markdown file at the root of a site that tells a model where things are. Jeremy Howard proposed it in September 2024, and the proposal's own page now says the files "are used most heavily for software documentation, where coding agents follow them to find API references and tutorials."

For search, the record points one way:

  • Google's guide says Google Search "ignores them", and that publishing one "will neither harm nor help".
  • OpenAI's and Anthropic's crawler documentation doesn't mention llms.txt.
  • Search Engine Land added one in March 2025. Semrush reported that from mid-August to late October 2025, GPTBot, PerplexityBot and ClaudeBot didn't request it once.
  • Google's John Mueller wrote in June 2025: "FWIW no AI system currently uses llms.txt."

Some of Google's developer documentation sites publish llms.txt files, which confuses the picture. Asked in January 2026 whether that was an endorsement, Mueller answered "to be direct, no", as Search Engine Roundtable reported.

Shruwd publishes an llms.txt too, so here's what it's for. It points coding agents at the markdown version of the docs, so an assistant setting up a brand through the API or the MCP server can read them without parsing HTML. It's documentation plumbing, and I don't expect it to earn a citation in AI search.

How to tell whether any of it worked

Change one thing, then ask the same questions again, several times each. AI answers differ from run to run, so a single before-and-after check can't separate your change from noise.

For Google's own AI features, Google's guide points to the Generative AI performance report in Search Console. The same guide says: "No third-party tool has access to our internal ranking or AI systems." That includes Shruwd. It can ask the questions repeatedly, count what comes back and read your logs. It can't see inside the model.

When you mark a fix as applied, Shruwd measures again after 14 days, at nine repetitions per prompt. It calls the result movement only when the ranges before and after don't overlap and the change is at least five percentage points.

Frequently asked questions

What is SEO for LLMs called?

LLM SEO, LLMO (large language model optimization), GEO (generative engine optimization) and AEO (answer engine optimization) all describe it. Wikipedia covers them in one article, and its "answer engine optimization" page redirects there. The work is mostly the same under each name.

How do you do SEO for LLMs?

Start with access: make sure robots.txt, your CDN and Search Console aren't blocking AI crawlers. Put your content in the raw HTML. Keep your pages indexed and ranking, in Bing as well as Google. Earn mentions on the pages that AI answers cite. Write direct answers with real facts. Then ask the same questions repeatedly to see what moved.

Is SEO still worth it in 2026?

Yes. AI answers are assembled from pages that search engines have crawled and indexed. Google states that a page must be indexed and eligible for a snippet to appear as a supporting link in AI Overviews or AI Mode.

All posts
Share

Find out what the answers say about you.

One measurement, your own crawler data, and the three access findings that matter most. Free, no card, about two minutes.

Start free