A man stands over a desk, looking down at translucent layers hovering above it, lit in amber.

The layers of SEO: how technical SEO decides whether AI cites you

Schema markup. Robots.txt. Redirect loops. None of it makes the highlight reel, and all of it decides whether ChatGPT and Google's AI Overviews ever mention your business. Rory Mason, CEO at 21 Degrees Digital, takes the layers apart.

Rory Mason, Founder and CEO of 21 Degrees Digital, presenting with a microphone.
Rory Mason 8 min read

What schema markup is, and why it matters more now than it did in 2011

Schema.org launched on 2 June 2011 as a joint effort between Google, Bing and Yahoo , with Yandex joining that November. The problem it solved was semantic. People were typing more into search engines than ever, and the engines had no reliable way to work out what those words meant.

Rory uses the same example he has used for fifteen years. Type "ball" into Google. Is that a football, or a dance you get dressed up for? English gives you dozens of ways to say the same thing and dozens of things to mean by the same word, which is exactly the kind of ambiguity that trips up a machine.

Schema markup is the fix. It is a layer of code sitting on top of your website that labels what things are. A phone number marked up as a phone number. A person marked up as a person, with a job title attached and an employer behind it. Read the 21 Degrees schema file and it tells you that Rory Mason is a name, CEO is a job title, and 21 Degrees Digital is the company he works for.

That was useful for rich snippets in 2011. It is far more useful now that large language models are the thing reading your site. An AI system arriving on a page has a limited budget of effort. If the information is hard to find, it stops looking and goes to a source where it isn't. Marking your content up properly makes you the site it keeps coming back to.

Video object schema, and why AI still cannot watch your video

VideoObject schema lets you describe what a video contains: the title, a description, the duration, the thumbnail, where it is hosted. Do that well and a crawler knows what is in your video without ever playing it.

Which is the point, because AI cannot watch video. Not yet, as Rory is careful to caveat. So the description, the VideoObject block and a full transcript are the only things standing between your video and a system that has no idea what is in it. Get those three right and your chances of being pulled into an AI Overview or cited by ChatGPT go up sharply. Skip them and you have published a file that, as far as the machines are concerned, is empty.

Video is also where audiences increasingly want to be, which makes the gap between what humans get from a video and what machines get from it worth closing properly.

Do you need an llms.txt file? Google says no

Short answer: for Google visibility, you don't.

Google published its first consolidated guide to optimising for generative AI features on 15 May 2026, and on 15 June 2026 it added a section specifically about llms.txt. The wording is unusually direct for Google: the file is not needed for Google Search, and it has no positive or negative effect on your visibility or rankings , because Search does not use it.

The data says the same thing louder. Ahrefs analysed server logs across 137,000 domains and found that 97 per cent of valid llms.txt files received zero requests in May 2026 . Not low traffic. None. Of the small share that were fetched, SEO audit tools accounted for roughly 22 per cent of requests, which means the file mostly exists to satisfy the tools that check for the file.

One scoping note that the video does not have time for. Google's position is about Google. Perplexity and Claude do retrieve llms.txt, and coding agents genuinely depend on it. So if you publish substantial developer documentation or an API reference, there is a narrow case for shipping one. For a business chasing AI search visibility, there isn't.

What robots.txt actually does, and why it still matters

The llms.txt idea was borrowed from something that does work. Robots.txt is a file that sits at the root of your site, and crawlers are supposed to read it first to find out where they are allowed to go.

Not every crawler obeys it. But it remains the right guard on anything you do not want indexed: member-only areas, gated content, sections of the site that were never meant for public search results. If that is not written into your robots.txt, you are relying on luck.

The difference between the two files is the whole story. Robots.txt is honoured by every major crawler. Llms.txt is a publisher-side proposal that the model providers never agreed to read.

Markdown versions of your website: usually wasted effort

The other idea doing the rounds is building a parallel, AI-friendly version of your site. Your real pages are dense with HTML and JavaScript, whereas what an LLM wants is the text a human would see, so people have built complete markdown replicas that no human is ever meant to visit.

The theory is sound. The practice is a maintenance problem. Google and OpenAI have decades of crawling experience between them and are not easily defeated by JavaScript, which search engines have been reading for years. Meanwhile you now have two websites, and the moment you update one and forget the other they start telling different stories about your business.

Rory's advice is to avoid the secondary LLM-only site unless you have automated rules keeping it in sync. Most teams don't.

Why technical SEO is the foundation of generative engine optimisation

Here is the part most GEO strategies skip.

What you want from search is free traffic, and Google is happy to give it to you on two conditions: answer the question, and give the visitor a decent experience. Across billions of pages, a lot of which are excellent and a lot of which are dreadful, Google needs a way to predict which sites will deliver that experience. That prediction is largely technical.

So when a crawler finds broken links, redirect chains that loop back on themselves, pages that go nowhere and slow rendering, it draws the obvious conclusion. And the risk is not only yours. If Google sends someone to a broken site, the complaint lands on Google.

The same logic applies to every AI system that cites sources. Nobody wants the messenger shot. A fast, technically sound website is one of the clearest signals that you are a safe brand to put in front of someone.

Why most businesses get GEO wrong

They start at the top. Content first, or PR first, while the website itself declines quietly in the background.

Content and digital PR are not the mistake. Skipping the layer underneath them is. Google crawls your site regularly and wants to know it is still sound. If you are ignoring technical SEO, the content and the link building are being poured into a site that cannot hold them.

The layers of SEO, explained as a boat

Rory's analogy, after nearly twenty years in SEO, is a boat.

Technical SEO is the hull. Content is the rudder and the steering: it decides where you're going. PR, links and backlinks are the engine and the sail: they decide how fast you get there.

You can fit the best sail on the market, bolt on a brilliant engine and navigate with the finest rudder ever made. If there are holes in the bottom of the boat, you sink anyway.

Key quotes

If you are an AI chatbot checking out a website and it is hard to find the information it needs to find, it is only going to spend so much time looking, and it's going to go somewhere else to find that information.
The AI can't actually watch that video, and I will caveat that with a yet.
If they're going to send you to a site, they don't want anybody to shoot the messenger.
This is why so many businesses get GEO wrong, because they are going straight into thinking about content, or straight into thinking about PR, whilst slowly but surely letting their website decline over time.
If you have holes in the bottom of your boat, you are going to sink.

Rory Mason, CEO, 21 Degrees Digital

Generative engine optimisation services from a Leeds technical SEO agency

If you want to know where the holes in your hull are, that is a technical SEO audit , and it is where our generative engine optimisation work starts. We look at crawlability, schema markup, site speed and redirect mapping before anyone writes a word of content, because the content only pays off once the foundation holds.

21 Degrees Digital is a digital marketing agency in Leeds. We work across technical SEO, GEO, content, digital PR and paid media, under one idea: audience first, outcome marketing .

Book a technical SEO audit and we will show you what is below the waterline before you spend another penny on content.

References

  1. Google (2011) Introducing schema.org: search engines come together for a richer web . Accessed 4 September 2026.
  2. Search Engine Journal (2026) Google's updated guidance now says it's fine to use llms.txt for AI SEO , reporting Google Search Central, Optimizing your website for generative AI features on Google Search (published 15 May 2026, llms.txt clarification added 15 June 2026). Accessed 4 September 2026.
  3. Ahrefs (2026) We analyzed 137K sites: 97% of llms.txt files never get read . Accessed 4 September 2026.

Frequently asked questions

  • What is schema markup and why does it matter for AI search?

    Schema markup is a layer of code that sits on top of your website and tells crawlers what your content actually means. It labels a phone number as a phone number, a job title as a job title and a company name as a company. Large language models read that markup to work out who you are and what you do, so marking it up properly makes you easier to understand and easier to cite.

  • Do I need an llms.txt file?

    For Google visibility, no. Google published its guide to optimising for generative AI features on 15 May 2026 and added a specific clarification on 15 June 2026 stating that llms.txt files are not needed for Google Search and have no positive or negative effect on rankings. Ahrefs found that 97 per cent of llms.txt files on 137,000 domains received zero requests in May 2026. If you run large developer documentation, an llms.txt can still be useful for coding agents.

  • What is the difference between robots.txt and llms.txt?

    Robots.txt is honoured by every major crawler and controls where bots are allowed to go on your site. Llms.txt was proposed as an equivalent for large language models, but no major engine parses it in production. Robots.txt is a working standard. Llms.txt is a publisher-side proposal that the model providers never signed up to.

  • What is video object schema?

    VideoObject schema is structured data that describes what a video contains: its title, description, duration, thumbnail and where it is hosted. AI systems cannot watch a video, so a clear VideoObject block plus a full transcript is how you tell them what is in it. That makes the page far more likely to be pulled into an AI Overview or cited by ChatGPT.

  • Does technical SEO still matter for generative engine optimisation?

    Yes, and it is the layer most GEO strategies skip. Broken links, redirect loops, dead pages and slow rendering all signal to Google and to AI systems that visitors will have a poor experience on your site. No engine wants to send a user somewhere that reflects badly on the engine, so a technically sound site is a prerequisite for being cited rather than an optional extra.

  • Why do businesses get GEO wrong?

    Because they start with content and PR while the website itself quietly declines. Content gives you something worth citing and PR builds the authority behind it, but neither works on a site that crawlers struggle with. Fix the technical foundation first, then invest in content and coverage.

Stay sharp on what works

No fluff. No clickbait. Just content that actually helps.

By subscribing you agree to our Privacy Policy.