What makes a Drupal site readable to LLMs
CTOs, heads of digital, and marketing leads
Drupal LLM readiness is the work that makes your pages quotable by an assistant: structured data emitted from the content model, metadata that agrees with what a visitor reads, and a copy of the content a machine can take in whole.
The short version
- Assistants answer from text they can lift cleanly, so the shape of your markup decides how often you get quoted.
- A replatformed site arrives with the structured data already coming out of the content model: organisation, website, web page, article and service, generated per page from fields.
- Drupal's AI layer arrives configured against your own provider keys, with generated alternative text and per-content-type metadata drafting.
- The machine-readable pack arrives with it: an llms.txt map, a markdown twin of every page, read-only JSON per content item, and crawl rules that name the AI agents one by one.
- All four are live on the site we replatformed, so you can open them instead of believing me.
- You can check every one of these on your own site this afternoon, with a browser and nothing else.
Two years ago the question about a website was where it ranked. The question I get asked now is whether it gets quoted, and by what.
Those turn out to be different problems with a lot of shared plumbing. A search engine indexes your page and sends someone to it. An assistant reads your page, decides whether the passage it needs is clean enough to lift, and either cites you or moves to a source that was easier to read. The second one rewards a kind of tidiness that most sites never had a reason to care about.
This post is about that tidiness on a Drupal site: what a replatform lands with, how you can measure where you stand today, and the one piece we are still building.
Why does an assistant quote one site and skip another?
Because one of them made the answer easy to lift, and lifting is most of the job.
When an assistant builds an answer it is looking for a passage it can attribute with confidence: a clear question, a self-contained answer near it, a heading that says what the section is, and some machine-readable declaration that the page is what it appears to be. Sites that publish that way get pulled into answers often. Sites where the same information is spread across four hover states and an accordion get pulled in less, because the assistant has to reconstruct meaning that was carried by layout.
Three things move the needle most, and all three are content-model problems wearing a marketing hat.
The page says one thing. One h1, headings that describe their sections, an opening paragraph that defines the subject in plain words. That is the passage an assistant quotes, so it should be a definition and not a warm-up.
The markup agrees with the prose. Structured data that says "this is an article, published on this date, by this author, part of this organisation" gives the machine a frame for the text. When the structured data and the visible text disagree, the page gets trusted less, which is why generating it from fields beats pasting it per page.
The content is reachable as text. Content assembled in the browser after load is harder for a crawler to see than content that arrives in the HTML. This is where a lot of otherwise good sites quietly lose ground, and it is one of the strongest arguments for a server-rendered CMS.
What arrives configured on the delivered site today?
The structured data and the AI layer, both wired to the content model and both working on handover.
On the structured data side, the delivered site emits JSON-LD from the fields themselves: organisation, website, web page, article and service. Because it comes out of the content model, a change to a headline changes the markup in the same edit. What a machine reads and what a visitor reads stay in step, which is a Google requirement and, more usefully, the thing that keeps the markup honest a year later when nobody remembers it exists.
Alongside it, the SEO foundations arrive configured: metadata per content type, XML sitemaps, clean path patterns and a redirect map with logging.
On the AI side, the delivered CMS arrives with Drupal's AI module suite installed and pointed at your own provider keys. Three things it does on day one:
| Capability | What it does on the delivered site |
|---|---|
| Assistive authoring | Drafting and content suggestions inside the editor, working on typed fields |
| Media descriptions | Alternative text generated for images that arrived without any, which is an accessibility fix and a search fix in one pass |
| Metadata drafting | Titles, meta descriptions and social metadata generated per content type from the content itself, with an editor approving |
Two details in there matter more than the features. The keys are yours, read from the environment, so they never travel in a configuration export or a database dump, and the capability belongs to your account. And the layer logs which model was called, for which task and when, while leaving the content of prompts and responses out of the log, so you can account for spend and answer a privacy question with the same record.
What is llms.txt, and why does it sit beside robots.txt?
It is a plain-text file at the root of a site that tells an assistant what the site is and where its canonical pages live.
Think of the family it belongs to. robots.txt tells a crawler where it may go. sitemap.xml tells it what exists. llms.txt tells a model what the site is for and which URLs carry the substance, in prose a model can read in one pass. It is a young convention that the ecosystem is still settling, and it is cheap to publish, which is a good combination for something with this much upside.
The fuller version of the idea goes past one file. A markdown copy of each page gives a model the prose and the headings with the navigation, scripts and styling out of the way. A JSON representation of each content item, straight from Drupal's own API layer, makes the content consumable as data by an assistant, an agent or an internal tool. And explicit crawl directives per bot let you say yes to the LLMs you want to appear in while keeping the controls you already have.
What does the delivered site actually arrive with?
That whole machine-readable pack, and this section used to say the opposite, so it is worth being precise about what changed.
When I first wrote this post the pack was in build and I labelled it that way. It has since shipped, and it is not shipped in the sense of a demo on our own marketing site. It is on the Drupal site the harness produced. qed42.com ran on Webflow; it runs on Drupal 11 today, replatformed with X to Drupal, and the cutover means the four surfaces are serving from the address you would have visited anyway:
https://www.qed42.com/llms.txt: the map, in plain text.https://www.qed42.com/insights.md: the markdown twin of a page, with the navigation and styling gone. Every page has one.https://www.qed42.com/jsonapi: the content as data, exposed for reading only. Writes are not available and the accounts resource is switched off, which is the part most people forget to check when they open an API on a public site.https://www.qed42.com/robots.txt: the crawl rules, naming the AI agents individually rather than waving at them collectively.
Go and open them. That is the only reason I list the addresses rather than describe the capability: a claim you can check in four clicks does not need me to be persuasive about it.
One thing is still in build, and it keeps its label: a deeper automated SEO audit, meaning a standing check of sitemap coverage, the crawl surface and metadata quality. That is an extension of what is already on the delivered site rather than the foundation for it, so nothing above is waiting on it.
This site, the one you are reading, still publishes its own /llms.txt, an /llms-full.txt corpus with every article inline, and a markdown copy of every post at /blogs/<slug>.md. That started as a demonstration of a mechanism we had not yet delivered. It is now just consistency.
How can you check your own site this afternoon?
Four checks, a browser, about twenty minutes. None of them need a tool.
- View source on your best page and search for
application/ld+json. If nothing comes back, you have no structured data. If something comes back, read it and ask whether it matches the page you are looking at. - Turn JavaScript off and reload. What survives is roughly what a crawler sees first. If your main content disappears, that is the biggest single item on your list.
- Fetch
/llms.txtand/sitemap.xml. The second one most sites have. The first one almost nobody has yet, which is exactly why publishing it is worth an afternoon. - Read your first paragraph aloud and ask if it defines the subject. If it is a warm-up, the passage an assistant would quote is somewhere further down the page, and it may never get there.
The results tend to sort into two piles: things that are a content edit, and things that are a platform limit. The first pile you can clear next week. The second pile is the reason a replatform comes up at all.
Where does this fit in a replatform?
At the point where you are already paying to touch every page.
The awkward economics of AI readiness on its own is that it asks for changes across the whole site for a benefit that shows up over several quarters. That is a hard sell as a standalone project and an easy one as part of a move you were doing anyway, because the content model you build during a replatform is the thing that generates the markup afterwards. Model the content properly once, and the structured data, the metadata and the machine-readable copies all fall out of it. Skip that step, and you are back to pasting markup per page and watching it drift.
So the sequence I would argue for is: move to a platform that renders on the server and holds a real content model, get the structured data coming out of the fields, and let the machine-readable copies fall out of the same model. That order means every step is useful on its own, which is the test I apply to any roadmap I am asked to believe. When I first wrote it, the third step was a promise about a later pack. It is now the same delivery as the second, which is the version of that argument I always wanted to be making.
This post is part of the X to Drupal series, alongside the pillar, why replatform into Drupal and what X to Drupal provides, and the companions on moving our own site off Webflow, what your team gets afterwards, and how we check the result really matches.
Frequently asked questions
What is llms.txt and does my site need one?
It is a plain-text file at the root of a site that tells an assistant what the site is and lists its canonical URLs, sitting alongside robots.txt and sitemap.xml. It is a young convention and cheap to publish, so it is worth doing early on a content-heavy site where being quoted has real value.
Does structured data help with LLMs or only with Google?
Both. Structured data gives any machine a frame for the text on the page: what kind of thing it is, who published it, when, and how it relates to the rest of the site. Search engines use it for rich results and assistants use it to attribute a passage with confidence.
Which AI features arrive configured after a replatform to Drupal?
Drupal's AI module suite arrives installed and wired to your own provider keys, with assistive authoring in the editor, alternative text generated for images that arrived without it, and metadata drafted per content type. The layer logs the model, the task and the time, and leaves prompt and response content out of the log.
Is the machine-readable pack available today?
Yes. The llms.txt map, a markdown twin of every page, read-only JSON per content item and per-bot crawl rules all arrive configured on the delivered site, alongside the structured data and the AI layer. You can check all four on the Drupal site we replatformed: qed42.com serves llms.txt, a markdown twin of every page, a read-only JSON API and a robots.txt that names the AI crawlers individually. A deeper automated SEO audit is the one piece still in build, and we label it that way wherever it appears.
How do I tell whether my current site is readable by an assistant?
View source and search for application/ld+json to see whether structured data exists, turn JavaScript off and reload to see what a crawler gets first, fetch /llms.txt and /sitemap.xml, and read your opening paragraph to check that it defines the subject. Those four checks take about twenty minutes and need only a browser.
Why does server-side rendering matter for LLMs?
Content that arrives in the HTML is available to any crawler on the first request. Content assembled in the browser after load asks the crawler to execute scripts before it sees anything, and coverage there varies by bot. A server-rendered CMS puts the text in front of every reader, human and machine, on the same request.