# X to Drupal, full text > X to Drupal is the QED42 harness that modernises an existing website by replatforming it onto a real Drupal CMS. Modernising a website in 2026 means AI inside the CMS, digital sovereignty, an open architecture and readiness for AI agents, and the replatform is the step that gets a site there: the visual result preserved, the content model built the way a senior Drupal architect would model it, the delivered site wired for AI capabilities and LLM visibility, and the whole thing handed over so the content team can run it: a runbook generated in place of a training week, a Model Context Protocol connection that stages requested changes in the CMS as drafts, and the site's own brand captured as context. The result is an exportable Drupal CMS for any Drupal 11 host, with no vendor lock-in. The proof is qed42.com itself, migrated from Webflow to Drupal 11: 569 article bodies, 2,007 images, zero custom modules, and the whole site installs as a Drupal Recipe. No parity percentage is published; the side-by-side comparison is the claim. This file carries the complete text of every article on the site, in series order, for assistants that prefer one document to a crawl. Per-article markdown is also available at .md for any /blogs/ URL, and the site map for machines is at /llms.txt. - Site: https://x-to-drupal.qed42.com - Maker: QED42 (https://www.qed42.com/) - Contact: business@qed42.com - Articles included: 7 --- # Why replatform into Drupal and what X to Drupal provides > Why replatforming into Drupal is worth doing, and how X to Drupal reads your live site, derives an architect-grade content model and delivers an editor-ready CMS. - Canonical: https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides - Author: Souvik Pal - Published: 2026-07-28 - Part of: the X to Drupal series by QED42 (post 1 of 7) *Replatforming into Drupal hands your content back to the people who own it. It has also always been priced like a rebuild, because the design and the content model both got recreated by hand. X to Drupal is our harness for doing that work from the rendered site itself. This post covers why the move is worth making and what the harness provides.* > **The short version** > - Replatforming into Drupal is worth it because it turns your site into structured content your team can edit, reuse, govern and report on without a developer in the loop. > - It has always been quoted like a rebuild: the design gets re-implemented by hand and the content model gets invented from scratch. > - Your existing site already holds both answers. They sit in the rendered pages, waiting for something that can read them out at production quality. > - X to Drupal reads the rendered site, derives the content model to a senior architect's standard, and verifies the result against what actually renders. > - What you receive is standard Drupal: typed fields with native widgets, Views for listings, real taxonomies, and zero custom modules. I have sat in the meeting where a client asks why moving their site to a new platform costs about the same as building it the first time. It is a fair question. The design already exists. The content already exists. And yet the quote covers weeks of design implementation and weeks of content modelling, as if none of it had ever been done. The honest answer has two halves. The move is worth making, for reasons that compound every year you run a content-heavy site. And the price has been a rebuild price because, until now, a rebuild is what the work amounted to. This post covers both halves, and what our harness, X to Drupal, changes about the second one. ## Why replatform into Drupal at all? A site on a real CMS hands control back to the people who own the content. Your marketing team publishes without a release. Your editors build a campaign page from existing components on a Tuesday afternoon. Your content gets structured well enough to be reused, translated, scheduled and reported on. Governance stops being a spreadsheet and becomes roles, permissions and review states in the platform. Drupal's central idea is that content is structured data first and a page second. Content types and fields describe exactly what a case study is. Paragraphs let editors compose page narrative from typed components, so an editor sees an image upload and a link field with the widgets Drupal already provides. Views turns any list of content into a configured listing with filters and pagination, so a new insight article appears on the listing, in the related block and in the sitemap without anyone touching a template. Taxonomies give you real categorisation to filter and report on. Drupal 11 and Drupal CMS ship a great deal of this ready made. Recipes install and configure whole capabilities in one step. SEO tooling, media handling and scheduling are available on day one, already configured. A modern Drupal install starts much closer to finished than it did five years ago. There is a newer reason too. Structured content is what LLMs can actually read and cite. A site whose pages are typed entities with clean markup and real metadata is far easier for an assistant to quote accurately than a pile of rendered divs. The series covers this in [what makes a Drupal site readable to LLMs](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search). Site builders like Webflow and Framer are genuinely good at what they do, and for a marketing site with a small team they are often the right call. The trade flips at a predictable point: when your content becomes an asset you want to model, query and reuse, the CMS earns its keep. ## Why has replatforming always been priced like a rebuild? Strip out project management and QA, and a typical replatform breaks into four buckets: design implementation, content modelling, content migration, and everything else (menus, redirects, forms, search, SEO metadata, integrations, permissions). Three of those four buckets are transcription. The information already exists in a form a person can read. It just does not exist in a form the new platform can consume, so people have re-created it by hand. **The design gets rebuilt even though it already exists.** Every rendered page carries the design in full: the spacing, the type scale, the breakpoints, the colours, the layout behaviour. A developer rebuilding that page reads those values off a screen with their eyes and types them back in, which is slow and lossy. The creative decisions were made years ago. What remains is data entry with very high standards. **The content model gets invented from scratch, and it is the harder half.** A rendered page tells you what it looks like. It does not tell you that these three cards are a listing of a repeating content type, that this band is global chrome shared by every page, that this heading and image belong to one repeatable component. Those are architectural judgements, and they separate a site an editor can run from a pile of pages an editor can only look at. Get the model wrong and editors ask for a "small change" and get a developer ticket, a new landing page needs a deployment, and everyone slowly stops using the CMS. **The person who can make those judgements is the scarcest person in the building.** A senior architect who knows Drupal deeply can make the few hundred modelling decisions well, and every one of them is already committed to something. So migration work gets staffed with whoever is available, the visual layer comes out fine because it is easy to check with your eyes, and the content model comes out shaky because nobody sees it until an editor tries to use it six months later. That is why a replatforming quote has looked like a rebuild quote. The work really was a rebuild, performed by people transcribing information the site already contained. ## What X to Drupal provides For the price to change, the work has to change. Three things would have to happen: something reads the visual layer off the rendered site with real precision, something derives the content model to a senior architect's standard, and something verifies the result against what actually renders. We went looking for a system that does all three, found excellent tools for pieces of it, and nothing that takes an arbitrary live website and produces a Drupal site with both the visual result and an architect-grade content model. So we built X to Drupal. Here is what it provides. **It reads the design off the rendered site.** The harness works from the pages your visitors actually see. It studies the design system the way an architect would and reproduces it as components in the Drupal theme. The old site stays exactly as it is; it becomes the specification. **It derives the content model to an architect's standard.** Page narrative lives in ordered paragraph components with typed fields, so an editor sees a heading field, an image upload, a link field and a repeatable group, each with its native widget. Any list of existing content becomes a View. Global chrome such as the header, footer and menus lives in blocks placed in theme regions. Categorisation uses real taxonomy vocabularies with real terms. These are the same shapes a senior Drupal architect chooses on a hand build, because they come from how we build production Drupal sites. **It verifies against what renders.** Configuration being present is not proof, and markup existing in the DOM is not proof. The harness compares rendered pages against the source site, and it signs in to walk the editorial journeys your team will use in the delivered CMS. A migration that looks done in the config and wrong in the browser is not done, so the checking continues until the differences are boring. The series covers this in [how we check a replatformed site really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). **It delivers standard Drupal.** The shipped site runs on zero custom modules. The theme is small, split and lint-clean. The configuration export contains what the site uses. Anyone who knows Drupal can maintain it from day one, and content created after the migration behaves exactly like content that came through it. **A person stays in the loop where judgement matters.** An architect reviews every call the run hands to a person. The automation does the transcription; the judgement calls stay human. That is the honest division of labour, and it is why the result holds up to review. We proved it on ourselves first: [we moved our own site from Webflow to Drupal](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal) with the harness before pointing it at anything else. ## What this changes about the decision The design question disappears from the project. Keeping the design fixed removes an entire round of stakeholder review, and the visual layer arrives as components read from the site you already approved. The content model stops depending on who happened to be available. The harness applies the modelling rules consistently across every component and every page, at the standard the scarce architect would set, and the architect's time goes to review, where their judgement counts. And the economics move accordingly. When the transcription work stops being hand work, a replatform stops being priced like a rebuild. If you want to know what that would mean for your own site, [get in touch](https://x-to-drupal.qed42.com/#contact). The rest of the series tells the story behind the harness: [fourteen generations of getting this wrong then right](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system), [what the build taught us about working with AI](https://x-to-drupal.qed42.com/blogs/what-building-this-taught-us-about-working-with-ai), [the day we moved our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), [what your team actually gets afterwards](https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms), [what makes the delivered site readable to LLMs](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### Why replatform into Drupal instead of staying on my current platform? Because structured content pays compounding returns on a content-heavy site. Drupal gives you typed content, editorial workflow, granular permissions, multilingual, a real API layer and a large module ecosystem. Site builders remain a good fit for small marketing sites; the trade favours Drupal once your content becomes an asset you want to model, reuse and report on. ### What does X to Drupal actually deliver? A working Drupal site built from your live site: the visual result reproduced as theme components, page narrative in ordered paragraph components with typed fields, listings as Views, global chrome as blocks in theme regions, real taxonomy vocabularies, and zero custom modules. An architect reviews the judgement calls while the harness does the transcription. ### Why does moving a website to Drupal usually look like a rebuild? Because three of the four buckets of work, design implementation, content modelling and content migration, are transcription done by hand: people re-create information the site already contains. X to Drupal does that transcription in the harness run, from the rendered site itself, and the architect's time goes to review. Get in touch through the form on our home page for what replatforming would mean for your own site. ### Can you migrate a website to Drupal without redesigning it? Yes, and it is the path X to Drupal is built for. The harness treats your current design as the specification and reproduces it, which removes an entire round of stakeholder review and keeps the focus on the content model, where the long-term value sits. ### Is the result standard Drupal or something custom? Standard Drupal. The delivered site runs on zero custom modules, uses Drupal's native widgets and editorial tooling, and ships a clean configuration export. Any team that knows Drupal can maintain it without learning anything about the harness that produced it. ### Can content editors change layouts in Drupal without a developer? They can when the content model is built for it. Editors compose pages from paragraph components with typed fields and native widgets, and listings assemble from existing content through Views. That freedom is a direct product of the modelling decisions, which is why X to Drupal holds them to an architect's standard. ### What happens to my SEO when I move platforms? The risk sits in URLs, metadata and redirects. Keep the URL structure, carry the metadata across, and put a redirect map in place for anything that has to change. Drupal covers all three with its SEO tooling, and the harness's verification compares the rendered result, metadata included, against the source site. --- # Fourteen generations of a website to Drupal system > What fourteen rewrites of an automated website to Drupal migration system taught us. The wrong turns, the decisions that held, and the one that changed everything. - Canonical: https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system - Author: Souvik Pal - Published: 2026-07-28 - Part of: the X to Drupal series by QED42 (post 2 of 7) *Building an automated website to Drupal migration system took fourteen full generations in about three months. This is the version-by-version account of what broke, what we decided, and which decision mattered most.* > **The short version** > - Fourteen numbered generations between 11 May and 21 August 2026. > - Reading a website accurately turned out to be the easy half. Deciding what the content model should be was the hard half. > - Generation 6 was the turn that mattered: components stored as one blob rendered perfectly and were useless to an editor. > - Generation 12 produced the rule that made everything hold: a fix that lives in the output is not a fix. > - This is a build log, so it covers decisions and outcomes. It does not cover how the system is put together. Three months ago I ran a pipeline end to end for the first time against qed42.com and got something that half worked. Fourteen generations later it produces a Drupal site I am willing to put in front of an architect. The distance between those two sentences is the interesting part, and almost none of it went where I expected. The pattern that repeats through the whole log: every time I thought the problem was accuracy, the problem was actually architecture. Every time I thought the problem was the model, the problem was the contract I had given it. ## Generations 1 to 4: reading the site was never the hard part The first two generations read the site well, but much of what they learned was lost before the build. Fonts stayed as placeholders and never made it into the theme. Background images went missing. Setup errors blocked verification before anything could be judged. Generation 3 added a second, structurally very different reference site, and the same root causes appeared on both. That was the first genuinely useful signal. Two very different sites failing identically means the failure is in the system, so the fix belongs in the system. Generation 3 also produced a list I still find useful, of every moment a human had to step in during a run: - Somebody had to say "this site is video and animation heavy" - Somebody had to notice that a six-category mega menu had come through as a flat list - Somebody had to notice a newsletter modal was missing entirely - Somebody had to fix spacing and layout by hand Every one of those is a missing automated step wearing a human's clothes. We turned the list into work, and it became the rule for the rest of the project: a human tweak during a run is a defect report against the system. Generation 4 settled an architectural question that kept coming back, and it settled it in favour of consistency: the same input must give the same result on every run. Variance is the thing you can least afford when you are trying to establish whether a change helped. ## Generation 5: the ceiling nobody could polish away By generation 5 a clean run came within touching distance of the bar on first pass and then stopped. There was no mechanism to close the gap, so it stayed open. We gave the finish a clear bar and a clear point to stop. Left open-ended, polish will happily spend an afternoon making a footer imperceptibly better. ## Generation 6: the one that mattered Generation 6 is the turn the whole project pivots on, and it started as a complaint about a footer colour. While looking at that, I looked properly at what the backend had become. Every component was one paragraph type with a single long-text field holding the entire payload as JSON. The theme decoded the JSON at render time and fed it to the component. It rendered beautifully. Every visual check was happy. And an editor opening that page saw a textarea full of JSON. That defeats the entire point of moving a site to a CMS. Drupal's whole story is typed fields with native widgets: a text field for a heading, an image upload for an image, a link field for a link, a repeatable group for list rows. A JSON blob has none of that. We had built something that passed every check we had and delivered none of the value we were selling. So generation 6 rebuilt the content model around typed fields, the way a senior Drupal architect would build it by hand. From that point on, the bar stopped being "does it render" and became "does it render and can an editor work in it". Everything after generation 6 is downstream of that sentence. ## Generations 7 and 8: decisions, then checks Generation 7 fixed a drift problem. The same decision was being made in several places, with the model making a judgement call each time, so it came out differently from site to site. That is why fixes never seemed to stick. We now make each decision once, and the rest of the run follows it. If I could keep one lesson from this log, it would be that one. Decide once. The model still does plenty of work, and all of it now follows a decision that is already made. Generation 8 added automated checks early in the run, so a structural problem is caught before any effort goes into the finish. The reason was simple. The system had been converging incident by incident. Every run surfaced a new unhandled shape and we patched reactively, which is a treadmill. At that point the measured difference was worst at tablet, close behind at mobile and smallest at desktop, and all three sat well outside the bar we had set. Uncomfortable readings, and useful ones, because they were finally honest and repeatable. ## Generations 9 to 11: measuring the right thing Generations 9 and 10 were about accuracy levers, and they produced a finding that reframed the remaining work. The remaining differences came from content and structure. Finish work only closes a gap when the content is already there. We had been trying to fix a missing-content problem with styling. Generation 11 fixed something I am slightly embarrassed about. Multi-page support was, in practice, theatre. Pages were enumerated one at a time, the sitemap was never read, and per-page changes to the header and footer were being lost. Fixing that properly meant treating the whole site as the unit of work. ## Generation 12: the rule that made everything hold Generation 12 ran across five sub-versions and produced the discipline that finally stopped the regressions. Fix after fix had appeared to land and then quietly disappeared on the next clean run. The reason was always the same: the fix had been applied to the output of a run, so the thing that generated the output still produced the old result. Every clean run overwrote it. Two rules came out of it, and they are now the ones I care most about: **A fix belongs in the generator, never in the output.** If you cannot reproduce the corrected result from a completely clean run, you have not fixed anything. You have decorated one run. **Verify what renders, never what is present.** Config being correct is not proof. Markup existing in the DOM is not proof. The only acceptable evidence is a rendered page compared against the source. We had several diagnoses go stale between sessions because they were verified against a hand-patched state that a clean run never produces. Generation 12 also caught something worth saying out loud: some of our own earlier diagnoses were wrong. When we re-verified four reported defects independently, one had been diagnosed accurately and three had been mis-located. Re-verifying load-bearing claims before acting on them became standard, and it has saved more time than it costs. ## Generations 13 and 14: the architect bar, then the package Generation 13 rebuilt the Drupal backend to the shape a senior architect would choose. The decisions, in the order they matter: - Page narrative lives in ordered paragraph components on the node, in source order - Global chrome such as header, footer and menus lives in blocks placed in theme regions, because it is genuinely global - Any list of existing entities becomes a View. Landing-page narrative is paragraphs; listings are Views - One flexible page type covers one-off and landing pages, with distinct content types reserved for genuinely repeating templates - Real taxonomy vocabularies with real terms, so filters and reporting work Then we audited it hard, across the whole system. The audit confirmed 22 gaps and refuted five. The meta-finding was the useful part: the designs lined up with Drupal 10 and 11 community consensus, and the failures were implementation and verification gaps. We were not wrong about what to build. We were wrong about how thoroughly we had built it. We also deleted a governance layer that had grown around the project, and distilled roughly 110,000 words of accumulated decision history into a single short reference that people actually read. Process accumulates like debt, and it needs paying down on the same schedule as code. Generation 14 tackled the thing nobody had been looking at: the package we hand over. The architecture was right, and the package was heavier than it needed to be. So we trimmed it to a lean install list, configuration that is all the site's own, and a theme in small lint-clean files. The editor-experience check became a required step, the last of the site's own data moved into real Drupal configuration, and the shipped site reached zero custom modules. Editorial dynamics like read time now derive at render from real fields, so content created after the migration behaves the same as content that came through it. ## What I would tell someone starting this Three things. Decide what "done" means before you start optimising, and make it something a person outside the project would accept. Ours moved from "does it render" to "does it render and can an editor work in it", and that change was worth more than any accuracy work. Put the checks early. A wrong result caught early costs minutes. The same result caught at the end costs a run. Expect the architecture problems to arrive disguised as accuracy problems. Every serious turn in this log started as a visual complaint and ended as a modelling decision. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [what this taught us about working with AI](https://x-to-drupal.qed42.com/blogs/what-building-this-taught-us-about-working-with-ai), [moving our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), [what your team gets afterwards](https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### Can AI handle a full website migration on its own? It can do a great deal of the work when the work is bounded. In our experience the model performs well when the decision is already made and it only has to execute it. It performs poorly when it is asked to re-derive a decision that was already made upstream, because the answer drifts between runs. ### Why did it take fourteen generations? Because the first several generations solved the visible problem, which was visual accuracy, while the real problem was the content model. Generation 6 reset the definition of done and much of the earlier work had to be redone against the new bar. ### Does the migrated site come with a proper content model or just pages? A content model. Page narrative sits in ordered paragraph components with typed fields, any listing of entities is a View, global chrome is block content in theme regions, and categorisation uses real taxonomy vocabularies with real terms. ### What was the single most useful engineering rule you found? A fix belongs in the generator, never in the output of a run. If a completely clean run does not reproduce the corrected result, the fix is decoration. This one rule ended a long run of regressions that had been costing days. ### How do you verify a migration is correct? By comparing what renders against the source site. Configuration being present, or markup existing in the DOM, is a weaker signal that produced several false positives for us. Rendered comparison is the only evidence we now accept. ### Is a system like this site-agnostic, or tuned to one site? Ours is built to be site-agnostic. We develop it against several very different sites, so a fix that suits only one of them shows up as a failure on the others. --- # What building this taught us about working with AI > AI code generation reliability comes from clear decisions, honest checks and rendered verification. Eight rules from three months of building with AI. - Canonical: https://x-to-drupal.qed42.com/blogs/what-building-this-taught-us-about-working-with-ai - Author: Souvik Pal - Published: 2026-07-28 - Part of: the X to Drupal series by QED42 (post 3 of 7) *AI code generation reliability comes from the constraints you put around the model, and these are the eight rules that survived three months of building a system that replatforms websites onto Drupal.* > **The short version** > - The model does its best work on a small, clear job whose decision is already made, and drifts when asked to make that decision again. > - One input, one truth. Drift is usually a clarity problem wearing a model problem's clothes. > - A fix that lives in the output is not a fix. It has to live in the thing that produces the output. > - Rendered result is the only evidence. Configuration being present proved to be a weak signal that cost us days. > - Three of four AI diagnoses we re-checked pointed at the wrong place. Re-verifying load-bearing claims pays for itself. I spent about three months building a system that takes a live website and produces a Drupal site from it, with the visual result preserved and the content model built properly. Fourteen generations in all. Most of the interesting failures had nothing to do with the model being weak. They came from me giving it a job with unclear edges. These are the rules I would give myself on day one. They apply to any large AI-assisted build, and none of them are about prompting. ## What does AI actually do well on work like this? Small, clear jobs with the decision already made. That is the sweet spot, and it is a bigger spot than it sounds. When a job had clear edges and nothing left to decide, the results were consistently good, and they stayed good run after run. The model is also genuinely strong on domain best practice. When we audited our own architecture decisions across the whole system, the finding that surprised me was that the designs lined up with Drupal 10 and 11 community consensus. The model knew the right shape. What it had not done was prove that the right shape had actually shipped, which is a different skill and the subject of most of the rules below. ## Why does the same input produce different results? Because somewhere downstream, something is being re-decided. Our worst drift period came from the model exercising judgement on the same question again and again. Fixes appeared to land and then behaved differently on the next site. It looked like model inconsistency. It was a design fault: we had given the same decision to several places and let each of them answer it. The fix was making each decision once and letting everything after it follow that decision. The model still does plenty of work, and every bit of it now follows a decision that is already made. One input, one truth. If you are seeing variance, look for the decision you accidentally asked for twice. Related, and worth stating separately: whenever a design choice traded consistency for cleverness, we chose consistency. The same input must give the same result on every run, because variance is the one thing you cannot afford while you are trying to establish whether a change helped. ## Why do AI fixes stop working? Because the model fixes the thing in front of it, and the thing in front of it is usually the output. This cost us more time than any other single mistake. A defect gets reported, the model investigates, finds the wrong value, corrects it, and the page renders correctly. Everybody is satisfied. Then the next clean run regenerates everything and the defect is back, because the generator that produced the wrong value was never touched. The rule that ended it: **a fix belongs in the generator, never in the output.** The acceptance test is a completely clean run reproducing the corrected result. Anything short of that is decoration. This is not really an AI problem, it is an old engineering discipline. AI just makes it much easier to fall into, because fixing the symptom is fast, satisfying, and looks identical to fixing the cause right up until the next run. ## Why is "it is in the config" not proof? Because presence is not behaviour. For a long stretch we verified fixes by checking that configuration existed, or that markup appeared in the DOM. Both are cheap to check and both produced false positives. Config can be correct while a naming mismatch stops the value ever reaching the template. Markup can exist in the DOM while rendering empty. We now accept one kind of evidence: the rendered page, compared against the source. It is slower to check and it is the only check that has never lied to us. The general form of this rule is worth keeping: when you ask a model to prove something worked, be specific about what counts as proof, because it will optimise for whatever you named. Name a weak proof and you get a weak proof, delivered confidently. ## Should you trust an AI diagnosis? Treat it as a hypothesis with a good hit rate, and re-verify anything you are about to build on. The sharpest example: four rendering defects were reported and diagnosed across earlier sessions. When we independently re-verified each one before acting, one had been diagnosed accurately and three had been mis-located. Not imagined, the defects were real. The proposed fix sites were wrong. There is a specific trap here. A diagnosis can be correct on Monday and stale by Wednesday, because the code moved underneath it. AI-assisted investigation produces confident, well-written, internally consistent explanations, and confidence is not correlated with freshness. We now re-confirm load-bearing claims at the moment we act on them, and the practice has saved far more time than it costs. ## Where should the checks go? Early in the run, and at a fixed bar. For several generations our system converged incident by incident. Each run surfaced something new, we patched it, and the next run surfaced something else. That is a treadmill, and no amount of model quality gets you off it. What got us off it was checking earlier in the run, and holding the bar steady whenever a check found something. Under pressure, people widen a pass condition until it stops complaining. A bar keeps its meaning only while it stays where you set it. The second half of this rule: make the important checks impossible to skip. One of our important checks could be skipped, so eventually it was, and the gaps it would have caught reached review before we closed them. A check with an off switch is a suggestion. ## What does a human tweak during a run tell you? That you have a missing feature, and you just found it for free. Early on I kept a list of every moment a person had to step in mid-run. Somebody had to mention that a site was animation heavy. Somebody had to notice a mega menu had flattened. Somebody had to give feedback on whitespace. Each of those felt like helpful collaboration at the time. Every one was a defect report. Reading the list that way turned a vague sense that the system needed babysitting into a concrete backlog, and closing that backlog is most of what made later runs quiet. If you are building anything AI-assisted and you find yourself nudging it in the same place twice, stop and write the nudge down. That is your next feature. ## What about the process that grows around the build? It needs pruning on the same schedule as the code. By generation 13 we had accumulated roughly 110,000 words of decision history across 73 documents, plus a whole governance ritual that had made sense when we introduced it and had quietly stopped earning its place. Nobody could read it, which means in practice nobody did, which means the decisions in it were not being honoured. We deleted the ritual and distilled the history into a single short reference that people actually read. That was as valuable as any code change in the same generation. AI-assisted work generates artefacts fast, and documents are artefacts. Left alone they become a layer that looks like rigour and functions as noise. ## What this means for how you scope AI work Three things I now believe more strongly than I did in May. The model is rarely the bottleneck. Clarity is. Almost every failure in my log traces back to a job with unclear edges, a decision asked twice, or a proof standard that was too easy to satisfy. Verification is the product. The generation is the easy half. What separates a demo from something you would put in front of a client is that somebody defined an uncomfortable bar and then made it impossible to route around. Expect architecture problems to arrive disguised as accuracy problems. Every serious turn in our build started as a visual complaint and ended as a modelling decision. If you only measure what is easy to see, you will optimise the wrong half for a surprisingly long time. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [fourteen generations of building it](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system), [moving our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), [what your team gets afterwards](https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### Can AI handle a large build on its own? It handles bounded work against a clear specification very well, and it needs a human to set the bar and design the verification. In our build the generation half went smoothly once each job had clear edges. The judgement about what "correct" means, and the checks that prove it, came from people. ### What is the most common cause of inconsistent AI output? A decision being made in more than one place. If two parts of a build both get to derive the same thing, they will eventually disagree. Making each decision once removed most of our variance. ### How do you stop AI from fixing the symptom? Define acceptance as a clean run reproducing the corrected result. If a fix cannot survive regenerating everything from scratch, it lives in the output and it will disappear. Making that the standing bar changed the behaviour immediately. ### How much should you trust an AI diagnosis of a bug? Treat it as a strong hypothesis and re-verify before you build on it. When we re-checked four diagnosed defects independently, one had been located correctly and three had not. The defects were real, the proposed fix sites were wrong. ### What kind of quality checks work best on AI-generated work? Checks that run early and hold a fixed bar. Make the checks that protect user-facing value impossible to skip, because a check with an off switch tends to be switched off. ### Did AI get the architecture right? Largely yes, and that surprised me. When we audited the design decisions against Drupal 10 and 11 community consensus they lined up well. The gaps were in implementation and verification, so the lesson was to spend the effort on proving the design shipped. ### What would you do differently from day one? Define what "done" means in terms a person outside the project would accept, before optimising anything. Ours started as "does it render" and became "does it render and can an editor work in it". That change was worth more than any accuracy work we did. --- # We moved our own site from Webflow to Drupal > A Webflow to Drupal migration on qed42.com, our own site. 569 articles and 2,007 images moved across, zero custom modules, and the visual result kept intact. - Canonical: https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal - Author: Souvik Pal - Published: 2026-07-28 - Part of: the X to Drupal series by QED42 (post 4 of 7) *A Webflow to Drupal migration, run on qed42.com with X to Drupal, the QED42 harness that replatforms an existing website onto a real Drupal CMS. 569 articles and 2,007 images moved across, and the visual result kept intact.* > **The short version** > - qed42.com ran on Webflow. It runs on Drupal 11 today, replatformed with X to Drupal, and the cutover is done: the main domain now serves the new site. > - The design came across intact. We did not redesign anything as part of the move. > - 569 insight article bodies migrated with their images: 2,007 images moved across. > - The backend is a real content model: paragraph components with typed fields, Views for every listing, real taxonomies, and zero custom modules carrying the site. > - Honest scope: we show you the comparisons and quote no single parity percentage, and animation-heavy sections are the hardest thing to reproduce exactly. We built X to Drupal to move other people's sites onto Drupal. The first site we pointed it at in anger was our own. That was deliberate. If a system claims to replatform a website faithfully, the fastest way to find out whether the claim survives contact with reality is to bet your own homepage on it. Our site is content-heavy, it has real listings with filters, a large insights library, case studies, forms, and a menu structure that matters. It is exactly the kind of site that shows up every weakness. ## Why move off Webflow at all? Webflow did a good job for us. It let a small team ship a strong-looking marketing site quickly, and it kept the design tight. For plenty of companies that is the right answer for a long time. Our reason to move was specific. QED42 builds Drupal for a living, and our own content operations were running somewhere else. That gap showed up in small ways every week: an editor wanting to restructure a listing, a campaign page needing a layout that did not exist, our insights library growing into something we wanted to model, filter and reuse properly. We wanted our own site to work the way we tell clients a content platform should work. Structured content, editors composing pages from real components, listings assembling themselves, workflow and permissions in the platform. ## What came across, and what it looks like The design came across. That is the part people want to see first, so here it is at desktop and mobile. The content came across too, which on a site like ours is the larger job. Our insights library holds hundreds of articles with rich bodies, inline images, captions, code blocks, tables and bylines. All of it moved, and the images moved with it into Drupal's own media and file handling so Drupal manages them as real assets with their own metadata. The numbers on the content side: | What | Count | |---|---| | Insight article bodies migrated | 569 | | Images moved across | 2,007 | | Custom modules in the shipped site | 0 | | Drupal version | 11 | That zero is the one I am most pleased with. It is easy to make a migrated site render correctly by putting the awkward parts in a custom module and moving on. A site carried by a bespoke module is a site the next team has to reverse engineer. Ours runs on Drupal core, contributed modules and configuration, which means anyone who knows Drupal can pick it up. ## What the listings do now This is where a replatform earns its money, and it is invisible in a screenshot of a homepage. Our insights listing is a Drupal View, which means Drupal assembles the list from the content itself. It filters by category, sorts by date, and pages properly. When somebody publishes an article, it appears on the listing, in the related content block, in the "from the team" block on other pages, and in the sitemap. Nobody edits a template. Nobody pastes a card into four places. The same pattern runs the work section. The listing works the same way, each case study is a content type with proper fields, and the related-work and next-case-study blocks assemble themselves from the same content, filtered by whatever you are reading at the time. Add a case study and everywhere it should appear, it appears. Editorial details that used to be baked into the page now derive from the content. Read time is calculated from the article. Categories come from taxonomy terms. That distinction matters because it decides whether content created after the migration behaves like content that came through it. On our site, an article written next month gets its read time and its listing placement exactly like the 569 that came across. ## What an editor sees The test I care about most. Open a page for editing and look at what you get. Each page is an ordered set of paragraph components. Each component has typed fields with native Drupal widgets: a text field for a heading, an image upload for an image, a link field for a call to action, a repeatable group for list rows. An editor can reorder components, add a new one, change an image, and publish. Getting to that was the hardest single decision in the whole project, and the story behind it is in [the version-by-version log](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system). An early generation stored every component as one JSON payload in a single field. It rendered perfectly and it handed editors a textarea full of code. Fixing that reset what we considered finished. ## What was genuinely hard Worth being straight about, because a replatforming story with no rough edges is not a real one. **Animation and video-heavy sections.** A hero with a video background or a scroll-driven animation is intrinsically hard to compare against a static reference, and it is the place where remaining differences cluster. We treat content and structure as non-negotiable and visual reproduction of motion as best effort. **Content that arrived dirty.** Content coming out of any site builder carries markup conventions from that builder. Getting it into clean, structured Drupal content, so the platform's own text formats and filters do the sanitising, took a dedicated pass. Skipping that pass gets you a site that renders correctly and carries a mess in the database. **Details that look like content and are not.** Our homepage had category chips that were part of a filtering mechanism rather than editorial content. Captured naively they become hard-coded content an editor cannot change. Telling those apart is a judgement call, and it is one we now handle deliberately. **Putting one number on parity.** We show comparisons and quote no single parity percentage, and that is a deliberate choice. Parity varies by page, by viewport and by how much motion a section carries, so any average flatters the easy pages and hides the hard ones. A screenshot pair at the same viewport and scroll position tells you more in half a second than a percentage does, and you get to judge it yourself. ## What we would tell someone considering the same move Keep the design fixed. A replatform where the design also changes turns into two projects and a much longer review cycle. Moving as-is and redesigning later, from a platform where redesign is cheap, is the easier sequence. Care about the content model more than the pixels. The pixels are what stakeholders check and the model is what your team lives in. A site that looks right and models badly will frustrate people for years. Ask what happens to content created after the migration. It is the question that separates a genuine platform move from an expensive snapshot. If new content does not behave like migrated content, something has been hard-coded that should have been modelled. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [fourteen generations of building it](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system), [what it taught us about working with AI](https://x-to-drupal.qed42.com/blogs/what-building-this-taught-us-about-working-with-ai), [what your team gets afterwards](https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms), [what makes the delivered site readable to LLMs](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### How long does a Webflow to Drupal migration take? It depends far more on the size of the content library and the number of distinct page templates than on how the design looks. Our own site is content-heavy with several listing types, and the harness run did the transcription work on it in days. For a timeline on your site, the template count and the content volume are the two numbers to send us. ### Will my site look the same after moving to Drupal? That is the design goal, and on our own site the visual result came across intact. Content and structure are treated as non-negotiable. Video backgrounds and scroll-driven animation are the hardest things to reproduce exactly, so those are where any remaining differences sit. ### Do you have to redesign when you replatform? No, and keeping the design fixed is usually the cheaper, lower-risk path. It removes a round of stakeholder review and keeps the project focused on the content model, which is where the long-term value sits. ### What happens to content published after the migration? It behaves like everything else. Read time calculates from the article, listing placement comes from the content type and taxonomy, and related content blocks pick it up automatically. Getting this right is the difference between a platform move and a snapshot. ### Does the migrated site need custom modules to work? Ours does not. The shipped site runs on Drupal core, contributed modules and configuration, with zero custom modules carrying it. That matters because a site held together by bespoke code is a site the next team has to reverse engineer. ### What happens to SEO on a move like this? The risks are URLs, metadata and redirects, and all three are planning work. We kept the URL structure, carried metadata across, and Drupal's SEO tooling covers sitemaps, metadata and path patterns natively. ### Can I see the site? Yes, and it is the site you would have visited anyway. qed42.com serves the replatformed Drupal 11 site today, built with X to Drupal from the Webflow original. The machine-readable surfaces the migration configures are open on it too, so you can check https://www.qed42.com/llms.txt and https://www.qed42.com/jsonapi rather than take our word for any of this. --- # What your team gets when the site becomes a Drupal CMS > What a Drupal CMS gives content teams after a replatform: editor independence, listings that build themselves, SEO foundations, Recipes, and where the work goes. - Canonical: https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms - Author: Souvik Pal - Published: 2026-07-28 - Part of: the X to Drupal series by QED42 (post 5 of 7) *A Drupal CMS for content teams changes who can change the website. This is what lands on day one after a replatform with X to Drupal, where the work goes, and what it does not cover.* > **The short version** > - Your editors compose pages from typed components and publish without a developer or a deployment. > - Drupal 11 and Drupal CMS ship media, scheduling, permissions and SEO foundations configured, so a large part of the site arrives finished. > - The delivered site installs on a fresh Drupal 11 as a Recipe, in one step, so the whole thing is reproducible. > - The transcription work moves into the harness run, so the project's time goes to the judgement calls. The section below says where the work goes. > - Drupal's AI layer arrives configured against your own provider keys, with generated alternative text and metadata drafted per content type. > - The machine-readable surfaces for LLMs arrive with it, and the media library arrives already holding your images rather than empty. > - Deeper AI authoring and a deeper automated SEO audit are the two things still in build, and we label them that way. The pixels are what stakeholders look at when a replatform lands. What decides whether the project was worth doing is what your team can do on the Monday after. So this post covers the after. What changes for the people who own the words, what arrives already built, where the work actually goes, and where the boundaries are. ## What changes on day one for your content team? Editors get to change the site. That sounds like the lowest possible bar and it is the thing most migrated sites fail at. On a site with a real content model, an editor opens a page and sees its sections as a list they can reorder, edit and add to. Each section has typed fields with the right widget: a text field for a heading, an image upload for an image, a link field for a call to action, a repeatable group for list rows. The practical effect shows up in the requests that stop arriving. A campaign page for next week's webinar gets built on Tuesday afternoon by the person running the campaign. Swapping a hero image takes a minute. Reordering a homepage does not need a release. The second change is that listings look after themselves. Any list of content on your site becomes a Drupal View: your insights listing, your case studies, the related-articles block, the "latest from the team" strip on a landing page. Publish an article and it appears in all of them, in the right order, filtered correctly, with the right card layout. Nobody maintains a list of links. The third change is governance. Roles, permissions and scheduled publishing are platform features. An editor writes the piece, schedules it, and it goes live at 9am on Thursday. That is configuration doing the job a process document used to do badly. A review step before anything publishes can be added as part of a custom scope. ## What arrives already built? More than most people expect, and this is the part that has changed most in Drupal over the last few years. A current Drupal install starts close to finished. The site we deliver comes with these working and configured: | Capability | What you get | |---|---| | Scheduling | Scheduled publishing on every content type | | Media management | A media library with focal-point cropping, image styles, responsive images, remote video | | SEO foundations | Metadata management, XML sitemaps, clean URL patterns, redirects | | Structured data | Organisation, website, web page, article and service markup emitted as JSON-LD from the content model | | AI in the CMS | The AI module suite wired to your own provider keys, with assistive authoring, generated alternative text and metadata drafted per content type | | Forms | A full form builder with spam protection, so marketing builds its own forms | | Admin experience | A modern admin theme with a dashboard and quick navigation | | Performance | Page and render caching with an object cache behind them, WebP conversion, focal-point cropping | | Privacy | Consent management for cookies and embeds | | Permissions | Role-based access down to individual fields | | Security | An enforcing content security policy, strict transport security carried over where your current site already sends it, and an editor role that is not an administrator role | None of that is bespoke work on your project. It is the ecosystem, configured. The practical significance is straightforward: capabilities that used to be separate pieces of project work are now setup steps. ## Where does the work actually go? I am going to describe this without putting a number on it. That is a deliberate choice. How it plays out depends on your content library, your template count, and how much of the current site survives review, so any single figure I gave you would travel further than it deserves and be wrong for most readers. What I can describe precisely is where the work goes. A replatform breaks into four buckets of work: re-implementing the design, deriving the content model, moving the content, and everything else. Three of those four are transcription. The information already exists in the current site and a person is retyping it into a new system with high standards. That is where the hours sit, and that is the part X to Drupal removes. What stays is the work that needs a person: deciding whether the current information architecture is still right, handling the integrations, planning the redirect map and the cutover, quality assurance, and the judgement calls about content that looks editorial and is actually mechanism. So the shape of the change is that a project stops being mostly transcription with some judgement, and becomes mostly judgement. The other effect worth naming: the senior architect whose time was going into modelling a site that already exists gets to spend it on what you are building next. ## How does SEO come out of a move like this? A platform move is the moment SEO risk concentrates, so it gets treated as a workstream and not a checkbox. Three things carry the risk: URLs, metadata and redirects. We keep the URL structure where we can, carry metadata across with the content, and put a redirect map in place for anything that has to change. Drupal covers all three natively: path patterns, metadata management per content type, and a redirect module that logs what is being hit. The delivered site arrives with the SEO foundations already configured: metadata management per content type, XML sitemaps, clean path patterns and redirects. It also emits structured data from the content model, so organisation, website, web page, article and service markup comes out of the fields and stays in step with what a visitor reads. The surfaces an LLM reads arrive with them, which the [post on LLM readiness](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search) goes into properly. A deeper automated audit is still in build, and the section below is honest about where that has got to. ## What is a Recipe and why should you care? This is the plug-and-play part, and it is the most underrated thing in the delivery. A Drupal Recipe is a packaged set of configuration and content that applies to a Drupal site in one step. The site we deliver comes as a Recipe, so it installs onto a fresh Drupal 11 with a single command and produces the site: content types, fields, paragraph components, Views, taxonomies, menus, text formats, the theme, and the content. What that gives you in practice: **Reproducibility.** Any developer can stand up the whole site from scratch, today or in two years. There is no undocumented state and no "ask the person who built it". **A clean handover.** Whether we keep working with you or your in-house team takes it on, what changes hands is a standard Drupal site with its configuration in version control. **Extension in one step.** The wider Recipe ecosystem works the same way, so adding a capability later, an events section, a case-study type, a privacy pack, takes an install. Getting the Recipe right took real work. A good Recipe holds only what your site needs, so it installs in one step and stays easy to read. That is what ships now. ## What ships with the AI layer today? The suite arrives installed and pointed at your own provider keys, so the capability belongs to your account and your CMS. What that gives your editors on day one: assistive authoring and content suggestions inside the editor, working on fields with a known purpose; alternative text generated for images that arrived without any; and titles, meta descriptions and social metadata drafted per content type from the content itself, with an editor approving. More than one provider is installed, so you choose which one answers. Two details behind it are worth a line each. Keys are read from the environment, so they never travel in a configuration export or a database dump. And the layer logs which model was called, for which task and when, while leaving the content of prompts and responses out of the log, which answers a spend question and a privacy question with the same record. The governance point is the one I would lead with in an internal pitch. An AI-drafted piece goes through the same roles and the same permissions as anything a person wrote, because it is the same content model underneath. ## What arrived since this post was written? This section used to name three capability packs and say all three were roadmap. One of them has shipped, and a few things arrived that were not on the list at all, so it is worth updating rather than quietly leaving stale. **The machine-readable pack now arrives with the site.** An `llms.txt` map, a markdown copy of every page, read-only JSON per content item from Drupal's own API layer, and crawl rules that name the AI agents individually. It is on the site we replatformed, so you can open `qed42.com/llms.txt` or `qed42.com/jsonapi` rather than take it on faith. [The LLM readiness post](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search) covers what each surface does and why the structured data is the foundation they sit on. **The media library arrives holding your images.** Every image the site uses is in the library on day one, ready for an editor to find and reuse. On our own site that is 1,334 media items. This is the change your content team notices first. **The image handling is proven before it is relied on.** A server can accept an image format and still be unable to produce a cropped copy of it, which fails quietly: the visitor sees the original, nobody sees an error, and the cropping you specified never happens. The delivery now tests that the server can produce what the site asks for, and says so plainly if it cannot. **Editor tooling is served from your own site.** An enforcing security policy and an editing widget that loads its code from an outside network are in direct conflict, and the widget loses. Those libraries now come from the site itself, so no admin screen depends on an outside network being reachable. **Verification signs in.** The checks used to run as an anonymous visitor, which cannot see a widget that only breaks once you log in. They now cover the admin screens your team will actually use, and every width your design declares rather than a fixed set. **Still in build:** deeper AI content authoring past the drafting and suggestions above, meaning expansion at length, summaries, taxonomy suggestions across a whole library and translation help; and a deeper automated SEO audit covering the metadata that came across, sitemap and redirect coverage, structured data and the crawl surface. Both arrive the way the site does, as Drupal Recipes, so adopting one is an install on a site you already have. On timing I will point you at a conversation. If either matters to your decision, ask us where it has got to, because a date on a blog post ages badly. ## Where does this fit best? It fits well when you have a site that basically works and a team frustrated by how hard it is to change. Content-heavy marketing sites, insight and resource libraries, case-study driven sites, multi-brand setups where the same components repeat. It fits well when you want to keep the design. A move that holds the design fixed removes a round of stakeholder review and keeps the focus on the model. It fits less well when you already know the site needs a redesign and a rethink of its information architecture. In that case a move gets you a faithful reproduction of something you were about to change, and starting from the design work is the better sequence. Worth saying plainly, because the honest answer is sometimes "not yet". It also has boundaries worth knowing about up front. Sites built heavily on scroll-driven animation and video backgrounds are the hardest to reproduce exactly. Content behind authentication, single-page applications and infinite-scroll interfaces need a conversation before anyone quotes anything. ## What to ask any vendor doing this Four questions that separate a platform move from an expensive snapshot. They apply to us too. 1. **What happens to content published after the migration?** If new content does not behave like migrated content, something got hard-coded that should have been modelled. 2. **Can an editor change this section without a developer?** Ask them to open the edit form and show you. This is where a JSON payload in a textarea gets found. 3. **How many custom modules is the site relying on?** The lower the number, the easier your site is for anyone else to maintain. Ours ships at zero. 4. **Can you rebuild the whole site from scratch, right now?** If yes, the configuration is real. If it needs a database somebody is protecting, it is not. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [fourteen generations of building it](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system), [what it taught us about working with AI](https://x-to-drupal.qed42.com/blogs/what-building-this-taught-us-about-working-with-ai), [moving our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), [what makes the delivered site readable to LLMs](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### Can content editors change layouts in Drupal without a developer? Yes, when the content model is built for it. Editors see each page as an ordered list of components with typed fields and native widgets, so they can reorder sections, add new ones and change content directly. A weak content model takes that freedom away, which is why the modelling decisions matter more than they look. ### What is a Drupal Recipe? A packaged set of configuration and content that applies to a Drupal site in one step. The site we deliver comes as a Recipe, so it installs onto a fresh Drupal 11 with a single command and produces the whole site: content types, fields, components, Views, taxonomies, menus, theme and content. ### Where does the work go with this approach? The transcription work moves into the harness run: re-implementing the design, deriving the content model, and moving content. The judgement work stays with people: whether the information architecture is still right, the integrations, the redirect map and the cutover, and quality assurance. ### Do the AI features ship today? Drupal's AI module suite arrives installed and wired to your own provider keys, with assistive authoring in the editor, alternative text generated for images that arrived without it, and metadata drafted per content type. The machine-readable surfaces an LLM reads arrive configured with it: an llms.txt map, a markdown twin of every page, read-only JSON per content item, and crawl rules naming the AI agents individually. Deeper authoring work such as long-form expansion, summaries and taxonomy suggestions across a library is still in build. For timing on that part, ask us where it has got to. ### Does an SEO audit come with the migration? The SEO foundations arrive configured on the delivered site: metadata per content type, XML sitemaps, path patterns, redirects, and structured data emitted from the content model. The machine-readable surfaces an LLM reads arrive with them, so the site is legible to an assistant on day one. A deeper automated audit covering sitemap coverage and the crawl surface is still in build, so that part is roadmap today. ### Who is a good fit for this? Teams with a content-heavy site that basically works, who are frustrated by how hard it is to change, and who want to keep the current design. Insight libraries, case-study sites and multi-brand setups with repeating components fit especially well. ### When is a replatform the wrong move? When you already know the site needs a redesign and a rethink of its information architecture. Moving first gets you a faithful copy of something you were about to change, so starting from the design work is the better sequence. --- # What makes a Drupal site readable to LLMs > Drupal LLM readiness after a replatform: the structured data, AI layer and machine-readable surfaces that arrive configured, and how to check your own site today. - Canonical: https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search - Author: Souvik Pal - Published: 2026-08-14 - Part of: the X to Drupal series by QED42 (post 6 of 7) *Drupal LLM readiness is the work that makes your pages quotable by an assistant: structured data emitted from the content model, metadata that agrees with what a visitor reads, and a copy of the content a machine can take in whole.* > **The short version** > - Assistants answer from text they can lift cleanly, so the shape of your markup decides how often you get quoted. > - A replatformed site arrives with the structured data already coming out of the content model: organisation, website, web page, article and service, generated per page from fields. > - Drupal's AI layer arrives configured against your own provider keys, with generated alternative text and per-content-type metadata drafting. > - The machine-readable pack arrives with it: an llms.txt map, a markdown twin of every page, read-only JSON per content item, and crawl rules that name the AI agents one by one. > - All four are live on the site we replatformed, so you can open them instead of believing me. > - You can check every one of these on your own site this afternoon, with a browser and nothing else. Two years ago the question about a website was where it ranked. The question I get asked now is whether it gets quoted, and by what. Those turn out to be different problems with a lot of shared plumbing. A search engine indexes your page and sends someone to it. An assistant reads your page, decides whether the passage it needs is clean enough to lift, and either cites you or moves to a source that was easier to read. The second one rewards a kind of tidiness that most sites never had a reason to care about. This post is about that tidiness on a Drupal site: what a replatform lands with, how you can measure where you stand today, and the one piece we are still building. ## Why does an assistant quote one site and skip another? Because one of them made the answer easy to lift, and lifting is most of the job. When an assistant builds an answer it is looking for a passage it can attribute with confidence: a clear question, a self-contained answer near it, a heading that says what the section is, and some machine-readable declaration that the page is what it appears to be. Sites that publish that way get pulled into answers often. Sites where the same information is spread across four hover states and an accordion get pulled in less, because the assistant has to reconstruct meaning that was carried by layout. Three things move the needle most, and all three are content-model problems wearing a marketing hat. **The page says one thing.** One h1, headings that describe their sections, an opening paragraph that defines the subject in plain words. That is the passage an assistant quotes, so it should be a definition and not a warm-up. **The markup agrees with the prose.** Structured data that says "this is an article, published on this date, by this author, part of this organisation" gives the machine a frame for the text. When the structured data and the visible text disagree, the page gets trusted less, which is why generating it from fields beats pasting it per page. **The content is reachable as text.** Content assembled in the browser after load is harder for a crawler to see than content that arrives in the HTML. This is where a lot of otherwise good sites quietly lose ground, and it is one of the strongest arguments for a server-rendered CMS. ## What arrives configured on the delivered site today? The structured data and the AI layer, both wired to the content model and both working on handover. On the structured data side, the delivered site emits JSON-LD from the fields themselves: organisation, website, web page, article and service. Because it comes out of the content model, a change to a headline changes the markup in the same edit. What a machine reads and what a visitor reads stay in step, which is a Google requirement and, more usefully, the thing that keeps the markup honest a year later when nobody remembers it exists. Alongside it, the SEO foundations arrive configured: metadata per content type, XML sitemaps, clean path patterns and a redirect map with logging. On the AI side, the delivered CMS arrives with Drupal's AI module suite installed and pointed at your own provider keys. Three things it does on day one: | Capability | What it does on the delivered site | |---|---| | Assistive authoring | Drafting and content suggestions inside the editor, working on typed fields | | Media descriptions | Alternative text generated for images that arrived without any, which is an accessibility fix and a search fix in one pass | | Metadata drafting | Titles, meta descriptions and social metadata generated per content type from the content itself, with an editor approving | Two details in there matter more than the features. The keys are yours, read from the environment, so they never travel in a configuration export or a database dump, and the capability belongs to your account. And the layer logs which model was called, for which task and when, while leaving the content of prompts and responses out of the log, so you can account for spend and answer a privacy question with the same record. ## What is llms.txt, and why does it sit beside robots.txt? It is a plain-text file at the root of a site that tells an assistant what the site is and where its canonical pages live. Think of the family it belongs to. `robots.txt` tells a crawler where it may go. `sitemap.xml` tells it what exists. `llms.txt` tells a model what the site is for and which URLs carry the substance, in prose a model can read in one pass. It is a young convention that the ecosystem is still settling, and it is cheap to publish, which is a good combination for something with this much upside. The fuller version of the idea goes past one file. A markdown copy of each page gives a model the prose and the headings with the navigation, scripts and styling out of the way. A JSON representation of each content item, straight from Drupal's own API layer, makes the content consumable as data by an assistant, an agent or an internal tool. And explicit crawl directives per bot let you say yes to the LLMs you want to appear in while keeping the controls you already have. ## What does the delivered site actually arrive with? That whole machine-readable pack, and this section used to say the opposite, so it is worth being precise about what changed. When I first wrote this post the pack was in build and I labelled it that way. It has since shipped, and it is not shipped in the sense of a demo on our own marketing site. It is on the Drupal site the harness produced. qed42.com ran on Webflow; it runs on Drupal 11 today, replatformed with X to Drupal, and the cutover means the four surfaces are serving from the address you would have visited anyway: - `https://www.qed42.com/llms.txt`: the map, in plain text. - `https://www.qed42.com/insights.md`: the markdown twin of a page, with the navigation and styling gone. Every page has one. - `https://www.qed42.com/jsonapi`: the content as data, exposed for reading only. Writes are not available and the accounts resource is switched off, which is the part most people forget to check when they open an API on a public site. - `https://www.qed42.com/robots.txt`: the crawl rules, naming the AI agents individually rather than waving at them collectively. Go and open them. That is the only reason I list the addresses rather than describe the capability: a claim you can check in four clicks does not need me to be persuasive about it. One thing is still in build, and it keeps its label: a deeper automated SEO audit, meaning a standing check of sitemap coverage, the crawl surface and metadata quality. That is an extension of what is already on the delivered site rather than the foundation for it, so nothing above is waiting on it. This site, the one you are reading, still publishes its own `/llms.txt`, an `/llms-full.txt` corpus with every article inline, and a markdown copy of every post at `/blogs/.md`. That started as a demonstration of a mechanism we had not yet delivered. It is now just consistency. ## How can you check your own site this afternoon? Four checks, a browser, about twenty minutes. None of them need a tool. 1. **View source on your best page and search for `application/ld+json`.** If nothing comes back, you have no structured data. If something comes back, read it and ask whether it matches the page you are looking at. 2. **Turn JavaScript off and reload.** What survives is roughly what a crawler sees first. If your main content disappears, that is the biggest single item on your list. 3. **Fetch `/llms.txt` and `/sitemap.xml`.** The second one most sites have. The first one almost nobody has yet, which is exactly why publishing it is worth an afternoon. 4. **Read your first paragraph aloud and ask if it defines the subject.** If it is a warm-up, the passage an assistant would quote is somewhere further down the page, and it may never get there. The results tend to sort into two piles: things that are a content edit, and things that are a platform limit. The first pile you can clear next week. The second pile is the reason a replatform comes up at all. ## Where does this fit in a replatform? At the point where you are already paying to touch every page. The awkward economics of AI readiness on its own is that it asks for changes across the whole site for a benefit that shows up over several quarters. That is a hard sell as a standalone project and an easy one as part of a move you were doing anyway, because the content model you build during a replatform is the thing that generates the markup afterwards. Model the content properly once, and the structured data, the metadata and the machine-readable copies all fall out of it. Skip that step, and you are back to pasting markup per page and watching it drift. So the sequence I would argue for is: move to a platform that renders on the server and holds a real content model, get the structured data coming out of the fields, and let the machine-readable copies fall out of the same model. That order means every step is useful on its own, which is the test I apply to any roadmap I am asked to believe. When I first wrote it, the third step was a promise about a later pack. It is now the same delivery as the second, which is the version of that argument I always wanted to be making. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [moving our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), [what your team gets afterwards](https://x-to-drupal.qed42.com/blogs/what-your-team-gets-when-the-site-becomes-a-drupal-cms), and [how we check the result really matches](https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches). --- ## Frequently asked questions ### What is llms.txt and does my site need one? It is a plain-text file at the root of a site that tells an assistant what the site is and lists its canonical URLs, sitting alongside robots.txt and sitemap.xml. It is a young convention and cheap to publish, so it is worth doing early on a content-heavy site where being quoted has real value. ### Does structured data help with LLMs or only with Google? Both. Structured data gives any machine a frame for the text on the page: what kind of thing it is, who published it, when, and how it relates to the rest of the site. Search engines use it for rich results and assistants use it to attribute a passage with confidence. ### Which AI features arrive configured after a replatform to Drupal? Drupal's AI module suite arrives installed and wired to your own provider keys, with assistive authoring in the editor, alternative text generated for images that arrived without it, and metadata drafted per content type. The layer logs the model, the task and the time, and leaves prompt and response content out of the log. ### Is the machine-readable pack available today? Yes. The llms.txt map, a markdown twin of every page, read-only JSON per content item and per-bot crawl rules all arrive configured on the delivered site, alongside the structured data and the AI layer. You can check all four on the Drupal site we replatformed: qed42.com serves llms.txt, a markdown twin of every page, a read-only JSON API and a robots.txt that names the AI crawlers individually. A deeper automated SEO audit is the one piece still in build, and we label it that way wherever it appears. ### How do I tell whether my current site is readable by an assistant? View source and search for application/ld+json to see whether structured data exists, turn JavaScript off and reload to see what a crawler gets first, fetch /llms.txt and /sitemap.xml, and read your opening paragraph to check that it defines the subject. Those four checks take about twenty minutes and need only a browser. ### Why does server-side rendering matter for LLMs? Content that arrives in the HTML is available to any crawler on the first request. Content assembled in the browser after load asks the crawler to execute scripts before it sees anything, and coverage there varies by bot. A server-rendered CMS puts the text in front of every reader, human and machine, on the same request. --- # How we check a replatformed site really matches > A website migration parity check past screenshots: side by side at three widths, every published address accounted for, and content traced unit by unit. - Canonical: https://x-to-drupal.qed42.com/blogs/how-we-check-a-replatformed-site-really-matches - Author: Souvik Pal - Published: 2026-08-14 - Part of: the X to Drupal series by QED42 (post 7 of 7) *A website migration parity check is the evidence that the new site carries everything the old one did: the same pages at the same widths, every published address answering, and every heading, paragraph, link label and image accounted for by name.* > **The short version** > - We show the two sites side by side at 1440, 768 and 375 pixels wide, the widths we publish. > - We publish no parity percentage. An average flatters the easy pages and hides the hard ones, so you get the pictures and your own judgement. > - Every address the old site advertises is checked against everything the new site can serve, redirects included. On qed42.com that list ran to 670 addresses. > - Content is traced unit by unit, so anything that did not arrive is named with a reason. > - A second pass compares the new site with the old one as it stands today, because the source keeps publishing while you migrate. The worst moment in a migration is three weeks after launch, when someone in the content team searches for a page they wrote in 2022 and it is gone. Nobody signed that off. Nobody noticed. It was in a category listing that quietly stopped existing, and the sign-off looked at nine pages out of six hundred. So the question this post answers is the one I would ask any vendor, us included: what did you actually check, and how would I know if you had missed something? ## Why do screenshots on their own fall short as proof? They prove one claim very well and stay silent on everything else. A pair of full-page captures at the same width, old on the left and new on the right, settles the visual question in about half a second. Your eye is better at this than any tool: you see the wrong heading weight, the card that lost its border, the section that came out in the wrong order. That is why the comparison is the centrepiece of how we present a migration, and why we show whole pages and let you scroll them. What that pair says nothing about is whether the other 600 pages came across, whether the addresses people have bookmarked still answer, or whether the third paragraph of a 2019 article survived the move. Those need counting, and counting is where the value of automated verification actually sits. So the parity story has two halves: the pictures, which you judge, and three checks behind them, which produce lists. ## Why do we publish no parity percentage? Because a single number is the most confident and least useful thing we could put on the page. Fidelity varies by page, by viewport and by how much motion a section carries. A site's contact page will match almost exactly. Its homepage, with a background video and a scroll-driven reveal, will be the hardest thing in the project. Average those and you get a figure that is flattering and technically defensible, and that leaves the decision you are actually making no better informed. There is a second reason, and it is the one I feel more strongly about. Publishing a score invites everyone to optimise the score. The moment a number goes on a slide, the incentive shifts from fixing the hard section to finding a measurement that treats it kindly. Showing the pictures keeps the incentive where it belongs. Capturing those pictures fairly turns out to take its own discipline. Background video, motion and listings that keep loading as you scroll all need care, so that both sides of every pair show the same moment and the same length of page. Getting that right is what lets a pair prove something. ## How do you know every page came across? You start from the old site's own list of what it publishes, and you make the new site answer for all of it. The check is boring and it is the one that catches real defects. Take the full set of addresses the old site advertises, from its sitemaps, its navigation, its listings and its internal links. Ask the new site for every one of them. Count a redirect as an answer, because a moved page that lands correctly is a moved page. Anything left over goes on a defect list with its address on it. On qed42.com that list ran to 670 addresses, and the check found 14 category listings that had no route on the new site. We built them. Nobody had noticed they were missing, and no visual comparison would have surfaced them, because they were pages that nobody had thought to look at. Two things make this check worth its cost. It runs against a machine-readable list, so it is complete by construction. And its output is a worklist with addresses on it, which is a thing a team can finish, unlike a score. ## What does tracing content unit by unit mean? Every heading, paragraph, link label and image reference gets followed from the reading of the old page through to the rendered new one. Sampling is the usual approach and it has an obvious hole: you find the classes of problem your sample happened to contain. Tracing every unit finds the ones that only affect a handful of pages, which are exactly the ones that survive to launch. The two that show up most often are the old article with an embedded table that only exists on eleven pages, and the image whose caption lived in a text field that had no equivalent in the new model until someone looked. When a unit does not arrive, it goes on a list with its name and a reason: the source had it inside a component with no counterpart, the source had it as a hard-coded block, the source has since dropped it. A named list with reasons is something a content lead can work through in an afternoon, which is why the check reports names and reasons and leaves the summing up to you. ## What happens to content published after you started? It comes back as a worklist, because the source site keeps working while the migration runs. This is the part most migration QA skips, and it causes more post-launch surprises than anything else on this page. A migration of a content-heavy site takes weeks. During those weeks the marketing team publishes, retitles, unpublishes and reorganises, all on the old site, all after the snapshot everything else was verified against. So there is a second pass, run close to cutover, that compares the new site with the old one as it stands that day, page by page. What comes out is three lists: published since the snapshot, changed since the snapshot, and gone since the snapshot. Your content team works those lists, which is far cheaper than re-reading the whole site and much safer than assuming nothing moved. ## What should you ask any vendor about verification? Five questions, and they apply to us. The answers tell you whether verification is a workstream or a slide. 1. **Where did your list of pages come from?** If it came from the old site's own published addresses, the check can be complete. If it came from a spreadsheet somebody typed, it is a sample with ambitions. 2. **Show me a page that failed.** A verification pass that found nothing found nothing because it was not looking. Ask to see the defect list and what happened to each item. 3. **What are your widths?** Ask which viewports were verified at and see the captures at each. A desktop-only sign-off on a site where most traffic is mobile is a sign-off of the wrong thing. 4. **How are redirects counted?** A moved page that lands correctly is a pass. A page that 200s with the wrong content is a defect wearing a pass, and only a content check will find it. 5. **When was the last comparison run?** If the answer predates the last month of publishing on the source site, ask for one more pass before cutover. The honest summary of all this is that verification is mostly bookkeeping, and bookkeeping is what makes the difference between a migration that lands and one that leaks. The pictures earn the meeting. The lists earn the launch. This post is part of the X to Drupal series, alongside the pillar, [why replatform into Drupal and what X to Drupal provides](https://x-to-drupal.qed42.com/blogs/why-replatform-into-drupal-and-what-x-to-drupal-provides), and the companions on [fourteen generations of building it](https://x-to-drupal.qed42.com/blogs/fourteen-generations-of-a-website-to-drupal-system), [moving our own site off Webflow](https://x-to-drupal.qed42.com/blogs/we-moved-our-own-site-from-webflow-to-drupal), and [what makes the delivered site readable to LLMs](https://x-to-drupal.qed42.com/blogs/what-makes-a-drupal-site-readable-to-ai-search). --- ## Frequently asked questions ### How do you verify a website migration? With a visual comparison and three counting checks. The two sites go side by side as full pages at 1440, 768 and 375 pixels wide. Every address the old site advertises is checked against everything the new site can serve, redirects included. Content is traced unit by unit. A final pass compares the new site with the old one as it stands that day. ### Why do you not publish a parity percentage? Fidelity varies by page, by viewport and by how much motion a section carries, so an average flatters the easy pages and hides the hard ones. Publishing a score also shifts the incentive towards a kinder measurement. The side-by-side comparison is the claim, and you are free to judge it harshly. ### What happens to pages that have no equivalent on the new site? They go on a defect list with their addresses, and they get built or redirected before launch. On qed42.com the address check ran to 670 addresses and found 14 category listings with no route on the new site, which we then built. ### Do redirects count as a page arriving? Yes. A moved page that lands correctly on the right content is a pass, and the check counts it as one. A page that returns a 200 with the wrong content is a defect, which is why the content trace runs alongside the address check. ### How do you handle content published while the migration is running? A second pass close to cutover compares the new site with the source as it stands that day and produces three lists: published since the snapshot, changed since it, and gone since it. Your content team works those lists, which is much cheaper than re-reading the whole site. ### Which parts of a site are hardest to reproduce? Sections built on scroll-driven animation and background video are the hardest, because their appearance depends on motion and a still comparison can only show the resting state. We show those sections in the comparison and say so, and content behind authentication or inside a single-page application needs a conversation before anyone quotes.