13 min read

Fourteen generations of a website to Drupal system

Engineers, tech leads, and the Drupal community

The journey
ShareXLinkedIn

Building an automated website to Drupal migration system took fourteen full generations in about three months. This is the version-by-version account of what broke, what we decided, and which decision mattered most.

The short version

  • Fourteen numbered generations between 11 May and 21 August 2026.
  • Reading a website accurately turned out to be the easy half. Deciding what the content model should be was the hard half.
  • Generation 6 was the turn that mattered: components stored as one blob rendered perfectly and were useless to an editor.
  • Generation 12 produced the rule that made everything hold: a fix that lives in the output is not a fix.
  • This is a build log, so it covers decisions and outcomes. It does not cover how the system is put together.

Three months ago I ran a pipeline end to end for the first time against qed42.com and got something that half worked. Fourteen generations later it produces a Drupal site I am willing to put in front of an architect. The distance between those two sentences is the interesting part, and almost none of it went where I expected.

The pattern that repeats through the whole log: every time I thought the problem was accuracy, the problem was actually architecture. Every time I thought the problem was the model, the problem was the contract I had given it.

Generations 1 to 4: reading the site was never the hard part

The first two generations read the site well, but much of what they learned was lost before the build. Fonts stayed as placeholders and never made it into the theme. Background images went missing. Setup errors blocked verification before anything could be judged.

Generation 3 added a second, structurally very different reference site, and the same root causes appeared on both. That was the first genuinely useful signal. Two very different sites failing identically means the failure is in the system, so the fix belongs in the system.

Generation 3 also produced a list I still find useful, of every moment a human had to step in during a run:

  • Somebody had to say "this site is video and animation heavy"
  • Somebody had to notice that a six-category mega menu had come through as a flat list
  • Somebody had to notice a newsletter modal was missing entirely
  • Somebody had to fix spacing and layout by hand

Every one of those is a missing automated step wearing a human's clothes. We turned the list into work, and it became the rule for the rest of the project: a human tweak during a run is a defect report against the system.

Generation 4 settled an architectural question that kept coming back, and it settled it in favour of consistency: the same input must give the same result on every run. Variance is the thing you can least afford when you are trying to establish whether a change helped.

Generation 5: the ceiling nobody could polish away

By generation 5 a clean run came within touching distance of the bar on first pass and then stopped. There was no mechanism to close the gap, so it stayed open.

We gave the finish a clear bar and a clear point to stop. Left open-ended, polish will happily spend an afternoon making a footer imperceptibly better.

Generation 6: the one that mattered

Generation 6 is the turn the whole project pivots on, and it started as a complaint about a footer colour.

While looking at that, I looked properly at what the backend had become. Every component was one paragraph type with a single long-text field holding the entire payload as JSON. The theme decoded the JSON at render time and fed it to the component. It rendered beautifully. Every visual check was happy.

And an editor opening that page saw a textarea full of JSON.

That defeats the entire point of moving a site to a CMS. Drupal's whole story is typed fields with native widgets: a text field for a heading, an image upload for an image, a link field for a link, a repeatable group for list rows. A JSON blob has none of that. We had built something that passed every check we had and delivered none of the value we were selling.

So generation 6 rebuilt the content model around typed fields, the way a senior Drupal architect would build it by hand. From that point on, the bar stopped being "does it render" and became "does it render and can an editor work in it". Everything after generation 6 is downstream of that sentence.

Generations 7 and 8: decisions, then checks

Generation 7 fixed a drift problem. The same decision was being made in several places, with the model making a judgement call each time, so it came out differently from site to site. That is why fixes never seemed to stick. We now make each decision once, and the rest of the run follows it.

If I could keep one lesson from this log, it would be that one. Decide once. The model still does plenty of work, and all of it now follows a decision that is already made.

Generation 8 added automated checks early in the run, so a structural problem is caught before any effort goes into the finish. The reason was simple. The system had been converging incident by incident. Every run surfaced a new unhandled shape and we patched reactively, which is a treadmill.

At that point the measured difference was worst at tablet, close behind at mobile and smallest at desktop, and all three sat well outside the bar we had set. Uncomfortable readings, and useful ones, because they were finally honest and repeatable.

Generations 9 to 11: measuring the right thing

Generations 9 and 10 were about accuracy levers, and they produced a finding that reframed the remaining work. The remaining differences came from content and structure. Finish work only closes a gap when the content is already there. We had been trying to fix a missing-content problem with styling.

Generation 11 fixed something I am slightly embarrassed about. Multi-page support was, in practice, theatre. Pages were enumerated one at a time, the sitemap was never read, and per-page changes to the header and footer were being lost. Fixing that properly meant treating the whole site as the unit of work.

Generation 12: the rule that made everything hold

Generation 12 ran across five sub-versions and produced the discipline that finally stopped the regressions.

Fix after fix had appeared to land and then quietly disappeared on the next clean run. The reason was always the same: the fix had been applied to the output of a run, so the thing that generated the output still produced the old result. Every clean run overwrote it.

Two rules came out of it, and they are now the ones I care most about:

A fix belongs in the generator, never in the output. If you cannot reproduce the corrected result from a completely clean run, you have not fixed anything. You have decorated one run.

Verify what renders, never what is present. Config being correct is not proof. Markup existing in the DOM is not proof. The only acceptable evidence is a rendered page compared against the source. We had several diagnoses go stale between sessions because they were verified against a hand-patched state that a clean run never produces.

Generation 12 also caught something worth saying out loud: some of our own earlier diagnoses were wrong. When we re-verified four reported defects independently, one had been diagnosed accurately and three had been mis-located. Re-verifying load-bearing claims before acting on them became standard, and it has saved more time than it costs.

Generations 13 and 14: the architect bar, then the package

Generation 13 rebuilt the Drupal backend to the shape a senior architect would choose. The decisions, in the order they matter:

  • Page narrative lives in ordered paragraph components on the node, in source order
  • Global chrome such as header, footer and menus lives in blocks placed in theme regions, because it is genuinely global
  • Any list of existing entities becomes a View. Landing-page narrative is paragraphs; listings are Views
  • One flexible page type covers one-off and landing pages, with distinct content types reserved for genuinely repeating templates
  • Real taxonomy vocabularies with real terms, so filters and reporting work

Then we audited it hard, across the whole system. The audit confirmed 22 gaps and refuted five. The meta-finding was the useful part: the designs lined up with Drupal 10 and 11 community consensus, and the failures were implementation and verification gaps. We were not wrong about what to build. We were wrong about how thoroughly we had built it.

We also deleted a governance layer that had grown around the project, and distilled roughly 110,000 words of accumulated decision history into a single short reference that people actually read. Process accumulates like debt, and it needs paying down on the same schedule as code.

Generation 14 tackled the thing nobody had been looking at: the package we hand over. The architecture was right, and the package was heavier than it needed to be.

So we trimmed it to a lean install list, configuration that is all the site's own, and a theme in small lint-clean files. The editor-experience check became a required step, the last of the site's own data moved into real Drupal configuration, and the shipped site reached zero custom modules. Editorial dynamics like read time now derive at render from real fields, so content created after the migration behaves the same as content that came through it.

What I would tell someone starting this

Three things.

Decide what "done" means before you start optimising, and make it something a person outside the project would accept. Ours moved from "does it render" to "does it render and can an editor work in it", and that change was worth more than any accuracy work.

Put the checks early. A wrong result caught early costs minutes. The same result caught at the end costs a run.

Expect the architecture problems to arrive disguised as accuracy problems. Every serious turn in this log started as a visual complaint and ended as a modelling decision.

This post is part of the X to Drupal series, alongside the pillar, why replatform into Drupal and what X to Drupal provides, and the companions on what this taught us about working with AI, moving our own site off Webflow, what your team gets afterwards, and how we check the result really matches.


Frequently asked questions

Can AI handle a full website migration on its own?

It can do a great deal of the work when the work is bounded. In our experience the model performs well when the decision is already made and it only has to execute it. It performs poorly when it is asked to re-derive a decision that was already made upstream, because the answer drifts between runs.

Why did it take fourteen generations?

Because the first several generations solved the visible problem, which was visual accuracy, while the real problem was the content model. Generation 6 reset the definition of done and much of the earlier work had to be redone against the new bar.

Does the migrated site come with a proper content model or just pages?

A content model. Page narrative sits in ordered paragraph components with typed fields, any listing of entities is a View, global chrome is block content in theme regions, and categorisation uses real taxonomy vocabularies with real terms.

What was the single most useful engineering rule you found?

A fix belongs in the generator, never in the output of a run. If a completely clean run does not reproduce the corrected result, the fix is decoration. This one rule ended a long run of regressions that had been costing days.

How do you verify a migration is correct?

By comparing what renders against the source site. Configuration being present, or markup existing in the DOM, is a weaker signal that produced several false positives for us. Rendered comparison is the only evidence we now accept.

Is a system like this site-agnostic, or tuned to one site?

Ours is built to be site-agnostic. We develop it against several very different sites, so a fix that suits only one of them shows up as a failure on the others.

Written by Souvik Pal, QED42.

ShareXLinkedIn