12 min read

What building this taught us about working with AI

Engineers, tech leads, and architects putting AI on real production work

AI learnings
ShareXLinkedIn

AI code generation reliability comes from the constraints you put around the model, and these are the eight rules that survived three months of building a system that replatforms websites onto Drupal.

The short version

  • The model does its best work on a small, clear job whose decision is already made, and drifts when asked to make that decision again.
  • One input, one truth. Drift is usually a clarity problem wearing a model problem's clothes.
  • A fix that lives in the output is not a fix. It has to live in the thing that produces the output.
  • Rendered result is the only evidence. Configuration being present proved to be a weak signal that cost us days.
  • Three of four AI diagnoses we re-checked pointed at the wrong place. Re-verifying load-bearing claims pays for itself.

I spent about three months building a system that takes a live website and produces a Drupal site from it, with the visual result preserved and the content model built properly. Fourteen generations in all. Most of the interesting failures had nothing to do with the model being weak. They came from me giving it a job with unclear edges.

These are the rules I would give myself on day one. They apply to any large AI-assisted build, and none of them are about prompting.

What does AI actually do well on work like this?

Small, clear jobs with the decision already made. That is the sweet spot, and it is a bigger spot than it sounds.

When a job had clear edges and nothing left to decide, the results were consistently good, and they stayed good run after run.

The model is also genuinely strong on domain best practice. When we audited our own architecture decisions across the whole system, the finding that surprised me was that the designs lined up with Drupal 10 and 11 community consensus. The model knew the right shape. What it had not done was prove that the right shape had actually shipped, which is a different skill and the subject of most of the rules below.

Why does the same input produce different results?

Because somewhere downstream, something is being re-decided.

Our worst drift period came from the model exercising judgement on the same question again and again. Fixes appeared to land and then behaved differently on the next site. It looked like model inconsistency. It was a design fault: we had given the same decision to several places and let each of them answer it.

The fix was making each decision once and letting everything after it follow that decision. The model still does plenty of work, and every bit of it now follows a decision that is already made.

One input, one truth. If you are seeing variance, look for the decision you accidentally asked for twice.

Related, and worth stating separately: whenever a design choice traded consistency for cleverness, we chose consistency. The same input must give the same result on every run, because variance is the one thing you cannot afford while you are trying to establish whether a change helped.

Why do AI fixes stop working?

Because the model fixes the thing in front of it, and the thing in front of it is usually the output.

This cost us more time than any other single mistake. A defect gets reported, the model investigates, finds the wrong value, corrects it, and the page renders correctly. Everybody is satisfied. Then the next clean run regenerates everything and the defect is back, because the generator that produced the wrong value was never touched.

The rule that ended it: a fix belongs in the generator, never in the output. The acceptance test is a completely clean run reproducing the corrected result. Anything short of that is decoration.

This is not really an AI problem, it is an old engineering discipline. AI just makes it much easier to fall into, because fixing the symptom is fast, satisfying, and looks identical to fixing the cause right up until the next run.

Why is "it is in the config" not proof?

Because presence is not behaviour.

For a long stretch we verified fixes by checking that configuration existed, or that markup appeared in the DOM. Both are cheap to check and both produced false positives. Config can be correct while a naming mismatch stops the value ever reaching the template. Markup can exist in the DOM while rendering empty.

We now accept one kind of evidence: the rendered page, compared against the source. It is slower to check and it is the only check that has never lied to us.

The general form of this rule is worth keeping: when you ask a model to prove something worked, be specific about what counts as proof, because it will optimise for whatever you named. Name a weak proof and you get a weak proof, delivered confidently.

Should you trust an AI diagnosis?

Treat it as a hypothesis with a good hit rate, and re-verify anything you are about to build on.

The sharpest example: four rendering defects were reported and diagnosed across earlier sessions. When we independently re-verified each one before acting, one had been diagnosed accurately and three had been mis-located. Not imagined, the defects were real. The proposed fix sites were wrong.

There is a specific trap here. A diagnosis can be correct on Monday and stale by Wednesday, because the code moved underneath it. AI-assisted investigation produces confident, well-written, internally consistent explanations, and confidence is not correlated with freshness. We now re-confirm load-bearing claims at the moment we act on them, and the practice has saved far more time than it costs.

Where should the checks go?

Early in the run, and at a fixed bar.

For several generations our system converged incident by incident. Each run surfaced something new, we patched it, and the next run surfaced something else. That is a treadmill, and no amount of model quality gets you off it.

What got us off it was checking earlier in the run, and holding the bar steady whenever a check found something. Under pressure, people widen a pass condition until it stops complaining. A bar keeps its meaning only while it stays where you set it.

The second half of this rule: make the important checks impossible to skip. One of our important checks could be skipped, so eventually it was, and the gaps it would have caught reached review before we closed them. A check with an off switch is a suggestion.

What does a human tweak during a run tell you?

That you have a missing feature, and you just found it for free.

Early on I kept a list of every moment a person had to step in mid-run. Somebody had to mention that a site was animation heavy. Somebody had to notice a mega menu had flattened. Somebody had to give feedback on whitespace. Each of those felt like helpful collaboration at the time.

Every one was a defect report. Reading the list that way turned a vague sense that the system needed babysitting into a concrete backlog, and closing that backlog is most of what made later runs quiet.

If you are building anything AI-assisted and you find yourself nudging it in the same place twice, stop and write the nudge down. That is your next feature.

What about the process that grows around the build?

It needs pruning on the same schedule as the code.

By generation 13 we had accumulated roughly 110,000 words of decision history across 73 documents, plus a whole governance ritual that had made sense when we introduced it and had quietly stopped earning its place. Nobody could read it, which means in practice nobody did, which means the decisions in it were not being honoured.

We deleted the ritual and distilled the history into a single short reference that people actually read. That was as valuable as any code change in the same generation. AI-assisted work generates artefacts fast, and documents are artefacts. Left alone they become a layer that looks like rigour and functions as noise.

What this means for how you scope AI work

Three things I now believe more strongly than I did in May.

The model is rarely the bottleneck. Clarity is. Almost every failure in my log traces back to a job with unclear edges, a decision asked twice, or a proof standard that was too easy to satisfy.

Verification is the product. The generation is the easy half. What separates a demo from something you would put in front of a client is that somebody defined an uncomfortable bar and then made it impossible to route around.

Expect architecture problems to arrive disguised as accuracy problems. Every serious turn in our build started as a visual complaint and ended as a modelling decision. If you only measure what is easy to see, you will optimise the wrong half for a surprisingly long time.

This post is part of the X to Drupal series, alongside the pillar, why replatform into Drupal and what X to Drupal provides, and the companions on fourteen generations of building it, moving our own site off Webflow, what your team gets afterwards, and how we check the result really matches.


Frequently asked questions

Can AI handle a large build on its own?

It handles bounded work against a clear specification very well, and it needs a human to set the bar and design the verification. In our build the generation half went smoothly once each job had clear edges. The judgement about what "correct" means, and the checks that prove it, came from people.

What is the most common cause of inconsistent AI output?

A decision being made in more than one place. If two parts of a build both get to derive the same thing, they will eventually disagree. Making each decision once removed most of our variance.

How do you stop AI from fixing the symptom?

Define acceptance as a clean run reproducing the corrected result. If a fix cannot survive regenerating everything from scratch, it lives in the output and it will disappear. Making that the standing bar changed the behaviour immediately.

How much should you trust an AI diagnosis of a bug?

Treat it as a strong hypothesis and re-verify before you build on it. When we re-checked four diagnosed defects independently, one had been located correctly and three had not. The defects were real, the proposed fix sites were wrong.

What kind of quality checks work best on AI-generated work?

Checks that run early and hold a fixed bar. Make the checks that protect user-facing value impossible to skip, because a check with an off switch tends to be switched off.

Did AI get the architecture right?

Largely yes, and that surprised me. When we audited the design decisions against Drupal 10 and 11 community consensus they lined up well. The gaps were in implementation and verification, so the lesson was to spend the effort on proving the design shipped.

What would you do differently from day one?

Define what "done" means in terms a person outside the project would accept, before optimising anything. Ours started as "does it render" and became "does it render and can an editor work in it". That change was worth more than any accuracy work we did.

Written by Souvik Pal, QED42.

ShareXLinkedIn