11 min read

How we check a replatformed site really matches

Engineers, tech leads, and the people signing off a migration

Verification
ShareXLinkedIn

A website migration parity check is the evidence that the new site carries everything the old one did: the same pages at the same widths, every published address answering, and every heading, paragraph, link label and image accounted for by name.

The short version

  • We show the two sites side by side at 1440, 768 and 375 pixels wide, the widths we publish.
  • We publish no parity percentage. An average flatters the easy pages and hides the hard ones, so you get the pictures and your own judgement.
  • Every address the old site advertises is checked against everything the new site can serve, redirects included. On qed42.com that list ran to 670 addresses.
  • Content is traced unit by unit, so anything that did not arrive is named with a reason.
  • A second pass compares the new site with the old one as it stands today, because the source keeps publishing while you migrate.

The worst moment in a migration is three weeks after launch, when someone in the content team searches for a page they wrote in 2022 and it is gone. Nobody signed that off. Nobody noticed. It was in a category listing that quietly stopped existing, and the sign-off looked at nine pages out of six hundred.

So the question this post answers is the one I would ask any vendor, us included: what did you actually check, and how would I know if you had missed something?

Why do screenshots on their own fall short as proof?

They prove one claim very well and stay silent on everything else.

A pair of full-page captures at the same width, old on the left and new on the right, settles the visual question in about half a second. Your eye is better at this than any tool: you see the wrong heading weight, the card that lost its border, the section that came out in the wrong order. That is why the comparison is the centrepiece of how we present a migration, and why we show whole pages and let you scroll them.

What that pair says nothing about is whether the other 600 pages came across, whether the addresses people have bookmarked still answer, or whether the third paragraph of a 2019 article survived the move. Those need counting, and counting is where the value of automated verification actually sits.

So the parity story has two halves: the pictures, which you judge, and three checks behind them, which produce lists.

Why do we publish no parity percentage?

Because a single number is the most confident and least useful thing we could put on the page.

Fidelity varies by page, by viewport and by how much motion a section carries. A site's contact page will match almost exactly. Its homepage, with a background video and a scroll-driven reveal, will be the hardest thing in the project. Average those and you get a figure that is flattering and technically defensible, and that leaves the decision you are actually making no better informed.

There is a second reason, and it is the one I feel more strongly about. Publishing a score invites everyone to optimise the score. The moment a number goes on a slide, the incentive shifts from fixing the hard section to finding a measurement that treats it kindly. Showing the pictures keeps the incentive where it belongs.

Capturing those pictures fairly turns out to take its own discipline. Background video, motion and listings that keep loading as you scroll all need care, so that both sides of every pair show the same moment and the same length of page. Getting that right is what lets a pair prove something.

How do you know every page came across?

You start from the old site's own list of what it publishes, and you make the new site answer for all of it.

The check is boring and it is the one that catches real defects. Take the full set of addresses the old site advertises, from its sitemaps, its navigation, its listings and its internal links. Ask the new site for every one of them. Count a redirect as an answer, because a moved page that lands correctly is a moved page. Anything left over goes on a defect list with its address on it.

On qed42.com that list ran to 670 addresses, and the check found 14 category listings that had no route on the new site. We built them. Nobody had noticed they were missing, and no visual comparison would have surfaced them, because they were pages that nobody had thought to look at.

Two things make this check worth its cost. It runs against a machine-readable list, so it is complete by construction. And its output is a worklist with addresses on it, which is a thing a team can finish, unlike a score.

What does tracing content unit by unit mean?

Every heading, paragraph, link label and image reference gets followed from the reading of the old page through to the rendered new one.

Sampling is the usual approach and it has an obvious hole: you find the classes of problem your sample happened to contain. Tracing every unit finds the ones that only affect a handful of pages, which are exactly the ones that survive to launch. The two that show up most often are the old article with an embedded table that only exists on eleven pages, and the image whose caption lived in a text field that had no equivalent in the new model until someone looked.

When a unit does not arrive, it goes on a list with its name and a reason: the source had it inside a component with no counterpart, the source had it as a hard-coded block, the source has since dropped it. A named list with reasons is something a content lead can work through in an afternoon, which is why the check reports names and reasons and leaves the summing up to you.

What happens to content published after you started?

It comes back as a worklist, because the source site keeps working while the migration runs.

This is the part most migration QA skips, and it causes more post-launch surprises than anything else on this page. A migration of a content-heavy site takes weeks. During those weeks the marketing team publishes, retitles, unpublishes and reorganises, all on the old site, all after the snapshot everything else was verified against.

So there is a second pass, run close to cutover, that compares the new site with the old one as it stands that day, page by page. What comes out is three lists: published since the snapshot, changed since the snapshot, and gone since the snapshot. Your content team works those lists, which is far cheaper than re-reading the whole site and much safer than assuming nothing moved.

What should you ask any vendor about verification?

Five questions, and they apply to us. The answers tell you whether verification is a workstream or a slide.

  1. Where did your list of pages come from? If it came from the old site's own published addresses, the check can be complete. If it came from a spreadsheet somebody typed, it is a sample with ambitions.
  2. Show me a page that failed. A verification pass that found nothing found nothing because it was not looking. Ask to see the defect list and what happened to each item.
  3. What are your widths? Ask which viewports were verified at and see the captures at each. A desktop-only sign-off on a site where most traffic is mobile is a sign-off of the wrong thing.
  4. How are redirects counted? A moved page that lands correctly is a pass. A page that 200s with the wrong content is a defect wearing a pass, and only a content check will find it.
  5. When was the last comparison run? If the answer predates the last month of publishing on the source site, ask for one more pass before cutover.

The honest summary of all this is that verification is mostly bookkeeping, and bookkeeping is what makes the difference between a migration that lands and one that leaks. The pictures earn the meeting. The lists earn the launch.

This post is part of the X to Drupal series, alongside the pillar, why replatform into Drupal and what X to Drupal provides, and the companions on fourteen generations of building it, moving our own site off Webflow, and what makes the delivered site readable to LLMs.


Frequently asked questions

How do you verify a website migration?

With a visual comparison and three counting checks. The two sites go side by side as full pages at 1440, 768 and 375 pixels wide. Every address the old site advertises is checked against everything the new site can serve, redirects included. Content is traced unit by unit. A final pass compares the new site with the old one as it stands that day.

Why do you not publish a parity percentage?

Fidelity varies by page, by viewport and by how much motion a section carries, so an average flatters the easy pages and hides the hard ones. Publishing a score also shifts the incentive towards a kinder measurement. The side-by-side comparison is the claim, and you are free to judge it harshly.

What happens to pages that have no equivalent on the new site?

They go on a defect list with their addresses, and they get built or redirected before launch. On qed42.com the address check ran to 670 addresses and found 14 category listings with no route on the new site, which we then built.

Do redirects count as a page arriving?

Yes. A moved page that lands correctly on the right content is a pass, and the check counts it as one. A page that returns a 200 with the wrong content is a defect, which is why the content trace runs alongside the address check.

How do you handle content published while the migration is running?

A second pass close to cutover compares the new site with the source as it stands that day and produces three lists: published since the snapshot, changed since it, and gone since it. Your content team works those lists, which is much cheaper than re-reading the whole site.

Which parts of a site are hardest to reproduce?

Sections built on scroll-driven animation and background video are the hardest, because their appearance depends on motion and a still comparison can only show the resting state. We show those sections in the comparison and say so, and content behind authentication or inside a single-page application needs a conversation before anyone quotes.

Written by Souvik Pal, QED42.

ShareXLinkedIn