What Is Visual Regression Testing? A Practical Guide

Visual regression testing is the practice of automatically comparing screenshots of a website or app before and after a change, so you catch unintended visual differences (broken layouts, overlapping elements, missing images, shifted text) before your users do.

A regression is something that used to work and no longer does. Functional tests catch regressions in behaviour: a form that stops submitting, a login that fails. Visual regression testing catches the other kind, the ones that only show up when someone looks at the page. Those are the bugs that get reported by clients and customers, because nothing in your logs or your test suite goes red.

Why visual bugs slip through

Most visual regressions are side effects. You update a plugin and it ships CSS that loads on every page. You change a shared component and a template you forgot about uses it differently. A new font loads slower and pushes the header onto two lines at one screen width.

Three things make these hard to catch by hand:

  • They happen far from the change. The edit was in the footer; the bug is on the checkout page.
  • They only appear in some combinations. The grid breaks at 768px but not at 1024px, or only on pages without a hero image.
  • The numbers multiply fast. 100 pages × 5 breakpoints × 2 browsers is 1,000 layouts. Nobody clicks through 1,000 layouts after every update.

A tool doesn’t get tired, doesn’t skip pages and notices a two-pixel shift. That is the whole case for automating it. Our list of 12 common visual bugs automated testing catches shows what this looks like in practice.

How visual regression testing works

Every visual regression testing tool, whatever it’s called, follows the same four steps.

1. Capture a baseline

You take screenshots of the pages you care about while they are known to be correct. This set is the baseline: the reference every later run is compared against.

2. Capture again after a change

After a deploy, an update or a pull request, you take the same screenshots again: same pages, same widths, same browser.

3. Compare

The tool compares each new screenshot with its baseline and highlights what changed. How it compares matters a lot (more on that below).

4. Review and approve

A person looks at the differences. Intended changes are approved; unintended ones are bugs to fix. Some tools let you promote approved screenshots to become the new baseline.

A visual regression comparison in Diffy: screenshot thumbnails with the changed areas highlighted

What you compare against

The “before” doesn’t have to be the same site at an earlier time. The common setups are:

  • Before and after on one environment. Screenshot, update, screenshot again. The usual workflow for CMS maintenance.
  • Staging vs production. Screenshot both environments and compare them, so you see what a release will change before it goes live.
  • Pull request preview vs a baseline. Run the comparison automatically for every pull request in CI. See how this works for Upsun preview environments or Pantheon deployments from GitHub.
  • Scheduled monitoring. Screenshot production on a schedule and compare it with a fixed baseline, so changes nobody deployed (auto-updates, third-party scripts, content edits) still get caught. We explain the trade-offs in monitoring vs pull request baselines.

What visual regression testing catches

  • Layout breaks: columns wrapping, grids collapsing, elements overlapping
  • Spacing and alignment changes after CSS or framework updates
  • Missing or wrongly sized images, icons and fonts
  • Responsive bugs that only appear at certain widths
  • Browser-specific rendering differences (for example Chrome vs Safari)
  • Components that changed everywhere they are used: headers, footers, forms, buttons
  • Content that disappeared or moved after a migration or template change

It does not tell you whether a form submits, whether a link goes to the right place or whether the page is fast. Pair it with functional and performance testing; it complements them rather than replacing them.

How screenshots get compared

The comparison method decides how useful the results are.

Pixel-by-pixel. Every pixel in the new screenshot is compared with the same pixel in the baseline. It’s precise and simple, but noisy: if a notice bar pushes the page down by 40px, everything below it is flagged as changed, even though nothing else moved relative to anything else.

Pixel comparison with a threshold. Same idea, but small differences below a set percentage are ignored. This cuts noise from anti-aliasing, and also risks hiding a small real bug.

Shift-aware comparison. The tool recognises that a block of content has only moved down the page and highlights just the element that caused the shift. Diffy uses its own algorithm that works this way, with pixel-perfect comparison available when you want every pixel checked.

Switching between Diffy's shift-aware comparison and pixel-perfect comparison

AI-assisted review. Newer tools add a layer that groups and summarises changes so a reviewer can start with the most important ones. It helps with review speed; it doesn’t replace the comparison itself. We wrote an honest piece on what AI actually does in visual testing.

False positives: the main reason teams give up

A visual test that flags changes on every run gets ignored within a week. Almost all false positives come from content that differs between page loads:

  • Carousels, sliders and animations caught mid-motion
  • Embedded videos, maps and ads
  • Cookie consent banners and chat widgets
  • Dates, “latest posts” lists and randomly ordered content
  • Fonts or images that haven’t finished loading

Our screenshot testing guide covers the capture side of this in detail. The fix is to stabilise the page before the screenshot: freeze animations with a small script, mask embeds, set the consent cookie, replace dynamic text with fixed placeholder text and wait for assets to load. Our guide to fixing false positives in visual regression testing covers each technique, and stabilising Elementor sites and screenshotting tabs, menus and sliders show worked examples.

Breakpoints and browsers

Most visual bugs are responsive bugs, so one desktop screenshot isn’t enough. Test at least mobile, tablet and desktop widths, and add the widths where your CSS changes layout. Our responsive design testing guide lists the widths worth testing and why.

Browsers matter too. Chrome and Safari use different rendering engines (Blink and WebKit), and the same CSS can render differently in each. If a lot of your visitors use iPhones, test WebKit as well as Chrome.

When to run visual regression tests

  • Before and after CMS updates. Plugin, theme, module and core updates are the most common source of visual breakage on WordPress and Drupal sites. See what breaks during WordPress and Drupal updates and visual regression testing for WordPress.
  • On every pull request. Catch the problem while the change is still fresh and cheap to fix.
  • Before a release. Compare staging with production and review exactly what will change.
  • During migrations. Compare the old site with the new one page by page. Our website migration testing checklist walks through it.
  • After CSS refactors and design system changes. One changed token can affect hundreds of pages. See CSS regression testing.
  • On a schedule. Catch changes that didn’t come from your team.

Types of visual regression testing tools

There are three broad approaches. Which one fits depends on who runs the tests and what you’re testing.

ApproachExamplesHow you use itBest for
Code-based, open sourcePlaywright toHaveScreenshot(), BackstopJS, Nightwatch VRTWrite tests or config in your repo; run locally or in CI; store baselines as image filesDevelopers who want full control and no screenshot limits
Component-basedChromatic with StorybookScreenshot components or stories rather than full pagesTeams with a mature component library in Storybook
Hosted, URL-basedDiffy, Percy, Applitools, Ghost InspectorGive the tool URLs; it captures, compares and hosts the results for reviewTesting whole websites, CMS maintenance, teams where non-developers review changes

Our round-up of visual regression testing tools goes through each of these with current pricing and free tiers. We’ve also compared Diffy with most of them in detail: Playwright, BackstopJS, Nightwatch, Chromatic, Percy, Applitools, Ghost Inspector, Sauce Labs, Visualping and Pagescreen. Each comparison includes a section on where the other tool is the better fit.

How to start in five steps

  1. Pick your pages. One or two examples of every template, plus your most important pages: homepage, pricing, checkout, contact.
  2. Pick your widths. Mobile, tablet and desktop to start.
  3. Take a baseline while the site is in a known good state.
  4. Stabilise before you trust it. Run the comparison twice with no changes in between. Anything flagged is noise; mask, freeze or mock it until two runs match.
  5. Make it routine. Run it before and after every update, on every pull request, or on a schedule, whatever matches how your site changes.

With Diffy you can do all of this in a browser tab: paste your URLs (or import them from your sitemap), choose breakpoints and run the first comparison. The free plan includes 500 screenshots a month with no credit card.

Try Diffy free

Frequently asked questions

What is visual regression testing in simple terms?

It is automated before-and-after screenshot comparison. You capture how your pages look when they are known to be correct, capture them again after a change, and a tool highlights every difference so you can decide whether each one is intended or a bug.

Is visual regression testing the same as screenshot testing?

Mostly, yes. Screenshot testing is the plain-language name for the same idea. Visual regression testing emphasises the goal: catching regressions, meaning things that used to look right and no longer do.

Does visual regression testing replace functional tests?

No. Functional tests check that things work, such as a form submitting or a login succeeding. Visual tests check that things look right. A button can pass every functional test while sitting behind an image on mobile, so most teams need both.

How many pages and breakpoints should I test?

Start with one or two examples of every template (homepage, landing page, article, listing, product, form) at three widths: mobile, tablet and desktop. A broken template usually breaks every page built on it, so this catches most regressions for a fraction of the screenshots.

Why do visual tests produce false positives?

Anything that changes between page loads shows up as a difference: carousels, animations, ads, embedded videos, cookie banners, dates and randomly ordered content. Freeze, mask or mock those elements before the screenshot is taken and most false positives disappear.

Can I do visual regression testing on a WordPress or Drupal site?

Yes. URL-based tools such as Diffy screenshot the rendered page, so the CMS, theme or page builder does not matter. The most common use is comparing a site before and after plugin, theme, module or core updates.