SEO testing

·

How to pick a control group when pages aren't identical

Your test and control pages do not need to be identical. They need to be comparable before the test and likely to respond similarly if nothing changes.

4 min read

You will almost never find two sets of SEO pages that are identical.

One page gets more traffic. Another ranks for more queries. Some have been live longer, target slightly different topics, or sit in different positions.

That does not prevent you from building a useful control group.

The goal is to create two groups that behaved similarly before the test and are exposed to the same major factors during it. Then change one group and leave the other alone.

You are matching expected behavior, not identical pages.

What makes SEO test and control pages comparable?

Comparable pages should be similar on the variables most likely to affect the metric you are testing.

Depending on the experiment, those may include:

  • Page type or template: Product pages should generally be compared with product pages, articles with articles, and so on.

  • Baseline traffic or impressions: A group of very high-volume pages behaves differently from a group receiving almost no search demand.

  • Ranking range: Pages around position 3 have different upside and CTR behavior from pages around position 30.

  • Query or intent type: Informational, branded, commercial, and transactional searches can respond differently.

  • Historical trend: Pages growing rapidly before the test should not all end up in one group while declining pages end up in the other.

  • Seasonality: Pages affected by different demand cycles can make the groups diverge even without the treatment.

You do not need every page to match on every characteristic.

You need the groups as a whole to be similar enough that one provides a reasonable estimate of what might have happened to the other without the change.

How do you build comparable test and control groups?

Start with pages eligible for the same experiment.

Remove pages that clearly should not be compared because they use a different template, serve a different intent, have extremely different traffic levels, or are affected by a known event.

Then match or balance the remaining pages using the variables that matter most for the test.

For a title-tag test, for example, you might prioritize:

  1. page type

  2. baseline impressions

  3. average position

  4. historical CTR

  5. pre-test trend

You can pair similar pages and randomly assign one to treatment and one to control, or create balanced groups and randomly assign within those groups.

Randomization is useful after you have identified comparable pages. It does not make fundamentally different pages comparable.

What does a good SEO control group look like?

Say you want to test a new title-tag format across 40 evergreen guide pages.

Before assigning groups, remove pages with unusual conditions: a recently published guide, a highly seasonal page, and several pages with almost no impressions.

For the remaining pages, compare:

  • previous-period impressions

  • average position

  • CTR

  • historical trend

Pair pages with reasonably similar baseline behavior. Then randomly assign one page from each pair to treatment and the other to control.

You end up with 18 treatment pages and 18 control pages.

The individual pages are not identical. One treatment page may rank at 8.7 while its paired control ranks at 9.4.

What matters is that the groups have similar distributions and behaved similarly before the test.

How do you know whether your control group is good enough?

Compare the groups using data from before the test.

Check:

  • total impressions or traffic

  • average or median ranking position

  • CTR

  • the distribution of high- and low-volume pages

  • historical trend

Do not look only at averages.

Two groups can have the same average traffic while one contains a few enormous pages and the other contains many medium-sized ones.

Plot the groups over the pre-test period if you can. You want to see whether they generally move in the same direction before anything changes.

If the groups already behave very differently before treatment, they are unlikely to give you a clean comparison afterward.

Should test and control groups have parallel trends?

Ideally, they should show reasonably similar movement before the test.

Suppose treatment traffic increased 15% during the month before launch while control traffic declined 8%. Even if their average traffic is identical on launch day, they are not starting from the same trajectory.

That does not mean every daily movement has to match.

You are looking for evidence that the groups respond similarly to the forces affecting both of them: demand, seasonality, algorithm changes, and normal search volatility.

If the pre-test trends are materially different, rebuild the groups before launching.

Matching the starting number is not enough. Match the behavior leading into the test.

What if you cannot build a credible control group?

Some SEO changes genuinely do not have a useful holdout.

Examples include:

  • a required sitewide template change

  • a migration

  • a technical fix that must apply everywhere

  • a legal or brand requirement

  • a page type with too few comparable URLs

Do not manufacture a weak control simply so you can call the analysis a test.

Use the strongest comparison available and be explicit about the limitation. That might mean historical trends, unaffected page types, external benchmarks, or a before-and-after analysis.

Those approaches can still provide useful evidence. They just support weaker causal claims.

A weak control does not make an analysis stronger simply because it has been labeled “control.”

Can you change the control group after the test starts?

Not because you dislike the result.

Define treatment and control groups before the experiment begins and keep the assignment fixed.

After seeing the result, it may be tempting to remove pages that behaved unexpectedly or move them between groups.

Doing that based on their performance introduces bias directly into the comparison.

Pages should only be excluded according to rules defined before the test, such as:

  • page became unavailable

  • implementation failed

  • tracking broke

  • a predefined eligibility condition was violated

Record exclusions and the reason for each one.

Do not improve the control group after you know what happened.

How many pages do you need in an SEO control group?

There is no universal number.

What matters is how much usable data the groups generate and how variable that data is.

Ten high-volume pages may provide more information than 100 pages receiving almost no impressions.

The sample also needs to be large enough that one unusual page does not determine the result.

Before launch, check how sensitive the group is to individual pages. Remove the largest page from the calculation and see how much the baseline changes.

If one URL can completely change the result, the group is probably too concentrated.

The sample-size calculation from your test plan should ultimately determine how much data you need to detect the effect you care about.

Should you match pages one-to-one?

Not necessarily.

Pairing pages can be a useful way to build balanced groups, especially when you have clear variables such as page type, traffic, and ranking range.

But the final analysis is usually about how the treatment group performs relative to the control group, not whether page A perfectly matches page B.

One-to-one matching is a method for creating comparable groups. It is not the goal itself.

Optimize for balanced groups, not perfect twins.

What should you check before launching the SEO test?

Before changing anything, confirm:

  1. treatment and control contain comparable page types

  2. traffic and ranking distributions are reasonably balanced

  3. pre-test trends move similarly enough to support the comparison

  4. no single page dominates either group

  5. assignment was made before seeing test-period performance

  6. exclusion rules are defined in advance

  7. the groups generate enough data for the effect you want to detect

Then lock the groups and start the experiment.

The SEO Split Test Planner can help build and validate treatment and control groups from Search Console data. The guide on how long an SEO test should run covers sample size and measurement duration, and how to read a test result you don’t like covers what happens after the data comes in.

Your control pages do not need to look exactly like your treatment pages. They need to give you a credible estimate of what would have happened without the change.

Get new SERPish things, when they’re ready.

Subscribe

Get new SERPish things, when they’re ready.

Subscribe