Diagnose-first geo holdout testing for analysts: pick SCM or panel DML, run power checks with GeoLift or Google Conversion Lift, and use an agency checklist.

Diagnose first: geo holdout testing for analysts, SCM vs panel DML

Treatment and holdout markets mapped for testing

A geo holdout test measures the causal incrementality of advertising by exposing some geographic markets to ads while keeping comparable markets ad-free. The outcome is a measured lift, or the absence of one, that you can use to reallocate spend with confidence. Geo holdouts are the right call when you need offline conversion data, privacy-safe measurement, or cross-device attribution that user-level tests cannot reliably deliver.


TL;DR:

  • Larger budgets are typically necessary for geo holdout tests because they need more conversions and longer durations to detect meaningful lift.
  • Selecting markets based on pre-test sales trends and grouping by natural boundaries improves the accuracy of representing true treatment effects.
  • Synthetic control methods are suitable when markets respond quickly and follow linear trends, while panel DML reduces bias in cases of nonlinear trends or shocks.
  • Proper diagnostics, such as high pre-period correlation and checks for contamination, are essential to validate that results reflect actual campaign impact.
  • Using open-source tools like GeoLift is recommended for feasibility testing, while platform services are better suited for scaled, production-grade measurement studies.

Moormarketing
Turn Measurement Into Growth
Moormarketing helps eCommerce brands develop data-driven digital strategies focused on overcoming obstacles and growing revenue.

Table of Contents

What geo holdout testing does and when it’s the right choice

A geo holdout test splits markets into two groups: one sees the campaign, the other does not. Analysts then compare the outcomes, usually sales, store visits or sign-ups, aggregated at the market level rather than tracked per user. This sidesteps the identity resolution problems that plague individual-level experiments.

That aggregation is also the method’s main privacy advantage. Because geo tests never need to match an ad exposure to a specific person, they keep working as cookies disappear and device graphs get patchier. This is a big part of why platforms have pushed geo-based measurement, including Conversion Lift based on geography, which compares exposed and control geographic groups to estimate causal impact and reports feasibility statuses such as “not enough data” or “no significant lift” when a result cannot be trusted.

Geo holdouts suit a specific set of problems well:

  • Brand campaigns where the outcome (awareness, store foot traffic) never touches a pixel.
  • Offline or omnichannel sales where online and in-store purchases both matter.
  • Marketplace sellers who cannot deploy user-level tracking on a platform they do not control.
  • Businesses that need a defensible, on-the-record incrementality number for budget conversations.

The trade-offs are real. Geo tests typically need larger budgets than user-level experiments to reach a detectable effect, they run longer because markets move slower than individual users, and the geography itself is a coarser lens: you learn what happened in a region, not to whom.

Design checklist: selecting markets, sizing holdouts and running power calculations

A geo holdout test lives or dies on its design. Follow this sequence before you launch anything.

  1. Select markets using pre-period similarity. Cluster candidate regions by historical sales or conversion trends so treatment and control markets track each other closely before the test starts.
  2. Think in Google Marketing Areas (GMAs) or commuting zones, not postcodes. Grouping by natural mobility boundaries keeps people who live in one market and shop in another from muddying the result.
  3. Set the holdout as a share of total conversions, not a share of markets. GeoLift’s own documentation frames holdout as a percentage of control-market conversions (for example, a range like holdout = c(0.5,1)), which directly ties the setting to statistical power, as shown in the GeoLift open-source toolkit.
  4. Run a power calculation before committing budget. Check the minimum detectable effect (MDE) at your planned spend and duration: if the MDE is larger than any lift you would realistically expect, the test is not worth running yet.
  5. Set a test duration that matches the buying cycle, and add a cooldown period if conversions lag the ad exposure by weeks rather than days.
  6. Build in contamination controls. Exclude border regions between treatment and control markets, use larger geographic aggregates where mobility is high, and where available, apply contamination modelling or mobility data to flag markets at risk of cross-exposure.

Pro Tip: Run the power calculation twice, once at your expected budget and once at 1.5 times that budget, so you know in advance whether a modest spend increase would meaningfully shorten the test.

Platforms that offer in-platform geo experiment setup build several of these steps in directly, including GMA-based unit selection and contamination modelling, which removes a layer of manual clustering work but still leaves budget, duration and holdout share decisions to the analyst.

Choosing an estimator: synthetic control versus panel DML

Two families of methods dominate geo lift analysis, and picking the wrong one for your data pattern will bias the result. Synthetic Control Methods (SCM), popularised through the open-source GeoLift toolkit, build a weighted combination of control markets that mimics the treatment market’s pre-period trend, then measure the gap that opens up once the campaign starts. It is well documented, comes with power calculators and plotting built in, and is the most common starting point for analysts prototyping a design.

Panel-style Double Machine Learning (DML) estimators take a different approach, modelling the outcome and treatment assignment separately across the full panel of markets and combining the two to isolate the causal effect. They demand more statistical setup than SCM but are less dependent on a single, cleanly matched synthetic market.

A 2025 simulation study found that Augmented Synthetic Control (ASC) methods can show severe bias under nonlinear trends, uneven lag structures, or shocks concentrated in treated markets, while panel DML variants reduced that bias and kept confidence interval coverage closer to nominal levels. The same study on estimator robustness recommends a diagnose-first framework rather than defaulting to one method.

That framework runs in three steps:

  • Check whether the pre-period trend is linear or nonlinear across candidate markets.
  • Check whether markets respond to the campaign at different lags.
  • Check whether any shock, a promotion, a stockout, a competitor launch, hit treated markets disproportionately.

Nonlinear trends, uneven lags or treatment-biased shocks all point towards a panel DML approach. Clean, parallel pre-trends with no major shocks make SCM/GeoLift a reasonable, faster option. Either way, validate the chosen estimator with placebo tests before trusting the headline number, a separate analysis of panel-aware estimator robustness reaches a similar conclusion, favouring DML specifically when simple augmentation of the synthetic control fails to match pre-period behaviour.

Reading results without fooling yourself

A geo holdout result is only as good as the diagnostics behind it. Before acting on any lift number, work through these checks:

  • Pre-period correlation between treatment and control markets should be high. Low correlation before the test started means the model had a weak baseline to compare against, which widens confidence intervals and makes a real lift harder to detect.
  • Check for contamination between adjacent markets. If people in a control market are regularly exposed to the campaign through travel, shared media markets or cross-border commuting, the true lift gets understated.
  • A “not enough data” status means conversion volume fell below the threshold needed for a reliable read, not that the campaign failed. Google’s guidance on geo lift diagnostics points to raising budget, broadening the conversion definition, or checking campaign targeting as the standard fixes.
  • “No significant lift” is a real finding, not a data problem, once feasibility checks pass: it means the campaign did not move the needle enough to distinguish from noise at the tested budget.

When a result comes back weak or ambiguous, the fix is rarely to rerun the same design. Increase budget to shift the minimum detectable effect, add conversion proxies (site visits, add-to-carts) alongside the primary metric to boost volume, or drop to segment-level analysis to find pockets where the campaign is working even if the aggregate is flat. Broadening the conversion window in particular can rescue an underpowered test without touching spend, a point echoed in general conversion rate optimisation guidance on widening what counts as a measurable outcome.

Practical toolset: what to prototype with and what to run at scale

Most teams end up choosing between an open-source prototype and a platform-managed study, and the two serve different stages of the same project.

  • GeoLift is free, script-based ®, and gives you market selection, power calculators and result plotting in one package. It is the fastest way to test whether a design is even feasible before spending on a live campaign, and its walkthrough vignette shows worked examples of holdout settings, power calculations and cooldown windows.
  • Conversion Lift based on geography runs inside Google Ads itself, handling GMA-based market selection and contamination modelling without custom code, at the cost of needing a live campaign and platform-level budget commitments.
  • Use GeoLift or a Python equivalent when you want to test feasibility cheaply and privately. Move to a platform-managed study once the design looks sound and you need production-grade reporting.
  • Either path needs the same inputs: clean historical conversion data by market, a defined test and cooldown window, and campaign-level settings that can actually be gated by region.

An agency checklist for a geo holdout pilot

Running a geo pilot end to end takes the same discipline as any paid media build: clean data, a realistic timeline, and clear deliverables. A typical pilot needs 12 to 18 months of market-level conversion history, a market selection pass using pre-period correlation, a holdout sized against a power calculation, and a test window long enough to cover the buying cycle plus cooldown.

Moor Marketing runs this kind of measurement work alongside campaign delivery on Google Ads rather than as a separate reporting exercise, so the media plan and the test design are built together from the start. Clients coming out of a pilot typically get a lift estimate, a shortlist of markets worth scaling first, and a handover pack their internal team can run against in future quarters.

An agency checklist for a geo holdout pilot — overview diagram

When to brief an agency versus run the pilot internally

Run it internally if you have clean market-level data, a statistician comfortable with SCM or DML, and time to wait out a full test window. Brief an agency when any of those is missing, or when the campaign is already live and you cannot afford a slow, self-taught first attempt.

A good brief specifies data access, exact KPI definitions, and what counts as success before the test starts. Judge agency performance on whether the design passed its own diagnostics, not just on whether it reported a lift.

— Liza

How Moor Marketing helps you run a geo holdout pilot

Moor Marketing builds the campaign and the measurement plan together, so your Google Advertising setup is structured for a clean holdout from the first briefing rather than retrofitted for one later.

Moormarketing

A typical engagement starts with a diagnostic call, moves through market selection and campaign build, and finishes with a result and a handover document your team can reuse. If you want a fast, hands-on start, the 12 Week DOUBLE Your Revenue Challenge is built for exactly that kind of rapid, structured engagement. You can also see the full service range on the Moor Marketing site and book a call to scope your own pilot.

Sources

FAQ

What is holdout testing and what is its purpose?

Holdout testing withholds a treatment, such as an ad campaign, from a defined group so their outcomes can serve as a baseline for comparison. Its purpose is to isolate the causal effect of the treatment by comparing exposed and unexposed groups under otherwise similar conditions.

What is geo testing?

Geo testing is a form of holdout testing that runs at the market level rather than the individual level, comparing conversion outcomes in exposed regions against comparable ad-free regions. It is commonly used to measure advertising lift when user-level tracking is unreliable or unavailable, as with Conversion Lift based on geography.

What is incremental testing?

Incremental testing measures the additional outcome caused specifically by a marketing action, beyond what would have happened anyway. A geo holdout test is one way to run incremental testing, using the ad-free control markets as the counterfactual for what sales would have looked like without the campaign.

What are the three main types of test marketing?

Test marketing approaches generally fall into standard test markets, controlled test markets run through a research partner, and simulated test markets that model outcomes before a live launch. Geo holdout testing sits within the standard test market category, since it uses real, live regions rather than a simulated or fully managed panel.

How long should a geo holdout test run?

A geo holdout test should run long enough to cover at least one full buying cycle for the product category, plus a cooldown period if conversions typically lag the ad exposure. Google’s geo lift guidance notes that geo studies are generally more accurate once the full study period completes rather than read partway through.

Share:

More Posts

Get strategies direct to your inbox every Tuesday

Contact us today
and let’s grow your
business together