1. What is GeoX?
Meridian GeoX is Google's open-source library for geo experiments in Python. The idea behind a geo experiment is simple. You change your marketing spend in part of the country: you switch it off, hold part of it back, or actually spend more. You leave the rest of the country as it was. Then you compare revenue between the two groups.
The hard part is the question: what would have happened in the test regions if you hadn't changed anything? You never know that for sure. GeoX estimates it with statistical models (Time-Based Regression, Synthetic Control and Synthetic Difference-in-Differences). The difference between that estimate and what you actually see is your incremental effect. In other words: the revenue you wouldn't have had without that marketing.
The result is a lift number you can defend in a budget meeting. Or one you can use to improve your Media Mix Model (MMM). More on that later.
Marc (Turntwo): To be honest, we expected a bit more. GeoX reads like GeoLift in Python: same steps, faster engine. A lot of familiar challenges aren't tackled in GeoX.
2. How does a geo experiment work?
Say you're a Dutch webshop with €40 million in annual revenue. You spend €180,000 a month on Meta and YouTube. The marketing manager wants to know what that really delivers. The platforms themselves say: a lot. But they always say that.
A geo experiment goes like this:
- Split the country into regions. In the Netherlands, this is usually done at city level: for example 30 to 40 cities or urban areas, defined by postcode where possible. Part of them becomes the test group, the rest the control group.
- Look back. In the weeks before the test, both groups need to behave similarly. If revenue in Utrecht and Eindhoven hasn't historically moved up and down together, you can't use Eindhoven as a control for Utrecht. We always calculate the best test and control groups up front.
- Start the experiment. For example, switch off your Meta and YouTube spend in the test regions for four weeks. Or hold back 30%.
- Estimate what would have happened otherwise. Based on the control regions, you predict revenue in the test regions without the intervention.
- Take the difference. Expected: €1,200,000 in revenue in the test regions. Measured: €1,080,000. That means €120,000 in revenue came from the channels you switched off. You can calculate a return on that.
Sounds straightforward. In practice, things rarely go wrong at step 4 or 5, but they often do at steps 1 and 2. Fortunately, GeoX has thought about that too.

3. What GeoX does well: setting up the test
GeoX puts a lot of emphasis on test design, before you switch anything off. The library generates several possible splits of regions into test and control, and scores them on three points:
- How well does the control group predict the test group before the test? A poor prediction means an unreliable result.
- How small can the effect be that you're still able to detect? This is called the Minimum Detectable Effect (MDE). If it's larger than the effect you expect, you're better off not starting, or changing the setup.
- How much budget are you putting on the line? How much spend do you switch off or hold back, and does that match what the organisation finds acceptable?
This is exactly the part that often gets skipped in practice. Teams start a test because "Rotterdam and The Hague are pretty similar anyway", without checking that in the data. Or they run for four weeks and find out afterwards that the effect could never have risen above the noise.
Marc (Turntwo): You don't win a geo experiment in the analysis, but in the setup. If the setup isn't right, the model mostly shows you noise afterwards, or can't say anything concrete about the experiment.
GeoX can also run several tests at once with one shared control group. In our example: Meta off in one set of regions, YouTube off in another, and one control group for both. That saves time and budget, but it does ask more of your data.
This kind of design step was already in GeoLift, and has recently been added to Python packages like diff-diff as well. The latter also offers more statistical methods, but is quite a bit more complex. GeoX can be a good first step to get started with incrementality testing.
4. The link with Meridian MMM
Google presents GeoX as a complement to Meridian, its Media Mix Model. An MMM estimates the effect of each channel on revenue based on historical data. But that's an estimate, and it can be well off if spend and revenue happened to move together in the past.
A geo experiment gives you a measured benchmark. You can feed that back into the model, so the estimate for that channel is right. Meridian can also tell you which channels would benefit most from a test like this.
Useful, but not new. Other tools like Robyn and PyMC-Marketing do the same. And diff-diff has MMM optimisation features too. The advantage of GeoX plus Meridian is that it's one package and works together out of the box.
5. The real bottleneck: data per region
This is our actual point, and it isn't about GeoX or any particular method.
Every geo experiment needs two things. Per region, per day:
- Your revenue (or other KPI) from your own systems. Not from Google Analytics, because that misses part of it due to consent. Not from the ad platform, because that's exactly the party you want to check. But from your order system, CRM or data warehouse.
- Your marketing spend at that same regional level. This is often the hardest part. Platforms report spend per campaign. Campaigns usually run nationally. And when they are set up regionally, Meta, Google and your own backend rarely use the same regional split. Meta often doesn't go deeper than, for example, region, which in the Netherlands means provinces.
In our webshop example, that means: linking orders to a city via postcode, getting spend from three platforms onto those same cities, and putting both into one table every day. If your data warehouse can't do that already, it'll take you a long time before you even get to the test itself.
This is where most teams get stuck. Not on the choice between statistical methods, but on "we don't have reliable marketing spend per city at daily level". A new library (unfortunately) doesn't change that.
The good news: you only have to do this work once. You need the same table (revenue and spend, per region, per day) for a (regional) MMM, for alerts on regional anomalies and for dashboards that look beyond national totals. Once you've laid that foundation properly, you can use whichever geo library fits best at the time.
6. Where do you start?
If you want to get started with GeoX (or another package), this is the order we'd follow:
- Check your data first. Can you pull revenue per region per day from your backend today, going back 12 months? And spend per region? If not, that's step one. And probably the biggest.
- Choose a regional split that works across all sources. Better 30 cities that match in all your sources than 100 postcode areas that only exist in your backend.
- Run the design step before you switch anything off. Let GeoX, GeoLift or diff-diff generate possible splits and look at the MDE. Is the smallest measurable effect larger than what you expect? Then change the setup or postpone the test.
- Start with one channel and one question. "What does our YouTube spend add?" is a test. "What works in our media mix?" is a research programme.
- Do something with the result. Feed it into your MMM if you have one. And otherwise, just into your budget discussion.
We'll keep testing GeoX over the coming weeks, alongside GeoLift and diff-diff. We're not expecting miracles, but we are expecting a nicer design process. If it turns out differently, you'll read about it here.
Stuck on that first point, revenue and spend data per region that don't come together? That's where most of the work is for most of our clients, and probably also the quickest win. Feel free to drop us a message.





