Skip to content

Writing · Analysis · attribution, measurement, experimentation

The ad platform says it drove the sale. Did it?

Attributed conversions and incremental impact are different numbers. A meal kit launch used matched cities and a holdout to tell channel effects from local demand.

Every ad platform reports attributed conversions: sales it was willing to take credit for, under rules the platform wrote. What a business wants to know is incremental impact: sales that would not have happened without the ad. Those are different numbers, and the gap between them is where budgets go wrong.

Where attributed credit comes from

The rules are generous by design.

  • Attribution windows. A click from 30 days ago still counts toward today's sale.
  • View-through. They saw the ad, didn't click, bought anyway. Counted.
  • Brand search. Someone who already knew you searched your name, clicked the ad above the free result, and bought. Counted. Would they have bought anyway? Almost certainly.
  • Overlap across platforms. Google and Meta each claim the same sale. Add the dashboards together and you sold more than you sold.

None of this is fraud. It is what "attributed" means. The dashboard is answering "which of our ads touched this buyer," not "did our ads cause this purchase."

What a measurement looks like

A measurement compares what happened to what would have happened without the ad. That needs a control: a group of people, or places, that didn't get the ad, chosen so that they look like the ones who did.

The cleanest version for most businesses is geographic. Pick markets, hold some out, run the campaign in the rest, and compare.

A launch that tested before it spent

In 2021, a meal kit company, the kind where a box arrives and you cook the meal yourself, was going to market with a well-known name attached and a small test budget. The question wasn't "how much should we spend." It was "on what." Billboards, audio, paid social: each had an internal champion and none had evidence.

Instead of picking one and going nationwide, we built an experiment.

Matching the cities. We assembled clusters of similar cities using public data: household income and demographics from the U.S. Census, and existing interest in meal delivery from Google Trends. The goal was groups of cities that looked alike on the things that predict whether someone buys a meal kit, so that a difference between groups could be read as the campaign and not as the city.

Assigning the channels. One cluster was the holdout: no advertising at all. Each of the other clusters got one channel, and only one.

  • Billboards (out-of-home)
  • Audio (streaming, Spotify-style placements)
  • Paid social

Reading the result. With a holdout in place, the question for each channel is no longer "how many conversions did the platform attribute to us." It is "how much did this cluster outperform the cluster that got nothing." That is the number the dashboards can't give you, because the dashboards never see the holdout.

What it showed. Two things, and only one of them was about channels.

Audio was the strongest performer for the money. Streaming audio lifted the treated cities clearly over the holdout at a cost that made it the best value of the three. Paid social also performed well against the holdout. Neither result would have been visible in click attribution, because nobody clicks a podcast ad.

The second finding was about place. The Pacific Northwest cluster responded far more than the others. Those cities had an affinity for the personality behind the product, and a concentration of well-paid tech workers who fit the meal kit buyer. The test hadn't been designed to find a region. It found one anyway, because matched clusters let you see where the response is coming from, not just whether it exists.

So the launch plan that came out of the test had a channel answer and a geography answer, both from a small budget, and both from comparing against cities that saw nothing.

Why this matters more at launch than later

Once a company is spending nationally on three channels, every market has seen every channel, and there is nothing left to compare against. You can still run holdouts, but you have to turn something off to do it, and turning things off is a hard sell internally.

At launch, the control is free. Nobody has seen anything yet. A company that tests in matched clusters before going wide learns which tactics actually move people and builds its media mix from evidence. A company that launches nationally on day one has to reverse-engineer the same answer from attributed conversions, which is the question we started with.

What to do Monday

Keep the platform dashboards for relative decisions: this ad versus that ad, this audience versus that one. They are fine at that.

Stop using them for absolute decisions: whether a channel is worth its budget at all. For that, run a holdout on your biggest channel once a year. If you are launching something new, run the geo test first, while the control is still free.

Methods and assumptions

How the cities were matched. Cities were grouped using U.S. Census variables (household income, household size, age distribution) and Google Trends interest in meal delivery and meal kit terms over the prior year. Clusters were formed so that each group had a similar mix on those variables and a similar baseline of interest. No cluster was chosen because it "felt right."

What the design controls for. Differences between cities in income, size, and pre-existing demand. Seasonal effects, since all clusters ran over the same weeks.

What it does not control for. Local events during the test window, differences in creative quality across channels, and the possibility that a channel works better in some kinds of cities than others. With a small budget the clusters were small, so the result is directional, not precise.

What would change the conclusion. If the holdout cluster had moved as much as the treated clusters, the campaign, not the channel, would be the wrong variable. If one cluster had moved for reasons unrelated to the test, the matching would have been the weak point.