Notes · The numbers · July 9, 2026

Incrementality, explained for people who've been nodding along (like me).

Somewhere around 2024, "incrementality" became the word in DTC. Every podcast. Every LinkedIn thread. Every agency deck. "Yeah, but is it incremental?" someone would say, and everyone would nod gravely, and I'd nod along with them.

Here's my confession: I nodded for a long time before I actually knew what it meant. I thought I knew. I assumed incremental sales were the extra ones, the halo, the sales your ads cause that the dashboard can't see. An increment is a little bit added on top, right? So incrementality must be the bonus bit on top of your ad results. That's what I had in my head, and nothing about the way people used the word ever forced me to check it.

Then one day I put my ego to one side and actually pinned the thing down. And it turned out incrementality wasn't what I thought at all. It's closer to the opposite. Most of the time it's not a bonus on top of your dashboard number. It's a smaller number hiding underneath it.

If you've been nodding along too, this note is for you. No shame in it. The word itself sets the trap.

The whole thing in three numbers.

Picture a shop with all its ads switched off. This week it sells 100 orders. Nobody advertised. Those sales came from regulars, word of mouth, people who already knew the brand. That's the baseline: what happens on its own.

Now switch the ads on. The shop sells 130.

The ads added 30. Those 30 are the incremental sales.

That's the entire concept. And notice "increment" means exactly what you already think it means: the added bit on top. The trap is on top of what. The increment is measured on top of the ads-off world (the 100), not on top of your ad account's results.

That's where I had it backwards. I made the ads the starting point and went looking for the extra bit above them. But the ads can't be the starting point, because the ads' own report is the thing we don't trust. The starting point has to be the world without them.

The same idea as a rooster.

A rooster crows every morning, right before the sun comes up. If he wanted to, he could take full credit: look what I did, I raised the sun. Every morning, same result. His track record is flawless.

Incrementality is the simple test that calls his bluff. You keep the rooster quiet one morning and check whether the sun comes up anyway.

Your ad platforms are the rooster. They crow right before the sale and claim they caused it. Testing incrementality means keeping them quiet for a bit, for some people or some cities, and watching whether the sales show up regardless. The sun coming up on its own? That's the 100. The rooster's actual contribution? That's the question.

Why your dashboard number is too big.

When Meta or Google reports your ROAS, they count every sale their ads "touched." Clicked an ad this week? Meta's sale. Just saw an ad scroll past yesterday? Depending on your settings, also Meta's sale.

But touched is not caused. Your most loyal customer was buying this month no matter what. She walked past a retargeting ad on her way to a checkout she was already heading for, and the platform booked the sale as its win. She's one of the 100, not one of the 30. Multiply that across thousands of orders and the dashboard looks brilliant while a chunk of the spend bought nothing you weren't already getting.

I've sat in a version of this meeting more than once. A founder pulls up the ad account, points at the retargeting campaign, ROAS of 6, best performer in the account, and asks why I'd even question it. So we run the boring test: pause it for two weeks. Revenue doesn't move. Not a dip, not a wobble. The campaign was a toll booth on a road people were already driving down. And the awkward part isn't the wasted spend. It's that the dashboard had been telling that story for a year, and everyone believed it because it was green.

The ads weren't driving the sales. They were standing next to them, taking a bow.

The three kinds of sales in your dashboard.

Every sale your platform claims falls into one of three buckets:

  1. Caused. The ad found someone who didn't know you, or wasn't going to buy, and made the sale happen. One of the 30. This is what you think you're paying for.
  2. Nudged. They were probably going to buy eventually; the ad moved it forward or tipped a wobbler. Worth something, but not the full sticker price.
  3. Claimed. They were buying anyway. One of the 100, with an ad standing nearby taking a bow. Worth nothing, billed at full price.
"Revenue from ads", as the platform reports it Caused Nudged Claimed the ad made it happen moved it forward was happening anyway All three billed to you as one flattering number.
The mix is the question. Testing is how you find out yours.

Your platform ROAS blends all three into one flattering number. Testing is how you find out your mix. And the mix is usually worse than you'd like: the campaigns that look best on the dashboard are often the ones fishing where the fish were already biting.

This isn't theory. The big tests are famous.

The companies with the best data teams in the world ran the rooster test, and the results made careers uncomfortable.

eBay ran the classic one. They turned off their paid search ads on their own brand name and measured what happened. Almost nothing. People who typed "eBay" into Google clicked the free listing sitting right under the paid one they'd removed. They had been paying for clicks from people already on the way in.

Airbnb got their version by accident. When 2020 gutted travel, they switched off essentially all performance marketing to save cash. Traffic held at around 95% of what it was. They never turned most of it back on, and said so publicly.

Uber found it the hard way. They paused a huge chunk of app-install spend expecting installs to crater. The line barely moved. Tens of millions of dollars had been buying installs that were coming anyway.

Three different businesses, same lesson: nobody knew their real 30 until they kept the rooster quiet. Your brand isn't eBay, and that's exactly the point. They could afford to burn money for years before finding out. You can't.

How you actually run the test.

The cleanest way to know if something works is the way medicine knows if a drug works: hold some people out. Show the ads to one group, hide them from another, compare. The gap between the groups is your 30.

One audience, split at random Group A sees the ads Group B no ads (the holdout) SALES the gap = the incremental effect the sales the spend actually caused Group B still buys. Just less.
The entire test in one picture: same audience, split at random, one difference between them.

In practice it runs three ways, from most rigorous to most blunt:

Holdout tests. Split the audience, keep the ads away from one slice, compare how each group buys. Meta and Google both offer versions of this (they call them conversion lift studies), though note who's running the lab. Best when your spend and volume are big enough for the math to mean something.

Geo tests. The workhorse for most brands. Pick a set of similar regions, cut (or raise) spend in half of them for three to four weeks, leave the others alone, and watch what actually changes. Turn Meta prospecting off in five states: if total revenue there drops 9% while the other states hold flat, that 9% is what Meta was actually driving. Not what it claimed. What it caused. Crude, honest, and no platform gets to grade itself.

The pause test. The bluntest tool: switch a channel off for two weeks and watch total revenue. If you pause a campaign with a claimed ROAS of 5 and total revenue doesn't move, you have your answer, and it cost you nothing but two weeks of spend you were wasting anyway. One caution: compare against your forecast or last year, not just last week, so a seasonal dip doesn't frame an innocent campaign, or a promo acquit a guilty one.

Even the blunt version tells you more than a quarter of dashboard screenshots.

iROAS, with the math done in front of you.

iROAS is just ROAS done honestly: not the revenue the platform claims, but the revenue the test proved the spend caused, divided by the spend.

Worked example. Your retargeting campaign spends $10,000 a month and Meta credits it with $50,000 in revenue. Platform ROAS: 5.0. Looks like your best campaign. Then you run a holdout: people who saw the retargeting ads bought at 5.5%, and the held-out group bought at 5.0% anyway. Almost everyone it "converted" was converting regardless. The genuinely added slice works out to about $9,000 of that $50,000.

Same campaign. Same $10,000 of spend. $50,000 What Meta claims "ROAS 5.0" $9,000 What the test proved iROAS 0.9
You're paying a dollar to cause ninety cents. Only one of these numbers shows up in your bank account.

iROAS: 0.9. You're paying a dollar to cause ninety cents.

Same campaign. The platform says 5.0, the test says 0.9, and only one of those numbers shows up in your bank account. A channel with a worse platform ROAS can easily be a better use of money once you know what each one actually causes. That reshuffles budgets, which is the whole point.

Where the over-claiming is worst.

You don't need to test everything at once. The over-claiming follows a pattern, so start where the lying is likeliest:

Notice the pattern: claimed performance and real performance run in roughly opposite order. The dashboard flatters exactly the spend that deserves it least. The lie even comes with published numbers now: 2026 test medians put Meta acquisition nearly honest at around 1.1x, retargeting at around 0.6x, and branded search at around 0.27x, so if you discount a platform number, discount it per channel, never with one blanket haircut.

And now, the thing I originally thought incrementality was: the halo.

Here's where my old wrong definition finally fits in, because the halo is real. It's just not what incrementality means. It's a piece of it.

Go back to the shop's 30 added sales. Say 22 of them bought on your website, where the platform might track them. The other 8 saw your ad, didn't click, and bought on Amazon three days later. Or searched your name and came in through an "organic" visit. Or picked you up in Target. Those 8 are still part of your 30, the ad caused them, but the platform can't follow people off its own property, so it never takes credit for them. That invisible slice is the halo.

So the platform's number is wrong in both directions at once. It counts sales it didn't cause (the rooster problem, pushing the number up), and it misses sales it did cause (the halo, pushing it down). Which error wins depends on the campaign: for retargeting and branded search, the over-claim dominates and the real number is far below the claimed one. For big cold prospecting on a brand that also sells on Amazon or in retail, the halo can be so large that the real number is above what the platform claims.

That's why the honest read happens at the level of the whole business. Blended MER, total revenue against total ad spend, catches both errors at once, and it's the whole reason operators steer on it. And it's why the tests above watch total revenue rather than the platform's scoreboard: a geo test doesn't care where the sale was tracked, only whether it happened.

When this matters, and when it doesn't.

Honesty runs both ways. If you're spending $20k a month, formal incrementality testing is overkill. Your blended numbers, watched properly, will tell you most of the truth, and the two-week pause test is free.

The discipline earns its keep as spend scales, because the gap between claimed and caused gets expensive fast. At $100k a month, a campaign that's 40% claimed-not-caused is a six-figure annual leak. The bigger the budget, the more the flattery costs.

The principle holds at any size, though: treat platform attribution as a witness with a stake in the verdict, not the judge. Cross-examine it against your total revenue, your contribution margin, and the occasional deliberate test.

How to run your first test without a data team.

You can do version one of this in a spreadsheet:

  1. Pick the suspect. Your highest-ROAS retargeting or branded-search campaign, the one you'd defend hardest. That's exactly the one to test.
  2. Pick the method. Under roughly $50k a month total spend: pause test. Above it, or if you can't stomach a full pause: geo test, cut the campaign in a few matched states only.
  3. Set the window. Two weeks minimum, three to four is better. Long enough to outlast attribution windows and normal noise.
  4. Decide the pass mark before you start. Write down what the platform claims the campaign drives per week. If total revenue drops by less than half of that when you pause it, the campaign was mostly claiming, not causing.
  5. Watch total revenue, not the platform. All sources, against your forecast. The platform's own reporting sits out this exam.
  6. Act on the answer. Mostly claimed? Move the money to prospecting or keep it as margin. Mostly caused? Turn it back on with actual confidence, which is worth something too.

The first test is the hardest, because a number you've been proud of might not survive it. Every test after that is just hygiene. And it's a lot cheaper than nodding along for another year. Ask me how I know.

What to ask whoever runs your growth.

Three questions separate the operators from the screenshot merchants:

Confident, specific answers are a good sign. If the answer is a dashboard, you have your answer too. And if whoever you ask just nods gravely, well. I know that nod.


If you've never tested whether your spend is causing sales or just standing next to them, that's usually where I'd start looking.

Book a 15-minute call

← Back to Notes