Measure marketing incrementality with a control group
Picture the Friday refill email. It goes to the customers whose thirty day supply is running low, and by Monday your dashboard credits it with, say, forty thousand pounds. Lovely number. Now the harder question. How much of that would those customers have bought anyway, email or no email?
That gap, between what a campaign gets credit for and what it actually caused, has a name. Incrementality. And more marketers are finally chasing it. 52% of US brand and agency marketers now run incrementality tests, according to a July 2025 EMARKETER and TransUnion survey. The rest are still reading a last click dashboard and calling it the truth.

Last click answers a different question than the one you actually care about. It asks which touch to hand the sale to. Incrementality asks whether the sale would have happened at all without you. One flatters your busiest campaigns. The other tells you which of them are carrying their weight.
Everything else on the dashboard is correlation in a nice suit.
What is marketing incrementality?
Incrementality is the extra revenue a campaign causes that would not have happened without it. EMARKETER puts it plainly: it measures whether a campaign "caused outcomes (sales, conversions, new customers) that would not have occurred without the ad exposure" (EMARKETER, 2026). Everything else on the dashboard is correlation in a nice suit.
You cannot see incrementality by staring at a single campaign, because you only ever observe the version of the world where you sent it. To measure the lift, you need the version where you did not. That is what a control group gives you. A slice of your audience, chosen at random, that gets nothing, so their behaviour becomes your baseline.
Why does last click attribution overstate your best campaigns?
Because it hands the whole sale to the final email, even when the customer had already decided to buy. Your abandoned cart flow and your branded automations look like heroes precisely because they fire at the people sitting closest to a purchase anyway.
Think about who is in an abandoned cart audience. They added the thing to the basket. Some meaningful share of them were coming back with or without your nudge. Last click gives your flow full credit for every one of those sales. The flow might be brilliant. It might also be taking a bow for revenue it never moved. From the dashboard alone, you genuinely cannot tell the difference.
Last click, one sale
- Hands the whole sale to the final email
- The cart flow takes a bow
- Credit for revenue it never moved
Control group, same sale
- Revenue per customer, held back versus sent
- Only the lift the send actually caused
- A 2X winner or a quiet leak, told apart
This is not an argument to switch everything off. It is an argument to stop trusting the credit and start measuring the cause. On the best campaigns the two are miles apart.
How do you run a control group test on your email campaigns?
You hold a random slice of the audience back, send to everyone else, then compare what each group is worth per person. Rejoiner builds its control group from "a randomly selected set of customers that represents 10% of the sample size" and reads the result after about 90 days on revenue per customer (Rejoiner). If the marketed group is worth more per head, that difference is your real lift. Not the credited number. The caused one.
The steps are not complicated, and you can run the first one this week:
- Pick one campaign or flow you want the truth about. Start with an offer or a discount, where the stakes are highest.
- Before it sends, randomly hold back around 10% of the eligible audience. Random is the whole game. If the two groups differ in any way other than "got the email", the test is worthless.
- Send to the other 90%. Change nothing else.
- Wait a full purchase cycle. Rejoiner reads at about 90 days so the repeat buyers have time to come back.
- Compare revenue per customer, held back versus marketed. The difference, positive or negative, is your incremental lift.
What you find will not always flatter you. Rejoiner shows a cart abandonment campaign generating "almost 2X the amount of revenue versus not sending it at all", the good case. It also shows the other kind, where a discounted send is quietly "throwing away" two pounds per purchasing customer, which at a hundred thousand customers "really starts to eat away at your revenue and profits" (Rejoiner). Same dashboard. Opposite reality. Only the control group tells them apart.
How big should the control group be, and how long should you wait?
Around 10% of the audience is the common starting point, and you wait a full purchase cycle before you read it, which for most repeat purchase brands means roughly 90 days. Big enough that the two groups behave alike, small enough that you are not starving revenue to run the experiment.
Two things quietly decide whether the answer is real. Size and time. The control group has to be large enough that the difference you see is signal, not the noise of a few big spenders landing on one side. And the window has to be long enough to catch the delayed purchase, because a refill or a replenishment often lands weeks after the send. Read it too early and you will underrate a flow that pays off slowly. Rejoiner recommends running these tests quarterly for anything with an offer attached, so the picture stays current as your list and your margins move.
One honest caution. A control group measures the campaign you tested, on the audience you tested it on. It does not tell you why, and it does not transfer. The Friday refill lift for a skincare brand says nothing about your welcome series. Test the things that cost you money or carry a discount first, and let the rest wait their turn.

Where this leaves your calendar
Most retention calendars are built on credited revenue, which means they are built on a number that overstates the winners and hides the leaks. A control group swaps opinion for evidence, one campaign at a time. It is slower than a dashboard. It is also the only version that survives contact with your P and L.
This measurement discipline is the mechanic underneath how we build PilotX. Four agents work each customer around their own timing, and everything those agents do is measured against a control group the brand sets, so the lift on the board is the brand's, not a story we tell you. The economics we quote, a modelled six to twenty times return, are modelled category maths, not a result we have measured on your account. The control group is how you would hold us to it, the same way you would hold any campaign to it.
You do not need us to start, though. Hold back 10% of one flow this week and read it in ninety days. If you want a faster sense of where the money is leaking first, our free Revenue Leak Audit models it against your own numbers, and if you would rather have the control group running on every Klaviyo send instead of one test at a time, that is what our Klaviyo control group setup and the founder pilot are for. Either way, stop trusting the credit. Measure the cause.
