Agentic Marketing ROI You Can Model on Your Own Numbers
It's Monday. You open Klaviyo, click into flows, and there it is. A handful of automated messages you set up months ago are quietly doing most of the work. In Klaviyo's 2026 benchmarks, flows drive about 41% of email revenue off just 5.3% of sends (Eightx, on Klaviyo 2026 data). The other 95% of what you send, the campaigns, the newsletters, the weekend promo, splits the rest.

Now look at the customers behind those sends. Across 156,110 shoppers, the average repeat purchase rate lands at 18.8%. Which means 81% buy once and never come back (BS&Co, February 2026). And of the ones who do return, half order again inside 30 days.
So the picture is this. Most of your revenue hides in a sliver of messages that fire at the right moment. Most of your customers leave after one order. And the window to save them is short. Your calendar cannot work that. A person scheduling campaigns cannot watch every customer and pick the right next move for each one, at the moment it matters.
That gap, between what your data could do per customer and what a schedule actually does, is what agentic marketing sets out to close. Instead of you building campaigns and hoping, software agents watch each customer, decide the next move for that person, and act on it, on the leash you set. That is the category. The question that matters for a founder is not whether it sounds clever. It is what it is worth on your numbers, and whether you can prove it.
The calendar
- One promo to the whole list
- Fired on your schedule
- Most customers gone after one order
The agents
- Next best move per customer
- Fired at the moment it matters
- Reaches the returner inside 30 days
What counts as agentic marketing ROI?
It is the extra revenue the agents produce above what your current lifecycle would have made on its own, and nothing more. Not the revenue a message happened to touch. The lift over a baseline. McKinsey's benchmark for getting personalisation right is a revenue lift of 5% to 15%, with acquisition costs down by as much as 50% and marketing ROI up 10% to 30% (McKinsey, summarised by Shopify Enterprise).
Those are the ceilings good targeting has reached for years. Agentic ROI is the same idea, measured the same honest way, with the decisions made per customer instead of per segment. So when someone quotes you a multiple, the only version worth trusting is the one framed as lift over a baseline you can see. Everything else is a bigger number describing the same revenue you would have earned anyway.
How do you model the return on your own numbers?
Start with the money already sitting in your base, then ask what a few points of retained revenue is worth. On a brand doing £8M a year at an 18.8% repeat rate, three more points of repeat purchase is comfortably into six figures of revenue, revenue you already paid to acquire. You do not need a data team for the first pass. You can do the whole thing on the back of a napkin:
- Take your revenue over the last twelve months.
- Find your repeat purchase rate, orders from returning customers over total orders.
- Pick a lift you would actually believe. Start low, two to three points.
- Multiply. That is revenue on the table before you spend another penny on ads.
- Divide by what the tool would cost you in a year. That is your modelled return.
Do it low on purpose. If the maths only works at a heroic lift, it does not work. If two or three points already pays for the whole thing several times over, you have found something real. And weight the model toward the flows and the moment, not the size of the send. Half of second orders happen within 30 days, so the return lives in reaching the right person inside that window, not in blasting a bigger list. If you want the two inputs pulled straight off your own store and lifecycle, our revenue leak audit does it on your data in a few minutes.

How do you prove the lift instead of trusting the dashboard?
Set a control group. Hold back a slice of customers the agents never touch, let the rest get the agentic treatment, and count only the difference between them. That difference is the real number. Everything else is a story your attribution wants to tell you.
And attribution does want to tell you a story. Platform default attribution can overstate email's contribution by 15% to 40% against a proper incrementality test (Eightx). So the flow that looks like it earned £30 a recipient may have earned far less that the customer would not have spent anyway. A control group strips that out. You stop counting revenue you would have kept regardless, and you start counting only what the agents actually added.
Make the tool prove it against a control group before you believe a word of the return. Including ours.
This is why so much AI spend disappoints. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, on escalating cost and unclear business value (Gartner, via Prefactor). Unclear business value is a measurement problem as much as a technology one. If you cannot separate what the agents did from what would have happened without them, you cannot defend the spend, and eventually someone kills it. The control group is how you keep a working system alive long enough to compound.
Where does a real modelled return actually land?
Higher than legacy personalisation, if your foundations are ready, and only against a control group. This is the frame we build PilotX on, running on your Shopify and Klaviyo data. Four agents work each customer, Discovery, Decision, Delivery and a Supervisor, and you choose how much they run on their own, from steering a couple of channels yourself to letting the confident moves fire on rules you set.
The modelled return climbs with how ready your data and lifecycle are. A floor around +19% when you are steering, roughly +30% in the middle, up to +50% at the top, always measured against a control group you choose. That top end is close to three times what the best legacy personalisation reaches, which McKinsey puts at 5% to 15%.
I will be plain about the status of those figures. They are modelled, not measured. They come from the economics run against a control group, not from a shelf of case studies. Which is exactly why the way we start leads with proof and not a pitch. We find the biggest gap in your lifecycle, build the fix on your own products before you connect anything, and show you a replay on your real customers you can watch in ten minutes. If it holds up, a short pilot runs it for real at a few cents per decision, measured against your control group, so the first number you ever see is your own.
What is the honest first step?
Do not start with us. Start with your own two numbers. Open your flows and write down the share of revenue they drive. Pull your repeat purchase rate. Put them side by side and you will see where the return is already hiding, and how much of it your schedule is missing. When you want a second read on those numbers, run the audit on your own data, or we will model the gap with you.
Either way, make the tool prove it against a control group before you believe a word of the return. Including ours.
