Agentic marketing

Run an AI Marketing Pilot Without Betting the Whole Quarter

A marketer I know runs email for a skincare brand. She told me she nearly handed the whole programme to an AI tool in one weekend. Every flow, every audience, live by Friday. Her boss had seen the demo and loved it.

A lifecycle marketer working at a laptop in a warm cafe
The marketer in the opening runs email for a skincare brand, and her name would sit next to any pilot that went wide too fast.

Then she stopped. If it went wrong it would land in the quarter, in her numbers, with her name next to it.

She was right to stop.

Gartner reckons more than 40% of agentic AI projects will be cancelled by the end of 2027, worn down by cost, unclear business value and thin controls, drawn from a poll of over 3,400 organisations already spending on it (Gartner forecast, reported by MarTech). MIT looked at the wider enterprise picture and found 95% of generative AI pilots delivered no measurable return at all (MIT GenAI Divide, via Legal.io). That is a lot of tools quietly switched off, and a lot of quiet embarrassment.

40%+agentic AI projects set to be cancelled by 2027Gartner
95%generative AI pilots returned nothing measurableMIT
~130vendors making agentic claims are the real thingGartner

Here is what I take from it. The tools mostly work. What breaks is how people bet on them. Big, broad, all in one go, with no honest way to tell whether the lift was real or would have happened anyway.

There is a calmer way to do this, and it is not slower where it counts.

The category behind all this now has a name. Agentic marketing. Software that watches each customer, decides the next move and acts, for each person, at the right time. It is genuinely capable, which is the exact reason you should give it one job first and not your whole engine.

Why do so many AI marketing projects get cancelled?

Because most of them start too broad and can never prove they earned their keep. That is the core of Gartner's finding: over 40% of agentic AI projects are set to be scrapped by the end of 2027, and the reasons are rarely the technology. They are escalating cost, unclear value and weak governance (Gartner, via MarTech).

There is a second trap. Gartner calls it agent washing, older chatbots and automation dressed up as agents, and it estimates only around 130 of the thousands of vendors making agentic claims are the real thing. So a chunk of those cancelled projects were betting on something that was never much of an agent to begin with.

The lesson underneath both numbers is simple. Scope decides survival. MIT found that the small group who did get a return, roughly 5%, won by embedding AI into one real workflow rather than sprinkling it thinly across everything.

The all in bet

  • Every flow live at once
  • Thin across the whole account
  • No honest read on the lift

The narrow pilot

  • One leaking flow, pointed at
  • Deep on a single workflow
  • A control group proves it

What should your first AI marketing pilot do?

One job. Pick the single place in your lifecycle that leaks the most money and point the pilot only there. Not the whole account. One gap.

For most consumer brands that gap is a flow you built a while back and never went back to. A winback that stopped converting. A quiet stretch after someone's first order where nothing useful happens. If you run on Shopify and Klaviyo, that is usually where the money is sitting, and it is where a narrow pilot pays for itself fastest.

First orderOne purchase, delighted
The quiet stretchNothing useful happens
The gapPoint the pilot here
WinbackOtherwise fires too late
A skincare serum bottle
A serum from the kind of skincare brand in the opening; point the first pilot at one quiet flow, not the whole programme.

Narrow is not timid. It is how the 5% in MIT's study actually won, by going deep on one workflow instead of thin across many. Pick a leak you can measure and you give yourself a fair test. If you are not sure which gap costs you most, that is worth an hour with your own numbers before you switch anything on, which is exactly what a revenue leak audit is for.

How do you know an AI marketing pilot actually worked?

You hold a control group. A slice of your audience, say 5% to 10%, keeps getting your normal experience while the pilot runs on everyone else. Then you compare revenue per person between the two. The gap is your real lift.

This matters more than it sounds. As the team at Rejoiner put it, opens and clicks never tell you what customers would have done if you hadn't emailed them. Plenty of the revenue your reports credit to a campaign would have arrived anyway. A control group is the only honest way to separate the two.

A control group is the only honest way to tell real lift from the revenue that would have arrived anyway.

The upside is real once you measure it properly. Admetrics puts the incremental revenue lift for brands running these tests well at 18% to 26%, off a control group as small as 5% held back for a month. A small price for a clear answer.

One thing on language. Call it a control group, not a holdout. Same mechanic, but the words matter: a control group frames it as a fair test you are running, which is what it is.

When should you widen an AI marketing pilot?

When the lift holds against the control group for two or three cycles, not on the first good week. One strong send can be luck. A pattern that survives a control group is a result you can build on.

Then widen slowly, and think of it as loosening a leash rather than flipping a switch. Start by steering one channel yourself. Move to approving each move the agents suggest before it goes out. Only once you trust the pattern do you let the confident, low risk moves run on the rules you set. You add scope as you earn confidence, so a bad call costs you a corner and never the quarter.

1
SteerYou drive one channel yourself
2
ApproveAgents suggest, you sign off each move
3
AutopilotConfident, low risk moves run on your rules

This is roughly how we run it at PilotX, so I will show my hand. Each customer gets four agents. One to discover what is happening, one to decide the next move, one to deliver it and one to supervise the rest. You set the control group, and you set how much freedom they get, from steering a couple of channels yourself, to approving every move, to letting the safe moves run on their own.

Measured against the control group you set, the modelled return climbs with how ready your data and foundations are. A floor around 19% more revenue when you are steering, a middle around 30%, and a ceiling up to 50% when the confident moves are running. Roughly three times where legacy personalisation tends to top out. On the fee itself, the model runs up to 55 times what you pay.

I want to be straight about those figures. They are modelled against a control group, not a measured result I am dressing up as proof. It is early, and I would rather you see it on your own data than take my numbers on faith.

Which is the whole idea behind how we start. We find the biggest gap in your lifecycle and build the fix on your own products, free, before you connect anything. You get a ten minute replay on your real customers, so you can see the move we would make and why. If it is worth trying, a two week pilot runs it for a few cents per decision, measured against the control group you set.

You do not need any of that to begin, though. Take one flow that used to work and has gone quiet. Hold back a tenth of the audience. Run your best fix on the rest for two weeks and compare. Whatever you learn is yours, and it costs almost nothing to find out.

Start narrow. One flow, one control group, two weeks. If the lift holds you will know what to widen next, and you will be holding proof instead of hoping the demo was real.

All articlesBook a demo