System1 Introduces AI Ad Testing Worth the Wait
When I stepped into the CGO role at System1, there was one genuinely daunting task on the list: apply the same rigor and behavioral science we’ve always brought to ad testing, but in a world where AI now exists. Today that becomes real, as we launch our first AI predictive tool. Given we’ve been measuring ad effectiveness for over two decades, why did it take so long?
It’s a huge step for the industry, learning how to use AI to understand how ads are actually working. And, truthfully, there’s a lot of trash out there. We’ve been through the wringer working out how to do this responsibly. It’s been knackering, so I’m sharing what we’ve learned along the way.
Two Approaches to AI Creative Testing
There are, broadly, two kinds of AI ad-testing tools on the market. Before you hand over cash, you should know which one you’re buying. I’ll call the first one Improv AI.
You prompt a general-purpose language model, ask it to imagine what a “35-year-old mother in Ohio” might think about your ad, and it does what improv performers do. Says yes, and, then makes something up on the spot that sounds plausible.
It’s fast, but underneath the confident delivery there’s no fixed methodology. Ask it the same question twice, on different days, and it can change its answer completely. You can’t know where the answer came from, or how it decided. You honestly have no idea if the model even looked at the ad, or whether it’s guessing from the title and the brand name alone.
Staying on the acting theme, the second type of AI creative testing is Method AI. The model has done the work, trained on a company’s own large historical database of human responses to real ads, learning the patterns well enough to predict how new creative is likely to score. It’s the far more defensible camp, and it’s the one System1’s model sits in. Because the last thing we wanted to launch was an expensive bingo machine.
It’s also not remotely our invention. Some of the biggest names in the industry have been doing versions of this for years. But the more you learn about how these models actually work, something doesn’t add up. Whatever you throw at them, they always have an answer.
What Differentiates System1’s AI Ad Testing?
Every prediction a model like ours makes is, underneath the score, a probability. It has looked at a new ad, compared it against everything structurally similar in the past, and produced its best estimate of where that ad would land with real people. Like a weather forecast. You don’t know it’s going to rain at 4:12 pm, but you know the odds well enough to grab an umbrella on your walk. Somewhere along the way, AI ad testing decided it wasn’t going to talk like a weather forecast. It decided to pretend it would rain at 4:12 pm. So you get a single, suspiciously precise number, whether or not the model has the faintest idea what it’s talking about.
Our model doesn’t do that. Ask Test Your Ad Screen how something will land and you get a band: low, medium or high, not a suspiciously exact 78 out of 100. It’s a less impressive slide. It’s also the honest answer.
Then there’s a separate issue: what it’s trying to predict in the first place. Open a typical AI ad-testing report and you’ll usually find a long list of metrics. It looks thorough. You might even get away with using it to convince your boss. But what do any of these numbers truly mean? Most of what’s on that list has never been shown to predict anything a business actually cares about. “4.4% feel inspiration”? Right.
So we’ve done our best to make something honest and useful, which took a while. About seven years ago, System1’s founder came up with the idea to test every US and UK TV ad. It now means we were able to train a model on a database of creativity that genuinely represents what’s out in the market.
I won’t pretend to have had anything to do with the training itself, but our engineers did good. They spent the early years working with Warwick University to properly understand how AI could predict human emotion. Three years later, the model matches real human responses nine times out of ten. What a world we live in.
But here’s the part I’m proudest of. It knows when an ad is different enough from anything made before that it can’t honestly predict the human response. When that happens, it doesn’t force a response. It lets you know it can’t produce a reasonable prediction. No fake certainty, in either direction.
So yes, it took years to build something smart enough to know when it’s wrong.
How to Use Test Your Ad Screen for AI Ad Testing
How should you actually use System1’s Test Your Ad Screen platform?
Screen the long tail of long-form ads you’d never have tested otherwise, and get a read back in a couple of minutes, not a couple of days.
Use the honest answers to work out which need a proper human look with System1’s Test Your Ad Pro, and which deserve more media spend than they’re currently getting.
This is the next stage of effectiveness. With more and more assets needed, marketers who know a little bit more at scale can responsibly grow brands through advertising.
Want to unlock greater efficiency and effectiveness with Test Your Ad Screen? Book a demo with our experts to see how the platform works and how you can use it to quickly and easily screen finished creative assets.