
TL;DR
Most ad testing programs produce a score that ranks ad concepts but arrives after the decision window has closed: by the time a preference score comes back, the shoot is done, and the ad campaign budget is committed.
This article provides a decision framework for choosing the right ad testing method, plus a five-stage diagnostic workflow that turns a creative effectiveness score into a repair brief.
Marketing teams that apply this framework arrive at a defensible creative brief while the direction is still open, avoiding costly mistakes that only surface once production has wrapped.
Best for: insights and CMI teams running pre-testing or pre-launch ad testing who need findings their creative and brand stakeholders can act on.
What is ad testing and why it matters
Ad testing is the practice of gathering feedback from a target audience on advertising creative before or during an ad campaign, to learn whether a message lands, confuses, or loses people before significant media spend is committed. It connects raw creative ideas to advertising effectiveness while there is still time to act on what you find, and it works at different stages of the process: concept, script, animatic, rough cut, finished film.

Pre-testing versus in-market testing
The two modes serve different purposes.
Pre-testing runs while creative direction is still open. A claim can be rewritten, a tone can shift, a scene can be cut.
In-market testing runs after the shoot, the edit, and the budget commitment. At that point, the only decision left is whether to run the creative or pull it.
Pre-testing exists to catch costly mistakes while they are still cheap to fix.
Why the timing gap costs you the insight
A preference score can indicate that concept B outperformed concept A. It cannot show the exact moment a viewer's attention slipped, or name the line of dialogue that triggered skepticism.
When a participant says "I almost chose the competitor," a static questionnaire has no way to follow that thread. The switching driver, the objection, the specific frame that nearly cost the sale, goes uncaptured.
The root cause sits in timing and method rather than in the research process itself. Testing early enough to change the work is what separates a program that improves campaign performance from one that merely documents it, and separates effective ads from creative that tested well but landed flat.
Ad testing methods: A decision framework
Choosing a research methodology for creative testing comes down to three variables: what question you are trying to answer (diagnose, predict, or optimize), how much stakeholder risk sits behind the decision, and how much time you have before the creative locks.
Method | Best question type | Stakeholder risk | Timeline |
|---|---|---|---|
Survey ratings | Optimize / rank | Low to medium | Days |
AI prediction models | Predict | Medium | Hours to days |
Neuromarketing | Diagnose attention/emotion | Medium to high | Weeks |
Qualitative interviews | Diagnose comprehension/credibility | High | Days (AI-moderated) to weeks (traditional) |
In-platform A/B testing | Optimize live creative | Low | Days to weeks (post-launch) |
Survey ratings
Best for ranking ads when a fast, comparable score is needed. An ad testing survey delivers preference rankings, Likert scores, and word clouds within days, and such quantitative data is straightforward to defend in a planning meeting.
The tradeoff is well understood: surveys identify the winning ad, without explaining why the losing concept failed. Survey design also caps what you can learn, because a fixed questionnaire can only measure the reactions someone anticipated while writing it. If the brief is "pick one from three," surveys do the job. If the brief is "fix this before production," they fall short.
AI prediction models
Best for forecasting how ads perform in market against normative databases, particularly when a team needs a projected lift score before committing media spend. These models return outputs in hours to days.
The limitation teams report is credibility at the stakeholder level: an ad effectiveness score without source evidence is difficult to defend in a creative review, and brand teams in particular are wary of outputs they can't trace back to a real person's reaction.
Neuromarketing (eye-tracking, EEG, facial coding)
Best for measuring attention and emotional engagement moment by moment, producing heatmaps and arousal curves that show where in a spot attention collapsed or emotion spiked. It can provide insights a self-report survey never will, because participants are poor narrators of their own attention.
The tradeoff is structural: specialized hardware, higher cost, and timelines that typically run to weeks. Neuromarketing also leaves comprehension failures and credibility issues in claims unexplained. It shows that something went wrong at second 14, but it doesn't identify what the participant misunderstood.
Qualitative interviews (AI-moderated, focus groups, or one-on-one)
Best for diagnosing the underlying cause of a creative problem: why a claim felt inauthentic, what drove switching hesitation, or where comprehension broke down. Open-ended questions let participants provide feedback in their own words, and probe-based findings with video evidence give stakeholders qualitative insights they can act on and trace back to a real participant. This is also the only setting where copy testing moves beyond whether a line is liked to whether it is understood and believed.
Traditional agency-run qualitative ad tests, whether focus groups or individual depth interviews, typically take six to twelve weeks from briefing to debrief. AI-moderated interviews run the same depth of inquiry in days, without sacrificing the video evidence or the verbatim record.
Two approaches that sit outside the framework
In-platform A/B testing on Google Ads, LinkedIn Ads, and similar channels measures how ads perform against live campaign performance data across multiple ads once spend is already running, which makes it an optimization tool rather than a pre-launch one.
Social media monitoring tools track unprompted reactions to your own and to competitors' ads. Useful for context and for spotting a backlash early, though they can't give you the controlled comparison a creative decision needs.
Combining methods
Most teams need a combination rather than a single method: surveys to produce quantitative data and quickly rank ad concepts, qualitative interviews to diagnose what isn't working and why, and predictive models to forecast before final media commitment. Treating these as substitutes rather than complements is where ad testing budgets get wasted.
How to get from "which ad won" to "what to change next"
A preference score answers a selection question: go with ad B. It leaves the repair question open: ad B's main claim read as implausible to participants who already use a competing product, and the brand linkage collapsed in the final five seconds.
One output closes a decision. The other informs the next ten, which is why the teams that get the most from ad testing treat it as an input to future ad campaigns rather than a gate on the current one. A single ad rarely fails for a single reason.
The five-stage diagnostic workflow

1. Establish baseline behavior before showing any creative.
Ask what participants currently use in the category, what they trust, which competitors' ads they can recall unprompted, and what last triggered a purchase. Open-ended questions at this stage surface the real competitive set, which is often different from the one the brand assumes it has, and map the switching friction creative will need to overcome.
2. Probe for moments that broke comprehension.
After viewing, locate the exact point at which a participant became confused or disengaged. Asking "what exactly made you pause?" or "what did you think that claim meant?" pinpoints the frame or line where comprehension failed and where viewer engagement dropped. These moments tend to cluster around jargon, implied price anchors, and benefit claims that outpace category familiarity.
3. Assess claim credibility.
Tone shifts, facial reactions, and pauses during viewing explain why a claim reads as inflated. Asking "what would make that offer worth it?" converts a vague "seems too good to be true" reaction into a specific credibility threshold: the brief for the next creative iteration. This is where a moderator has to dig deeper than the first answer, because an initial reaction is usually a summary and the reason sits one question further down.
4. Diagnose brand linkage.
Determine whether participants can recall which brand the ad was for without prompting, and whether the creative reinforces or contradicts what they already believe about the brand. Weak ad recall and weak brand recognition are among the most common failure modes in ad testing, and they are almost always correctable at the execution level, but only if the team knows the problem exists.
5. Test differentiation against current options.
Use displacement questions, such as "what would you stop using or doing if this existed?", to quantify switching friction rather than liking. An ad that feels appealing yet communicates no unique selling proposition tends to generate positive scores but fails to drive sales, because it prompts substitution within a repertoire rather than incremental demand. The same probes reveal whether your key messages land or get absorbed as generic category claims.
What makes the workflow diagnostic rather than descriptive
The difference is adaptive probing. A static script can't follow up when a participant pauses at a price point or reacts visibly to a brand claim. Conveo's AI research assistant probes based on what participants actually say, asking follow-ups a human moderator would in a live session, which also makes it practical to see how different groups respond to the same creative without commissioning a separate study per segment.
Teams running this workflow through Conveo see the output as annotated video: a handful of frames from the ad with verbatim quotes overlaid and arrows linking specific reactions to thematic findings. Each annotation maps to a decision: revise the claim, strengthen the brand cue, reframe the offer. The result is a repair brief and a set of creative ideas grounded in what participants actually said, rather than a ranking.
See it in action: how AI-moderated video interviews actually work.
Enterprise-grade rigor in ad testing
The speed-versus-rigor tension is real, and it surfaces every time an insights leader or procurement gatekeeper evaluates a new ad testing platform. Faster timelines should still keep the methodology intact and produce findings that withstand a senior stakeholder asking where they came from. The answer is to build rigor into the infrastructure itself rather than to slow the work down.
Every participant is grounded in a real person.
No synthetic participants or avatars stand in for consumer reactions. When a participant's expression shifts at a specific frame, or their tone changes when a brand claim lands flat, the recording captures that moment and traces it to the person who said it. Stakeholders asking for source documentation receive timestamped video clips and verbatim quotes, along with a clear chain of custody for the consumer insights that end up in a stakeholder deck.
Conveo's AI research assistant is built by researchers.
It adapts to what participants actually say rather than following a rigid script, probing hesitations a fixed question list would miss. Multimodal analysis reads speech, tone, and facial cues together, so the research team can see the emotional engagement a creative generates with more precision than a post-exposure survey alone would surface.
Human researchers retain control over study design and interpretation, where methodological judgment is irreplaceable. The platform handles the operational work, including the data analysis that would otherwise consume the bulk of a fieldwork cycle.
Full auditability.
Every study is designed and documented so that the research methodology can withstand a compliance review: who designed it, which participants took part, what was asked, and how findings were derived. This applies to every study in the program, including the ones that never go to review.
Compliance infrastructure built for procurement.
Conveo is SOC 2 Type II certified and GDPR compliant, with EU hosting (Belgium), SSO, and customer data deletion upon request. Having these in place means a study can move forward instead of sitting in a vendor security queue.
A searchable insight library.
Every clip, theme, and finding from an ad study connects to prior work, so nothing gets researched twice. Stakeholders who want to verify a conclusion can pull the source clip directly rather than debate an analyst's summary, which is part of why procurement accepts AI-generated findings that come from Conveo.
Who this approach isn't for
This diagnostic approach is built for marketing teams and insights functions that need to understand *why* a creative choice is working or failing, beyond which option scored higher. A simpler method is the better call when:
The only decision is picking a winning ad from a small set of finished concepts, with no time or budget to revise. A straightforward ad-testing survey answers that question faster and more cheaply.
The findings need to inform a single ad or a one-off decision, rather than carry over into previous or future campaigns.
The creative is locked, and the media plan is committed, so the only remaining choice is whether to run it.
Cross-market ad testing: Consistency and comparability
Running ad testing across multiple markets creates a comparability problem most programs underestimate until they are staring at conflicting results. The risk goes well beyond a tagline translating poorly.
The underlying construct being tested, whether a claim feels credible or whether a tone reads as premium or aggressive, carries different weight in different cultural contexts. A benefit framed around individual achievement may land strongly in one market and feel tone-deaf in another, and the only reliable way to find out is to observe how different groups respond to the same set of creative.
Challenge | How it's addressed | Outcome |
|---|---|---|
Cultural interpretation risk | Adaptive probing plus human researcher review | Findings reflect market-specific meaning instead of assumed equivalence |
Translation effects on claims | Transcription and translation in 50+ languages, with human review | Meaning shifts caught before analysis |
Comparability across markets and waves | Standardized probe strategy, multimodal analysis, and a shared insight library | Global creative decisions grounded in consistent, connected evidence |
A hesitation in Germany gets probed differently from a hesitation in Brazil, because the AI research assistant follows what each participant actually says rather than applying a fixed sequence uniformly. Human researchers then review findings for cultural nuance before synthesis, and the same review catches translation shifts across 50+ languages before they reach analysis.
For global research operations teams, this is also a governance question. When a CMO asks why brand perception differs between France and the UK, the answer needs to be grounded in comparable audience response data, gathered from a representative sample of the right audience in each market under a shared protocol.
A standardized probe strategy and a shared insight library mean a wave run in the UK in Q1 can be meaningfully compared to a wave run in the US in Q3 without rebuilding context from scratch.
"The pace, responsiveness, and research expertise of the Conveo team, on top of the top AI-moderated qual platform, have been invaluable to us in scaling brand advertising internationally."
— Matt Harris, Research & Insights Lead, Canva
Creative elements to test and core metrics
A well-scoped ad testing program starts with deciding exactly which ad elements to present to participants.
Five creative elements to test
Creative elements fall into five categories, each capable of moving campaign performance in a different direction.

Visuals. Often the first thing participants notice and the first thing they misread. Color, imagery, and composition can signal the wrong brand values before a single word registers.
Copy. Carries the argument. Copy testing involves checking how headline framing, tone, and the specific language used to describe a benefit align with participants' perceptions of whether the message is for them.
Call to action. Frequently undertested. CTA wording influences perceived urgency more than the surrounding creative, and a weak CTA can undercut an otherwise strong concept.
Format. Ad formats read differently by context. The same creative behaves differently depending on whether it's a 15-second pre-roll on Google Ads, a static social card, or an audio ad during a podcast break.
Sound. Particularly in video and TV ads, audio shapes emotional response in ways participants often can't articulate without being probed.
Six key metrics to diagnose
Thorough ad testing tracks six metrics that together describe ad effectiveness.
Metric | What it tells you |
|---|---|
Recall | Whether participants attribute the ad to the right brand. Consistent ad recall across waves is what helps build brand recognition over time. |
Attention | Where interest spikes and where it drops. |
Emotion | The feeling the creative actually creates, which often differs from the response the brief intended. |
Purchase intent | The most direct commercial signal, though it reads most accurately when paired with its reasoning. |
Relevance | Whether the message lands for the right audience rather than for audiences in general. |
Differentiation | Whether the ad communicates a unique selling proposition or blends into category noise. The most commonly neglected metric, and the one that turns creative effectiveness into a lasting competitive edge. |
Weighting metrics to the campaign objective
The key metrics that matter most depend on the campaign's purpose, and mapping them to business outcomes is what keeps the program credible with finance and brand leadership.
Brand-building work should weight recall and differentiation.
Performance campaigns built to drive sales should weight purchase intent and relevance.
In a representative scenario, a team testing two ad concepts might find that a survey shows one concept scoring higher on stated preference. The diagnostic question is why, and which ad elements to change as a result. That is where a qualitative pass earns its place, turning creative testing into a repeatable contributor to campaign success.
How Conveo runs ad testing at enterprise scale
For enterprise creative and marketing teams, the gap between a preference score and a repair instruction is where ad spend is wasted. Survey-only testing can tell a team which concept was preferred; it leaves unanswered why the other concept lost viewers at the six-second mark, or what made a brand claim feel implausible.

The workflow runs in five steps.
1. Study design. Teams start around the creative stage: early-direction conversations while ad concepts are still malleable, rather than post-production validation when changing course is expensive.
2. Recruitment. Participants come through Conveo's integrated panel network, or from a team's own list via CSV upload or QR code, so the sample matches the target audience the campaign was built for rather than whoever was easiest to reach.
3. AI-moderated interviews. The AI research assistant conducts video interviews with adaptive probing, following what participants actually say rather than cycling through a fixed script. When someone says, "I'm not sure I believe that," the next question digs deeper.
4. Multimodal analysis. As sessions close, multimodal data analysis integrates speech, tone, and facial cues. The research team can see the exact moment a viewer became skeptical or tuned out: a brow furrow at the price point, a vocal hesitation when the brand is introduced, a visible drop in viewer engagement seconds into the payoff, each timestamped and linked to the surrounding verbatim quote.
5. Reporting and reuse. Stakeholders don't have to take the researcher's word for it; they can watch the clip themselves. Every finding flows into a searchable insight library, connecting the current creative round to previous campaigns and prior brand work so the next brief starts with institutional memory rather than a blank page.
Teams report cutting this cycle from weeks to days. Over successive rounds, that library lets qualitative insights compound into a competitive edge rather than reset each quarter, and it makes the line between creative decisions and business outcomes visible.
Frequently Asked Questions
What is ad testing in marketing?
What is System 1 ad testing?
What does "AD test" mean in a medical context?
What is ad testing software and how does it work?
Can ad testing be done online, and what are the tradeoffs?









