Ad testing: From preference scores to creative direction

A decision framework for choosing an ad testing method, plus a five-stage diagnostic workflow that turns a creative effectiveness score into a repair brief.

Articles

Five connected white pills on a beige background reading Establish, Probe, Assess, Diagnose and Test, the five stages of ad testing

In this article

Qualitative insights at the speed of your business

Conveo automates video interviews to speed up decision-making.

TL;DR

  • Most ad testing programs produce a score that ranks ad concepts but arrives after the decision window has closed: by the time a preference score comes back, the shoot is done, and the ad campaign budget is committed.

  • This article provides a decision framework for choosing the right ad testing method, plus a five-stage diagnostic workflow that turns a creative effectiveness score into a repair brief.

  • Marketing teams that apply this framework arrive at a defensible creative brief while the direction is still open, avoiding costly mistakes that only surface once production has wrapped.

  • Best for: insights and CMI teams running pre-testing or pre-launch ad testing who need findings their creative and brand stakeholders can act on.

What is ad testing and why it matters

Ad testing is the practice of gathering feedback from a target audience on advertising creative before or during an ad campaign, to learn whether a message lands, confuses, or loses people before significant media spend is committed. It connects raw creative ideas to advertising effectiveness while there is still time to act on what you find, and it works at different stages of the process: concept, script, animatic, rough cut, finished film.

Definition card titled Ad testing: gathering feedback from a target audience on advertising creative before or during an ad campaign

Pre-testing versus in-market testing

The two modes serve different purposes.

  • Pre-testing runs while creative direction is still open. A claim can be rewritten, a tone can shift, a scene can be cut.

  • In-market testing runs after the shoot, the edit, and the budget commitment. At that point, the only decision left is whether to run the creative or pull it.

Pre-testing exists to catch costly mistakes while they are still cheap to fix.

Why the timing gap costs you the insight

A preference score can indicate that concept B outperformed concept A. It cannot show the exact moment a viewer's attention slipped, or name the line of dialogue that triggered skepticism.

When a participant says "I almost chose the competitor," a static questionnaire has no way to follow that thread. The switching driver, the objection, the specific frame that nearly cost the sale, goes uncaptured.

The root cause sits in timing and method rather than in the research process itself. Testing early enough to change the work is what separates a program that improves campaign performance from one that merely documents it, and separates effective ads from creative that tested well but landed flat.

Ad testing methods: A decision framework

Choosing a research methodology for creative testing comes down to three variables: what question you are trying to answer (diagnose, predict, or optimize), how much stakeholder risk sits behind the decision, and how much time you have before the creative locks.

Method

Best question type

Stakeholder risk

Timeline

Survey ratings

Optimize / rank

Low to medium

Days

AI prediction models

Predict

Medium

Hours to days

Neuromarketing

Diagnose attention/emotion

Medium to high

Weeks

Qualitative interviews

Diagnose comprehension/credibility

High

Days (AI-moderated) to weeks (traditional)

In-platform A/B testing

Optimize live creative

Low

Days to weeks (post-launch)

Survey ratings

Best for ranking ads when a fast, comparable score is needed. An ad testing survey delivers preference rankings, Likert scores, and word clouds within days, and such quantitative data is straightforward to defend in a planning meeting.

The tradeoff is well understood: surveys identify the winning ad, without explaining why the losing concept failed. Survey design also caps what you can learn, because a fixed questionnaire can only measure the reactions someone anticipated while writing it. If the brief is "pick one from three," surveys do the job. If the brief is "fix this before production," they fall short.

AI prediction models

Best for forecasting how ads perform in market against normative databases, particularly when a team needs a projected lift score before committing media spend. These models return outputs in hours to days.

The limitation teams report is credibility at the stakeholder level: an ad effectiveness score without source evidence is difficult to defend in a creative review, and brand teams in particular are wary of outputs they can't trace back to a real person's reaction.

Neuromarketing (eye-tracking, EEG, facial coding)

Best for measuring attention and emotional engagement moment by moment, producing heatmaps and arousal curves that show where in a spot attention collapsed or emotion spiked. It can provide insights a self-report survey never will, because participants are poor narrators of their own attention.

The tradeoff is structural: specialized hardware, higher cost, and timelines that typically run to weeks. Neuromarketing also leaves comprehension failures and credibility issues in claims unexplained. It shows that something went wrong at second 14, but it doesn't identify what the participant misunderstood.

Qualitative interviews (AI-moderated, focus groups, or one-on-one)

Best for diagnosing the underlying cause of a creative problem: why a claim felt inauthentic, what drove switching hesitation, or where comprehension broke down. Open-ended questions let participants provide feedback in their own words, and probe-based findings with video evidence give stakeholders qualitative insights they can act on and trace back to a real participant. This is also the only setting where copy testing moves beyond whether a line is liked to whether it is understood and believed.

Traditional agency-run qualitative ad tests, whether focus groups or individual depth interviews, typically take six to twelve weeks from briefing to debrief. AI-moderated interviews run the same depth of inquiry in days, without sacrificing the video evidence or the verbatim record.

Two approaches that sit outside the framework

In-platform A/B testing on Google Ads, LinkedIn Ads, and similar channels measures how ads perform against live campaign performance data across multiple ads once spend is already running, which makes it an optimization tool rather than a pre-launch one.

Social media monitoring tools track unprompted reactions to your own and to competitors' ads. Useful for context and for spotting a backlash early, though they can't give you the controlled comparison a creative decision needs.

Combining methods

Most teams need a combination rather than a single method: surveys to produce quantitative data and quickly rank ad concepts, qualitative interviews to diagnose what isn't working and why, and predictive models to forecast before final media commitment. Treating these as substitutes rather than complements is where ad testing budgets get wasted.

Not sure which method fits your question, timeline, and stakeholder risk? Talk to us about the study you have coming up.

Not sure which method fits your question, timeline, and stakeholder risk? Talk to us about the study you have coming up.

How to get from "which ad won" to "what to change next"

A preference score answers a selection question: go with ad B. It leaves the repair question open: ad B's main claim read as implausible to participants who already use a competing product, and the brand linkage collapsed in the final five seconds.

One output closes a decision. The other informs the next ten, which is why the teams that get the most from ad testing treat it as an input to future ad campaigns rather than a gate on the current one. A single ad rarely fails for a single reason.

The five-stage diagnostic workflow

The five-stage diagnostic workflow on an orange gradient: establish baseline, probe comprehension, assess credibility, diagnose brand linkage, test differentiation

1. Establish baseline behavior before showing any creative.

Ask what participants currently use in the category, what they trust, which competitors' ads they can recall unprompted, and what last triggered a purchase. Open-ended questions at this stage surface the real competitive set, which is often different from the one the brand assumes it has, and map the switching friction creative will need to overcome.

2. Probe for moments that broke comprehension.

After viewing, locate the exact point at which a participant became confused or disengaged. Asking "what exactly made you pause?" or "what did you think that claim meant?" pinpoints the frame or line where comprehension failed and where viewer engagement dropped. These moments tend to cluster around jargon, implied price anchors, and benefit claims that outpace category familiarity.

3. Assess claim credibility.

Tone shifts, facial reactions, and pauses during viewing explain why a claim reads as inflated. Asking "what would make that offer worth it?" converts a vague "seems too good to be true" reaction into a specific credibility threshold: the brief for the next creative iteration. This is where a moderator has to dig deeper than the first answer, because an initial reaction is usually a summary and the reason sits one question further down.

4. Diagnose brand linkage.

Determine whether participants can recall which brand the ad was for without prompting, and whether the creative reinforces or contradicts what they already believe about the brand. Weak ad recall and weak brand recognition are among the most common failure modes in ad testing, and they are almost always correctable at the execution level, but only if the team knows the problem exists.

5. Test differentiation against current options.

Use displacement questions, such as "what would you stop using or doing if this existed?", to quantify switching friction rather than liking. An ad that feels appealing yet communicates no unique selling proposition tends to generate positive scores but fails to drive sales, because it prompts substitution within a repertoire rather than incremental demand. The same probes reveal whether your key messages land or get absorbed as generic category claims.

What makes the workflow diagnostic rather than descriptive

The difference is adaptive probing. A static script can't follow up when a participant pauses at a price point or reacts visibly to a brand claim. Conveo's AI research assistant probes based on what participants actually say, asking follow-ups a human moderator would in a live session, which also makes it practical to see how different groups respond to the same creative without commissioning a separate study per segment.

Teams running this workflow through Conveo see the output as annotated video: a handful of frames from the ad with verbatim quotes overlaid and arrows linking specific reactions to thematic findings. Each annotation maps to a decision: revise the claim, strengthen the brand cue, reframe the offer. The result is a repair brief and a set of creative ideas grounded in what participants actually said, rather than a ranking.

See it in action: how AI-moderated video interviews actually work.

Enterprise-grade rigor in ad testing

The speed-versus-rigor tension is real, and it surfaces every time an insights leader or procurement gatekeeper evaluates a new ad testing platform. Faster timelines should still keep the methodology intact and produce findings that withstand a senior stakeholder asking where they came from. The answer is to build rigor into the infrastructure itself rather than to slow the work down.

Every participant is grounded in a real person.

No synthetic participants or avatars stand in for consumer reactions. When a participant's expression shifts at a specific frame, or their tone changes when a brand claim lands flat, the recording captures that moment and traces it to the person who said it. Stakeholders asking for source documentation receive timestamped video clips and verbatim quotes, along with a clear chain of custody for the consumer insights that end up in a stakeholder deck.

Conveo's AI research assistant is built by researchers.

It adapts to what participants actually say rather than following a rigid script, probing hesitations a fixed question list would miss. Multimodal analysis reads speech, tone, and facial cues together, so the research team can see the emotional engagement a creative generates with more precision than a post-exposure survey alone would surface.

Human researchers retain control over study design and interpretation, where methodological judgment is irreplaceable. The platform handles the operational work, including the data analysis that would otherwise consume the bulk of a fieldwork cycle.

Full auditability.

Every study is designed and documented so that the research methodology can withstand a compliance review: who designed it, which participants took part, what was asked, and how findings were derived. This applies to every study in the program, including the ones that never go to review.

Compliance infrastructure built for procurement.

Conveo is SOC 2 Type II certified and GDPR compliant, with EU hosting (Belgium), SSO, and customer data deletion upon request. Having these in place means a study can move forward instead of sitting in a vendor security queue.

A searchable insight library.

Every clip, theme, and finding from an ad study connects to prior work, so nothing gets researched twice. Stakeholders who want to verify a conclusion can pull the source clip directly rather than debate an analyst's summary, which is part of why procurement accepts AI-generated findings that come from Conveo.

Who this approach isn't for

This diagnostic approach is built for marketing teams and insights functions that need to understand *why* a creative choice is working or failing, beyond which option scored higher. A simpler method is the better call when:

  • The only decision is picking a winning ad from a small set of finished concepts, with no time or budget to revise. A straightforward ad-testing survey answers that question faster and more cheaply.

  • The findings need to inform a single ad or a one-off decision, rather than carry over into previous or future campaigns.

  • The creative is locked, and the media plan is committed, so the only remaining choice is whether to run it.

Cross-market ad testing: Consistency and comparability

Running ad testing across multiple markets creates a comparability problem most programs underestimate until they are staring at conflicting results. The risk goes well beyond a tagline translating poorly.

The underlying construct being tested, whether a claim feels credible or whether a tone reads as premium or aggressive, carries different weight in different cultural contexts. A benefit framed around individual achievement may land strongly in one market and feel tone-deaf in another, and the only reliable way to find out is to observe how different groups respond to the same set of creative.

Challenge

How it's addressed

Outcome

Cultural interpretation risk

Adaptive probing plus human researcher review

Findings reflect market-specific meaning instead of assumed equivalence

Translation effects on claims

Transcription and translation in 50+ languages, with human review

Meaning shifts caught before analysis

Comparability across markets and waves

Standardized probe strategy, multimodal analysis, and a shared insight library

Global creative decisions grounded in consistent, connected evidence

A hesitation in Germany gets probed differently from a hesitation in Brazil, because the AI research assistant follows what each participant actually says rather than applying a fixed sequence uniformly. Human researchers then review findings for cultural nuance before synthesis, and the same review catches translation shifts across 50+ languages before they reach analysis.

For global research operations teams, this is also a governance question. When a CMO asks why brand perception differs between France and the UK, the answer needs to be grounded in comparable audience response data, gathered from a representative sample of the right audience in each market under a shared protocol.

A standardized probe strategy and a shared insight library mean a wave run in the UK in Q1 can be meaningfully compared to a wave run in the US in Q3 without rebuilding context from scratch.

"The pace, responsiveness, and research expertise of the Conveo team, on top of the top AI-moderated qual platform, have been invaluable to us in scaling brand advertising internationally."

— Matt Harris, Research & Insights Lead, Canva

Creative elements to test and core metrics

A well-scoped ad testing program starts with deciding exactly which ad elements to present to participants.

Five creative elements to test

Creative elements fall into five categories, each capable of moving campaign performance in a different direction.

Five creative elements to test, shown as numbered steps on an orange gradient: visuals, copy, call to action, format, and sound
  1. Visuals. Often the first thing participants notice and the first thing they misread. Color, imagery, and composition can signal the wrong brand values before a single word registers.

  2. Copy. Carries the argument. Copy testing involves checking how headline framing, tone, and the specific language used to describe a benefit align with participants' perceptions of whether the message is for them.

  3. Call to action. Frequently undertested. CTA wording influences perceived urgency more than the surrounding creative, and a weak CTA can undercut an otherwise strong concept.

  4. Format. Ad formats read differently by context. The same creative behaves differently depending on whether it's a 15-second pre-roll on Google Ads, a static social card, or an audio ad during a podcast break.

  5. Sound. Particularly in video and TV ads, audio shapes emotional response in ways participants often can't articulate without being probed.

Six key metrics to diagnose

Thorough ad testing tracks six metrics that together describe ad effectiveness.

Metric

What it tells you

Recall

Whether participants attribute the ad to the right brand. Consistent ad recall across waves is what helps build brand recognition over time.

Attention

Where interest spikes and where it drops.

Emotion

The feeling the creative actually creates, which often differs from the response the brief intended.

Purchase intent

The most direct commercial signal, though it reads most accurately when paired with its reasoning.

Relevance

Whether the message lands for the right audience rather than for audiences in general.

Differentiation

Whether the ad communicates a unique selling proposition or blends into category noise. The most commonly neglected metric, and the one that turns creative effectiveness into a lasting competitive edge.

Weighting metrics to the campaign objective

The key metrics that matter most depend on the campaign's purpose, and mapping them to business outcomes is what keeps the program credible with finance and brand leadership.

  • Brand-building work should weight recall and differentiation.

  • Performance campaigns built to drive sales should weight purchase intent and relevance.

In a representative scenario, a team testing two ad concepts might find that a survey shows one concept scoring higher on stated preference. The diagnostic question is why, and which ad elements to change as a result. That is where a qualitative pass earns its place, turning creative testing into a repeatable contributor to campaign success.

How Conveo runs ad testing at enterprise scale

For enterprise creative and marketing teams, the gap between a preference score and a repair instruction is where ad spend is wasted. Survey-only testing can tell a team which concept was preferred; it leaves unanswered why the other concept lost viewers at the six-second mark, or what made a brand claim feel implausible.

Conveo logo above four checked steps: study design, recruitment, AI-moderated interviews, and multimodal analysis

The workflow runs in five steps.

1. Study design. Teams start around the creative stage: early-direction conversations while ad concepts are still malleable, rather than post-production validation when changing course is expensive.

2. Recruitment. Participants come through Conveo's integrated panel network, or from a team's own list via CSV upload or QR code, so the sample matches the target audience the campaign was built for rather than whoever was easiest to reach.

3. AI-moderated interviews. The AI research assistant conducts video interviews with adaptive probing, following what participants actually say rather than cycling through a fixed script. When someone says, "I'm not sure I believe that," the next question digs deeper.

4. Multimodal analysis. As sessions close, multimodal data analysis integrates speech, tone, and facial cues. The research team can see the exact moment a viewer became skeptical or tuned out: a brow furrow at the price point, a vocal hesitation when the brand is introduced, a visible drop in viewer engagement seconds into the payoff, each timestamped and linked to the surrounding verbatim quote.

5. Reporting and reuse. Stakeholders don't have to take the researcher's word for it; they can watch the clip themselves. Every finding flows into a searchable insight library, connecting the current creative round to previous campaigns and prior brand work so the next brief starts with institutional memory rather than a blank page.

Teams report cutting this cycle from weeks to days. Over successive rounds, that library lets qualitative insights compound into a competitive edge rather than reset each quarter, and it makes the line between creative decisions and business outcomes visible.

See how this workflow runs on your next ad testing study.

See how this workflow runs on your next ad testing study.

Frequently Asked Questions

Ad testing is the process of evaluating creative assets, such as TV ads, digital video, social ads, or static display, with a target audience before or after a campaign launches. The goal is to measure whether the creative communicates the intended key messages, elicits the desired response, and is likely to drive the business outcomes it was designed to achieve.

System 1 ad testing is a methodology developed by System1 Group that predicts advertising effectiveness by measuring second-by-second emotional response from real consumers. The platform assigns ads a Star Rating intended to predict long-term brand growth potential, alongside metrics for short-term sales activation, and is used for pre-testing ad creative at different stages, from early scripts and storyboards through finished ads, across TV, digital, and social ad formats.

In a medical context, "AD test" most commonly refers to audiometry, a clinical hearing test that evaluates hearing sensitivity across frequencies and volumes. It's unrelated to advertising or marketing research; a primary care provider or audiologist is the right place to start for hearing evaluations.

Ad testing software gives marketing teams a structured way of gathering feedback on creative before it goes to market. Most platforms recruit a representative sample of target consumers, show them the same set of ads in a controlled environment, and capture responses through surveys, emotional-reaction tools, or moderated discussion. Where traditional testing captures only what participants say they think, platforms that include AI-moderated qualitative interviews also capture the reasoning behind a reaction, which is what separates a ranking of effective ads from an explanation of why they work.

Yes, online ad testing is now the standard approach for most enterprise teams. Participants access the study on their own devices and provide feedback in their own contexts, removing the logistical costs of in-person sessions and making multi-market recruiting practical. The tradeoff is context: a participant watching a video ad on a laptop doesn't replicate the experience of seeing it mid-scroll on a phone, in a LinkedIn Ads feed, or hearing audio ads between podcast segments. Teams that need channel-specific findings typically run separate studies or use platforms that can segment audience response by device and viewing context.

Qualitative insights at the speed of your business

Conveo automates video interviews to speed up decision-making.

Your next read.

Success stories

Canva brings the voice of the consumer into every decision with Conveo

A study launched at 6:15 p.m. Results before breakfast. See how Canva uses Conveo to run research at the speed decisions actually happen.

Rómulo Rejón

Head of Customer Marketing

Success stories

Trend or fad? NRG validates cultural shifts by running qual at scale with Conveo

Hollywood has spent decades telling dads how to be dads. NRG wanted to know which version they actually recognize. So they ran a qual study at quant scale that wasn't possible before.

Rómulo Rejón

Head of Customer Marketing

Success stories

Ninth Seat partners with Conveo to understand every consumer in the moment

Four conversations with the same consumer, moderated in the moment. How a 40-year insights agency uses AI smartly, keeps research human, and wins more work because of it.

Rómulo Rejón

Head of Customer Marketing