What Is MaxDiff? A Practical Guide to Best-Worst Scaling

Rating scales let participants call everything important. MaxDiff forces a trade-off every round, so the ranking separates. Here is how a MaxDiff question works, how to design a valid study and how to read the scores.

Articles

White MaxDiff label with a cursor and three sparkles on an orange-to-pink gradient, the cover of this guide to best-worst scaling

In this article

Qualitative insights at the speed of your business

Conveo automates video interviews to speed up decision-making.

TL;DR

  • MaxDiff (best-worst scaling) is a survey method that forces participants to trade items off against each other instead of rating them, producing a clear preference ranking.

  • The problem: Traditional rating scales let participants rate every item highly, since nothing forces a choice.

  • The fix: MaxDiff makes participants trade items off against each other, exposing what actually matters.

  • Best for: Insights teams needing a defensible cut list built from importance scores that actually separate.

  • The outcome: Real priorities, carried forward wave to wave, as actionable data instead of guessing what to build or drop. Conveo runs MaxDiff analysis as a continuous, AI-moderated program that builds with every wave.

What Is MaxDiff?

MaxDiff, short for maximum difference scaling, is a best-worst scaling survey method used across market research that asks participants to pick the best and worst options from a small set, repeated across several item sets. Each MaxDiff survey question forces a genuine trade-off rather than a rating, which is what separates it from an ordinary numeric scale. That forced-choice mechanism removes rating-scale bias: no clustering everything at "important," no cultural variance in how people use a 7-point scale. The output is a statistical analysis of relative preference: a clean ranking and trade-off view showing how each item compares against the others.

White card on an orange gradient defining MaxDiff as a best-worst scaling method where participants pick the best and worst options from small sets

Why MaxDiff Beats Rating Scales

A features list where every item scores 8 or higher out of 10 on a numeric scale tells you nothing about which three to keep when the budget allows only three. This is the scale-use bias traditional rating scales are prone to: participants rate every product attribute, service attribute, or messaging line equally, so nothing stands out. MaxDiff solves this by forcing a choice: best versus worst, item against item, every round. That forced trade-off is what produces a real priority order, a measure of relative preference, instead of a wall of high scores. This is why Conveo builds MaxDiff as a forced-choice exercise from the start.

Dimension

Rating scale

MaxDiff

Task

Rate each item on its own

Pick the best and worst item in each small set

What it allows

Every item scored as important

One best and one worst per set, every round

Typical output

Clustered high scores

A preference ranking with scores that separate

Main bias

Scale-use bias, cultural differences in scale use

Depends on a balanced design and a fixed item list

How a MaxDiff Question Works

A MaxDiff question works in three steps:

Three numbered steps on an orange gradient, show a small set, force best and worst, repeat across sets, the flow of one MaxDiff question
  1. Show a small set. Four items appear at once, say four specific features of a product.

  2. Force best and worst. Respondents choose the single best and single worst instead of rating each one.

  3. Repeat across sets. This MaxDiff example repeats with different combinations of items across a typical 10 to 15 sets (also called rounds), meaning each respondent works through multiple MaxDiff questions before the session ends, so every feature appears a set number of times against different alternatives.

That balanced exposure is what lets researchers calculate a reliable preference score instead of a one-off guess, producing more reliable data than a single rating pass ever could.

When to Use MaxDiff

MaxDiff earns its place when stakeholders must weigh multiple attributes against each other and make a forced choice. That's true whether the decision is an internal roadmap call or a question about what consumers actually want most, beyond what scores highest on a customer satisfaction survey.

These are the classic forced-tradeoff decisions where rating scales flatten everything into ties:

  • Feature roadmaps needing to prioritize specific features under limited engineering capacity

  • Message testing where only three claims can headline packaging

  • Pricing strategies that need to know which price points actually move a decision

  • Benefit prioritization for a repositioning brief

Ranking questions built this way surface the important factors that budget, shelf space, or a launch deadline will demand anyway, instead of a wall of four-out-of-five scores.

When Not to Use MaxDiff

Where It Doesn't Fit

MaxDiff assumes you already know your item list, since ranking questions like these depend on having a fixed set of options to trade off. Exploratory questions, where the goal is to discover what to test rather than rank known options, need a different method first, whatever market research tools eventually run the trade-off exercise. Undefined categories, early-stage concepts, and open-ended "what matters to you" questions all fall outside its design.

The Gap Conveo Closes

The trade-off is simple: MaxDiff tells you what participants prioritize and leaves the reasoning out. Traditional MaxDiff stops at the ranking, leaving the reasoning to a separate follow-up question or a standalone study.

MaxDiff tells you what people prefer. Conveo's MaxDiff tells you why. After the choice tasks, the AI research assistant asks each participant why their best and worst items landed where they did, so the score and the story behind it come from the same participant.

Get the MaxDiff ranking and the reasoning from the same participants:

Get the MaxDiff ranking and the reasoning from the same participants:

Designing a Valid MaxDiff Study

Rankings are only as stable as the items and tasks behind them, and that stability is what separates reliable data from a directionless chart. Before trusting output, run this checklist:

Four checked boxes on cream for a valid MaxDiff study: item construction, cognitive load, design balance and quality flags

1. Item Construction

Cut ambiguous MaxDiff items and double-barreled statements that bundle multiple attributes into one score. Flag translation risks early: a phrase that reads neutral in English can carry unintended weight in German or Japanese, distorting the preference measurement itself.

2. Cognitive Load

More than four to five items per set, or too many tasks in one sitting, produces fatigue and the scale-use bias MaxDiff exists to avoid. Too much cognitive load can lead respondents to satisfice, and the survey data may look decisive when it is actually noise. Keeping sets tighter preserves the greater discrimination between items that MaxDiff is designed to produce.

3. Design Balance

Every item needs roughly equal exposure and pairing across the MaxDiff study, the experimental design decision that determines whether the output can be trusted at all. Unequal exposure skews estimates before hierarchical Bayes modeling ever runs.

4. Quality Flags

Watch for inconsistent survey responses across repeated tasks, speeding, and straight-lining. Any of these means the underlying MaxDiff data isn't yet actionable data, regardless of how clean the output chart looks.

Interpreting MaxDiff Scores

What the Scores Mean

A ranked list is not an interpretation. MaxDiff measures relative importance on a shared scale: utilities from a hierarchical Bayes model place an item at 0.35 above one at 0.20, but the gap between adjacent importance scores rarely justifies a confident stakeholder claim on its own. Converting those importance scores into preference shares (the modeled likelihood that one item wins head-to-head against another) often tells a cleaner story to stakeholders than the raw score alone. Segmentation often reveals that a flat aggregate score is hiding two groups of consumer preferences pulling in opposite directions.

What the Scores Don't Tell You

MaxDiff scores also can't explain themselves. A high utility might mask substitutability between similar items, and a low one might reflect confusion rather than rejection, both of which can distort MaxDiff results if taken at face value. This is where Conveo's video-based reasoning probe helps: because the follow-up conversation happens on camera, tone and expression sit right next to the ranking whenever a stakeholder asks why a particular item won or lost.

MaxDiff in Practice: What It Looks Like in Conveo

MaxDiff measures are only as valuable as the program that runs them repeatedly. Here's what that looks like inside Conveo:

  1. Conveo pairs MaxDiff with a follow-up conversation that probes the reasoning behind each choice, so the results explain which items win or lose with consumers and why.

  2. Conveo runs MaxDiff in a continuous program, so each wave builds on what the last one learned instead of re-establishing everything each time a decision comes up.

"Conveo has definitely changed what we can pitch to clients, because it allows us to bridge the gaps between quantitative and qualitative methodologies and tell a more complete end-to-end story. It lets us close the loop: from the archetypes we've identified in the qualitative, which of them actually resonate the most? It feels more authoritative, more backed up by the data."

— Fergus Navaratnam-Blair, VP Trends and Futures, National Research Group (NRG)

While the Decision Is Still Open

Conveo interviews are AI-moderated, so participants complete the choice tasks and the follow-up conversation on their own schedule, across markets and time zones. The ranking and reasoning arrive together while the roadmap or pack decision is still open, and each wave builds on what the last one learned instead of starting from scratch. In practice, that can look like 100 interviews in 3 days. See the step-by-step setup guide for exact configuration.

Run your next MaxDiff with the reasoning captured in the same session:

Run your next MaxDiff with the reasoning captured in the same session:

Frequently asked questions

In a pack-claims test for a snack brand, a respondent might see four items, each a distinct claim: "high protein," "no artificial flavors," "resealable pack," and "30% less sugar." Respondents choose "high protein" as most important and "resealable pack" as least important, effectively identifying the best and worst options in that set. This choice repeats across 10 to 15 rounds with different item combinations, building a statistically robust ranking of every claim relative to the others. Feature prioritization studies work the same way, with product features standing in for claims.

MaxDiff and conjoint analysis solve related but different problems, and TURF analysis sits alongside both. Both MaxDiff and conjoint force participants to make real trade-offs rather than rate everything highly, which is why researchers trust the output more than traditional rating scales alone. MaxDiff uses a lighter choice task and needs a smaller sample than conjoint, since conjoint estimates more parameters per participant; Conveo sizes each study's sample individually rather than applying a fixed number.

Yes, with limits. MaxDiff can rank concepts by customer preference, but it won't tell you why a concept won or whether it can displace current behavior. A concept can win a preference exercise, even score well on customer satisfaction measures, and still fail to change what someone actually buys or uses. Two practices help close that gap: use sequential monadic exposure when testing multiple concepts, a separate design choice that avoids the anchoring and order effects that distort comparative choices, and cap concept sessions at three to four concepts, since a tighter set produces greater discrimination between options. Then ask about displacement directly, for example "What would you stop using or ignore to make room for this?", to test whether a concept changes behavior rather than only earning approval. Conveo's MaxDiff runs the ranking and the follow-up probing in the same AI-moderated session, so the preference score, the resulting preference shares and the reasoning behind them come from the same participant.

A typical MaxDiff study includes 12 to 20 items. A list under eight items won't create meaningful trade-offs, since participants have too little to weigh against each other. More than 25 items pushes session length beyond 25 minutes and risks cognitive overload, degrading the quality of the resulting survey data: more data collected badly is worse than less data collected well. Add a "missing item" check at the end, because a fixed list can omit the factors that actually drive decisions.

Plan for roughly 100 participants for a stable MaxDiff ranking, gathered across multiple MaxDiff questions per session so each item is traded off enough times. That is larger than the eight to 12 participants typical for pure qualitative interviews, and smaller than the 200+ often required for conjoint analysis. If you plan to segment by persona, market or use case, plan for a larger sample; Conveo's research team sizes it per study rather than applying a fixed rule of thumb.

Yes, with caution. MaxDiff works across markets, but translation and cultural norms create validity risks that a single-market study never has to account for. Item wording that scales consumer preferences cleanly in one language can shift meaning entirely in another, distorting the comparisons the method depends on. Back-translate items to catch meaning shifts, then run a pilot in each market to confirm consistent interpretation before fielding the full experimental design at scale. Be cautious with direct pricing, hypothetical preference and emotion questions in multi-market studies, since cultural norms can inflate or suppress survey responses in ways that look like real differences in customer preferences.

Qualitative insights at the speed of your business

Conveo automates video interviews to speed up decision-making.

Your next read.

Articles

MaxDiff analysis: how to interpret scores accurately

Utility scores rank every item in a MaxDiff study, but close gaps often sit inside the margin of error. Here is how to tell signal from noise and pair each score with the reasoning behind it.

Headshot of Alex de Hemptinne

Alex de Hemptinne

Head of Customer Success

News

Greenbook - Conveo AI Demo Recording - From Questions to Insights in Minutes

Generative AI is transforming market research, but the real breakthrough isn’t just faster data processing, it’s the way AI can work alongside humans to uncover deeper, more actionable insights. Together with Greenbook, Conveo recently organized a webinar where Niels, Head of Research, demonstrated how the company’s AI Insights Platform streamlines every stage of the research process, from study design to real-time moderation to instant, multi-layered analysis, using a global “coffee rituals” study to showcase its ability to blend speed, scale, and human empathy.

Headshot of Hendrick Van Hove

Hendrik Van Hove

Founder & CPO

Success stories

Canva brings the voice of the consumer into every decision with Conveo

A study launched at 6:15 p.m. Results before breakfast. See how Canva uses Conveo to run research at the speed decisions actually happen.

Rómulo Rejón

Head of Customer Marketing