
TL;DR
MaxDiff (best-worst scaling) is a survey method that forces participants to trade items off against each other instead of rating them, producing a clear preference ranking.
The problem: Traditional rating scales let participants rate every item highly, since nothing forces a choice.
The fix: MaxDiff makes participants trade items off against each other, exposing what actually matters.
Best for: Insights teams needing a defensible cut list built from importance scores that actually separate.
The outcome: Real priorities, carried forward wave to wave, as actionable data instead of guessing what to build or drop. Conveo runs MaxDiff analysis as a continuous, AI-moderated program that builds with every wave.
What Is MaxDiff?
MaxDiff, short for maximum difference scaling, is a best-worst scaling survey method used across market research that asks participants to pick the best and worst options from a small set, repeated across several item sets. Each MaxDiff survey question forces a genuine trade-off rather than a rating, which is what separates it from an ordinary numeric scale. That forced-choice mechanism removes rating-scale bias: no clustering everything at "important," no cultural variance in how people use a 7-point scale. The output is a statistical analysis of relative preference: a clean ranking and trade-off view showing how each item compares against the others.

Why MaxDiff Beats Rating Scales
A features list where every item scores 8 or higher out of 10 on a numeric scale tells you nothing about which three to keep when the budget allows only three. This is the scale-use bias traditional rating scales are prone to: participants rate every product attribute, service attribute, or messaging line equally, so nothing stands out. MaxDiff solves this by forcing a choice: best versus worst, item against item, every round. That forced trade-off is what produces a real priority order, a measure of relative preference, instead of a wall of high scores. This is why Conveo builds MaxDiff as a forced-choice exercise from the start.
Dimension | Rating scale | MaxDiff |
|---|---|---|
Task | Rate each item on its own | Pick the best and worst item in each small set |
What it allows | Every item scored as important | One best and one worst per set, every round |
Typical output | Clustered high scores | A preference ranking with scores that separate |
Main bias | Scale-use bias, cultural differences in scale use | Depends on a balanced design and a fixed item list |
How a MaxDiff Question Works
A MaxDiff question works in three steps:

Show a small set. Four items appear at once, say four specific features of a product.
Force best and worst. Respondents choose the single best and single worst instead of rating each one.
Repeat across sets. This MaxDiff example repeats with different combinations of items across a typical 10 to 15 sets (also called rounds), meaning each respondent works through multiple MaxDiff questions before the session ends, so every feature appears a set number of times against different alternatives.
That balanced exposure is what lets researchers calculate a reliable preference score instead of a one-off guess, producing more reliable data than a single rating pass ever could.
When to Use MaxDiff
MaxDiff earns its place when stakeholders must weigh multiple attributes against each other and make a forced choice. That's true whether the decision is an internal roadmap call or a question about what consumers actually want most, beyond what scores highest on a customer satisfaction survey.
These are the classic forced-tradeoff decisions where rating scales flatten everything into ties:
Feature roadmaps needing to prioritize specific features under limited engineering capacity
Message testing where only three claims can headline packaging
Pricing strategies that need to know which price points actually move a decision
Benefit prioritization for a repositioning brief
Ranking questions built this way surface the important factors that budget, shelf space, or a launch deadline will demand anyway, instead of a wall of four-out-of-five scores.
When Not to Use MaxDiff
Where It Doesn't Fit
MaxDiff assumes you already know your item list, since ranking questions like these depend on having a fixed set of options to trade off. Exploratory questions, where the goal is to discover what to test rather than rank known options, need a different method first, whatever market research tools eventually run the trade-off exercise. Undefined categories, early-stage concepts, and open-ended "what matters to you" questions all fall outside its design.
The Gap Conveo Closes
The trade-off is simple: MaxDiff tells you what participants prioritize and leaves the reasoning out. Traditional MaxDiff stops at the ranking, leaving the reasoning to a separate follow-up question or a standalone study.
MaxDiff tells you what people prefer. Conveo's MaxDiff tells you why. After the choice tasks, the AI research assistant asks each participant why their best and worst items landed where they did, so the score and the story behind it come from the same participant.
Designing a Valid MaxDiff Study
Rankings are only as stable as the items and tasks behind them, and that stability is what separates reliable data from a directionless chart. Before trusting output, run this checklist:

1. Item Construction
Cut ambiguous MaxDiff items and double-barreled statements that bundle multiple attributes into one score. Flag translation risks early: a phrase that reads neutral in English can carry unintended weight in German or Japanese, distorting the preference measurement itself.
2. Cognitive Load
More than four to five items per set, or too many tasks in one sitting, produces fatigue and the scale-use bias MaxDiff exists to avoid. Too much cognitive load can lead respondents to satisfice, and the survey data may look decisive when it is actually noise. Keeping sets tighter preserves the greater discrimination between items that MaxDiff is designed to produce.
3. Design Balance
Every item needs roughly equal exposure and pairing across the MaxDiff study, the experimental design decision that determines whether the output can be trusted at all. Unequal exposure skews estimates before hierarchical Bayes modeling ever runs.
4. Quality Flags
Watch for inconsistent survey responses across repeated tasks, speeding, and straight-lining. Any of these means the underlying MaxDiff data isn't yet actionable data, regardless of how clean the output chart looks.
Interpreting MaxDiff Scores
What the Scores Mean
A ranked list is not an interpretation. MaxDiff measures relative importance on a shared scale: utilities from a hierarchical Bayes model place an item at 0.35 above one at 0.20, but the gap between adjacent importance scores rarely justifies a confident stakeholder claim on its own. Converting those importance scores into preference shares (the modeled likelihood that one item wins head-to-head against another) often tells a cleaner story to stakeholders than the raw score alone. Segmentation often reveals that a flat aggregate score is hiding two groups of consumer preferences pulling in opposite directions.
What the Scores Don't Tell You
MaxDiff scores also can't explain themselves. A high utility might mask substitutability between similar items, and a low one might reflect confusion rather than rejection, both of which can distort MaxDiff results if taken at face value. This is where Conveo's video-based reasoning probe helps: because the follow-up conversation happens on camera, tone and expression sit right next to the ranking whenever a stakeholder asks why a particular item won or lost.
MaxDiff in Practice: What It Looks Like in Conveo
MaxDiff measures are only as valuable as the program that runs them repeatedly. Here's what that looks like inside Conveo:
Conveo pairs MaxDiff with a follow-up conversation that probes the reasoning behind each choice, so the results explain which items win or lose with consumers and why.
Conveo runs MaxDiff in a continuous program, so each wave builds on what the last one learned instead of re-establishing everything each time a decision comes up.
"Conveo has definitely changed what we can pitch to clients, because it allows us to bridge the gaps between quantitative and qualitative methodologies and tell a more complete end-to-end story. It lets us close the loop: from the archetypes we've identified in the qualitative, which of them actually resonate the most? It feels more authoritative, more backed up by the data."
— Fergus Navaratnam-Blair, VP Trends and Futures, National Research Group (NRG)
While the Decision Is Still Open
Conveo interviews are AI-moderated, so participants complete the choice tasks and the follow-up conversation on their own schedule, across markets and time zones. The ranking and reasoning arrive together while the roadmap or pack decision is still open, and each wave builds on what the last one learned instead of starting from scratch. In practice, that can look like 100 interviews in 3 days. See the step-by-step setup guide for exact configuration.
Frequently asked questions
What is an example of a MaxDiff question?
How is MaxDiff different from conjoint analysis?
Can MaxDiff be used for concept testing?
How many items can you include in a MaxDiff study?
What sample size do you need for MaxDiff?
Can you run MaxDiff in multiple markets?










