Analyze what participants say, how they say it and what's on camera
Conveo reads words, tone of voice and what the camera shows in video interviews, and links every cue to the moment it happened.
Emily KavanaghAt a glance
- Verbal, vocal and visual cues detected in recorded video interviews
- Brands, products and objects in frame picked up without being named
- Every cue links to its moment in the recording
- Emotion charts per question, with coverage shown alongside
- Filter an interview's cues by emotion, category and intensity
Release details
- Shipped
- Area
- Analysis & AI
- Kind
- Improved
- Author
- Emily Kavanagh
Ask someone which of two brands they prefer and they may say they don't mind. Watch the same answer on video and you might see them slow down at the cheaper one, smile at the other, and sit in front of a cupboard stocked with it. A transcript keeps the words and drops the rest. Checking tone and body language by hand works for five interviews and falls apart at fifty.
Multimodal analysis, how it works and what it changes, no sound needed.
How it works
Three layers of signal. Conveo analyzes recorded video interviews for verbal cues (words, phrasing, stated intent), vocal cues (tone, pace, pitch, hesitation) and visual cues (facial expressions, gestures, behavior).
Context in the frame. Brands, products and objects visible around the participant are picked up, even when nobody names them.
Every cue has a timestamp. Select a cue and the recording jumps to that moment, so you can replay the exchange around it.
Patterns across the study. Emotion charts in Question coding show how cues vary from question to question, with coverage shown alongside.
Check the moment behind every reading
Each interview has an emotion view. A timeline with question markers shows what the participant was answering when a cue appeared, and each cue is labeled as visual, vocal or verbal. Filter by emotion, category and intensity, then compare what you see with the participant's own words before you quote it.
Research teams asked for it in five situations
These come from teams we work with in consumer goods, food and drink, flavors and ingredients, home appliances, and advertising and research agencies. We kept them anonymous.
Taste and product tests. The first reaction to a flavor or a new product, next to the rating given afterwards.
Ad and concept testing. How people respond to creative moment by moment, alongside what they say about it.
In-home ethnography. Cooking, cleaning or loading the dishwasher, filmed while participants talk it through.
What sits in the pantry. Brands and pack sizes spotted on camera without the participant having to name them.
Shop-alongs. Reactions to shelf, packaging and pricing while participants shop live on camera.
What this means for your research
The say-do gap becomes something you can find in the data and replay. You can scan a study for hesitation or delight, then go straight to the clips that back it up.
Some limits to know. Cues are automated interpretations of what was recorded. They do not tell you what someone felt inside, and they do not prove that a stimulus caused a reaction. People express themselves differently, so a still face is no verdict, and an interview without usable video has no visual reading. Treat cues as leads for qualitative review and check them across several participants.
Read the emotional analysis guide.


