AI Moderated Interviews: How to Use Them for Concept Testing
How to run concept tests with AI moderated interviews so you learn why an idea lands, not just which concept scores highest.

Most concept tests tell you which idea scored higher. They rarely tell you why. A respondent rates a pack a 6 out of 10, ticks "somewhat unique," and moves on. You still do not know whether the name confused them, the benefit felt generic, or they would never buy this category at that price.
That missing "why" is the job of an interview. The problem is that interviews have never fit a concept sprint. You cannot put 200 people through a live moderator in a week, and a 12-person focus group is too small to trust when you are choosing between three directions.
AI moderated interviews sit in that gap. Each person sees the concept, answers in their own words, and gets follow-up questions in real time — at survey volume, without a researcher in the room.
Why concept tests stall on scores
A standard concept testing survey is good at ranking. Comparison tests, monadic tests, and sequential monadic tests all produce clean numbers: purchase intent, uniqueness, relevance, believability. Those numbers are useful. They are also incomplete.
Three things tend to go missing:
- Comprehension. People rate ideas they did not actually understand. A high uniqueness score on a misunderstood concept is a false positive.
- The reason for rejection. "Unlikely to buy" is a decision, not a diagnosis. You need to know whether to kill the idea, rewrite the benefit, or change the audience.
- The unexpected thread. The most useful finding is often the one you did not put on the Likert scale: a use case you had not named, a competitor they mentally compared it to, a word in the headline that made them distrust the brand.
Focus groups recover some of that language, then introduce groupthink and a sample too small to choose a winner. Human-led interviews recover it well, then blow the timeline. AI moderation is the method that keeps the probe without giving up the N.
What an AI-moderated concept test actually looks like
Each session is a loop: show the stimulus, ask a question, listen, probe if the answer is thin or interesting, then move on. Respondents speak, type, or use video. Interviews run in parallel, so 150 people can finish in a day or two instead of a month of calendar slots.
The research design still looks like a concept test you already know:
- Screen for category users, rejectors, or the segment you actually need.
- Expose one concept (monadic) or a short sequence (sequential). Keep exposure fair if you are comparing.
- Capture a few closed items so you still have scores: appeal, uniqueness, relevance, intent.
- Interview the why with a short discussion guide. This is where AI moderation earns its keep.
- Read themes and quotes against the scores, then decide: build, rewrite, or kill.
Voiceform is built for this loop. You can attach packs, product pages, feature lists, Figma files, or a paragraph of copy, keep the probes consistent across cells, and try a concept testing interview to see how a session feels from the respondent side.
Write a discussion guide, not a questionnaire
The follow-ups are where the depth comes from, so the guide should be short. Five to seven core questions is plenty. If you ask twenty, the AI has no room to probe and you have reinvented a long survey.
Use this as a starting structure. Adjust the wording to the stimulus; do not skip the jobs underneath.
- First reaction. "You just saw this idea. What went through your mind?" Probe anything vague: "interesting," "fine," "not for me."
- Comprehension. "In your own words, what is this offering, and who is it for?" If they cannot explain it, later scores are noise.
- Benefit. "What would this do for you that you cannot already do?" Probe the gap versus what they use today.
- Friction. "What would stop you from trying this?" Price, trust, hassle, and "I do not get it" are different problems.
- Fit. "When would you actually use this?" Listen for occasions, not abstract praise.
- Choice. If they saw more than one: "Which would you pick, and what made the others lose?" Force a reason, not just a favorite.
- Missing. "If you could change one thing before this launched, what would it be?"
For broader interview prompts you can steal from, customer interview questions is a useful companion. The difference here is that every question should serve a concept decision, not a general relationship interview.
Set probing so it stays useful
A well-configured AI moderator does not follow up on everything. Exhausted respondents give worse answers, and a probe on every rating turns a 12-minute test into a 30-minute one.
Concentrate depth on the questions that change the decision:
- Confusion on the concept
- Low appeal or low intent
- A claim they do not believe
- A comparison they make to a competitor you did not mention
Let lighter items stay light: demographics, category usage, a simple uniqueness rating. You still want those data points. You do not need three follow-ups on them.
Pilot with five to ten people before you field the full sample. You will find questions the AI probes in an unhelpful direction, a stimulus that does not load, or a phrase respondents consistently misread. Fixing that after five responses is cheap. After three hundred it is not.
Monadic, sequential, or comparison
AI moderation does not replace those designs. It sits on top of them.
- Monadic. Each person sees one concept. Best when you need uncontaminated reactions and enough sample per cell to compare scores. Pair each cell with the same discussion guide so the "why" is comparable.
- Sequential monadic. Each person sees more than one, in rotated order. Efficient, but fatigue and contrast effects creep in. Cap it at two or three concepts and keep the interview portion short.
- Comparison. People see options together and choose. Good for a forced pick. Worse for understanding each idea on its own. If you use it, still interview the loser: why it lost is often more actionable than why the winner won.
The rule of thumb: if the decision is "which concept goes forward," you need both a score you can compare and language that explains the gap. Do not run a comparison test with no probes, and do not run 40 interviews with no closed items.
Who to recruit, and how many
AI moderation is worth using from around fifty participants. Below that, a skilled human interviewer is usually the better spend.
For concept testing specifically:
- One concept, directional. 50–75 category users. Enough to see whether the idea is confused, appealing, or dead.
- Two or three concepts, a decision. 100–150 per cell for monadic, or 150–300 total for sequential, depending on how small a difference you need to detect.
- Creative or pack work with segments. Size the cells to the segments that matter, not the overall market. A winning score among non-users can hide a miss among current buyers.
Screen tightly. Concept tests fail quietly when they fill with people who would never buy the category. If you need a panel, Voiceform respondents or partners like Prolific keep recruitment from becoming the long pole.
How to read the output
You will have scores, transcripts, and themes. Use them together.
Start with the closed items so you know which concept is ahead and among whom. Then read the interview themes as a diagnosis:
- High intent, confused explanation. The idea has heat and a communication problem. Rewrite before you build.
- Clear explanation, low intent. They got it and still would not buy. That is a concept problem, not a copy problem.
- Low uniqueness, frequent competitor mentions. You are landing in an occupied slot. Change the claim or kill it.
- A use case you did not write. That thread is often the real product. Follow it in a second, tighter test.
Do not treat a handful of colorful quotes as the finding. Look for the same reason showing up across dozens of interviews. The method's advantage is sample size. Use it.
If you need a refresher on collecting qualitative data at scale, what is a qualitative survey covers the fundamentals. For the method itself, including where AI moderation is a poor fit, see what is an AI moderated interview.
Common mistakes
- Too many concepts per person. Three is already a lot once each one gets a spoken debrief. Four or five produces noise.
- A questionnaire pretending to be a guide. If every question is closed plus an optional comment box, you did not run an interview.
- Probing on every item. Pick the two or three questions that decide the round.
- Leading the moderator. "Did you love the bold new look?" will get you the bold new look. Ask what they noticed first.
- Skipping the pilot. The first fielding always teaches you something about the guide.
- Reading quotes instead of themes. One vivid respondent is not a segment.
Frequently asked questions
Do I still need rating questions?
Yes. Scores tell you which concept is ahead. Interviews tell you whether to trust that lead and what to change. Run both in the same study.
Can people react to a prototype, not just a board?
Yes. Send them to a Figma file or staging URL, ask them to talk through it, then debrief. You get behavior plus explanation, which a static board cannot give you.
How long should the interview be?
Ten to fifteen minutes is the useful range for a concept test. Long enough for a first reaction, comprehension, and friction. Short enough that sequential exposure still works.
Will people criticize a concept they know is unfinished?
Often more honestly than they would with a human moderator in the room. Social pressure drops. You still have to ask for criticism directly, or you will collect polite compliments.
When should I not use AI moderation?
Sensitive or distressing subject matter, very small expert samples, and live co-creation still want a person. Early pack and product concepts with a few hundred category users are a strong fit.
The short version
Concept testing has always needed two outputs: a ranking and a reason. Surveys delivered the ranking. Interviews delivered the reason, too slowly. AI moderated interviews let you keep both in the same fieldwork window — provided you write a real discussion guide, probe the questions that change the decision, and read themes against the scores.
If you want to run one, try Voiceform free, try a concept testing interview, or book a demo.


