Vipul Nair

Product, Recruitment, Delivery: What Actually Separates One AI Research Partner From Another

Written by
Portrait of Vipul Nair.
Vipul NairCo-founder, Echovane
Illustrated future city viewed from a green hillside at dusk.

I have now sat on both sides of this conversation, and the pattern is consistent. The first ten minutes of an AI research demo look more or less identical wherever you go. There is an AI that talks, it asks a follow-up, and there is a dashboard with themes on it.

Then you run a study, and the differences turn out to have been substantial and almost entirely invisible in the demo.

This article is my attempt to say where those differences sit and what to ask in order to see them beforehand. I work at one of these companies, so read it with that in mind. I have tried to write the version I would have wanted when I was the one being pitched to.

Three Layers, and Most Evaluations Only Test One

Product is what a demo is designed to exercise, and it does matter. But a flawless moderator interviewing the wrong people produces a beautifully executed irrelevance, and a study that is methodologically perfect and lands as a transcript archive three weeks after the decision has been taken has not helped anybody either.

Let me take each layer in turn.

Diagram showing product, recruitment, and delivery as three layers of AI research partner evaluation.
Most demos test the product layer. Real study quality also depends on whether the right people are recruited and whether the work reaches stakeholders in a usable form.

Layer One: What the Moderator Can Do in a Live Conversation

It Should Behave Like a Call, Not Like a Form

Our moderator is conversational over video and voice. Participants do not type, do not read a wall of text to see the question, and do not press a button in order to speak. It works the way a video call works.

This sounds like an interface detail and it is not. Every element sitting between a participant and their own thinking costs you depth. Somebody composing a written answer is editing themselves as they go. Somebody reading a question on screen is processing it as a form, not as something a person asked them. Removing that friction is a substantial part of why long AI-moderated interviews hold together at all.

How to test it: take an interview yourself, as a participant, without a script prepared for you. Give it a lazy answer and see whether it lets you get away with it.

Length, Memory, and What Happens at Minute Seventy

We regularly run ninety-minute interviews, and the platform will go to two and a half hours, though we do not recommend it. The harder achievement is holding context across it without the conversation degrading, and our response times typically stay under a second and do not increase as context accumulates.

In research terms the payoff is concrete. If a participant says something at minute five and contradicts it at minute seventy, we remember and probe the contradiction. Those moments are frequently the most valuable thing in a transcript, because they are where the self-presentation and the actual behaviour come apart.

How to test it: ask what the average completed interview length is across their studies, and what happens to response time at minute sixty. Vague answers to that second question are informative.

Branching Into Sub-Topics, Not Just Follow-Ups

Nearly every platform in this category advertises contextual follow-up, so it is worth being precise about what that means.

Contextual follow-up deepens the answer and returns to the planned route. Sub-topic branching changes the route itself, based on what the participant said.

We do both, and the second is the one that changes study design. A participant says three of eight features matter to them, and the guide opens a different set of questions built around those three. A participant picks two ads out of a cluttered set, and the guide brings those two back and goes deeper on them specifically. Two participants in the same study can answer meaningfully different guides, and that is the intended behaviour rather than an inconsistency to be apologised for.

How to test it: ask to see a discussion guide with a real branch in it, rather than a follow-up being described as one.

Diagram comparing contextual follow-up with sub-topic branching in an AI-moderated interview.
Follow-up deepens an answer. Branching changes the route of the study based on what the participant says.

Observed Behaviour, Not Only Claimed Behaviour

This is the layer I would weight most heavily if I were the one evaluating, because it is the hardest thing to fake in a demo and the most consequential in a study.

Echovane is more than an AI-moderated interview platform. We run self-ethnographies and digital shop-alongs, and our visual intelligence analyses what participants are doing, not only what they say about it. It also lets us blend methods within a single study. Pairing a shop-along with an AI-moderated interview is one of the combinations we see most often, because the observation supplies the specifics and the interview supplies the reasoning.

How to test it: ask for a case where the system noticed something the participant never mentioned, and probed on it.

Pre and Post Tasks That Feed the Interview

Participants can complete tasks before and after the session, answering over voice, video, text and more. They can share content they like, write a brand love letter, or share something they are comfortable sharing from their own social feeds or AI chat apps.

What actually matters is the wiring between the two. A pre-task can be set up so that what happens inside it changes the probes in the interview. A participant completes a pantry audit, and the interview then asks about what was observed in that specific pantry, not about pantries in general.

How to test it: ask whether pre-task content is available to the moderator during the session, or only to the analyst afterwards. The difference is larger than it sounds.

Stimulus Handling, and Structure Inside Qual

For ad and concept work we sequence audio, video, image and text stimuli, support monadic and sequential monadic designs, rotate across concept sets, randomise, and apply the other standard approaches for removing order bias.

We also support closed-ended questions inside the same study, including grid questions where the design calls for them, for the moments when a stakeholder needs a number attached to whatever the qualitative has just surfaced.

How to test it: describe your actual concept test design and ask them to walk you through how it would be configured.

Layer Two: Whether the Right People End Up in the Study

This layer deserves an article of its own and has one, so I will keep it brief here.

The headline numbers are ninety countries, sixty-five languages and upwards of twenty million respondents, with coverage extending across fifty industries, a hundred job titles and a thousand skills. For niche recruits we integrate with specialist qualitative panels and run our own recruitment. Vetting is layered: identity checks, behavioural screening, fraud detection, and then the interview itself as a final quality gate, because a participant can pass every check at the door and still fail to sustain sixty minutes of specific conversation about a category they do not take part in.

Several cohorts can also be routed off a single screener, each carrying its own discussion guide, so that segments are defined by one instrument in one fielding window rather than three.

How to test it: give them your hardest specification rather than your easiest, and ask what the real feasibility and timeline are.

Layer Three: What Reaches a Stakeholder's Desk, and When

This is the layer that gets evaluated least and determines satisfaction most.

Full Service, and Where the Line Sits

We are done for you. You will never have to open the platform during the research unless you want to. We help with screeners and discussion guides, though the final decisions are always yours, and you open the platform once fieldwork is complete, because that is the point at which a researcher can properly mine the transcripts with the analysis tooling.

The line sits there deliberately. Running fieldwork is not the work insights teams are hired to do. Interpretation is.

The Analysis Surfaces

The full interview corpus is searchable by chat, for data tables, verbatims and more. The feature researchers tend to like most, though, is asking for insight clips. Something along the lines of "show me moments where the participant spoke about speciality coffee" returns the clips themselves, and a few of them can be assembled into a highlight reel for a presentation in a handful of clicks.

That last part carries more weight than it might appear to, because the gap between having the evidence and being able to show the evidence is where a lot of qualitative work loses its force in a boardroom.

What Stakeholders See

A client dashboard, which is a read-only version of the insights dashboard the researcher prepares. Stakeholders are spared the internal mechanics of the study and see the end insights. PowerPoint reports beyond the dashboard, which we ask you to flag at the start of a study rather than at the end of one. And Echovane study links can be appended to the end of quantitative studies where that is useful.

On security, we are SOC 2 compliant and ISO 27001 certified, with end-to-end data protection, granular access controls and a continuously managed security programme. In regulated categories this tends to be the first conversation rather than the last.

How to test it: ask to see the client-facing view rather than the researcher view. That is what your stakeholders will judge the work by.

Diagram showing the client-facing dashboard, analysis chat, PowerPoint reports, and study links as delivery surfaces.
The delivery layer determines whether evidence turns into something a stakeholder can actually use.

The Questions I Would Actually Ask

If I had one evaluation call and wanted to see past the demo, these are what I would use.

  • Can I take an interview as a participant, right now, without a script prepared for me?
  • What is the average completed interview length across your studies, and what happens to response time at minute sixty?
  • Show me a discussion guide with a real sub-topic branch in it.
  • Show me a case where the system observed something the participant never said, and probed it.
  • Here is my hardest recruitment specification. What is the real feasibility and timeline?
  • How is the screener written, and who writes it?
  • What does the client-facing deliverable look like, as opposed to the researcher view?
  • Which parts of this study will my team be expected to run?
  • What does this system not do well?

The last one is the most useful question on that list. Any company claiming its approach is right for every research question is not being straight with you, and the answer tells you more about how the partnership will go than any capability slide will.

Where We Are Not the Right Answer

So, in that spirit.

If your question is about group dynamics, about how social influence shapes the way opinions form, or about how people build on each other's ideas, a well-moderated focus group is still the right instrument and we are not trying to replace it.

If you need a handful of interviews with a small, senior audience you already have relationships with, the operational advantages of what we do compress to almost nothing. You would be better off picking up the phone.

And where the value of the research rests on a relationship between a particular moderator and a particular participant, built over years on a deeply sensitive topic, the best human moderators still bring something we do not.

Where This Leaves Us

The three layers are not equally visible, and that is the whole of the problem. Product is what gets demoed, recruitment is what gets assumed, and delivery is what gets discovered.

An evaluation that tests all three tends to reach a clear answer quickly. An evaluation that tests only the first tends to reach a confident answer that turns out to be wrong about six weeks later.

If you are running an evaluation and want to test all three layers rather than the first, I would be glad to take you through ours. Book a demo.