MPredict anythingMIROFISH 米罗鱼
Respondent labMethod Comparison

Synthetic Users vs. Real Customer Research

Aug 8, 202611 min readMiroFish Editorial
Respondent lab
  1. 01Question
  2. 02Synthetic screen
  3. 03Real study
  4. 04Calibration
  5. 05Decision
Evidence states
Grounded
Inferred
Unsupported
Needs real research
Direct answer

Direct answer

Synthetic users can help teams rehearse likely reactions, draft research questions, and inspect assumptions before spending research budget. Real customer research is still required for lived experience, prevalence, usability, willingness to pay, and behavior. Treat synthetic users as a planning layer that improves real studies, not as a replacement for real customers.

Research ledger

Synthetic vs. real research matrix

Use this matrix to choose the right evidence level for the decision rather than forcing every question through one method.

QuestionSynthetic users help withReal research needed forDecision limit
Concept clarityFind likely confusionConfirm real comprehensionRevise before launch tests
Market demandList hypothesesMeasure prevalence and intentDo not forecast demand
UsabilityPredict friction pointsObserve real interactionDo not replace task testing
PricingSurface pushback themesMeasure willingness to payDo not set price alone
Q01

What is the difference between synthetic users and real customer research?

Synthetic users are model-based simulations of possible customer responses, while real customer research observes actual people, choices, constraints, and behavior.

The difference matters because many customer questions depend on reality. A real buyer can misunderstand a feature, hesitate because of budget politics, abandon a task, or reveal an unexpected workaround. A synthetic user can only reason from the information and assumptions provided to the model.

Synthetic users are therefore best used upstream. They help teams prepare better research, not avoid research. The output can make a study sharper by naming hypotheses, segment risks, and questions that deserve real evidence.

Q02

What do synthetic users do well?

Synthetic users do well at structured brainstorming, objection mapping, wording checks, persona stress tests, and research-plan preparation.

A grounded synthetic panel can compare concept variants quickly. It can show that one message triggers trust concerns, another creates category confusion, and another lacks proof. That is useful before the team pays for recruiting or launches a larger survey.

The value comes from speed and breadth. Synthetic users can explore more variants than a small research team can interview in a day. The weakness is that this breadth is simulated, so the strongest outputs are hypotheses and research priorities.

  • Map objections
  • Draft better probes
  • Screen variants
  • Find missing segments
Q03

What does real customer research do that synthetic users cannot?

Real customer research observes actual experience, behavior, emotion, accessibility constraints, purchase friction, and segment prevalence.

Interviews reveal lived context. Surveys estimate prevalence. Usability tests show what people do with an interface. Field data shows behavior after a real choice. These forms of evidence answer questions a language model cannot settle by generating plausible responses.

Synthetic users can make those real studies better. They can suggest where to probe, which claims to test, and which segments to recruit. But the authority for customer truth should remain with observed customers.

Q04

How should teams combine synthetic users and real research?

Use synthetic users before real research to sharpen hypotheses, then compare their themes against interviews, surveys, usability tests, or behavior.

A practical sequence is: run a synthetic panel, code themes, write a real research guide, collect observed evidence, and update the synthetic workflow based on what matched or failed. This creates a calibration loop instead of a one-off demo.

The comparison should preserve misses. If synthetic users invent an objection that real customers never mention, that is useful calibration. If they miss a major accessibility concern, the workflow should be narrowed or redesigned.

Q05

Which decisions are safe for synthetic users?

Low-risk, reversible decisions such as wording cleanup, research planning, and variant shortlisting are safer than launch, pricing, eligibility, or compliance decisions.

The safer question is not whether synthetic users are accurate in general. It is what happens if they are wrong. If the cost is a better interview guide or an extra concept variant, the method can be proportionate. If the cost is excluding a customer group or setting a price, real evidence is required.

NIST's risk-management framing is useful here: increase evidence, review, and controls as consequence and irreversibility increase.

Q06

How should synthetic and real findings be reported together?

Report synthetic themes, real evidence, conflicts, unresolved gaps, and the decision each evidence layer is allowed to support.

Do not blend synthetic quotes and customer quotes as if they have the same status. Label the source of every theme. State which findings were confirmed by real customers and which remain hypotheses.

A clean report protects the team from overclaiming. It also makes synthetic research more useful because future runs can be calibrated against real findings instead of judged by surface plausibility.

Research notes

How should the comparison be used in practice?

Synthetic customer research should start with the decision, not with a panel prompt. A team should write the product question, the audience it wants to understand, the action it may take, and the evidence threshold that would make the result useful. Without that anchor, generated feedback can become a collection of plausible quotes. Plausible quotes may inspire a workshop, but they are not research evidence unless the workflow records what the synthetic panel was built from, what it was allowed to infer, and where real customer data would be needed before action.

The source packet is the practical difference between a synthetic respondent and a stereotype. It should include current product context, customer interviews, support themes, sales objections, usage data, survey findings, competitor claims, category language, pricing constraints, and any known segment differences. The packet should also mark gaps. If the team has no evidence for a segment, the synthetic panel can explore possible reactions, but it should not pretend to represent that segment. Missing evidence is a research task, not a prompt-writing problem.

A synthetic panel needs coverage logic. Decide which segments matter to the decision, which attributes are grounded, which are intentionally varied, and which attributes are excluded because they are irrelevant or unsupported. Persona names and demographic detail are less important than decision-relevant variables: job to be done, constraint, budget authority, prior awareness, trust source, current workaround, adoption risk, and reason to reject. A small grounded panel is often more useful than a large decorative panel that only varies surface-level biography.

The output should be coded into themes, objections, hypotheses, and evidence gaps. Raw transcripts are useful for inspection, but they should not be the final research artifact. A product manager needs to know which objections repeated across segments, which appeared only under one assumption, which response depended on unsupported source material, and which claim should be tested with real customers. Coding makes the simulation auditable and prevents one vivid quote from dominating the decision.

Validation is local. A synthetic respondent workflow that helps screen early messaging may fail for pricing, regulated categories, minority segments, accessibility needs, or high-stakes customer decisions. The validation record should state the domain, audience, source cutoff, model version, prompt version, comparison evidence, and intended use. If the workflow has not been compared with real interviews, survey data, usability tests, or observed behavior in a similar setting, label it exploratory. That label does not make the work useless; it keeps the claim honest.

Real customer research remains the standard for customer truth. Synthetic customer research can prepare the real study by surfacing hypotheses, draft questions, missing segments, confusing language, and likely objections. It can help teams spend research budget more deliberately. It should not replace interviews when empathy, lived experience, accessibility, legal risk, or purchase behavior matters. It should not replace surveys when the team needs prevalence. It should not replace usability testing when the question depends on interaction with a real interface.

Governance should match consequence. Low-risk concept exploration can use a lightweight protocol and a clear handoff. Pricing, health, finance, employment, public policy, or access-related decisions require stronger review, real respondent evidence, and explicit human accountability. NIST AI RMF is useful because it asks teams to map the context, measure risks, and manage the system within its intended use. A synthetic panel that sounds confident can still omit affected users, amplify bias, or turn weak evidence into a fluent recommendation.

A useful handoff memo separates four layers. First, what synthetic respondents said. Second, which themes were stable across repeated runs. Third, which source evidence supports those themes. Fourth, what real customer observation should come next. This format gives teams immediate value without overstating certainty. The memo can justify revising copy, narrowing a survey, recruiting a missing segment, or preparing a usability task. It cannot justify saying that customers will buy, churn, comply, or adopt at a specific rate unless comparable observed data supports that claim.

Source ledger

Evidence used

  1. R-01
    Using GPT for Market Research

    Harvard Business School Working Paper

    A market-research paper testing whether large language models can approximate human survey patterns in bounded settings.

  2. R-02
    Out of One, Many: Using Language Models to Simulate Human Samples

    Political Analysis

    A foundational study on simulating human samples with language models, useful for understanding both potential and sampling limits.

  3. R-03
    Generative Agent Simulations of 1,000 People

    Stanford University research team

    An interview-grounded agent simulation study evaluated against held-out behavioral measures.

  4. R-04
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

    A risk-management framework for mapping, measuring, and managing AI systems in their intended context of use.

  5. R-07
    AAPOR Code of Professional Ethics and Practices

    American Association for Public Opinion Research

    Ethics guidance for survey and public-opinion research, useful when synthetic respondents might be confused with real respondent evidence.

Related cluster
From panel to research plan

Use simulated customer reaction to prepare better real studies.

Ground a panel in your own source material, surface objections, and turn the output into a real research handoff.

Compare research paths