MPredict anythingMIROFISH 米罗鱼
Experiment protocolAgentic Experimentation

Can AI Agents Simulate A/B Test Outcomes?

Aug 6, 202615 min readMiroFish Editorial
Experiment strip
Protocol / pretest
  1. 01
    StageHypothesis
  2. 02
    Control ATreatment B
  3. 03
    StagePopulation
  4. 04
    StageRepeated runs
  5. 05
    DecisionLive-test gate
Direct answer

Direct answer

AI agents can simulate reactions to A/B variants for candidate screening, mechanism discovery, and hypothesis stress-testing. They cannot establish the causal effect that a live randomized experiment measures. Treat an agentic experiment as a calibrated pretest: ground agents in relevant evidence, repeat controlled runs, compare against held-out outcomes when available, and send consequential decisions through a real A/B test.

Protocol artifact

Agentic experiment evidence ladder

Move downward only when the previous layer is documented. Simulation evidence narrows the live test; it does not inherit the authority of randomization.

StageControlTreatmentDecision rule
QuestionCurrent experienceProposed changeName the behavior that could change
SimulationRepeated A runsRepeated B runsCompare coded mechanisms and outcomes
CalibrationKnown historical outcomeHeld-out historical outcomeQuantify error before prospective use
Live gateRandomly assigned usersRandomly assigned usersUse a prespecified causal metric
Q01

What does it mean to simulate an A/B test with AI agents?

An agentic A/B simulation exposes defined synthetic participants to control and treatment conditions, repeats the comparison, and records how outcomes and mechanisms differ inside the model.

The useful unit is not a persuasive transcript. It is an experiment specification: a decision, a control, a treatment, a target population, an outcome rubric, and a repeatable run configuration. Agents should receive only the information that a comparable real participant could know at that point in the journey. The resulting contrast describes the simulated system under those conditions.

This makes agentic testing different from asking one language model which headline it prefers. A preference prompt has no assignment protocol, population model, repeated trials, or independent outcome definition. A structured simulation can reveal objections, interpretation paths, interaction effects, and overlooked segments. Those are useful inputs to experimental planning even when the generated frequencies are not population estimates.

MiroFish can support this pre-launch rehearsal by turning source material into a multi-agent scenario that teams can inspect and rerun. It should be used to compare candidate narratives and surface mechanisms, not to claim that a simulated lift is the lift a production experiment will deliver.

Q02

What does current research show about agentic A/B prediction?

Early research suggests that calibrated agentic systems can recover useful directional signal in some historical experiments, but the evidence is preliminary, domain-specific, and not a license to replace live tests.

[Experimental] A 2026 arXiv preprint evaluated simulated predictions against 67 historical marketing A/B tests. The paper reports 0.70 baseline sign agreement and substantial error reduction after a two-phase pre-period calibration procedure. It also reports lower standard errors for a within-subject design. These are results from one research design and dataset, not a universal benchmark for AI agents or MiroFish.

The stronger lesson is methodological. The researchers did not assume that plausible agents were valid. They compared predictions with known outcomes, estimated systematic error, and changed the design to reduce noise. That sequence—historical backtesting, calibration, then prospective use—is more defensible than judging quality from realistic language alone.

Other generative-agent research uses interview evidence and held-out behavioral measures to evaluate synthetic representations. Together these studies support a cautious direction: rich grounding and independent evaluation matter, while transfer to a new population, product, or outcome must be demonstrated rather than presumed.

Q03

Which decisions are suitable for an AI-agent pretest?

Use agentic pretests for reversible, pre-launch decisions where the main value is ranking options, finding failure mechanisms, or improving the design of a later live experiment.

Good candidates include choosing which message variants deserve traffic, identifying likely objections before a concept test, rehearsing how social groups may transmit a claim, and checking whether a survey question is interpreted consistently. These decisions benefit from breadth and speed, and a wrong simulated ranking can still be caught in the live gate.

Avoid using synthetic outcomes as the sole basis for high-impact pricing, health, employment, credit, legal, or public-policy decisions. The model may omit affected groups, encode historical bias, or generate a stable story that has no causal connection to real behavior. NIST AI RMF recommends matching measurement and governance effort to the context and potential harm.

A practical rule is to ask what happens if the simulation is wrong. If the answer is 'we test a different headline next week,' a pretest may be proportionate. If the answer affects rights, safety, material access, or irreversible budget, require stronger independent evidence and accountable human review.

  • Screen many variants before spending traffic
  • Discover objections and causal hypotheses
  • Improve treatment wording and measurement
  • Never use generated lift as guaranteed business impact
Q04

How should teams design a valid agentic comparison?

Hold non-treatment inputs constant, define assignment and exposure rules before running, isolate the treatment contrast, and code outcomes with a rubric that does not change after results appear.

Start with the estimand in plain language: whose behavior, under which alternative experiences, over what period, and measured how? The Rubin potential-outcomes framework explains why causal effects require comparing outcomes that cannot both be observed for the same real unit at the same time. Synthetic agents can render both paths, but that convenience does not solve the identification problem in the real population.

Version the source packet, agent definitions, model identifier, prompts, sampling settings, tools, and timestamps. Run a baseline repeatedly before comparing variants so you can see natural system variation. Use the same coding rules for both arms and hide the arm label from human coders where practical.

Decide whether a within-subject or between-subject design better matches the question. Within-subject comparisons can reduce agent-level noise but risk memory, order, and contrast effects. Between-subject comparisons preserve isolation but require enough independent agents and runs to separate treatment signal from synthetic population variation.

Q05

How can simulated outcomes be calibrated before prospective use?

Backtest on comparable completed experiments, reserve outcomes from the build process, estimate directional and magnitude error, and define the confidence needed before simulation can influence prioritization.

Calibration requires examples where the real outcome is already known. Build the agentic experiment without exposing that outcome, record its prediction, then compare. Track sign agreement, rank correlation, absolute error, calibration by segment, and cases where the mechanism was wrong even when the direction was right. One successful example is not a validation set.

Pre-period data can help identify systematic differences between synthetic and observed baselines. CUPED is a related but distinct technique for live randomized experiments: it uses pre-experiment covariates to reduce variance. Do not label an informal prompt adjustment as CUPED or assume variance reduction removes bias.

If no historical experiments are comparable, label the output exploratory. The team can still use it to improve treatments and measurement, but it should not attach a probability or expected lift to the live outcome. Calibration is local to the domain, population, metric, and model configuration tested.

Q06

When must a simulated experiment hand off to a live A/B test?

Hand off when the decision depends on a real causal effect, the cost of error is material, or the simulated recommendation will change customer exposure, budget, access, or policy.

The handoff memo should state which variants were dropped, which mechanisms survived repeated runs, which segments remain uncertain, and what the live experiment must measure. It should also list the simulation's known gaps so the production test does not quietly inherit synthetic assumptions.

Prespecify the live primary metric, guardrails, randomization unit, analysis window, stopping rule, and minimum detectable effect. Use the simulation to sharpen these choices, not to peek at live results or redefine success after launch. The live test remains the source of causal evidence for the actual population and implementation.

After the test, compare observed mechanisms and outcomes with the simulation. Retain misses as carefully as hits. That feedback creates a domain-specific calibration record and tells the team when future agentic pretests deserve more—or less—weight.

Protocol notes

What should an evidence-ready agentic experiment record?

An evidence-ready record begins with provenance. Give the experiment a stable protocol ID and preserve the exact business question, decision owner, preparation date, source cutoff, and intended use. Attach the control and treatment artifacts exactly as agents received them, including surrounding context, timing, channel, and any interaction affordances. List every source used to construct the population and distinguish observed attributes from modeling assumptions. If an analyst cannot reconstruct which claim entered which agent, the result may still inspire discussion, but it is not ready to influence a consequential shortlist. Provenance also prevents a subtle form of hindsight: quietly replacing a weak source, refining a treatment, or changing eligibility after seeing outputs while continuing to describe the work as one experiment.

The population record should describe coverage, not merely persona count. Name the eligible population, segment definitions, evidence behind segment proportions, attributes intentionally held constant, and groups for which evidence is sparse. Record how individual agents were instantiated, whether histories were reused across arms, and how social relationships or network positions were assigned. A hundred generated profiles do not represent a market simply because they are different. Diversity in names and prose can coexist with identical incentives or stereotypes. Reviewers need to see which real dimensions of heterogeneity could affect the treatment and whether those dimensions were grounded, sampled, stress-tested, or omitted. When population proportions are unknown, run explicit alternative mixes and report whether the shortlist changes rather than choosing one convenient synthetic census.

The assignment record should make the contrast auditable. For a between-subject design, retain group allocation, stratification or matching variables, repeated assignment seeds where supported, and baseline balance summaries. For a within-subject design, retain exposure order, context-reset procedure, counterbalancing, and any washout or independent-state method. When agents interact, document the randomization level and possible spillover between control and treatment worlds. Individual assignment inside one shared network can contaminate the contrast because treated information reaches controls; world-level assignment estimates a different effect. Neither choice is automatically wrong, but the estimand must match it. If assignment is described only as 'we ran both prompts,' the experiment lacks the structure required to interpret a difference.

The measurement record translates generated behavior into evidence. Define one primary outcome tied to the decision, material guardrails, and a small set of exploratory mechanisms. Store the coding rubric, examples, counterexamples, missing-output rule, coder identity or model version, and agreement check. Free-form summaries should link back to coded runs so a reviewer can inspect minority and tail cases. Avoid scoring generic enthusiasm when the decision concerns action: an agent can praise a treatment while refusing to buy, share, comply, or continue. Report the full arm distributions, number of agents, number of repeated worlds, ambiguous cases, and exclusions. A percentage without its denominator and assignment structure invites readers to mistake synthetic frequency for estimated population prevalence.

The technical run manifest separates treatment signal from system drift. Preserve model provider and identifier, dated model version when available, system and user prompts, temperature or sampling settings, tool definitions and versions, retrieval configuration, timestamps, concurrency, retry rules, context limits, and random seeds where the stack supports them. Record failures rather than rerunning silently until a complete response appears. If a provider updates a model during a study, mark the boundary and treat it as a technical variant. Exact transcript reproduction may be impossible in stochastic hosted systems, but MIASE's minimum-information principle still applies: another team should be able to reconstruct the procedure, identify configuration differences, and compare outcome distributions under a meaningfully equivalent setup.

The analysis record should preserve prespecified and exploratory work separately. State how arm contrasts, paired differences, rank ordering, uncertainty, and sensitivity variants were calculated. If the team selected the best of many outcomes or treatments, record the multiplicity rather than reporting the final comparison as though it were the only question asked. Keep null and contrary findings in the same ledger as supporting findings. Generated explanations can suggest mechanisms, but reviewers should trace them to behavior and source evidence instead of rewarding fluency. When sample sizes are small or calibration is absent, use terms such as recurring, uncommon, or design-sensitive instead of attaching a probability that resembles a real-world rate.

The calibration record links simulation to observations without erasing misses. Before revealing a completed experiment's outcome, freeze the synthetic prediction and its confidence label. Then compare control baseline, direction, magnitude, rank, segments, and mechanisms. Record whether the case belonged to the build set, adjustment set, or untouched evaluation set. If a change is made after a miss, version it and evaluate the new procedure on different held-out cases. Do not let one aggregate score hide where the model fails: a workflow may rank message variants usefully but fail on prices, new markets, or treatments that depend on operational friction. Calibration should end in a narrow usage statement and an expiration condition tied to material model, prompt, evidence, or population changes.

Finally, the decision record connects the experiment to accountability. Name which candidates advance, which are revised or dropped, the reason under the prespecified rule, the unresolved risks, and the next source of evidence. For an advancing treatment, link the live protocol: eligible units, randomization, primary metric, guardrails, power basis, analysis window, stop conditions, owner, and rollback path. State explicitly that the simulation did not establish a live causal effect. After production results arrive, append them to the same record and compare the frozen expectations with observations. This final step turns agentic experimentation into a learning system. Without it, teams accumulate attractive simulations; with it, they accumulate local evidence about where simulation helps, where it fails, and how much weight future pretests deserve.

A mature program separates three kinds of uncertainty that are often collapsed into one confidence statement. World uncertainty concerns facts the team does not know about the market, affected groups, incentives, or future context. Model uncertainty concerns whether agents, prompts, interaction rules, and language-model behavior represent the intended mechanism. Sampling uncertainty concerns variation across assignments and stochastic runs. Each demands a different response: gather evidence for the world, redesign or calibrate the model, and repeat or rebalance the experiment for sampling. Adding more synthetic runs addresses only the third category. It cannot repair a missing stakeholder or an invalid treatment contrast. The protocol should label each material uncertainty, assign an owner, and state the observation or test that could reduce it.

Teams should also distinguish prediction from preparation. A simulated outcome may be operationally valuable even when it is not calibrated as a forecast. If several grounded worlds expose the same objection, the team can prepare an answer, instrument that objection in research, or add a guardrail to the live test without claiming the objection will occur at a particular rate. This is often the safest route to value: use generated behavior to expand the mechanisms considered, then let interviews, telemetry, experiments, and domain review determine which operate in reality. The decision memo should say whether an action is justified by expected impact, low preparation cost, reversibility, or calibrated evidence. These rationales are not interchangeable.

Governance should be proportionate but visible. A low-risk creative screen may need one analyst, a versioned protocol, and a live experiment owner. A pricing, eligibility, or public-policy scenario may require domain review, affected-group representation, legal or compliance approval, stronger held-out evidence, and an explicit prohibition on autonomous action. NIST AI RMF's govern-map-measure-manage cycle provides a useful structure: establish accountability, map context, measure performance and harm, and manage deployment and monitoring. The ledger can hold these controls without pretending that one checklist makes every use safe. Rigor, independence, and escalation should increase with the consequence of being wrong.

Schedule a dated review of the protocol and its sources; a pretest built for yesterday's population or product state should not silently govern tomorrow's experiment.

Source ledger

Evidence used

  1. E-01
    Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

    arXiv preprint

    [Experimental] A 2026 preprint evaluating agentic predictions against 67 historical marketing experiments. Its findings are preliminary and do not establish general predictive accuracy.

  2. E-02
    Designing Reliable Experiments with Generative Agent-Based Modeling

    arXiv preprint

    A research framework for experimental design with generative agents, including measurement, repeated trials, and design-validity concerns.

  3. E-07
    Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies

    Journal of Educational Psychology

    A foundational treatment of potential outcomes and the missing-counterfactual problem behind causal inference.

  4. E-04
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

    A risk-based framework for governing, mapping, measuring, and managing AI systems in their intended context of use.

  5. E-05
    Generative Agent Simulations of 1,000 People

    Stanford University research team

    A study of interview-grounded agents evaluated against held-out behavioral measures, illustrating both the value and limits of rich persona evidence.

  6. E-03
    Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data

    Microsoft Experimentation Platform

    The CUPED paper explains how pre-experiment covariates can reduce variance in randomized online experiments without changing the estimand.

Continue the experiment
From protocol to rehearsal

Stress-test the variants before production traffic decides.

Use your own source material to rehearse treatments, inspect mechanisms, and prepare a sharper live experiment.

Rehearse an experiment