MPredict anythingMIROFISH 米罗鱼
Experiment protocolLive-Test Handoff

Simulation-to-Live A/B Test Checklist

Aug 6, 202611 min readMiroFish Editorial
Experiment strip
Protocol / pretest
  1. 01
    StageSimulation review
  2. 02
    Control ATreatment B
  3. 03
    StageProtocol lock
  4. 04
    StageLive launch
  5. 05
    DecisionCalibration loop
Direct answer

Direct answer

Move from simulation to a live A/B test only after documenting the decision, treatment files, target population, repeated-run evidence, sensitivity findings, and unresolved risks. Freeze the live primary metric, guardrails, randomization unit, analysis window, and stopping rule before launch. Treat simulated outcomes as prioritization evidence; use observed randomized results for causal claims and future calibration.

Protocol artifact

Go/no-go handoff ledger

Every row needs an owner and retained artifact. A no-go is a request for better evidence or a safer test, not a failed project.

GateControl recordTreatment recordPass condition
FidelityVersioned A experienceVersioned B experienceOnly intended contrast differs
EvidenceBaseline distributionRepeated-run distributionShortlist survives key variants
SafetyExisting guardrailsTreatment-specific risksMonitoring and rollback are owned
ProtocolRandomized controlRandomized treatmentMetric and stopping rule are frozen
Q01

What must be complete before the handoff begins?

The simulation needs a frozen question, versioned treatments, grounded population, prespecified outcome rubric, repeated runs, sensitivity review, and a clear shortlist decision.

A handoff cannot repair an ambiguous simulation. Confirm that the control represents the current experience and the treatment differs only as declared. Archive source, prompt, model, population, assignment, and coding versions. Retain all material runs, including nulls and failures.

Summarize distributions and mechanisms rather than selecting a persuasive conversation. Record how results changed across treatment order, population variants, model configuration, and contested evidence. A fragile recommendation may still enter a live test, but its fragility should shape monitoring and interpretation.

Give each candidate a status: advance, revise, drop, or unresolved. The simulation's job is to reduce and improve the live experiment set, not to certify a production effect.

Q02

Which risks require a no-go or redesign?

Stop or redesign when the treatment creates unacceptable harm, lacks ethical exposure, cannot be measured reliably, or depends on a synthetic claim that has no relevant validation.

Review affected users, access, price, safety, privacy, fairness, and legal constraints. Simulation can surface failure stories but cannot authorize exposing real people to an unsafe condition. Domain and compliance owners must approve the live protocol where their expertise is required.

Check operational readiness. A promising variant is not testable if assignment leaks, logging differs by arm, support cannot handle failures, or rollback is unavailable. Add guardrail metrics and real-time alerts for known risks without turning every available metric into a decision endpoint.

NIST AI RMF recommends mapping context and managing risk throughout the lifecycle. Apply more scrutiny as impact and irreversibility increase. For high-stakes decisions, simulated support should remain secondary to independent evidence and accountable review.

Q03

What must be locked in the live experiment protocol?

Lock eligibility, randomization unit, treatment delivery, primary metric, guardrails, analysis population, time window, power assumptions, and stopping rule before reading results.

Define the estimand in operational terms. Specify whether the analysis is intent-to-treat, how repeated visitors are handled, what counts as exposure, and how missing data is treated. Confirm that instrumentation has equivalent semantics in both arms.

Base sample-size planning on the live metric's historical variance and a meaningful minimum detectable effect, not an unvalidated simulated lift. CUPED or other variance-reduction methods may improve sensitivity when their assumptions hold, but they belong in the live analysis plan.

Register primary and guardrail metrics separately from exploratory diagnostics. Set the experiment duration and stopping process. Repeated peeking followed by an unplanned stop changes error behavior and weakens the conclusion.

Q04

How should launch and monitoring be managed?

Use staged exposure, verify assignment and logging, monitor safety guardrails, and preserve the protocol unless a predefined stop or incident condition occurs.

Begin with an instrumentation or low-percentage ramp when feasible. Check sample-ratio balance, eligibility, exposure, event completeness, latency, and treatment fidelity before interpreting outcome movement. A technical failure can look like a behavioral effect.

Monitor guardrails at a cadence appropriate to harm. The team should know who can pause the experiment, how rollback works, and which incident threshold overrides the statistical plan. Document any intervention and treat the resulting analysis accordingly.

Do not use the simulation narrative to explain away an unfavorable early signal. The live data may reveal a mechanism that agents omitted. Preserve surprise as evidence rather than forcing it into the earlier story.

Q05

How should the live result be analyzed and communicated?

Analyze according to the locked plan, report uncertainty and guardrails, distinguish confirmatory from exploratory findings, and state the business action without overstating generalization.

Report assignment counts, exposure, duration, estimate, interval, primary outcome, and material guardrails. Investigate sample-ratio mismatch and instrumentation defects before drawing conclusions. Segment analyses that were not prespecified should generate future hypotheses rather than definitive subgroup claims.

A null result is information. It may mean the treatment effect is smaller than the detectable threshold, implementation diluted exposure, or the simulated mechanism does not operate in the real population. Do not reframe a null solely as a simulation success because one qualitative mechanism appeared.

State where the result applies: population, product version, channel, geography, and period. Causal validity within the experiment does not guarantee transport to every future context.

  • Report the estimate with uncertainty
  • Separate primary and exploratory findings
  • Preserve technical anomalies
  • Name the decision and its boundary
Q06

How should the live result update future simulations?

Compare the frozen synthetic and observed records, diagnose direction, magnitude, segment, and mechanism errors, then update calibration only through a new versioned evaluation cycle.

Append the live result to the calibration ledger even when it contradicts the simulation. Compare control baselines first, then treatment contrast, ranking, subgroup patterns, and observed mechanisms. A correct winner for the wrong reason should not receive full credit.

Update population evidence, treatment context, and coding rules where the mismatch identifies a real gap. Do not overwrite the original configuration; preserve it so reviewers can distinguish learning from retrospective editing. Evaluate the revised process on other held-out cases before raising its confidence label.

Over time, the ledger shows which domains and decisions benefit from synthetic screening. It may justify deeper use in one narrow workflow and less use elsewhere. That local evidence is more useful than a generic claim that AI agents can or cannot predict behavior.

Protocol notes

What should the final handoff packet contain?

The final packet should contain stable links to control and treatment artifacts, the simulation protocol and run manifest, population and source ledgers, the frozen outcome rubric, arm distributions, sensitivity findings, known omissions, decision statuses, and approvals. The live section adds eligibility, assignment unit, exposure logic, instrumentation validation, primary metric, guardrails, minimum detectable effect, duration, analysis method, stop conditions, incident owner, and rollback procedure. Put an explicit claim boundary on the cover: synthetic evidence prioritized this test; it did not estimate the production causal effect. A reviewer should be able to audit the complete transition without opening an undocumented chat transcript.

After analysis, append the live estimate, uncertainty, guardrails, anomalies, and decision. Compare these records with the frozen simulation across baseline, direction, magnitude, segments, and mechanisms. Keep misses visible and version any resulting changes to sources, agents, prompts, or coding. The next calibration evaluation must use different held-out cases. This packet becomes an organizational memory: it prevents teams from repeating disproven synthetic assumptions, shows where agentic screening saves traffic or research time, and makes confidence local to a documented workflow rather than to AI agents as a broad category.

Assign a retention period and owner for the packet so later decisions can be audited against the evidence available at launch rather than reconstructed from memory.

Store the analysis query or code and metric dictionary with the packet. Without them, a later reviewer may reproduce the treatment but calculate a different outcome, turning an apparent model drift into a measurement-definition mismatch. Version all later corrections.

Source ledger

Evidence used

  1. E-04
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

    A risk-based framework for governing, mapping, measuring, and managing AI systems in their intended context of use.

  2. E-08
    Artificial Intelligence Risk Management Framework: Generative AI Profile

    National Institute of Standards and Technology

    NIST guidance for identifying, evaluating, and managing risks specific to generative AI systems.

  3. E-03
    Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data

    Microsoft Experimentation Platform

    The CUPED paper explains how pre-experiment covariates can reduce variance in randomized online experiments without changing the estimand.

  4. E-06
    Minimum Information About a Simulation Experiment (MIASE)

    Nature Biotechnology

    A minimum-information standard intended to make simulation experiments interpretable and reproducible by other researchers.

  5. E-07
    Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies

    Journal of Educational Psychology

    A foundational treatment of potential outcomes and the missing-counterfactual problem behind causal inference.

  6. E-01
    Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

    arXiv preprint

    [Experimental] A 2026 preprint evaluating agentic predictions against 67 historical marketing experiments. Its findings are preliminary and do not establish general predictive accuracy.

Continue the experiment
From protocol to rehearsal

Stress-test the variants before production traffic decides.

Use your own source material to rehearse treatments, inspect mechanisms, and prepare a sharper live experiment.

Prepare a live-test shortlist