How to Design an AI-Simulated Experiment
- 01StageEstimand
- 02Control ATreatment B
- 03StageAgent frame
- 04StageRun manifest
- 05DecisionDecision gate
Direct answer
Design an AI-simulated experiment by defining the decision and estimand, writing control and treatment conditions, grounding the target population, choosing an exposure design, and prespecifying outcomes before running. Version every input, repeat the baseline and variants, test sensitivity, and document a live-test gate. The output should be a shortlist and risk map—not an unsupported causal or accuracy claim.
AI-simulated experiment canvas
Complete each row before launching runs. Empty cells are unresolved design choices, not details to infer after seeing results.
| Protocol field | Control A | Treatment B | Prespecified rule |
|---|---|---|---|
| Experience | Current message or flow | Single intended change | No hidden treatment differences |
| Population | Same eligibility frame | Same eligibility frame | Declare exclusions and segments |
| Outcome | Coded response distribution | Coded response distribution | Freeze rubric before runs |
| Gate | Retain baseline | Advance, revise, or drop | Live validation for material action |
How should the decision and estimand be defined?
State whose outcome would change, which two experiences are compared, over what horizon, and which action follows from the contrast.
Begin with a real decision: which two onboarding messages should enter a live test, which response plan needs preparation, or which concept requires another interview round. An estimand turns that decision into a comparison. It specifies the population, treatment, comparator, outcome, and time window.
Avoid broad goals such as 'predict customer behavior.' A workable specification is: among first-time self-serve buyers represented by this evidence packet, how does a proof-led landing-page message change coded trust objections relative to the current page during an initial evaluation conversation? This is still a simulated contrast, but it is inspectable.
Write what the result may change and what it may not. If the simulation can drop a clearly confusing treatment but cannot set a revenue forecast, record that boundary before results appear.
How should control and treatment conditions be constructed?
Keep the control faithful to the current experience, change only the intended treatment dimension, and give both arms equivalent context, format, and opportunity to respond.
A treatment bundle makes interpretation difficult. If the new variant changes headline, price framing, social proof, and call to action, a difference cannot be attributed to one element. Bundles can still be compared as product packages, but the claim must concern the package rather than a component.
Preserve fidelity. Summarizing the control in two sentences while giving the treatment a polished full brief biases exposure. Include the same channel constraints, timing, surrounding content, and interaction rules. Remove labels such as 'new' or 'improved' unless users would see them.
Archive exact treatment files with stable IDs. The report should point to A-01 and B-01 rather than relying on a later copy-and-paste reconstruction.
How should the synthetic population be grounded?
Use current evidence to define relevant traits, incentives, constraints, and segment proportions, while labeling modeled assumptions and underrepresented groups.
Useful sources include customer interviews, support themes, survey data, behavioral analytics, product research, market constraints, and policy requirements. Map each consequential agent attribute to evidence or label it as a modeling choice. Do not fill every gap with a confident demographic stereotype.
Deep interview-grounded research shows that richer context can improve behavioral representation on held-out measures, but it does not guarantee population validity. Keep a separate evaluation set where possible and avoid claiming that a small evidence packet represents an entire market.
Create an omission review. Ask which users experience the decision differently, who is absent from available data, and which stakeholder can block or amplify the outcome. Test plausible population variants when those gaps could reverse the shortlist.
Which outcomes and coding rules should be prespecified?
Choose outcomes connected to the decision, define observable coding rules, separate primary and exploratory measures, and freeze them before inspecting treatment results.
Agent output is unstructured, so the measurement layer needs discipline. A primary outcome might be the presence of a specified trust objection, completion of an intended action, time to escalation, or selection among defined options. Sentiment alone is rarely sufficient because positive language may coexist with refusal.
Write examples and counterexamples for every code. Use human review, structured model coding, or both, and audit agreement on a sample. Keep coders blind to treatment where possible. Record missing and ambiguous outputs rather than forcing every run into a clean category.
Exploratory mechanisms can generate hypotheses, but they should not replace the primary outcome after the favored treatment underperforms. This preregistration-lite discipline reduces narrative cherry-picking.
How should the experiment be run and repeated?
Version the complete run manifest, establish baseline variability, repeat both arms under matched configurations, and test the assumptions most capable of reversing the decision.
The manifest should contain source IDs, treatment IDs, agent definitions, model and prompt versions, tool versions, sampling settings, timestamps, seeds where supported, exposure order, assignment logic, and coding rules. MIASE provides a useful minimum-information principle even though LLM-agent systems require additional generative details.
Repeat the control before interpreting treatment differences. Then pair or balance runs so configuration changes do not coincide with arm changes. Report distributions, not one representative transcript. Small run sets should use descriptive language rather than probability claims.
Add sensitivity variants for contested evidence, population mix, model version, treatment order, and influential assumptions. A recommendation that reverses under a minor plausible change should advance only as a hypothesis requiring stronger live evidence.
- Freeze the outcome rubric
- Version evidence and treatments
- Repeat assignments and worlds
- Preserve null and failed runs
What belongs in the review and live-test gate?
The gate should summarize the shortlist, evidence, uncertainty, sensitivity, prohibited claims, live protocol, and named owner responsible for the next decision.
End each treatment in one of four states: advance, revise, drop, or unresolved. Give a reason tied to the prespecified outcome and risk register. A polished narrative is supporting material, not the decision rule. Preserve minority and tail failure modes that an average could hide.
State which real observation would invalidate the simulation. For an advancing treatment, define the live population, assignment unit, primary metric, guardrails, analysis window, and stopping rule. Do not use simulated effect magnitude as a production power assumption unless it has demonstrated local calibration.
Assign a human owner and review date. After real evidence arrives, compare it with the frozen simulation and update the calibration ledger. This closes the experiment loop instead of treating simulation as a one-time prediction artifact.
What should be included in a preregistration-lite record?
Before the first treatment run, save a short record containing the decision, estimand, population boundary, source cutoff, control and treatment IDs, exposure design, primary outcome, guardrails, coding rubric, expected run count, sensitivity variants, exclusion rules, and decision threshold. Add the model and prompt manifest plus the named owner. The record does not need academic infrastructure; a dated, immutable repository artifact is enough. Its purpose is to distinguish what the team intended to test from explanations invented after reading generated output. Exploratory observations remain welcome, but they must be labeled and routed into a later protocol.
Define completion as a decision artifact rather than a number of attractive transcripts. The package should include arm distributions, repeated-run count, ambiguous and failed runs, population coverage gaps, sensitivity reversals, minority failure modes, source ledger, and an advance/revise/drop/unresolved status for every candidate. If a treatment advances, attach the live-test question and list what simulation cannot establish. This minimum record supports review, reruns, and later calibration without making the process unnecessarily bureaucratic. It also exposes when the experiment cannot answer its stated question before additional runs consume time and model cost.
Preserve reviewer comments and protocol deviations beside the record so later teams can distinguish a deliberate choice from an accidental omission.
Give each deviation a date, owner, rationale, and expected analytical consequence for future review.
Evidence used
- E-02Designing Reliable Experiments with Generative Agent-Based Modeling
arXiv preprint
A research framework for experimental design with generative agents, including measurement, repeated trials, and design-validity concerns.
- E-06Minimum Information About a Simulation Experiment (MIASE)
Nature Biotechnology
A minimum-information standard intended to make simulation experiments interpretable and reproducible by other researchers.
- E-07Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
Journal of Educational Psychology
A foundational treatment of potential outcomes and the missing-counterfactual problem behind causal inference.
- E-05Generative Agent Simulations of 1,000 People
Stanford University research team
A study of interview-grounded agents evaluated against held-out behavioral measures, illustrating both the value and limits of rich persona evidence.
- E-04Artificial Intelligence Risk Management Framework (AI RMF 1.0)
National Institute of Standards and Technology
A risk-based framework for governing, mapping, measuring, and managing AI systems in their intended context of use.
- E-08Artificial Intelligence Risk Management Framework: Generative AI Profile
National Institute of Standards and Technology
NIST guidance for identifying, evaluating, and managing risks specific to generative AI systems.
Related experiment protocols
Stress-test the variants before production traffic decides.
Use your own source material to rehearse treatments, inspect mechanisms, and prepare a sharper live experiment.