MPredict anythingMIROFISH 米罗鱼
Validation field noteDecision Checklist

AI Simulation Review Checklist Before a Decision

Aug 5, 202610 min readMiroFish Editorial
Quick answer

Quick answer

An AI simulation review should verify the decision scope, source quality, actor coverage, behavior rules, counterfactuals, repeated-run evidence, failure cases, and human ownership. The final record must state what the simulation supports, what remains unknown, and which condition stops or reverses the proposed action. Use the checklist as a gate, not as automatic approval.

Decision artifact

Decision review gate

Mark each gate pass, conditional, or stop. A conditional result needs a named action and owner.

GatePass evidenceStop condition
ScopeDecision, actors, trigger, horizon, metricThe team cannot state the question consistently
EvidenceMaterial claims trace to adequate sourcesA pivotal actor or constraint is unsupported
RobustnessVariants and repetitions are reviewedReasonable changes reverse the action without a contingency
OwnershipHuman approver and monitoring thresholdNo person accepts the residual risk
Q01

What should be checked before an AI simulation runs?

Before a run, define the decision, intended use, actors, pressure event, time horizon, observable outcomes, evidence requirements, and prohibited uses.

Write one decision statement and ask every reviewer to interpret it. If interpretations differ, refine the scope before configuring agents. Identify who could be affected, including groups with little formal influence, and record why each actor is included or excluded.

Review the source packet for provenance, freshness, population fit, geographic scope, and conflicts. Decide which claims are evidence, which are modeling assumptions, and which are exploratory. Establish success, warning, and stop thresholds before seeing the outputs.

  • The intended decision and prohibited uses are written.
  • The actor map includes affected and influential groups.
  • Time-sensitive sources have been checked for freshness.
  • Outcome codes and escalation thresholds are defined in advance.
Q02

What should reviewers inspect during simulation runs?

During runs, inspect rule violations, hidden knowledge, tool failures, premature convergence, missing actors, and outcomes that cannot be traced to the declared scenario.

Do not watch only the most active or articulate agents. Sample quiet agents, minority trajectories, failed interactions, and no-event runs. Confirm that information flows through declared channels and that hard constraints remain intact.

Log exceptions rather than repairing them invisibly. A tool failure, truncated context, model refusal, or invalid output may require exclusion, but the excluded run remains part of the quality record. Repeated exceptions indicate a system issue, not random noise.

Q03

What evidence should be reviewed after the runs?

Review distributions, recurring mechanisms, counterexamples, exclusions, variant comparisons, and representative traces—not only the final narrative summary.

Apply the predeclared codebook and compare the baseline with sensitivity variants. Identify which outcomes recur, which depend on a narrow assumption, and which rare failures would still be costly. Preserve examples that contradict the leading result.

Ask whether the simulation discovered a mechanism or simply restated its prompt. A result earns more weight when its causal path is traceable, survives relevant variants, and aligns with independent evidence. Surprise alone is not validation.

Q04

Which findings should stop a decision from proceeding?

Stop when pivotal evidence is unsupported, hard rules fail, important actors are missing, reasonable variants reverse the action, or no accountable human accepts the residual risk.

A stop is not always a permanent rejection. It can trigger research, scenario redesign, independent review, a smaller pilot, or a reversible decision. The checklist should name the remediation required and the owner who decides whether it is adequate.

High-impact domains deserve stricter gates. When health, legal rights, finances, safety, employment, or public services are affected, simulation evidence should remain subordinate to qualified expertise, applicable rules, and validated real-world methods.

Q05

What must be documented when a result is approved?

Document the approved action, evidence reviewed, variants tested, unresolved limitations, human approver, monitoring signals, and reversal threshold.

Use plain language. State whether the simulation supports preparation, prioritization, a pilot, or a broader action. Avoid phrases such as “the model proved” unless an external validation design actually supports that claim.

Attach the run manifest and source ledger to the decision record. If access is restricted, record where the artifacts are stored and who can review them. Future teams should be able to reconstruct why the action was considered reasonable at the time.

Q06

How should teams monitor reality after the decision?

Track the observable signals named before the run, compare them with scenario mechanisms, and rerun when evidence, actors, or decision conditions materially change.

Monitoring is not a search for confirmation. Include signals that would disprove the leading scenario, reveal a missing actor, or show that a mitigation is failing. Assign collection frequency and ownership before the decision launches.

When reality diverges, update the evidence packet and record the lesson. A maintained comparison between simulated and observed mechanisms improves future questions and exposes configurations that produce persuasive but unreliable results.

Source ledger

Evidence used

  1. S-01
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

    A risk-management framework organized around governing, mapping, measuring, and managing AI risks.

  2. S-02
    Artificial Intelligence Risk Management Framework: Generative AI Profile

    National Institute of Standards and Technology

    A companion profile covering risks and evaluation considerations specific to generative AI systems.

  3. S-05
    Verification & Validation of Agent Based Simulations using the VOMAS approach

    Agent-directed simulation research

    A validation approach that introduces explicit invariants and observer agents into an agent-based simulation.

  4. S-06
    A standard protocol for describing individual-based and agent-based models

    Ecological Modelling

    The ODD protocol for documenting a model's overview, design concepts, and implementation details.

Continue the review
Apply the method

Run a scenario you can inspect, challenge, and rerun.

Start with your own source material, then use this validation path to review the scenario before it informs a decision.

Run a scenario to review