MPredict anythingMIROFISH 米罗鱼
Validation field noteRobustness Testing

Sensitivity Analysis for AI Scenario Simulations

Aug 5, 202611 min readMiroFish Editorial
Quick answer

Quick answer

Sensitivity analysis tests whether an AI scenario conclusion survives reasonable changes to its inputs and configuration. Vary high-impact assumptions, actor weights, evidence interpretations, model settings, and random seeds; then compare outcome distributions and decision reversals. The goal is not to find one perfect forecast. It is to distinguish stable conclusions from results that depend on a narrow setup.

Decision artifact

Scenario sensitivity matrix

Prioritize variants that are both uncertain and capable of changing the action.

Input familyVariant to testSignal to compare
EvidenceUse the strongest credible competing interpretationDoes the recommended action reverse?
ActorsAdd, remove, or reweight an affected groupWhich narrative or coalition changes?
MechanismDelay the trigger or weaken an incentiveDoes the same causal path still appear?
ExecutionChange model, sampling setting, or seedHow much output variance is technical?
Q01

What is sensitivity analysis for an AI simulation?

Sensitivity analysis measures how changes in uncertain inputs affect simulation outputs and the decision derived from them.

Traditional sensitivity methods attribute output variance to model inputs. LLM-agent systems add new sources of variation: natural-language interpretations, generated plans, interaction order, tool output, model version, and sampling. The same principle still applies—change inputs deliberately and observe which conclusions move.

Begin with the decision metric, not the most visible text. A launch simulation might track the prevalence of a trust objection, time to escalation, or effectiveness of a mitigation. Without a defined outcome, teams often compare wording differences while missing a decision reversal.

Q02

Which assumptions should be varied first?

Vary assumptions that combine high uncertainty with high decision influence: contested evidence, pivotal actors, strong incentives, timing, and technical configuration.

List the assumptions behind the conclusion and score each for uncertainty and leverage. A well-established product price may be low uncertainty but high leverage; a speculative audience weight may be high on both. Start with the assumptions in the high-high quadrant.

Include omissions. Adding an overlooked regulator, distributor, minority audience, or internal veto holder can matter more than changing an existing persona. The purpose is to test the structure of the scenario, not merely tune numeric settings.

  • Evidence interpretations that credible reviewers dispute.
  • Actors that can amplify, block, or reframe the pressure event.
  • Timing and information-access assumptions.
  • Model version, prompt version, sampling configuration, and seed.
Q03

How should a sensitivity test be designed?

Use a staged design: establish a repeated-run baseline, vary one input at a time, then test combinations among the inputs that materially move the result.

One-at-a-time variants are easy to explain and useful for screening. They can miss interactions, so follow with combined variants for influential inputs. For example, a skeptical audience may become decisive only when a trusted intermediary also delays its response.

Keep everything else fixed and versioned. Use the same output coding scheme across variants. If the model or prompt changes during the test, treat that as a separate technical variant rather than silently mixing runs.

Q04

How should teams compare sensitivity-test outcomes?

Compare decision direction, outcome frequency, timing, mechanism, and tail failures—not just the average sentiment or final narrative.

A result can be directionally stable while the path varies. If most variants produce the same objection through different actors, preparation may still be justified. Conversely, a similar average score can hide distinct failure modes that demand different mitigations.

Use tables or plots that show distributions by variant. Report how many runs support each conclusion and mark configurations that reverse the action. Avoid precise probability language unless the design and calibration support it; generated frequencies are properties of the simulation setup, not direct estimates of population prevalence.

Q05

What makes an AI scenario conclusion fragile?

A conclusion is fragile when small plausible changes reverse the decision, remove the mechanism, or shift the outcome beyond the team's tolerance.

Fragility is useful information. It can identify a fact to research, a stakeholder to interview, a mitigation to test, or a decision that should remain reversible. Do not hide a fragile result behind an average across incompatible scenarios.

Distinguish model fragility from world fragility. If changing only the language model reverses the result, the technical layer needs investigation. If competing but credible source assumptions reverse it across models, the real-world uncertainty may be the dominant issue.

Q06

How should sensitivity findings change a decision?

Stable findings can support preparation; fragile findings should trigger research, monitoring, staged action, or explicit contingency plans.

Tie each influential variable to an action. Monitor it if it will become observable, research it if evidence can reduce uncertainty, or build a contingency if neither is possible. A sensitivity report that does not change the decision process is only technical documentation.

Preserve the matrix with the decision record. When real signals arrive, compare them with the tested variants and update the scenario. This makes simulation a learning loop instead of a one-time forecast artifact.

Source ledger

Evidence used

  1. S-01
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

    A risk-management framework organized around governing, mapping, measuring, and managing AI risks.

  2. S-02
    Artificial Intelligence Risk Management Framework: Generative AI Profile

    National Institute of Standards and Technology

    A companion profile covering risks and evaluation considerations specific to generative AI systems.

  3. S-07
    Variance based sensitivity analysis of model output

    Computer Physics Communications

    A technical reference for attributing output variance to uncertain model inputs and their interactions.

  4. S-06
    A standard protocol for describing individual-based and agent-based models

    Ecological Modelling

    The ODD protocol for documenting a model's overview, design concepts, and implementation details.

Continue the review
Apply the method

Run a scenario you can inspect, challenge, and rerun.

Start with your own source material, then use this validation path to review the scenario before it informs a decision.

Compare scenario variants