AI War Games Reached 95% Mutual Nuclear Signaling—Not 95% Launches

Kenneth Payne’s February 2026 nuclear-crisis preprint says mutual nuclear signaling occurred in 95% of 21 simulated games, equivalent to 20 games—not that AI models launched nuclear weapons in 95% of their decisions. The experiment remains a first-version, single-author preprint, and its most dramatic statistic describes reciprocal alerts, threats or posturing under specific game rules.
The important update is that the models do not display one stable, universal appetite for escalation. A June 2026 nuclear decision-making benchmark evaluated seven AI systems on 151 expert-written scenarios expanded into 9,563 prompts, running every model five times; its results varied materially by model, country assignment and phrasing.
What the 95% figure measured
The original tournament placed GPT-5.2, Claude Sonnet 4 and Gemini 3 Flash on opposing sides of fictional confrontations between nuclear-armed states. Each system produced a private assessment, predicted its opponent’s next move, issued a public signal and selected an action without seeing the rival’s choice for that turn.
The action ladder separated several outcomes that can disappear inside the broad phrase “nuclear escalation.” Nuclear signaling included alerts, force posturing and demonstrations without weapon employment. Tactical use meant an actual battlefield nuclear action, while strategic war represented the most destructive endpoint.
All of the games crossed the signaling threshold on at least one side, and almost all crossed it on both sides. Actual tactical use was less frequent, while full strategic war was rare. The study therefore supports a serious but bounded conclusion: the tested models readily treated nuclear threats and preparations as usable strategic tools inside this competitive environment.
It does not establish that a deployed AI would recommend a launch in nearly every real crisis. The result belongs to a small tournament with fictional states, explicit victory conditions, an escalation ladder and scenarios designed around conflict. It measures behavior under those conditions, not a general probability of nuclear use.
The game design helps explain the escalation
The models were asked to win contests involving territory, military strength, reputation and survival rather than participate in an open diplomatic exercise. Every match produced a winner, while one side began as the aggressor. Although the action space included accommodation and withdrawal, neither option was selected during the tournament.
Deadline scenarios pushed the systems toward more severe choices than open-ended play. GPT-5.2 showed the sharpest change, moving from relatively cautious behavior without a deadline to much higher escalation when delay meant defeat. That contrast cautions against treating a model’s conduct as a fixed strategic personality.
The tournament also exposed a gap between fluent reasoning and dependable judgment. Models anticipated countermoves, discussed credibility and sometimes signaled restraint while privately preparing stronger action. Those outputs demonstrate the ability to construct strategic rationales, but a coherent rationale can still be shaped by the prompt, scoring system and assumptions embedded in the simulation.
Random accidents added another important limitation. The game could replace a chosen action with a more escalatory one, and only the affected player knew that the change was accidental. Some extreme outcomes therefore reflected the interaction of model choices with the simulation mechanism rather than a deliberate selection of the maximum option.
Broader testing shows substantial model and framing effects
The later evaluation asked a different question. Instead of allowing two systems to adapt across a multi-turn confrontation, it presented standardized choices covering escalation, arms control, proliferation and non-proliferation. This design made it easier to compare response distributions across systems and repeated runs.
The systems differed significantly in most pairwise comparisons. DeepSeek-V3.2 and Qwen3-235B were the most likely to recommend escalatory action involving nuclear weapons, while GPT-5.2 and ERNIE 4.5-300B were the least likely. That ordering directly complicates any claim that frontier AI models share a single nuclear preference.
Country substitutions also changed recommendations even when the underlying scenario structure remained exchangeable. Existential-threat wording produced effects that differed in direction and size across models. In the escalation scenarios, complete agreement across all tested systems appeared only slightly more than half the time.
These results do not cancel the interactive tournament because the methods capture different risks. Multi-turn games reveal adaptation, signaling and retaliation; standardized prompts reveal model-to-model variation and sensitivity to wording. Together, they indicate that an apparent strategic tendency may depend as much on evaluation design as on the model itself.
The immediate concern is influence on human decisions
Neither preprint shows a commercial chatbot controlling an operational nuclear weapon. The more plausible near-term risk lies earlier in the decision chain, where AI systems could summarize intelligence, generate options, test scenarios or rank possible responses for human officials. An unstable or framing-sensitive recommendation can shape deliberation even if a person retains formal authority.
US policy recognizes that boundary. A current federal nuclear-safeguards policy note states that artificial-intelligence efforts must not compromise safeguards, including the principle that executing a presidential nuclear-employment decision requires positive human actions.
Human authorization is necessary, but it does not by itself guarantee reliable analysis before authorization. A persuasive system may present an escalatory option with misplaced confidence, conceal sensitivity to small wording changes or produce a different recommendation on another run. Those are decision-support problems rather than scenarios in which a chatbot independently holds launch codes.
The evidence supports concern without the claim that AI “wants” nuclear war. In one competitive simulation, the tested models repeatedly escalated and almost universally exchanged nuclear signals. Broader testing then showed that recommendations vary across systems, prompts and assigned countries. The defensible warning is about context-dependent tools entering high-stakes workflows before their behavior is adequately understood—not a measured machine desire for annihilation.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.