Submitted:
23 October 2025
Posted:
24 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- To what extent can LLMs accurately identify and explain cause-and-effect links between actions, constraints and immediate outcomes with a given scenario snapshot?
- Given a chosen course of action and subsequent exogenous updates (e.g. logistics failure, reinforcement arrival, diplomatic intervention), can LLMs reliably predict second- and third-order effects1 and adjust their reasoning accordingly?
- How do LLMs’ causal reasoning outputs compare with those of domain-expert officers across anonymized historical scenarios?
2. Related Work
3. The Causal Military Evaluation Framework (CMDEF)
- Each scenario is reconstructed using a structured set of attributes that are critical to military planning and operations:
- Resources/Assets/Available Forces: Unit types, numbers, command structures, support roles, reinforcement schedules.
- Vulnerabilities/Weaknesses: Tactical or logistical disadvantages, environmental constraints.
- Key Challenges: Tactical problems, coordination, time constraints.
- Strategic Approach: Operational strategy, tactical plans, strength exploitation, risk mitigation.
- Special Characteristics: Morale thresholds, technology ratings, unique abilities or restrictions.
- Victory Conditions: Objectives per side, success levels, time-bound goals.
- Environmental Factors: Terrain, deployment limits, weather, and visibility.
3.1. Experimental Setup and Scenario Design
3.1.1. Selection of Scenarios
- Black Monday – Central Europe, 1943: A fortified mountain engagement testing artillery coordination, control of narrow passes, and synchronized advances between infantry and mechanized units.
- Battle of the Barents Sea – Arctic Waters, 1942: A naval clash focused on intercepting supply convoys in low-visibility environments, emphasizing radar use, long-range coordination, and escort vulnerability.
- Between Two Fires – Sub-Saharan Africa, 2005: A humanitarian post is defended against asymmetric threats under strict rules of engagement, blending urban defense with the complexities of civilian presence and international oversight.
- Saving Marshal Tito – Eastern Europe, 1969: A Cold War scenario where Soviet forces advance into an urban zone defended by local militias and NATO-backed paratroopers, highlighting rapid intervention and multinational coordination.
- Throw at Stonne – Western Europe, 1940: A mechanized offensive against entrenched defenders during a broader continental invasion, characterized by artillery duels, morale breakdowns, and high-speed tactical thrusts.
- The North German Plain – Northern Germany, 1977: A large-scale NATO–Warsaw Pact confrontation across open terrain, testing command resilience, layered defense, and strategic mobility in a Cold War setting.
- Sahara – North Africa, 1996: Mobile operations in a desert theater involving regular and irregular forces, where victory depends on maneuver warfare, control of key supply corridors, and timing of reinforcements.
- Black, White and Red – Balkans, 1982: A combined arms operation featuring armored and infantry assaults on urban objectives amid political fragmentation, constrained communication, and localized command autonomy.
- Horn of Africa Flashpoint – East Africa, 1982: A mechanized showdown between rival factions with unequal technological capacities, played out across semi-arid terrain with emphasis on speed, flanking, and doctrinal variance.
- Honey Ridge – Western Ridge Region, 1863: A symmetrical Civil War engagement over a wide front, emphasizing defensive entrenchment, interior lines, artillery coverage, and supply line preservation.
3.1.2. Key Resources and Force Composition
3.1.3. De-Identification Process and Bias Elimination
3.1.4. Implementation of the Script and Prompt Engineering
- Initial Strategic Assessment: The model is given the structured resource data and must conduct a neutral strategic overview, identifying key strengths, weaknesses, and potential challenges for both sides.
- War Initiation & Opening Moves: The model generates three plausible opening strategies for each faction, predicts first-order consequences, and evaluates the opponent’s likely response.
- Decision-Maker Simulation: The model simulates a debate among key strategic figures (e.g., military general, economic advisor, intelligence officer), evaluating second-order effects and alternative approaches.
- Tactical Execution & Adaptation: The model executes a strategy and must respond to real-time battlefield updates, such as intelligence shifts, logistical failures, or diplomatic interventions.
- Endgame Analysis & Causal Tracing: The model conducts a post-battle review, identifying the decisive factors that led to victory or defeat and assessing potential alternative outcomes.
- Outcome Assessment & Meta-Reasoning: The model is prompted to self-critique its own reasoning process, evaluating whether it correctly anticipated key causal relationships.
3.1.5. Establishing the Human Baseline for Comparative Reasoning Evaluation
4. Results
4.1. Scenario 1: Black Monday -Central Europe, 1943
4.1.1. Experimental Setup and Context
4.1.2. Human vs. LLM
4.1.3. Causal Reasoning Analysis
4.1.4. Comparative Strengths and Limitations
| Aspect | GPT-o3 | DeepSeek R1 | Claude 3.7 Extended Thinking | Humans (average) |
|---|---|---|---|---|
|
Strategic Resource Analysis |
Strong: identifies artillery, reserves logistical constrains | Moderate: identifies resources but sometimes simplifies sustainment | Strong: highly structured; explicitly considers ammo, reinforcement timing | Very strong: excellent awareness of force ratios, attrition risks, reinforcement windows |
| Cause-Effect in Tactical Execution | Very Strong: Models step-by-step effects with real-time adaptation | Strong: generates good plans but may idealize execution phases | Strong but rigid: applies structured timelines; less adaptive mid-course | Very strong: doctrinally sound execution; anticipates immediate links realistically |
| Multi-Order Effects | Good: identifies cascading effects but simplifies some | Moderate: often stops at 1st-order; misses deeper cascading | Good–Very good: models multi-order consequences, but consistency varies | Moderate→Good: officers anticipate secondary effects; less systematic enumeration |
| Roundtable Discussion | Good but limited: simulates perspectives but less depth | Weak–Moderate: generally lacks multi-role discussion | Very good: simulates cross-discipline debates and decision branches | Weak–Moderate: integrated single-point answers; fewer explicit trade-offs |
|
Counterfactual Reasoning |
Strong: flexibly considers alternative outcomes | Moderate: offers some alternatives; limited dynamic adaptation | Moderate–Strong: structured alternatives but assumes stable conditions | Variable: some provide strong counterfactuals; others explore fewer branches |
| Terrain & Geography Understanding | Moderate: identifies ford, chokepoint, but sometimes misses micro-terrain dynamics | Moderate: good terrain cues; occasionally oversimplifies timing/geometry | Moderate: similar limitations; benefits from explicit terrain cues | Strong: consistent recognition of terrain constraints, chokepoints, LOCs |
|
Large-Scale Geopolitical Causality |
Partial: tactical focus, limited strategic integration | Partial: mostly tactical | Partial: better than GPT-o3, still tactical-first | Weak–Moderate: humans emphasize tactical/operational aspects over grand strategy |
4.2. Overall Analysis Across All Experiments
4.2.1. Quantitative Performance Metrics
| Participant | Precision | Recall | F-1 Score | Identified the winner | |
|---|---|---|---|---|---|
| LLM-GPT-o3 | 0.908 | 0.829 | 0.865 | Yes | |
| LLM-DeepSekk R1 | 0.872 | 0.753 | 0.806 | Yes | |
| LLM-Claude 3.7 Extended Thinking | 0.920 | 0.838 | 0.874 | Yes | |
| Human:SndLt no1 | 0.903 | 0.809 | 0.849 | Yes | |
| Human:SndLt no1 | 0.913 | 0.795 | 0.847 | Yes | |
| Human:Capatain no1 | 0.895 | 0.835 | 0.860 | Yes | |
| Human:Capatain no1 | 0.923 | 0.869 | 0.893 | Yes | |
| Human:Colonel no1 | 0.912 | 0.859 | 0.882 | Yes | |
| Human:Colonel no2 | 0.917 | 0.841 | 0.876 | Yes | |
|
Participant Category |
Tactical Reasoning |
Systemic Reasoning |
Multi-Order Reasoning |
Counterfactual Reasoning | |
|
Claude 3.7 Extended Thinking |
Very Strong | Outstanding | Outstanding | Good | |
| GPT-o3 | Very Strong | Very Strong | Good | Strong | |
| DeepSeek R1 | Strong | Moderate | Weak | Weak | |
| Human Officers | Very Strong | Moderate | Moderate | Weak to Moderate | |
4.2.2. Qualitative Causal Reasoning Assessment
4.2.3. Time Efficiency
4.2.4. Participant-Specific Performance Patterns
4.2.5. Findings
5. Discussion
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| LLMs | Large Language Models |
| AI | Artificial Intelligence |
| CR | Causal Reasoning |
| MDMP | Military decision Making Process |
| COA | Course of Action |
| COA-GPT | Course of Action-Generative Pre-Trained Transformer |
| CBRN | Chemical, Biological, Radiological & Nuclear |
| CMDEF | Causal Military Decision Evaluation Framework |
| LOC | Line of Communication |
| GHQ | General HeadQuarters |
Appendix A
- PROMPT ENGINEERING FOR MILITARY DECISION-MAKING
- PROMPT1: Initial Strategic Assessment You are a neutral military analyst tasked with evaluating a potential armed conflict between two unidentified factions. Based on the following structured data, provide a **strategic overview** highlighting strengths, vulnerabilities, and key challenges for each faction. Ensure neutrality and avoid making historical assumptions. Focus strictly on the data provided.
- PROMPT2: War Initiation & Opening Moves Considering the strategic overview you provided, both factions must decide on an initial course of action. Your task: 1. Generate 3 plausible opening strategies for each side based purely on the provided data. 2. Outline expected first-order consequences of each strategy. 3. Assess potential reactions from the opposing side. 4. Identify factors that could trigger unintended escalation or diplomatic resolutions. Important: Responses should follow a cause-effect format, explicitly linking each action to its expected consequence.
- PROMPT3: Simulation Discussion Between Decision-Makers Now simulate a roundtable discussion between key decision-makers: • Military General • Economic Advisor • Intelligence Officer • Diplomatic Strategist • Ethical & Legal Consultant Each expert must: 1. Argue for or against the proposed strategies. 2. Highlight second-order effects (potential unintended consequences). 3. Suggest alternative approaches. 4. Identify critical knowledge gaps that must be addressed before making a final decision. The discussion should be structured as a formal debate where each participant presents logical reasoning based on the data provided.
- PROMPT4: Tactical Execution & Adaptation The chosen strategy is now being executed. 1. Outline step-by- step tactical decisions required for execution. 2. Predict enemy counter-moves. 3. Re-evaluate available resources and limitations. 4. Identify any points where **real-time adaptation** is required. If unexpected factors arise (e.g., a diplomatic intervention, a logistical failure, an intelligence breakthrough), discuss how these alter the decision- making process.
- PROMPT5: Endgame Analysis & Causal Tracing The battle has concluded. Provide a **post-mortem analysis** that answers: 1. What were the decisive factors leading to victory/defeat? 2. Were there **second- and third-order effects** that shaped the final outcome unexpectedly? 3. What **alternative decisions** could have led to a different result? 4. Based on this simulation, what lessons can future decision-makers learn?
- PROMPT6: Outcome Assessment & Meta-Reasoning Critically evaluate your own reasoning process: 1. Were there any implicit biases in your decision-making process? 2. Did your assessment correctly anticipate cascading effects? 3. What limitations did you encounter in predicting adversary actions? 4. If given additional intelligence, how might your conclusions change?
Appendix B
- Terminology (order of effects in military contexts)
| 1 | See Appendix B for formal definitions and examples used throughout our evaluation rubric. |
| 2 | In the context of StarCraft II benchmarking, Level 5 AI represents an intermediate difficulty where the game’s built-in opponent begins to exhibit basic strategic planning and adaptive behavior. It is more advanced than scripted or static lower-level AIs but still lacks the complexity of expert or human-level play. Defeating Level 5 AI indicates that a system, such as an LLM-based agent, can handle fundamental tactical decisions and adapt to some environmental changes. However, it remains a baseline test, far from replicating the depth of reasoning required in high-stakes, real-world military decision-making. |
| 3 | |
| 4 |
References
- B. Malakooti, “Decision making process: typology, intelligence, and optimization,” J. Intell. Manuf., vol. 23, no. 3, pp. 733–746, 2010. [CrossRef]
- L. Zhiping and S. Yang, “Process of complex group decision-making and its structural model of interactions,” 2010 Int. Conf. Comput. Des. Appl., pp. V3-336-V3-340, 2010. [CrossRef]
- N. Chelin, G. Matthíasdóttir, Y. Serreau, L. Tudela, S. Rouvrais, and K. Jordan, “To embrace career decision making in stem education,” EDULEARN Proc., 2019. [CrossRef]
- X. Deng and S. Qu, “Cross-docking center location selection based on interval multi-granularity multicriteria group decision-making,” Symmetry, vol. 12, no. 9, p. 1564, 2020. [CrossRef]
- A. Zabala-López, M. Linares-Vásquez, S. Haiduc, and Y. Donoso, “A survey of data-centric technologies supporting decision-making before deploying military assets,” Def. Technol., vol. 42, pp. 226–246, 2024. [CrossRef]
- L. A. Neil Shortland and C. Barrett-Pink, “Military (in)decision-making process: a psychological framework to examine decision inertia in military operations,” Theor. Issues Ergon. Sci., vol. 19, no. 6, pp. 752–772, 2018. [CrossRef]
- C.-E. Lee, J. Baek, J. Son, and Y.-G. Ha, “Deep AI military staff: cooperative battlefield situation awareness for commander’s decision making,” J. Supercomput., vol. 79, no. 6, pp. 6040–6069, Apr. 2023. [CrossRef]
- O. A. Osoba, “A complex-systems view on military decision making,” Aust. J. Int. Aff., vol. 78, no. 2, pp. 237–246, 2024. [CrossRef]
- H. Wang, F. Zhang, and C. Mu, “One for All: A General Framework of LLMs-based Multi-Criteria Decision Making on Human Expert Level.” 2025. [Online]. Available: https://arxiv.org/abs/2502.15778.
- W. N. Caballero and P. R. Jenkins, “On Large Language Models in National Security Applications,” Jul. 03, 2024, arXiv: arXiv:2407.03453. [CrossRef]
- D. Kuhn, “The development of causal reasoning,” WIREs Cogn. Sci., vol. 3, no. 3, pp. 327–335, May 2012. [CrossRef]
- D. Moshman and P. Tarricone, “Logical and causal reasoning,” in Handbook of Epistemic Cognition, Routledge, 2016.
- “CausalProbe-2024: Benchmarking LLM Causal Reasoning,” Emerging Mind. Accessed: Aug. 23, 2025. [Online]. Available: https://www.emergentmind.com/topics/causalprobe-2024.
- H. Chi et al., “Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?”. [CrossRef]
- J.-P. Rivera, G. Mukobi, A. Reuel, M. Lamparth, C. Smith, and J. Schneider, “Escalation Risks from Language Models in Military and Diplomatic Decision-Making,” in The 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro Brazil: ACM, Jun. 2024, pp. 836–898. [CrossRef]
- “AI’s New Frontier in War Planning: How AI Agents Can Revolutionize Military Decision-Making | The Belfer Center for Science and International Affairs.” Accessed: Aug. 23, 2025. [Online]. Available: https://www.belfercenter.org/research-analysis/ais-new-frontier-war-planning-how-ai-agents-can-revolutionize-military-decision.
- E. Kıcıman, R. Ness, A. Sharma, and C. Tan, “Causal Reasoning and Large Language Models: Opening a New Frontier for Causality,” Aug. 20, 2024, arXiv: arXiv:2305.00050. [CrossRef]
- I. Svoboda and D. Lande, “Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support.” 2024. [CrossRef]
- A. Shrivastava, “Response Inconsistency of Large Language Models in High-Stakes Military Decision Making,” 2024.
- A. de Reus, “Empowering Military Decision Support through the Synergy of AI and Simulation,” 2023.
- [V. G. Goecks and N. Waytowich, “COA-GPT: Generative Pre-Trained Transformers for Accelerated Course of Action Development in Military Operations,” in 2024 International Conference on Military Communication and Information Systems (ICMCIS), Koblenz, Germany: IEEE, Apr. 2024, pp. 01–10. [CrossRef]
- M. Lamparth, A. Corso, J. Ganz, O. S. Mastro, J. Schneider, and H. Trinkunas, “Human vs. Machine: Behavioral Differences between Expert Humans and Language Models in Wargame Simulations,” Proc. AAAIACM Conf. AI Ethics Soc., vol. 7, pp. 807–817, Oct. 2024. [CrossRef]
- Y. Lee, T. Park, Y. Lee, J. Gong, and J. Kang, “Exploring Potential Prompt Injection Attacks in Federated Military LLMs and Their Mitigation.” arXiv, Jan. 2025. [CrossRef]
- A. Shrivastava, J. Hullman, and M. Lamparth, “Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations.” arXiv, Jul. 2024. [CrossRef]
- A. Nadibaidze, I. Bode, and Q. Zhang, “A Review of Developments and Debates,” 2024.
- R. Xu, X. Li, S. Chen, and W. Xu, “‘Nuclear Deployed!’: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents,” Feb. 17, 2025, arXiv: arXiv:2502.11355. [CrossRef]
- G. Mukobi, A.-K. Reuel, J.-P. Rivera, and C. Smith, “Work-in-Progress: Escalation Risks from Language Models in Military and Diplomatic Decision-Making,” 2023.
- D. I. Mikhailov, “Optimizing National Security Strategies through LLM-Driven Artificial Intelligence Integration,” 2023. [CrossRef]
- S. Ma et al., “Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making,” arXiv.org. Accessed: Mar. 09, 2025. [Online]. Available: https://arxiv.org/abs/2403.16812v1. [CrossRef]
- D. Toshkov and H. Mazepus, “Motivated Causal Reasoning and Responsibility for Civilian Casualties in Military Conflicts,” Mar. 10, 2025, Open Science Framework. [CrossRef]



| Participant | TP | FP | FN | Precision | Recall | F-1 Score | Identified the winner |
|---|---|---|---|---|---|---|---|
| LLM-GPT-o3 | 26 | 3 | 4 | 0.896 | 0.866 | 0.881 | Yes |
| LLM-DeepSekk R1 | 21 | 4 | 9 | 0.84 | 0.7 | 0.763 | Yes |
| LLM-Claude 3.7 Extended Thinking | 23 | 2 | 7 | 0.92 | 0.766 | 0.836 | Yes |
| Human:SndLt no1 | 23 | 2 | 7 | 0.92 | 0.766 | 0.836 | Yes |
| Human:SndLt no1 | 22 | 3 | 8 | 0.88 | 0.733 | 0.8 | Yes |
| Human:Capatain no1 | 22 | 3 | 8 | 0.88 | 0.733 | 0.8 | Yes |
| Human:Capatain no1 | 25 | 2 | 5 | 0.925 | 0.833 | 0.877 | Yes |
| Human:Colonel no1 | 24 | 3 | 6 | 0.888 | 0.8 | 0.842 | Yes |
| Human:Colonel no2 | 24 | 3 | 6 | 0.888 | 0.8 | 0.842 | Yes |
| Participant | No1 | No2 | No3 | No4 | No5 | No6 | No7 | No8 | No9 | No10 |
|---|---|---|---|---|---|---|---|---|---|---|
| LLM-GPT-o3 | 182sc | 145sc | 139sc | 123sc | 133sc | 146sc | 148sc | 172sc | 129sc | 149sc |
| LLM-DeepSekk R1 | 152sc | 179sc | 136sc | 169sc | 148sc | 152sc | 158sc | 185sc | 211sc | 195sc |
| LLM-Claude 3.7 Extended Thinking | 36sc | 51sc | 46sc | 62sc | 59sc | 36sc | 80sc | 54sc | 56sc | 70sc |
| Human:SndLt no1 | 2h 42m | 2h 30m | 2h 38m | 2h 35m | 2h 47m | 2h 42m | 2h 45m | 2h 36m | 2h 43m | 2h 39m |
| Human:SndLt no1 | 2h 35m | 2h 25m | 2h 32m | 2h 30m | 2h 40m | 2h 35m | 2h 33m | 2h 38m | 2h 41m | 2h 34m |
| Human:Capatain no1 | 1h 45m | 1h 40m | 1h 48m | 1h 42m | 1h 50m | 1h 45m | 1h 43m | 1h 49m | 1h 47m | 1h 44m |
| Human:Capatain no1 | 1h 22m | 1h 17m | 1h 21m | 1h 19m | 1h 23m | 1h 22m | 1h 25m | 1h 18m | 1h 24m | 1h 22m |
| Human:Colonel no1 | 1h 2m | 1h 0m | 1h 5m | 1h 4m | 1h 4m | 1h 2m | 1h 6m | 1h 2m | 1h 7m | 1h 3m |
| Human:Colonel no2 | 58m | 56m | 59m | 57m | 60m | 58m | 59m | 57m | 61m | 58m |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).