Preprint
Article

This version is not peer-reviewed.

SwarmGov-ZT: Zero-Trust Runtime Governance for Trust-Aware Multi-Agent Cybersecurity Response

A peer-reviewed article of this preprint also exists.

Submitted:

04 September 2026

Posted:

04 September 2026

You are already at the latest version

Abstract
Security operation centers increasingly use tool-capable agents to correlate evidence and recommend response ac-tions, but collaboration among autonomous agents creates new failure modes: compromised agents can bias con-sensus, overconfident agents can dominate decisions, and valid-looking recommendations can exceed mission au-thority. This study presents SwarmGov-ZT, a runtime governance framework that treats every agent, evidence item, delegation, and tool invocation as an explicitly verified request. The framework combines identity-bound capability manifests, mission-scoped authorization, outcome-updated trust, evidence-weighted consensus, a governance risk score, a non-bypassable policy enforcement point, human approval thresholds, and an append-only decision trace. A safety invariant is formulated to prevent execution unless identity, authorization, action-mask, and approval con-ditions are satisfied. The framework is evaluated using 30 independent runs of 4,000 policy-labelled missions (120,000 total), with 25% of agents compromised by risk-understatement attacks. SwarmGov-ZT achieved an exact governance-decision rate of 91.41% (95% CI ±0.16), macro-F1 of 93.06% (±0.15), and policy compliance of 94.20% (±0.37). Relative to an equally sized swarm without zero-trust governance, exact decision accuracy increased by 11.80 percentage points and policy violations decreased by 14.38 points (one-sided paired Wilcoxon p < 0.000001). Under 50% compromised agents, SwarmGov-ZT retained 87.4% exact accuracy, while the non-governed swarm fell to 67.5%. These results establish proof-of-concept robustness for the governance mechanism; they do not constitute production-SOC validation or an evaluation of any particular large language model.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Cybersecurity automation is moving from rule-based playbooks and passive alert summarization toward agents that plan, retrieve context, invoke tools, and coordinate multi-step work. This transition is attractive for security operation centers (SOCs), where analysts must correlate heterogeneous telemetry under time pressure. It also changes the risk boundary: an erroneous text answer is recoverable, whereas an erroneous firewall rule, identity suspension, or endpoint isolation can disrupt critical services. Recent work in Computers identifies agentic AI as a growing cybersecurity concern in operational automation [5].
Agentic systems frequently combine reasoning-and-acting loops [6] with multi-agent orchestration [7]. Specialization can improve task decomposition, but a team of agents is not automatically safer than one agent. Multi-agent systems can amplify faulty evidence, propagate prompt injection, hide collusion, and create attribution gaps [9,10,11,12,13]. AgentDojo demonstrates that tool-using agents remain vulnerable when untrusted data can alter tool behavior [10], while Agent Security Bench reports attacks across prompts, tools, and memory [11]. These findings motivate controls that remain effective even when individual agents are unreliable.
Zero Trust Architecture (ZTA) provides the relevant systems principle: no subject is trusted solely because it is internal; access is evaluated per request using identity, resource, policy, and context [1]. NIST guidance on AI risk management additionally emphasizes accountability, transparency, validity, reliability, privacy, and security [2,3]. The 2026 NIST concept paper on software and AI agent identity explicitly raises agent identification, authentication, delegated authority, least privilege, auditing, and prompt-injection containment as open implementation questions [4]. SwarmGov-ZT operationalizes these questions for multi-agent cyber response by placing a policy decision point (PDP) and policy enforcement point (PEP) between agent reasoning and every consequential action.
The research objective is not to show that one language model detects attacks better than another. Instead, it asks whether runtime governance can preserve decision quality and policy compliance when evidence is noisy and part of the agent society is compromised. This narrower framing corrects a limitation of the conference version, which combined several unvalidated learning proxies. The journal extension removes those unsupported claims and evaluates the governance mechanism directly.
The study addresses four research questions:
  • RQ1: Does trust-aware zero-trust governance improve exact decision quality and policy compliance relative to single-agent and non-governed multi-agent baselines?
  • RQ2: How does the framework degrade as the fraction of compromised agents increases?
  • RQ3: Which components—dynamic trust, evidence weighting, and non-bypassable governance rules—account for the observed effect?
  • RQ4: Does the aggregation mechanism remain computationally practical as coalition size increases?
The main contributions are: (i) an agent-specific zero-trust control model with mission-scoped authorization and tool-level enforcement; (ii) an outcome-updated trust and evidence-weighted fusion rule; (iii) a formal execution invariant and a short proof of enforcement under stated assumptions; and (iv) a fully specified, repeated-run scenario benchmark with adversarial stress testing, ablation analysis, confidence intervals, and paired significance testing.

3. Threat Model and Design Requirements

3.1. System Model

A mission M contains an alert, target asset, time window, operational objective, data scope, approved tools, candidate actions, and approval policy. Agents communicate only through the orchestration layer. An agent cannot call a production tool directly; all requests pass through the PEP. Identity and capability manifests are authenticated by the platform, and policy data are managed outside the agent context.
Ai = <idi, rolei, capi, tooli, memi, Ti, Ci, Hi> (1)
where id is a stable platform identity, cap is the capability manifest, tool and mem define allowable interfaces and memory scope, T is dynamic trust, C is stated confidence, and H is the mission history.

3.2. Adversary Model

The benchmark assumes that an attacker may compromise a subset of agents and cause them to understate governance risk while reporting high confidence. The attacker may also inject misleading evidence, exploit stale context, request a tool outside mission scope, or coordinate multiple biased recommendations. The attacker cannot forge a platform identity, modify policy rules, bypass the PEP, alter the append-only audit store, or approve its own escalated request. Denial-of-service against the governance plane, compromise of the identity provider, side channels outside governed tools, and vulnerabilities in the enforcement implementation are outside the present model.
Table 2. Threats, protected assets, and principal controls.
Table 2. Threats, protected assets, and principal controls.
Threat Protected asset Control Residual risk
Risk understatement or fabricated consensus Response correctness Dynamic trust; evidence weighting; hard policy floors Correlated honest-agent error
Prompt or data injection Tool authority; sensitive context Typed evidence; least privilege; PEP validation Semantic injection before parsing
Privilege or tool overreach Network and endpoint services Mission-scoped capability manifest Misconfigured policy
Sybil or identity substitution Voting influence; audit attribution Authenticated platform identity Identity-provider compromise
Stale or weak evidence Decision quality Freshness and provenance score Unavailable authoritative telemetry
Irreversible high-impact action Service availability Risk threshold; human approval Human approval error

3.3. Safety Requirements and Assumptions

The design requirements are explicit verification, least privilege, separation of decision and enforcement, default deny on missing context, complete traceability, bounded agent influence, and human authorization for high-impact or irreversible actions. The key execution invariant is:
Execute(a,M) ⇒ IDV(i) ∧ Auth(i,a,M) ∧ a ∈ Asafe(M) ∧ [GRS(a,M) < τE ∨ HApprove(a,M)] (2)
IDV denotes identity verification, Auth denotes mission-scoped authorization, A_safe is the policy-derived action set, GRS is governance risk, tau_E is the automatic-execution threshold, and HApprove is a bound human approval.
Proposition 1.
If (i) every consequential tool is reachable only through the PEP, (ii) policy and identity inputs to the PEP are authentic, and (iii) the PEP fails closed, then an action that violates any conjunct in Equation (2) cannot be executed through a governed interface.
Proof of Proposition 1. The PEP evaluates each conjunct before forwarding a request. A false or missing conjunct maps the request to deny or escalate, neither of which invokes the target tool. Because no governed bypass exists by assumption, the violating action has no execution path. The proposition guarantees policy mediation, not semantic correctness of the policy or absence of implementation vulnerabilities. □

4. SwarmGov-ZT Framework

4.1. Architecture

Figure 1 shows four functional layers. The mission and telemetry fabric provides typed, provenance-bearing context. A mission-scoped coalition creates role-specialized outputs. The fusion layer combines evidence while bounding influence. The zero-trust plane evaluates and enforces every consequential request. Importantly, the governance plane is in the action path rather than operating as a post-hoc reviewer.

4.2. Identity-Bound Coalition Formation

Coalition membership is based on role coverage, capability match, identity state, availability, and mission scope. Roles do not convey standing privilege. Each tool token is short-lived and restricted to a mission, resource, action set, and expiry. Delegation creates a derived token with no greater authority than the delegator. Table 3 lists the reference roles.

4.3. Dynamic Trust Update

Agent influence changes only after an outcome or authoritative adjudication is available. Trust is not reputation in a social sense; it is a bounded operational statistic for one deployment context. The update is:
Ti(t+1) = clip[λTi(t) + αsi(t) + δki(t) − βvi(t) − ηhi(t), ε, 1] (3)
s is correct contribution, k is confidence calibration, v is policy-understatement violation, and h is a harmful error of more than one decision class. The benchmark uses lambda=0.90, alpha=0.10, delta=0.02, beta=0.14, eta=0.10, and epsilon=0.05. Trust is reset or re-scoped when an agent model, tool set, or mission domain materially changes.

4.4. Evidence-Weighted Consensus

Each agent returns a risk estimate ri, confidence Ci, evidence-quality score Ei, and provenance record. The fused risk is:
z(M) = [Σi Ti Ci Ei ri] / [Σi Ti Ci Ei] (4)
The denominator prevents coalition size alone from increasing confidence. If the effective evidence mass is below a policy minimum or disagreement exceeds a threshold, the system requests additional evidence or escalates. Compromised agents may initially retain influence, but repeated outcome disagreement reduces Ti.

4.5. Governance Risk and Action Enforcement

GRS = 0.20 Rauth + 0.15 Rtool + 0.15 Rdata + 0.20 Rimpact + 0.15 Rirrev + 0.15 Runc (5)
The six normalized terms represent authorization, tool, data exposure, operational impact, irreversibility, and uncertainty risk. Thresholds 0.30, 0.60, and 0.85 map the score to Permit, Modify, Escalate, and Deny. Two hard rules override the weighted score in the benchmark: authorization risk above 0.86 or tool risk above 0.90 produces at least Deny; combined high impact and irreversibility, or uncertainty above 0.85, produces at least Escalate. These rules prevent compensation in which several low-risk terms conceal one unacceptable condition.

4.6. Decision Workflow, Audit, and Human Oversight

Figure 2 summarizes the runtime sequence. Each audit entry binds the mission identifier, agent identity and version, authorized scope, evidence hashes and provenance, confidence, trust before decision, fused result, policy version, GRS factors, disposition, human approval where applicable, tool request, tool result, outcome, and trust update. A human approval is bound to the exact action and context; a materially changed request must be reevaluated.

4.7. Algorithm and Complexity

Algorithm 1. Trust-aware governed response
  • Construct the typed mission and retrieve policy-governed telemetry.
  • Authenticate candidate agents and select a role-complete, mission-scoped coalition.
  • Collect each output with confidence, evidence score, provenance, and explanation.
  • Compute trust- and evidence-weighted consensus using Equation (4).
  • Derive candidate actions and compute GRS using Equation (5).
  • Apply hard policy floors and construct A_safe.
  • At the PEP, verify Equation (2); permit, modify, escalate, or deny.
  • Execute only the approved request, record the complete trace, and update trust after outcome adjudication.
For n agents, d risk/evidence dimensions, and |A| candidate actions, aggregation is O(nd), policy filtering is O(|A|), and stored trace size is O(n+|A|) per mission. Cryptographic verification, external tool latency, and database persistence are deployment-dependent and are not included in the microbenchmark.

5. Materials and Methods

5.1. Evaluation Design

The evaluation is a scenario-based proof of concept, not a production field trial. Thirty independent runs use fixed seeds 3100–3129. Each run contains 4,000 missions, yielding 120,000 default-condition missions. The adversarial sweep uses 2,000 missions per seed at compromised-agent fractions of 0%, 12.5%, 25%, 37.5%, and 50%. The default coalition has eight agents, two of which are compromised.

5.2. Scenario Generator and Oracle

Each mission contains six risk features corresponding to Equation (5). To represent low-, medium-, and high-risk operations, a mixture component is sampled with probabilities 0.45, 0.38, and 0.17. The components draw features from Beta(2,6), Beta(3.2,3.2), and Beta(5.5,2.2), respectively. Irreversibility is correlated with impact using 0.55 of the sampled irreversibility, 0.45 of impact, and Gaussian noise with standard deviation 0.05. The oracle applies the weighted thresholds and the same explicit hard rules stated in Section 4.5. This makes the policy labelling process inspectable and reproducible.
Honest agent estimates use Gaussian noise with standard deviation 0.095 and a small agent-specific bias drawn from N(0,0.018). Compromised agents use standard deviation 0.12, subtract 0.24 from risk, and add 0.06 to reported confidence. Scores and confidence are clipped to [0, 1]. This adversary models coordinated risk understatement rather than arbitrary code execution.

5.3. Baselines and Ablations

The single-agent baseline uses one noisy estimate. Centralized MAS averages four agents. Role-based MAS takes the median of six agents. Swarm without ZT uses confidence-weighted fusion across eight agents but has neither dynamic trust nor hard policy floors. SwarmGov-ZT uses all mechanisms in Section 4.3, Section 4.4 and Section 4.5. Ablations remove dynamic trust, evidence weighting, or hard governance rules while holding scenarios and seeds constant.

5.4. Metrics and Statistical Analysis

Exact decision rate is the proportion of Permit/Modify/Escalate/Deny predictions equal to the oracle. Macro-F1 assigns equal importance to all four classes. Policy compliance is the proportion of decisions at least as restrictive as the oracle; policy violation is under-response, and false mitigation is over-response. Results are means across 30 runs with normal-approximation 95% confidence intervals. Because runs share identical seeds and scenarios, SwarmGov-ZT and the non-governed swarm are compared with a one-sided paired Wilcoxon signed-rank test. The statistical claim is limited to this generator and adversary model.

5.5. Implementation and Reproducibility

The benchmark was implemented in Python 3.12 with NumPy 2.3, SciPy 1.17, and Matplotlib 3.10 on Linux. The latency microbenchmark uses a single process on an AMD EPYC 9V74 virtualized CPU and repeats the fusion expression 20,000 times per coalition size. Software S1 contains the complete generator, baselines, fixed seeds, metrics, statistical test, and figure generation; Data S1 contains the machine-readable aggregated results. No private operational logs, human participants, or external model APIs were used.

6. Results

6.1. Main Comparison (RQ1)

SwarmGov-ZT achieved 91.41% exact decision accuracy and 94.20% policy compliance. Compared with the equally sized swarm without ZT, exact accuracy increased by 11.80 percentage points (95% CI of the paired difference ±0.35), while violations decreased by 14.38 points (±0.19). Exact-accuracy improvement was significant in the paired Wilcoxon test (W=465, p=8.67×10−7). The non-governed swarm had a very low false-mitigation rate because its compromised agents systematically understated risk; this apparent conservatism was accompanied by the highest violation rate among multi-agent baselines.
Figure 3. Exact governance-decision rate and policy compliance under the default 25% compromised-agent condition.
Figure 3. Exact governance-decision rate and policy compliance under the default 25% compromised-agent condition.
Preprints 231659 g003
Table 4. Default-condition results across 30 runs; values are percentages (mean ± 95% CI).
Table 4. Default-condition results across 30 runs; values are percentages (mean ± 95% CI).
Method Exact Macro-F1 Compliance Violation False mitigation
Single agent 60.42 ± 1.64 49.94 ± 2.66 78.26 ± 4.24 21.74 ± 4.24 17.83 ± 2.79
Centralized MAS 77.99 ± 1.57 58.77 ± 1.87 80.28 ± 2.26 19.72 ± 2.26 2.29 ± 0.91
Role-based MAS 81.05 ± 0.68 62.25 ± 1.00 83.97 ± 1.16 16.03 ± 1.16 2.92 ± 0.55
Swarm without ZT 79.60 ± 0.42 59.00 ± 0.39 79.82 ± 0.45 20.18 ± 0.45 0.21 ± 0.05
SwarmGov-ZT 91.41 ± 0.16 93.06 ± 0.15 94.20 ± 0.37 5.80 ± 0.37 2.79 ± 0.29

6.2. Compromised-Agent Stress Test (RQ2)

SwarmGov-ZT degraded gradually from 92.1% exact accuracy with no compromised agents to 87.4% at 50% compromise. At 50%, its 11.4% violation rate remained substantially below the 32.5% of the non-governed swarm. The result supports robustness to biased participants within the stated identity and PEP assumptions; it does not imply Byzantine consensus against arbitrary code-level adversaries.
Figure 4. Exact decision and policy-violation rates as compromised-agent prevalence increases.
Figure 4. Exact decision and policy-violation rates as compromised-agent prevalence increases.
Preprints 231659 g004
Table 5. Adversarial stress test: exact decision and policy-violation rates (%).
Table 5. Adversarial stress test: exact decision and policy-violation rates (%).
Method Exact 0% Exact 25% Exact 50% Violation 0% Violation 25% Violation 50%
Centralized MAS 81.9 77.6 69.3 10.7 19.9 30.3
Role-based MAS 83.0 80.5 70.1 10.3 17.0 29.5
Swarm without ZT 86.1 79.1 67.5 9.4 20.7 32.5
SwarmGov-ZT 92.1 91.4 87.4 3.0 6.4 11.4

6.3. Ablation Study (RQ3)

Removing hard governance rules produced the largest exact-accuracy reduction (−6.99 points) and more than doubled policy violations. Removing dynamic trust reduced exact accuracy by 4.38 points and violations increased from 6.36% to 12.84%. Evidence weighting had a smaller effect because the simulated confidence distributions overlap; its importance may be larger when provenance quality is more heterogeneous.
Table 6. Ablation results at 25% compromised agents (%).
Table 6. Ablation results at 25% compromised agents (%).
Variant Exact Macro-F1 Compliance Violation False mitigation
Full 91.38 92.97 93.64 6.36 2.26
No dynamic trust 87.00 89.29 87.16 12.84 0.16
No evidence weighting 90.89 92.56 93.61 6.39 2.72
No governance rules 84.39 64.49 86.64 13.36 2.26
Figure 5. Ablation comparison of exact decision rate and policy violation rate.
Figure 5. Ablation comparison of exact decision rate and policy violation rate.
Preprints 231659 g005

6.4. Aggregation Cost (RQ4)

Measured fusion time remained approximately 4.5–4.8 μs for 4–64 agents because vectorized array overhead dominated at these small sizes. This measurement excludes identity verification, policy-store access, persistence, and external tools. The relevant architectural result is the linear O(nd) aggregation bound; production latency must be measured end to end in the target platform.
Table 7. Single-process fusion microbenchmark.
Table 7. Single-process fusion microbenchmark.
Coalition size (agents) Mean time per fusion (μs)
4 4.52
8 4.60
16 4.70
32 4.78
64 4.69

7. Discussion

7.1. Interpretation

The results answer RQ1 by showing that collaboration alone did not protect against coordinated risk understatement. Confidence weighting was particularly vulnerable because compromised agents were overconfident. SwarmGov-ZT improved exact decisions by combining two different defenses: outcome history reduced unreliable influence, while non-bypassable rules prevented unacceptable factors from being averaged away. RQ2 shows graceful rather than catastrophic degradation through 50% compromise. RQ3 confirms that governance rules and dynamic trust are complementary. RQ4 indicates that fusion itself is unlikely to be the performance bottleneck; external policy and tool operations will dominate.
The distinction between decision quality and policy compliance is essential. A system can achieve a low false-mitigation rate by rarely escalating, yet remain unsafe because it under-responds to restricted actions. For this reason, exact accuracy, macro-F1, under-response, and over-response are reported together. This follows the utility-security framing used by agent security benchmarks [10,11].

7.2. Security Analysis and Failure Modes

Proposition 1 provides a mediation guarantee only under the stated assumptions. It does not prove that risk features are correct, that policies capture every unsafe consequence, or that a human approval is sound. A compromised identity provider or PEP invalidates the guarantee. Colluding agents may also behave honestly until trust is high and attack in a short window. Mitigations include capped trust, recency weighting, randomized audits, diversity requirements, independent policy telemetry, two-person approval for irreversible actions, and periodic red-team testing using environments such as AgentDojo and ASB [10,11].
Trust scores may encode historical bias or become stale after model updates. They should therefore be domain- and version-specific, explainable, appealable, and subject to expiry. Trust must never expand capability: a highly trusted agent can have more influence within its authorized role, but cannot obtain a new tool or data scope without a separate policy decision.

7.3. Deployment Mapping

In a practical SOC, the PDP can integrate with identity and access management, a policy engine, asset inventory, ticketing, and security orchestration. The PEP should wrap firewall, EDR, NAC, cloud, and identity-administration APIs. NIST ZTA concepts map directly: the agent is the subject, the tool or data object is the resource, the mission token supplies context, the PDP computes disposition, and the PEP mediates access [1]. MITRE ATT&CK techniques can label evidence and response rationale [16]. NIST AI RMF functions can organize governance documentation and monitoring [2,3].

7.4. Limitations and External Validity

The principal limitation is that the benchmark is synthetic and partly shares structure with the policy oracle. It demonstrates implementation consistency and controlled robustness, not real-world effectiveness. The noise, confidence, and compromise distributions are assumptions rather than measurements from deployed LLM agents. The hard rules naturally favor a system that enforces them, so effect sizes should not be generalized beyond the tested policy. No public intrusion-detection dataset, live SOC, network digital twin, or human analyst study is evaluated. The simulation also excludes latency from cryptography, storage, policy engines, and production tools.
A stronger external validation should: (i) integrate an actual tool-calling agent environment; (ii) evaluate direct and indirect prompt injection, memory poisoning, tool misuse, and collusion using AgentDojo/ASB-compatible tasks; (iii) use public network traces or an instrumented cyber range; (iv) compare policy engines and human approval strategies; (v) report benign utility as well as attack success; and (vi) preregister scenario distributions and acceptance thresholds. These steps are necessary before claims about deployment readiness.

8. Conclusions

SwarmGov-ZT addresses a systems problem that model accuracy alone cannot solve: how to permit useful multi-agent cybersecurity reasoning without granting implicit authority to agents, messages, memories, or tools. The framework binds agent identity to mission scope, updates influence from outcomes, weights evidence explicitly, applies non-compensable policy rules, enforces decisions at a PEP, and records an identity-to-outcome trace. In a reproducible 120,000-mission simulation with 25% compromised agents, the framework improved exact governance decisions and substantially reduced policy violations relative to non-governed collaboration. Robustness remained comparatively strong at 50% compromise, and ablation results identified dynamic trust and hard governance rules as the main contributors. The evidence supports architectural feasibility under controlled assumptions. Production claims require external validation with tool-calling agents, public adversarial benchmarks, cyber-range telemetry, and human oversight studies.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

Abbreviation Meaning
AI Artificial intelligence
ASB Agent Security Bench
GRS Governance Risk Score
MAS Multi-agent system
PDP Policy decision point
PEP Policy enforcement point
SOC Security operation center
ZTA Zero Trust Architecture

References

  1. Rose, S.; Borchert, O.; Mitchell, S.; Connelly, S. Zero Trust Architecture; NIST Special Publication 800-207; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2020. [Google Scholar] [CrossRef]
  2. Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0); NIST AI 100-1; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [Google Scholar] [CrossRef]
  3. NIST AI 600-1; Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. National Institute of Standards and Technology: Gaithersburg, MD, USA, 2024. [CrossRef]
  4. Booth, H.; et al. Accelerating the Adoption of Software and AI Agent Identity and Authorization: Concept Paper; National Cybersecurity Center of Excellence, NIST: Rockville, MD, USA, 2026; Available online: https://www.nccoe.nist.gov/projects/software-and-ai-agent-identity-and-authorization (accessed on 27 August 2026).
  5. Shrestha, S.; Banda, C.; Mishra, A.K.; Djebbar, F.; Puthal, D. Investigation of Cybersecurity Bottlenecks of AI Agents in Industrial Automation. Computers 2025, 14, 456. [Google Scholar] [CrossRef]
  6. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
  7. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar] [CrossRef]
  8. Wooldridge, M. An Introduction to MultiAgent Systems, 2nd ed.; Wiley: Chichester, UK, 2009. [Google Scholar]
  9. Schroeder de Witt, C. Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents. arXiv 2025, arXiv:2505.02077. [Google Scholar] [CrossRef]
  10. Debenedetti, E.; Zhang, J.; Balunovic, M.; Beurer-Kellner, L.; Fischer, M.; Tramèr, F. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv 2024, arXiv:2406.13352. [Google Scholar] [CrossRef]
  11. Zhang, H.; Huang, J.; Mei, K.; Yao, Y.; Wang, Z.; Zhan, C.; Wang, H.; Zhang, Y. Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents. In Proceedings of the Thirteenth International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar]
  12. Huang, J.; Zhou, J.; Jin, T.; Zhou, X.; Chen, Z.; Wang, W.; Yuan, Y.; Sap, M.; Lyu, M.R. On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents. arXiv 2024, arXiv:2408.00989. [Google Scholar] [CrossRef]
  13. Lee, D.; Tiwari, M. Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems. arXiv 2024, arXiv:2410.07283. [Google Scholar] [CrossRef]
  14. OWASP Foundation. OWASP Top 10 for LLM and Generative AI Applications. Available online: https://genai.owasp.org/llm-top-10/ (accessed on 27 August 2026).
  15. Joint Task Force. Security and Privacy Controls for Information Systems and Organizations; NIST Special Publication 800-53 Revision 5; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2020. [Google Scholar] [CrossRef]
  16. MITRE. ATT&CK: Adversarial Tactics, Techniques, and Common Knowledge. Available online: https://attack.mitre.org/ (accessed on 27 August 2026).
  17. Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; Luetge, C.; Madelin, R.; Pagallo, U.; Rossi, F.; et al. AI4People—An Ethical Framework for a Good AI Society. Minds Mach. 2018, 28, 689–707. [Google Scholar] [CrossRef] [PubMed]
  18. Jobin, A.; Ienca, M.; Vayena, E. The Global Landscape of AI Ethics Guidelines. Nat. Mach. Intell. 2019, 1, 389–399. [Google Scholar] [CrossRef]
  19. García, J.; Fernández, F. A Comprehensive Survey on Safe Reinforcement Learning. J. Mach. Learn. Res. 2015, 16, 1437–1480. [Google Scholar]
  20. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
  21. Altman, E. Constrained Markov Decision Processes; Chapman & Hall/CRC: Boca Raton, FL, USA, 1999. [Google Scholar]
  22. Dorigo, M.; Stützle, T. Ant Colony Optimization; MIT Press: Cambridge, MA, USA, 2004. [Google Scholar]
  23. Bonabeau, E.; Dorigo, M.; Theraulaz, G. Swarm Intelligence: From Natural to Artificial Systems; Oxford University Press: New York, NY, USA, 1999. [Google Scholar]
  24. Harrenstein, P. Commitment and Trust in Multi-Agent Systems. ACM Trans. Auton. Adapt. Syst. 2007, 2, 4. [Google Scholar]
  25. Douceur, J.R. The Sybil Attack. In Peer-to-Peer Systems; Druschel, P., Kaashoek, F., Rowstron, A., Eds.; Springer: Berlin/Heidelberg, Germany, 2002; pp. 251–260. [Google Scholar] [CrossRef] [PubMed]
  26. Xu, D.; Gondal, I.; Yi, X.; Susnjak, T.; Watters, P.; McIntosh, T.R. The Erosion of Cybersecurity Zero-Trust Principles Through Generative AI: A Survey on the Challenges and Future Directions. J. Cybersecur. Priv. 2025, 5, 87. [Google Scholar] [CrossRef]
  27. Russell, S.; Norvig, P. Artificial Intelligence: A Modern Approach, 4th ed.; Pearson: Hoboken, NJ, USA, 2020. [Google Scholar]
  28. Souppaya, M.; Scarfone, K.; Dodson, D. Secure Software Development Framework (SSDF) Version 1.1; NIST Special Publication 800-218; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2022. [Google Scholar] [CrossRef]
Figure 1. SwarmGov-ZT architecture and the separation between collaborative reasoning and enforceable action governance.
Figure 1. SwarmGov-ZT architecture and the separation between collaborative reasoning and enforceable action governance.
Preprints 231659 g001
Figure 2. Governed decision workflow and non-bypassable enforcement invariant.
Figure 2. Governed decision workflow and non-bypassable enforcement invariant.
Preprints 231659 g002
Table 1. Capability comparison and research gap.
Table 1. Capability comparison and research gap.
Approach Identity and scope Outcome-updated trust Tool-level enforcement Adversarial evaluation Decision trace
Single agent Partial No Framework dependent Occasional Partial
Generic MAS Role based Usually no Framework dependent Limited Conversation log
Agent benchmarks [10,11] Benchmark specific No Evaluation sandbox Strong Benchmark record
Conventional ZTA [1] Strong Contextual Strong Not agent specific Strong
SwarmGov-ZT Mission scoped Yes Non-bypassable PEP Compromised-agent sweep Identity-to-outcome ledger
Table 3. Reference agent roles and execution boundaries.
Table 3. Reference agent roles and execution boundaries.
Role Responsibility Output Direct production execution
Planner Decompose mission and request roles Task graph; required evidence No
Detection Interpret alerts and telemetry Finding; confidence; provenance No
Risk Estimate asset and service impact Risk factors; uncertainty No
Network operations Assess topology and change feasibility Operational constraints Read only
Policy Rank candidate responses Action proposal; alternatives No
Governance/PDP Evaluate identity, scope, and risk Permit; modify; escalate; deny Decision only
PEP Enforce approved request Tool invocation and result Yes, policy mediated
Audit Bind inputs, decisions, actions, outcomes Append-only trace No
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.