Submitted:
28 December 2025
Posted:
29 December 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We propose a novel dynamic trust assessment mechanism that continuously evaluates agent reliability using multi-dimensional metrics including behavioral consistency, decision accuracy, and collaboration efficiency.
- We develop an adversary-aware orchestration strategy that combines reinforcement learning with game-theoretic principles to proactively detect and mitigate various attack vectors including prompt injection and action perturbation.
- We introduce an adaptive collaboration topology that dynamically reconfigures agent communication structures based on real-time trust assessments and task requirements, reducing coordination overhead by 39.8%.
- We implement a comprehensive explainable decision tracing framework for complete audit chains, meeting regulatory requirements for high-risk applications.
- We demonstrate through extensive experiments that TrustOrch significantly improves system robustness against adversarial attacks while maintaining high performance in benign scenarios.
2. Related Work
2.1. Trust Management in Multi-Agent Systems
3. System Architecture
3.1. Overview
3.2. Dynamic Trust Assessment Mechanism
3.3. Adversary-Aware Orchestration Strategy
| Algorithm 1 Adversary-Aware Orchestration Training |
| Input: Initial orchestration parameters , learning rates |
| Output: Robust orchestration policy |
| 1: Initialize adversarial policy randomly |
| 2: for episode to K do |
| 3: // Adversarial policy update |
| 4: Generate trajectories using current |
| 5: Update using gradient ascent on |
| 6: // Orchestration policy update |
| 7: Simulate attacks using |
| 8: Update using policy gradient with robustness term |
| 9: // Trust assessment update |
| 10: Update trust scores based on agent behaviors |
| 11: end for |
| 12: return |
3.4. Adaptive Collaboration Topology
4. Security Architecture
4.1. Layered Defense Mechanism
4.2. Blockchain-Based Trust Verification
- Global Chain: Maintains agent identities and high-level aggregated trust scores
- Regional Chains: Record task-specific interactions and performance metrics
- Local Chains: Store detailed execution logs for audit purposes
5. Experimental Evaluation
5.1. Experimental Setup
- Static Trust (ST): Traditional static trust model with fixed topology
- ERNIE: Adversarial regularization framework [11]
- TrustChain: Blockchain-based trust management [8]
- MSR: Mean Subsequence Reduced algorithm for secure consensus
5.2. Performance Metrics
- Robustness Score (RS): Percentage of successful task completions under attack
- Communication Overhead (CO): Average messages per task
- Trust Accuracy (TA): Precision in identifying malicious agents
- Response Latency (RL): Average decision time in milliseconds
5.3. Results and Analysis
5.3.1. Robustness Against Adversarial Attacks
5.3.2. Communication Efficiency
5.3.3. Trust Assessment Accuracy
5.3.4. Scalability Analysis
5.4. Case Study: Autonomous Vehicle Coordination
6. Discussion
6.1. Key Insights
6.2. Limitations and Future Work
7. Conclusions
References
- MarketsandMarkets, "Multi-Agent Systems Market - Global Forecast to 2028," Market Research Report, Tech. Rep., 2023.
- L. Yuan, F. Chen, Z. Zhang et al., "Communication-robust multi-agent learning by adaptable auxiliary multi-agent adversary generation," Frontiers of Computer Science, vol. 18, 186331, 2024. [CrossRef]
- O. Ma, Y. Pu, L. Du et al., "SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems," in Proc. ACM SIGSAC Conference on Computer and Communications Security, 2024.
- J. Zhu, C. Lu, J. Li, and F.-Y. Wang, "Secure consensus control on multi-agent systems based on improved PBFT and Raft blockchain consensus algorithms," IEEE/CAA Journal of Automatica Sinica, vol. 12, no. 7, pp. 1407-1417, 2025. [CrossRef]
- A. Pattanaik, Z. Tang, S. Liu, and G. Bommannan, "Robust Deep Reinforcement Learning with Adversarial Attacks," in Proc. 17th International Conference on Autonomous Agents and MultiAgent Systems, pp. 2040-2042, 2018.
- W. Chen, Y. Su, J. Zuo et al., "Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents," arXiv preprint arXiv:2308.10848, 2023.
- D. Chen, K. Zhang, Y. Wang et al., "Multi-Agent Collaboration Mechanisms: A Survey of LLMs," arXiv preprint arXiv:2501.06322, 2025.
- S. Raza, R. Sapkota, M. Karkee, and C. Emmanouilidis, “TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-Based Agentic Multi-Agent Systems,” arXiv preprint arXiv:2506.04133, 2025.
- S. Malik et al., "A blockchain-enabled trust aware energy trading framework using games theory and multi-agent system in smart grid," Energy, vol. 255, 124452, 2022.
- S. Malik, V. Dedeoglu, S. S. Kanhere, and R. Jurdak, "TrustChain: Trust Management in Blockchain and IoT Supported Supply Chains," in IEEE International Conference on Blockchain, pp. 184-193, 2019.
- Anonymous, "Time-Exact Multi-Blockchain Architectures for Trustworthy Multi-Agent Systems," OpenReview, 2025.
- L. Yuan, J. Zhang, and F. Chen, "Adaptive Auxiliary Adversary Generation for Robust Multi-Agent Communication," in Proc. International Conference on Machine Learning, pp. 17534-17543, 2023.
- A. Bukharin, Y. Li, Y. Yu et al., "Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable Algorithms," in Advances in Neural Information Processing Systems 36, 2023.
- A. Sharif and D. Marijan, "Adversarial Deep Reinforcement Learning for Improving the Robustness of Multi-agent Autonomous Driving Policies," in Proc. 29th IEEE International Conference on Software Analysis, Evolution and Reengineering, 2023.
- O. Ma, X. Liu, and Y. Xia, "Detecting adversarial directions in deep reinforcement learning to make robust decisions," in Proc. 40th International Conference on Machine Learning, pp. 17534-17543, 2023.
- Zou, Y., Qi, N., Deng, Y., Xue, Z., Gong, M., & Zhang, W. (2025, July). Autonomous resource management in microservice systems via reinforcement learning. In 2025 8th International Conference on Computer Information Science and Application Technology (CISAT) (pp. 991-995). IEEE.
- Yao, G., Liu, H., & Dai, L. (2025). Multi-agent reinforcement learning for adaptive resource orchestration in cloud-native clusters. arXiv preprint arXiv:2508.10253.
- Li, Y. (2024). Differential Privacy-Enhanced Federated Learning for Robust AI Systems. Journal of Computer Technology and Software, 3(4).
- Sun, Y., Zhang, R., Meng, R., Lian, L., Wang, H., & Quan, X. (2025, July). Fusion-based retrieval-augmented generation for complex question answering with LLMs. In 2025 8th International Conference on Computer Information Science and Application Technology (CISAT) (pp. 116-120). IEEE.
- Zheng, J., Chen, Y., Zhou, Z., Peng, C., Deng, H., & Yin, S. (2025). Information-Constrained Retrieval for Scientific Literature via Large Language Model Agents.
- Pan, S., & Wu, D. (2025). Trustworthy summarization via uncertainty quantification and risk awareness in large language models. arXiv preprint arXiv:2510.01231.
- Hu, X., Kang, Y., Yao, G., Kang, T., Wang, M., & Liu, H. (2025). Dynamic prompt fusion for multi-task and cross-domain adaptation in LLMs. arXiv preprint arXiv:2509.18113.
- Wang, Y., Wu, D., Liu, F., Qiu, Z., & Hu, C. (2025). Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models. arXiv preprint arXiv:2511.03981.
- Liu, X., Qin, Y., Xu, Q., Liu, Z., Guo, X., & Xu, W. (2025). Integrating Knowledge Graph Reasoning with Pretrained Language Models for Structured Anomaly Detection.
- Lyu, S., Wang, M., Zhang, H., Zheng, J., Lin, J., & Sun, X. (2025). Integrating Structure-Aware Attention and Knowledge Graphs in Explainable Recommendation Systems. arXiv preprint arXiv:2510.10109.
- Li, J., Gan, Q., Liu, Z., Chiang, C., Ying, R., & Chen, C. (2025). An Improved Attention-Based LSTM Neural Network for Intelligent Anomaly Detection in Financial Statements.
- Ying, R., Lyu, J., Li, J., Nie, C., & Chiang, C. (2025). Dynamic Portfolio Optimization with Data-Aware Multi-Agent Reinforcement Learning and Adaptive Risk Control.
- Chang, W. C., Dai, L., & Xu, T. (2025). Machine Learning Approaches to Clinical Risk Prediction: Multi-Scale Temporal Alignment in Electronic Health Records. arXiv preprint arXiv:2511.21561.
- Liu, R., Zhang, R., & Wang, S. (2025). Graph Neural Networks for User Satisfaction Classification in Human-Computer Interaction. arXiv preprint arXiv:2511.04166.
- Xie, J., Wu, Y., Zhang, Y., Zhang, X., Xie, Y., & Qu, Y. (2025, October). PLATO-TTA: Prototype-Guided Pseudo-Labeling and Adaptive Tuning for Multi-Modal Test-Time Adaptation of 3D Segmentation. In Proceedings of the 33rd ACM International Conference on Multimedia (pp. 2226-2234).
- Song, X., Liu, Y., Luan, Y., Guo, J., & Guo, X. (2025). Controllable Abstraction in Summary Generation for Large Language Models via Prompt Engineering. arXiv preprint arXiv:2510.15436.





| Method | RS (%) | CO (msgs) | TA (%) | RL (ms) |
|---|---|---|---|---|
| Static Trust | 62.3 | 145.2 | 71.4 | 23.5 |
| ERNIE | 78.5 | 112.3 | 82.7 | 31.2 |
| TrustChain | 75.2 | 98.7 | 88.3 | 45.8 |
| MSR | 69.8 | 156.4 | 76.5 | 28.9 |
| TrustOrch | 91.7 | 87.3 | 94.2 | 34.6 |
| Predicted Malicious | Predicted Benign | |
|---|---|---|
| Actual Malicious | 188 (TP) | 12 (FN) |
| Actual Benign | 8 (FP) | 792 (TN) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.