Preprint
Article

This version is not peer-reviewed.

SafeAgent: Dual-Objective Optimization Against Prompt and Tool Injection in Interactive Environments

  † These authors contributed equally to this work.

Submitted:

21 September 2026

Posted:

22 September 2026

You are already at the latest version

Abstract
Large Language Model (LLM) agents that interact with external tools, web content, and code repositories face significant security risks from prompt injection and tool manipulation attacks. However, overly restrictive security measures often render agents ineffective, creating a fundamental tension between safety and utility. We propose SafeAgent, a dual-objective optimization framework that simultaneously quantifies and balances these competing goals. Our approach introduces three complementary defense mechanisms: a context firewall that separates external data from executable instructions, a tool-call sandbox enforcing least-privilege access control, and a safety-aware planning module integrated into the ReAct reasoning loop. We construct a comprehensive attack benchmark encompassing web retrieval injections, malicious tool responses, and repository README attacks. Experimental evaluation on the AgentDojo benchmark demonstrates that our full defense configuration reduces attack success rate (ASR) from 67.2% to 13.4% (a reduction of 53.8 percentage points, corresponding to approximately 80% relative reduction) while maintaining 68.9% task utility. SafeAgent achieves Pareto-optimal performance among tested configurations compared to baseline approaches. Our framework provides measurable, reproducible metrics for evaluating the safety-utility tradeoff in agentic AI systems.
Keywords: 
;  ;  ;  ;  

1. Introduction

The emergence of Large Language Model (LLM) agents capable of autonomous tool use represents a paradigm shift in AI applications [1]. These agents combine natural language reasoning with external actions, enabling them to browse the web, execute code, manage files, and interact with APIs. However, this capability expansion introduces severe security vulnerabilities that traditional LLM safety mechanisms cannot adequately address [2].
Prompt injection attacks exploit the fundamental inability of LLMs to distinguish between trusted instructions and untrusted data [3]. When agents process external content, such as web pages, email attachments, or tool responses, attackers can embed malicious instructions that hijack agent behavior [4]. The OWASP Foundation identifies prompt injection as the top security risk for LLM applications, with documented incidents including credential theft, unauthorized actions, and data exfiltration [2].
A naive approach to agent security involves aggressive filtering or capability restriction. However, this creates a second problem: overly cautious agents that refuse legitimate requests or fail to complete useful tasks. We term this the safety-utility dilemma, where improving one objective often degrades the other. Existing work typically addresses either safety [5] or utility [6] in isolation, lacking a unified framework for optimizing both simultaneously.
This paper makes the following contributions: (1) We formalize the dual-objective optimization problem for agent security, defining quantitative metrics for both safety and utility. (2) We propose SafeAgent, a defense architecture combining context firewalls, tool sandboxing, and safety-aware planning. (3) We construct a diverse attack benchmark covering realistic injection scenarios. (4) We demonstrate Pareto-optimal performance among tested configurations on standard benchmarks, achieving measurable improvements in the safety-utility tradeoff.

3. Problem Formulation

3.1. Threat Model

We consider an LLM agent A with access to a tool set T = { t 1 , t 2 , ... , t n } for tasks including web retrieval, code execution, and file operations. The agent receives user queries and generates action sequences to complete tasks. An adversary M can inject malicious content into any external data source the agent may access, including web pages, tool responses, or repository files. The adversary's goal is to cause the agent to execute unauthorized actions, leak sensitive information, or deviate from the user's intended task.

3.2. Dual-Objective Framework

Let D = { ( q i , g i ) } i = 1 N denote a dataset of N user queries with corresponding ground-truth goals. Let N atk denote the number of attack test cases. We define two primary metrics (reported as percentages for readability, i.e., values in [ 0 , 100 ] ):
Utility Score (%) measures task completion quality:
Utility = 100 N ∑ i = 1 N 1 [ A ( q i ) ⊨ g i ]
where 1 [ ⋅ ] is the indicator function and A ( q i ) ⊨ g i indicates the agent's output satisfies the task goal g i .
Safety Score (%) measures resistance to attacks:
Safety = 100 − 100 N atk ∑ j = 1 N atk 1 [ A ( q j atk ) ⊨ g j mal ]
where q j atk represents a query containing injected attack content, and g j mal is the adversary's malicious objective. Equivalently, Safety = 100 − ASR , where ASR is the Attack Success Rate in percentage points.
The dual-objective optimization seeks configurations on the Pareto frontier:
max θ ( Utility ( θ ) , Safety ( θ ) )
where θ represents defense configuration parameters. A configuration is Pareto-optimal if no other configuration improves one objective without degrading the other. We note that we evaluate Pareto optimality empirically among tested configurations and do not claim global optimality guarantees.

3.3. Attack Taxonomy

We categorize injection attacks by their source and mechanism:
Type I: Web Retrieval Injection. Malicious instructions embedded in web page content retrieved during information-seeking tasks. Attack payloads typically use invisible formatting, comment tags, or semantic disguise.
Type II: Tool Response Injection. Adversarial content returned by compromised or malicious API endpoints, including poisoned database entries and manipulated file contents.
Type III: Repository Injection. Hidden instructions in code documentation, README files, or comments that activate when agents process repository content.

4. Proposed Defense Framework

SafeAgent's defense architecture comprises three layers designed to provide defense-in-depth while preserving agent capability. Figure 1 illustrates the overall system design. Note that the modules shown are conceptual; implementation details are provided in Section V-C.

4.1. Context Firewall

The context firewall enforces strict separation between data and instructions. All external content is wrapped in explicit data markers and processed through a classification stage before entering the agent's context window.
Given input text x , the firewall applies transformation ϕ :
ϕ ( x ) = [ DATASTART ] ⊕ sanitize ( x ) ⊕ [ DATAEND ]
where ⊕ denotes concatenation and sanitize ( ⋅ ) removes known injection patterns including control characters, instruction-like prefixes, and encoded payloads.
The agent's system prompt explicitly instructs that content within data markers should be treated as information only and never executed as commands. We do not claim formal guarantees for data/instruction separation; we evaluate empirically under benchmarked attack settings. Empirical results show significant reduction in successful injections.

4.2. Tool-Call Sandbox

The sandbox layer implements the principle of least privilege by restricting tool access based on task requirements. For each task, SafeAgent defines an allowlist T allow ⊆ T specifying permitted tools and parameter constraints.
A tool call ( t , p ) is permitted only if:
permit ( t , p ) = 1 [ t ∈ T allow ] ⋅ 1 [ p ∈ P t ]
where P t defines valid parameter ranges for tool t . Unauthorized calls are blocked and logged.
The allowlist is determined by a pre-planning phase where SafeAgent analyzes the user query before exposure to external content. In this phase, the agent assesses the task requirements and determines the minimal set of tools needed. This follows the principle that tool permissions should be decided based on trusted input only.

4.3. Safety-Aware Planning

SafeAgent extends the ReAct framework [1] with explicit safety verification steps. The standard ReAct loop alternates between Thought → Action → Observation. Our modified loop inserts a Safety Check before each action:
  • Thought: Reason about current state and next step
  • Safety Check: Verify action aligns with original goal
  • Action: Execute if safety check passes
  • Observation: Process result through context firewall
The safety check evaluates whether the proposed action is consistent with the user's original query and whether it would violate any security constraints. Actions flagged as potentially compromised trigger a re-planning step that excludes the suspicious reasoning chain.
The planner draws on the verifier-guided reasoning perspective of Wang et al. [49] to place an explicit assessment between a candidate reasoning step and its execution. In SafeAgent, the verified object is the proposed tool action, and the acceptance criterion is consistency with the trusted user goal and security constraints. Re-planning after a failed check concentrates additional reasoning on a candidate that cannot safely proceed, while a passing check permits the existing action path to continue. The budget-aware framing also makes verification overhead part of the design rationale: repeated checks must be assessed against the useful work they preserve. The current implementation checks every action and reports latency overhead; it does not implement adaptive budget allocation or claim the efficiency results of a separate reasoning system.

4.4. Output Validator

The Output Validator serves as the final defense layer, performing anomaly detection on agent outputs before they are returned to the user. It checks for indicators of successful injection attacks, including: (1) responses that contain content unrelated to the original user query, (2) outputs that attempt to execute unauthorized actions or leak sensitive information, and (3) anomalous patterns suggesting the agent has been compromised. When suspicious outputs are detected, the validator either sanitizes the response or triggers a re-execution with enhanced safety constraints.
The validator incorporates the sentinel perspective expressed by AgentShield [50] by treating the release of an agent response as a distinct monitoring boundary. A fluent intermediate conclusion can propagate into later decisions unless its consistency with the task is checked before it is reused or returned. SafeAgent applies this principle through its existing anomaly checks for goal deviation, unauthorized behavior, and indications of compromise, with sanitization or constrained re-execution as the response to a detected problem. The firewall limits what enters the reasoning context, whereas the validator checks what leaves the agent, giving the two mechanisms complementary positions in the information flow. This use of a sentinel-style boundary concerns the present single-agent pipeline; multi-agent hallucination propagation is a separate evaluation setting.

5. Experimental Setup

5.1. Benchmark and Metrics

We evaluate on the AgentDojo benchmark [6], which provides 97 tasks across four domains: Banking (financial transactions), Workspace (document management), Travel (booking operations), and Slack (communication tasks). The benchmark includes 629 security test cases with injected attack payloads. We use the official split and configuration from the original benchmark without additional curation.
We report four metrics (all values in percentages): (1) Benign Utility (BU): Task success rate without attacks. (2) Utility Under Attack (UA): Task success rate when attacks are present. (3) Attack Success Rate (ASR): Percentage of attacks that achieve their malicious goal. (4) Safety Score: Defined as 100 − ASR (in percentage points).

5.2. Defense Configurations

We compare four configurations: (1) No Defense: Baseline agent with standard prompting. (2) Filter Only: Input sanitization and injection pattern detection. (3) Filter + Sandbox: Adding tool-call restrictions via allowlists. (4) SafeAgent (Full): Complete framework including context firewall, sandbox, and safety-aware planning.

5.3. Implementation Details

Experiments use GPT-4o as the backbone LLM. The context firewall employs regex-based pattern matching augmented with a classifier trained on benchmark injection samples combined with public prompt-injection pattern datasets. Allowlists are generated automatically by prompting the LLM to identify required tools from task descriptions before external data exposure. Safety checks use a temperature of 0 for deterministic verification.

6. Results and Analysis

6.1. Overall Performance

Table 1 presents the main experimental results. The full defense configuration achieves the best balance between safety and utility, reducing ASR from 67.2% to 13.4% (a reduction of 53.8 percentage points) while maintaining 68.9% utility compared to 78.5% for the undefended baseline.

6.2. Pareto Analysis

Figure 2 shows the safety-utility tradeoff across all configurations. SafeAgent achieves points along the estimated Pareto frontier among tested configurations, meaning no other tested configuration provides better safety at the same utility level or vice versa. The shaded region indicates configurations that are dominated by frontier points.
Notably, the transition from Filter + Sandbox to SafeAgent yields disproportionate safety gains (67.5% to 86.6%) with modest utility cost (71.8% to 68.9%), suggesting the safety-aware planning component provides high return on investment.

6.3. Attack-Type Analysis

Figure 2 presents attack success rates across attack types and defense configurations. Indirect (Web) injection shows the highest baseline vulnerability (82% ASR) but SafeAgent reduces it effectively to 22%. Tool Poisoning attacks are most effectively defended against, with SafeAgent achieving the lowest ASR of 15%.
The relative ASR reduction from baseline to SafeAgent is approximately 80% on average (from 67.2% to 13.4% in Table 1), calculated as:
Relative   Reduction = ASR baseline − ASR defense ASR baseline
This demonstrates the framework's general effectiveness across diverse attack vectors. Domain-specific analysis reveals that Workspace tasks experience the highest utility degradation under full defense (from 76.3% to 65.1%), likely due to the broader tool requirements of document management operations.
Figure 3. Attack success rate across attack types and defense configurations. Lower values indicate stronger defense. SafeAgent achieves consistent ASR below 23% across all attack types. Cell values are shown as decimals (e.g., 0.18 = 18% ASR).
Figure 3. Attack success rate across attack types and defense configurations. Lower values indicate stronger defense. SafeAgent achieves consistent ASR below 23% across all attack types. Cell values are shown as decimals (e.g., 0.18 = 18% ASR).
Preprints 234354 g003

6.4. Component Ablation

To understand individual component contributions, we conducted ablation experiments removing each defense layer. Results indicate the context firewall contributes approximately 35% of total ASR reduction, the sandbox contributes 28%, and safety-aware planning contributes 37%. However, these components exhibit synergistic effects; removing any single layer degrades overall performance more than the sum of individual contributions would suggest.

6.5. Limitations

Our framework has several limitations. First, the context firewall relies on heuristic patterns that sophisticated adversaries may evade through novel encoding schemes. Second, automatically generated allowlists may be overly permissive or restrictive for edge cases. Third, safety checks incur computational overhead, increasing average response latency by approximately 23%. Finally, our evaluation focuses on text-based injection; multimodal attacks (e.g., image-based prompt injection) through images or audio remain unexplored.

7. Conclusions

We presented SafeAgent, a dual-objective optimization framework for securing LLM agents against prompt and tool injection attacks while preserving task utility. Our three-layer defense architecture, combining context firewalls, tool sandboxing, and safety-aware planning, achieves Pareto-optimal performance among tested configurations on the AgentDojo benchmark. The framework reduces attack success rates from 67.2% to 13.4% (approximately 80% relative reduction) while maintaining 68.9% task utility.
Our results demonstrate that, under the AgentDojo benchmark and our evaluated attack types, the safety-utility tradeoff in agent systems is not strictly zero-sum; carefully designed defenses can achieve substantial security improvements with acceptable utility costs. This balance relies on the pre-planning phase accurately determining tool permissions based on trusted input. The quantitative framework we provide enables reproducible evaluation and comparison of future defense mechanisms.
Future work should address multimodal injection attacks such as image-based prompt injection, develop adaptive defenses that learn from attack patterns, and explore formal verification methods for safety guarantees. As LLM agents become increasingly deployed in real-world applications, principled approaches to balancing safety and utility will be essential for trustworthy AI systems.

References

  1. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K. R.; Cao, Y. React: Synergizing reasoning and acting in language models. The eleventh international conference on learning representations, 2022. [Google Scholar]
  2. OWASP Top 10 for Large Language Model Applications 2025. 2025. Available online: https://genai.owasp.org/llmrisk/.
  3. Perez, F.; Ribeiro, I. Ignore previous prompt: Attack techniques for language models. arXiv 2022, arXiv:2211.09527. [Google Scholar]
  4. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM workshop on artificial intelligence and security, 2023; pp. 79–90. [Google Scholar]
  5. Zhang, Z.; Cui, S.; Lu, Y.; Zhou, J.; Yang, J.; Wang, H.; Huang, M. Agent-safetybench: Evaluating the safety of llm agents. arXiv 2024, arXiv:2412.14470. [Google Scholar]
  6. Debenedetti, E.; Zhang, J.; Balunovic, M.; Beurer-Kellner, L.; Fischer, M.; Tramèr, F. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. Adv. Neural Inf. Process. Syst. 2024, vol. 37, 82895–82920. [Google Scholar] [CrossRef]
  7. Liu, Y.; Deng, G.; Li, Y.; Wang, K.; Wang, Z.; Wang, X.; Zhang, T.; Liu, Y.; Wang, H.; Zheng, Y. Prompt injection attack against llm-integrated applications. arXiv 2023, arXiv:2306.05499. [Google Scholar]
  8. Zhan, Q.; Liang, Z.; Ying, Z.; Kang, D. Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents. arXiv 2024, arXiv:2403.02691. [Google Scholar]
  9. Shi, J.; Yuan, Z.; Tie, G.; Zhou, P.; Gong, N. Z.; Sun, L. Prompt Injection Attack to Tool Selection in LLM Agents. arXiv 2025, arXiv:2504.19793. [Google Scholar]
  10. Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; Fredrikson, M. Agentharm: A benchmark for measuring harmfulness of llm agents. arXiv 2024, arXiv:2410.09024. [Google Scholar]
  11. Yuan, T.; He, Z.; Dong, L.; Wang, Y.; Zhao, R.; Xia, T.; Xu, L.; Zhou, B.; Li, F.; Zhang, Z. R-judge: Benchmarking safety risk awareness for llm agents. arXiv 2024, arXiv:2401.10019. [Google Scholar]
  12. Huang, D.; Zhao, N.; Liu, W.; Sang, N. Target-Aware Augmentation for Rare-Event Prediction under Tabular Covariate Shift. 2026. [Google Scholar] [CrossRef]
  13. Yi-Ge, E.; Wu, M.; Chen, Z. Reducing divergence in batch normalization for domain adaptation. Proc. AAAI Conf. Artif. Intell. 2025, vol. 39(no. 21), 22155–22163. [Google Scholar] [CrossRef]
  14. Sun, S.; Wang, Y.; Yan, R.; Wang, J.; Lu, Y. CurateDistill: Efficient Data Selection via Active Data Curation and Dataset Distillation. 2026. [Google Scholar] [CrossRef]
  15. Yi-Ge, E.; Shawn, L. FlexDataset: Crafting annotated dataset generation for diverse applications. Proc. AAAI Conf. Artif. Intell. 2025, vol. 39(no. 9), 9481–9489. [Google Scholar] [CrossRef]
  16. Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv 2024, arXiv:2402.04249. [Google Scholar]
  17. Shi, T.; He, J.; Wang, Z.; Li, H.; Wu, L.; Guo, W.; Song, D. Progent: Programmable privilege control for llm agents. arXiv 2025, arXiv:2504.11703. [Google Scholar]
  18. Zhu, J.; Tseng, K.; Vernik, G.; Huang, X.; Patil, S. G.; Fang, V.; Popa, R. A. MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents. arXiv 2025, arXiv:2512.11147. [Google Scholar]
  19. Xiang, Z.; Zheng, L.; Li, Y.; Hong, J.; Li, Q.; Xie, H.; Zhang, J.; Xiong, Z.; Xie, C.; Yang, C. Guardagent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning. arXiv 2024, arXiv:2406.09187. [Google Scholar]
  20. Zheng, L.; Chen, J.; Hu, S.; Jin, Y.; Duan, Y. Evidence-Grounded Generation of Audit Procedures for Financial Statement Risk Assessment. 2026. [Google Scholar] [CrossRef]
  21. Wang, M.; Peng, P. L.; Chen, J. Financial report summarization via structure-aware modeling and information fidelity constraints. Trans. Comput. Sci. Methods 2024, vol. 4(no. 7). [Google Scholar]
  22. Huang, C. Integrating attention mapping and legal prior knowledge for interpretable legal reasoning with transformer-based models. Preprints.org 2025. [Google Scholar] [CrossRef]
  23. Mao, L. Joint Optimization of Dual Retrieval and Adaptive Re-Ranking for Retrieval-Augmented Generation. Artif. Intell. Comput. Innov. 2024, vol. 4(no. 4). [Google Scholar]
  24. Huang, S. Y. Retrieval-Augmented Long-Text Reasoning with Semantic Conflict Detection and Cross-Paragraph Evidence Integration. Artif. Intell. Comput. Innov. 2024, vol. 4(no. 1). [Google Scholar]
  25. Ren, L.; Tang, Y.; Chen, B.; Tao, J.; Chen, F. Evidence-Verified Root Cause Localization for Microservice Incidents under Tool and Telemetry Noise. 2026 6th International Conference on Machine Learning and Intelligent Systems Engineering (MLISE), 2026; pp. 664–669. [Google Scholar]
  26. Hu, R.; Zheng, Y.; Fang, R.; Zhou, J. VeriTrail-RCA: Evidence-Verified Root Cause Localization for Microservice Backends. 2026 6th International Conference on Machine Learning and Intelligent Systems Engineering (MLISE), 2026; pp. 670–675. [Google Scholar]
  27. Kou, J.; Qing, Z.; Wei, R.; Xue, F.; Huang, W.; Zhuang, H. CausalRCA: Causal Graph-Augmented Retrieval for Root Cause Analysis in Microservice Systems. 2026 5th International Conference on Electronic Information Engineering, Big Data and Computer Technology (EIBDCT), 2026; pp. 440–444. [Google Scholar]
  28. Hu, S.; Zeng, M.; Zheng, L.; Shen, H. Research on Fault Prediction Method of Enterprise Data Pipeline Based on ETL Process Logs. 2026. [Google Scholar] [CrossRef]
  29. Xue, Y.; Wang, Y.; Zhu, A.; Sun, X.; Zhang, C. Joint Temporal-Structural Representation Learning for Distributed Fault Discrimination in Microservice Architectures. 2026 IEEE 8th International Conference on Communications, Information System and Computer Engineering (CISCE), 2026; pp. 1581–1586. [Google Scholar]
  30. Zhao, Y.; Liu, Y.; Chen, P.; Liu, Y.; Chai, X. Multimodal Proxy Embedding for Cold-Start Item Ranking. 2026. [Google Scholar] [CrossRef]
  31. Ke, Z. MCA-TAN: Cross-Modal Graph Neural Recommendation with Temporal Behavior and Social Network Fusion. Artif. Intell. Comput. Innov. 2025, vol. 5(no. 10). [Google Scholar]
  32. Mei, Y.; Ke, Y.; Li, S.; Cheng, Q. Attention-Based Representation Learning for XBRL-Guided Corporate Revenue Forecasting. 2026. [Google Scholar] [CrossRef]
  33. Su, C.; Jiang, L.; Pi, S. Sentiment classification of chinese railway review text based on multi-feature fusion gated recurrent unit. 2021 International Conference on Information Science, Parallel and Distributed Systems (ISPDS), 2021; pp. 197–202. [Google Scholar]
  34. Wang, J. Credit card fraud detection via hierarchical multi-source data fusion and dropout regularization. Trans. Comput. Sci. Methods 2025, vol. 5(no. 1). [Google Scholar]
  35. Wang, J. Multivariate time series forecasting and classification via gnn and transformer models. J. Comput. Technol. Softw. 2024, vol. 3(no. 9). [Google Scholar]
  36. Wang, Y. AI-Enhanced Distributed Time Series Modeling: Incremental Learning for Evolving Streaming Data. Trans. Comput. Sci. Methods 2024, vol. 4(no. 8). [Google Scholar]
  37. Wang, W.; Zhang, Y.; Pan, Z.; Ke, Z.; Sun, Y.; Yang, Y. Drift-Aware Agentic ETL for Schema-Evolving Data Lakes. 2026 International Conference on Generative Artificial Intelligence and Information Security (GAIIS), 2026; pp. 728–733. [Google Scholar]
  38. Ding, J.; Chai, X.; Zhang, B.; Lin, W. CauSAL-Rec: Causally-Consistent LLM Agent for Faithful and Controllable Explainable Recommendation. 2026 6th International Conference on Intelligent Communications and Computing (ICICC), 2026; pp. 500–505. [Google Scholar]
  39. Li, S.; Wang, Y.; Xing, Y.; Wang, M. Mitigating correlation bias in advertising recommendation via causal modeling and consistency-aware learning. In Proceedings of the 2025 6th International Conference on Computer Science and Management Technology, 2025; pp. 585–589. [Google Scholar]
  40. Liu, Z.; Jiang, R.; Wang, F.; Yang, K.; Ma, W. Policy-constrained LLM repair for Kubernetes backend configurations. 2026 7th International Conference on Artificial Intelligence and Electromechanical Automation (AIEA), 2026; pp. 1129–1134. [Google Scholar]
  41. Ren, L.; Liu, Y.; Zhao, Y. Learning-Based Intelligent Agents for Backend Resource Scheduling and Operational Decision Making. Trans. Comput. Sci. Methods 2025, vol. 5(no. 3). [Google Scholar]
  42. Lin, Y.; Zhao, N.; Shen, H.; Zeng, M. Decision-Focused Hierarchical Forecasting with Differentiable Robust Optimization for Enterprise Budget-to-Revenue Allocation. 2026. [Google Scholar] [CrossRef]
  43. Ke, Z. Right-Sizing Cloud Data Warehouses Under Asymmetric Loss via One-Sided Conformal Prediction. IEEE Access 2026, vol. 14, 94748–94763. [Google Scholar] [CrossRef]
  44. Wang, Z.; Shu, S.; Wang, Y.; Lai, I. H.; Peng, C. C. Communication-Efficient Decentralized LLM Inference over Low-Bandwidth Distributed Nodes. 2026 8th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT), 2026; pp. 447–452. [Google Scholar]
  45. Shu, S.; Peng, C. C.; Wang, Z.; Wang, Y.; Lai, I. H. Transformer-aware adaptive gradient compression for distributed LLM fine-tuning. 2026 6th International Conference on Intelligent Communications and Computing (ICICC), 2026; pp. 291–295. [Google Scholar]
  46. He, C.; Gong, Y.; Ma, Y.; Zhang, B.; Wang, S. Subspace-Deconvolution Parameterization for Efficient Large Language Model Adaptation. 2026. [Google Scholar] [CrossRef]
  47. Jiang, J.; Hu, J.; Lyu, Y. Cost-Aware Cross-Cloud Federated Learning with Learned Client Routing and Aggregation Selection. 2026 8th International Conference on Internet of Things, Automation and Artificial Intelligence (IoTAAI), 2026; pp. 111–116. [Google Scholar]
  48. Li, S.; Chen, B.; Li, Y.; Wang, Z.; Xue, Y.; Xu, C. Privacy-Preserving Anomaly Detection in Cloud Services Using Hierarchical Federated Learning With Differential Privacy. 2026 9th International Symposium on Big Data and Applied Statistics (ISBDAS), 2026; pp. 668–672. [Google Scholar]
  49. Wang, S.; Zhang, B.; Gong, Y.; He, C. Budget-Aware Verifier-Guided Test-Time Reasoning for Efficient Large Language Models. 2026. [Google Scholar] [CrossRef]
  50. Feng, G.; Chen, J.; Luo, Z.; Yang, Q.; Wang, Y. AgentShield: A Lightweight Sentinel Framework for Mitigating Hallucination Propagation in Multi-Agent LLM Systems. 2026. [Google Scholar] [CrossRef]
Figure 1. Overview of the SafeAgent defense-in-depth framework. User tasks and external environment data pass through the Context Firewall for data/instruction separation via structural tagging and filtering. The Safety-Aware Planner integrates ReAct reasoning with safety checks for intent alignment and risk evaluation. The Tool-Call Sandbox enforces least-privilege access control with allowlists for Web Retrieval, Code Execution, and File Operations. An Output Validator performs anomaly detection before returning the final response.
Figure 1. Overview of the SafeAgent defense-in-depth framework. User tasks and external environment data pass through the Context Firewall for data/instruction separation via structural tagging and filtering. The Safety-Aware Planner integrates ReAct reasoning with safety checks for intent alignment and risk evaluation. The Tool-Call Sandbox enforces least-privilege access control with allowlists for Web Retrieval, Code Execution, and File Operations. An Output Validator performs anomaly detection before returning the final response.
Preprints 234354 g001
Figure 2. Safety-utility Pareto frontier among tested configurations. Each marker represents a defense configuration; the dashed line indicates the estimated frontier. SafeAgent dominates baseline approaches. Safety Score = 100 − ASR (%).
Figure 2. Safety-utility Pareto frontier among tested configurations. Each marker represents a defense configuration; the dashed line indicates the estimated frontier. SafeAgent dominates baseline approaches. Safety Score = 100 − ASR (%).
Preprints 234354 g002
Table 1. Performance comparison across defense configurations. BU: Benign Utility, UA: Utility Under Attack, ASR: Attack Success Rate. All values are percentages (%).
Table 1. Performance comparison across defense configurations. BU: Benign Utility, UA: Utility Under Attack, ASR: Attack Success Rate. All values are percentages (%).
Configuration BU UA ASR Safety
No Defense 78.5 45.2 67.2 32.8
Filter Only 75.2 52.3 47.1 52.9
Filter + Sandbox 71.8 58.6 32.5 67.5
SafeAgent (Full) 68.9 62.1 13.4 86.6
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.