Submitted:
25 August 2025
Posted:
26 August 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Existing Evaluation Metrics
2.2. Existing Frameworks and Benchmarks
3. A Balanced Evaluation Framework for Agentic AI
3.1. Framework Overview
- Robustness&Adaptability — measures the agent’s resilience to changing conditions, adversarial inputs and unexpected events. Metrics include success rate under noisy inputs, recovery time from failures, ability to adapt to new goals and resilience to adversarial examples [3].
- Safety&Ethics — evaluates whether the agent avoids harmful actions, mitigates biases and adheres to ethical norms. Metrics encompass hallucination rate, harmful-content generation, fairness scores and compliance with regulatory requirements [4]. A harm-reduction index integrates hallucination, toxicity and fairness measures.
- Human-Centred Interaction — captures how users perceive and interact with the agent. Metrics include user satisfaction (CSAT/NPS), trust scores, transparency, explainability and cognitive load. Human-in-the-loop assessments use instruments like TrAAIT to measure trust [6].
- Economic&Sustainability Impact — examines cost–benefit trade-offs and long-term sustainability of deployment. Metrics include productivity gain, return on investment, carbon footprint of compute resources and alignment with organisational goals. This axis addresses the economic dimension often overlooked in technical benchmarks [2].
3.2. Measuring Complex Behaviours
3.3. Visualising the Framework

4. Proposed Experiments and Validation
4.1. Goal-Drift Evaluation
4.2. Robustness Under Perturbations
4.3. Safety and Ethical Auditing
4.4. Human-Centred Trust Evaluation
4.5. Economic and Sustainability Analysis
5. Case Studies: Agentic AI in the Wild
5.1. Legacy Application Modernisation
5.2. Market-Research Data Quality and Insight Generation
5.3. Credit-Risk Memo Generation
5.4. Cross-Case Analysis
6. Implications and Future Directions
6.1. Towards Balanced Benchmarks
6.2. Reproducibility and Open Evaluation
6.3. Human-Agent Collaboration and Trust
6.4. Policy and Governance
7. Conclusions
Acknowledgments
References
- R. Sapkota, G. Tambwekar, A. Crespi, A. Ramachandran and H. Long. “AI Agents vs Agentic AI: A Conceptual Taxonomy, Applications and Challenges.” arXiv preprint arXiv:2505.10468, 2025. [CrossRef]
- K. J. Meimandi, N. Arsenlis, S. Kalamkar and A. Talwalkar. “The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims.” arXiv preprint arXiv:2506.02064, 2025.
- C. Bronsdon. “AI Agent Evaluation: Methods, Challenges, and Best Practices.” Galileo, 2025. https://www.galileo.ai/blog/ai-agent-evaluation.
- QAwerk. “AI Agent Evaluation: Metrics That Actually Matter.” Blog, 2025. https://qawerk.com/blog/ai-agent-evaluation-metrics/.
- McKinsey&Company. “Seizing the Agentic AI Advantage: A CEO Playbook.” Report, 2025. https://www.mckinsey.com/featured-insights/mckinsey-technology-and-innovation/seizing-the-agentic-ai-advantage.
- A. F. Stevens, P. Stetson and colleagues. “Theory of trust and acceptance of artificial intelligence technology (TrAAIT): An instrument to assess clinician trust and acceptance of artificial intelligence.” Journal of Biomedical Informatics, 148:104550, 2023. (Cited for the TrAAIT model and trust instrument). [CrossRef]
- R. Arike, E. Donoway, H. Bartsch and M. Hobbhahn. “Technical Report: Evaluating Goal Drift in Language Model Agents.” arXiv preprint arXiv:2505.02709, 2025. (Cited for the goal-drift evaluation methodology.).
- C. Dilmegani. “Large Language Model Evaluation in 2025: 10+ Metrics & Methods.” AIMultiple, 2025. https://research.aimultiple.com/large-language-model-evaluation/. (Cited for the need to combine automated metrics with human and fairness evaluations.
| Case | Agentic approach | Reported impact |
|---|---|---|
| Legacy modernisation | Humans supervise squads of agents to document, code, review and integrate features | reduction in time/effort |
| Data quality & insights | Agents detect anomalies, analyse internal/external signals and synthesise drivers | productivity gain; annual savings |
| Credit-risk memos | Agents extract data, draft sections, generate confidence scores; humans supervise | 20–60 % productivity; 30 % faster decisions |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).