Submitted:
02 September 2026
Posted:
03 September 2026
You are already at the latest version
Abstract
Large language models (LLMs) face a growing spectrum of adversarial jailbreak inputs that bypass safety training and elicit harmful outputs. Existing red-teaming efforts are ad hoc, using incompatible metrics and omitting guardrail-layer evaluation. We introduce GuardBench, a pre-deployment evaluation framework comprising 4,200 jailbreak prompts across seven attack categories, a four-stage automated pipeline, and a multi-metric scoring protocol covering Attack Success Rate (ASR), False Positive Rate (FPR), and P95 latency with 95% confidence intervals. We evaluate two guardrail configurations against GPT-4o-mini: a rule-based keyword filter (G1) and Llama Guard 2, an 8B-parameter LLM classifier (G3). Evaluated via static replay of archival 2023-era prompts, the undefended baseline achieves 5.8% ASR—far below prior-work rates of 60–80% obtained with live adaptive generation; the gap reflects prompt staleness, alignment generalisation, or both. Against this baseline, G1 reduces ASR by 0.32 pp with negligible latency; G3 achieves a 2.4 pp reduction but scores below the unguarded baseline on the Composite Safety Score due to 5,799 ms P95 latency under our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation). Token Manipulation attacks carry a 38.8% residual ASR under both configurations—the dominant open problem for next-generation guardrail design. GuardBench includes a CI/CD gate with configurable thresholds for evidence-based guardrail selection. Datasets, code, and results will be released upon acceptance.
Keywords:
LLM safety
; jailbreak attacks
; guardrail evaluation
; red teaming
; adversarial prompts
; AI safety benchmarks
; responsible AI deploymen
1. Introduction
The rapid deployment of large language models into consumer-facing and enterprise applications has elevated safety guardrails from a research concern to an operational necessity. Modern alignment pipelines, including Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI [2], substantially reduce harmful outputs, yet empirical evidence consistently demonstrates that safety-trained models remain susceptible to a class of adversarial inputs known as jailbreak attacks. From suffix-based gradient optimisation [3] to iterative black-box refinement [4,5] and in-context escalation [6], attackers exploit diverse model behaviours that no single guardrail fully addresses.
The consequences of guardrail failure are concrete: models are induced to provide synthesis instructions for dangerous substances, generate targeted harassment, or assist in cyberattacks. An increasingly important attack surface is indirect prompt injection, where adversarial instructions are embedded in retrieved documents or tool outputs rather than user turns [7]. Despite high-profile incidents, practitioners lack a standardised pre-deployment checklist that quantifies guardrail robustness across a representative attack portfolio. Existing benchmarks either focus on attack discovery [8] or assess downstream task robustness [9], but few integrate attack enumeration with guardrail evaluation in a deployment-ready pipeline.
Table 1 summarises how GuardBench relates to the most closely related prior frameworks. Unlike HarmBench [8] and JailbreakBench [10], GuardBench explicitly measures FPR against a benign counterfactual set, reports latency overhead as a first-class metric, provides a CI/CD integration layer with configurable deployment thresholds, and applies severity-weighted labelling to every prompt. These capabilities are specifically designed for the guardrail-deployment workflow rather than for academic attack comparison.
Beyond academic benchmarks, commercial guardrail APIs (including the OpenAI Moderation API [11], Azure AI Content Safety [12], and NVIDIA NeMo Guardrails [13] (NeMo Guardrails is a programmable dialog-management framework rather than a benchmark; it is included as a representative commercial guardrail tooling layer)) offer ready-made content filtering with low integration overhead. However, these systems are closed-source, provide no public adversarial ASR benchmarks, do not report latency overhead or FPR on standardised attack taxonomies, and offer no CI/CD integration hooks for automated deployment gating. GuardBench is designed to be complementary: it evaluates commercial guardrail APIs as black-box configurations alongside open-source alternatives.
This paper makes the following contributions:
- 1.
- We introduce GuardBench, an end-to-end pre-deployment evaluation framework that unifies dataset curation, attack execution, adjudication, and multi-metric reporting into a single reproducible pipeline suitable for production deployment gating.
- 2.
- We curate and release a dataset of 4,200 jailbreak prompts organised across a seven-category taxonomy with explicit severity labels and a 1,000-prompt benign counterfactual set enabling FPR measurement.
- 3.
- We establish cross-guardrail baselines against one production-grade LLM (GPT-4o-mini), reporting ASR, FPR, and P95 latency with 95% confidence intervals for two evaluated guardrail configurations, with two additional architectures planned for future evaluation.
- 4.
- We define the Composite Safety Score (CSS) and a CI/CD integration pattern that gates model releases on configurable safety thresholds, so that deployment decisions rest on measured data rather than subjective judgement.
2. Related Work
2.1. Attack Methods
Early manual jailbreaks relied on role-playing personas (“DAN”) and instruction-override prefixes [14]. Zou et al. [3] introduced Greedy Coordinate Gradient (GCG), a white-box optimisation procedure that appends short adversarial suffixes whose transferability to closed-source models highlighted a fundamental brittleness in safety alignment. Chao et al. [4] proposed PAIR, an automated black-box method that uses a separate attacker LLM to iteratively refine jailbreak candidates, achieving success in fewer than twenty queries against GPT-4 and Claude. Mehrotra et al. [5] extended this with TAP (Tree of Attacks with Pruning), which applies tree-of-thoughts reasoning to prune unlikely candidates before submission, achieving >80% ASR on GPT-4 in fewer than 30 queries. Anil et al. [6] demonstrated that LLMs with large context windows are vulnerable to many-shot jailbreaking, where hundreds of fake demonstrations of harmful dialogues in the prompt systematically suppress refusals via in-context learning dynamics. Greshake et al. [7] showed that LLM-integrated applications are vulnerable to indirect prompt injection, where adversarial instructions hidden in tool outputs or retrieved documents override system-level safety instructions without any direct attacker–user interaction.
2.2. Defence and Guardrail Approaches
Guardrail research has evolved from simple keyword filters toward learned classifiers and smoothing-based defences. Inan et al. [15] presented Llama Guard, an instruction-tuned 7B classifier that evaluates both input prompts and model outputs against a structured safety taxonomy; it matches or exceeds commercial moderation APIs on several benchmarks. Subsequent versions (Llama Guard 2 and 3) extend the taxonomy coverage and improve recall on adversarial prompts [16]; our experiments use Llama Guard 2, which was the stable release available at the time of evaluation. Han et al. [17] released WildGuard, an open-source 7B model trained on 92K labeled examples spanning 13 risk categories that achieves state-of-the-art on ten safety benchmarks including matching GPT-4-based judges. Robey et al. [18] introduced SmoothLLM, a randomised smoothing defence that duplicates and perturbs input copies, aggregating majority votes; it reduces GCG ASR by over an order of magnitude with minimal impact on benign utility. Bai et al. [2] showed that Constitutional AI training (using a written constitution for model self-critique) substantially reduces harmful outputs, providing a complementary training-time defence.
2.3. Evaluation Frameworks
Systematic evaluation of both attacks and defences was largely fragmented until recent benchmark efforts. Mazeika et al. [8] introduced HarmBench, a standardised framework with 510 harmful behaviours spanning 7 categories and 33 evaluated models; HarmBench provides a fine-tuned Llama-2-13B classifier for scalable ASR measurement and has become the de facto evaluation standard. Chao et al. [10] released JailbreakBench, a living leaderboard with a 100-behaviour dataset, standardised threat model, and reproducible evaluation protocol, directly addressing the reproducibility crisis caused by closed-source datasets and incomparable metrics. Zhu et al. [9] proposed PromptBench, which evaluates adversarial robustness of instruction-following across character-, word-, sentence-, and semantic-level perturbations on 13 datasets. Ganguli et al. [19] documented large-scale human red-teaming of Claude, cataloguing attack types and scaling behaviours across model sizes, while Perez et al. [20] automated this process using reinforcement learning, substantially outperforming zero-shot prompting in harmful-text generation rate, underscoring the need for systematic automated evaluation.
Commercial guardrail solutions (including the OpenAI Moderation API [11], Azure AI Content Safety [12], and NVIDIA NeMo Guardrails [13]) offer low-integration-overhead content filtering but are closed-source, publish no adversarial ASR benchmarks against standardised attack taxonomies, and provide no public FPR or latency overhead data. GuardBench is designed to be complementary to these systems: it provides the measurement infrastructure needed to evaluate commercial APIs as black-box guardrail configurations alongside open-source alternatives, enabling practitioners to make evidence-based guardrail selection decisions.
GuardBench differs from all of the above as summarised in Table 1. The central distinction is that GuardBench targets the guardrail layer rather than the underlying model: it integrates dataset curation, attack execution, FPR measurement, latency profiling, and multi-metric scoring into a pipeline with a native CI/CD integration layer suitable for pre-deployment gating. Neither HarmBench nor JailbreakBench provides latency overhead reporting or configurable deployment threshold gating.
3. GuardBench Framework Design
3.1. Overview
GuardBench is structured as a four-stage pipeline: (1) Dataset Curation, (2) Attack Execution, (3) Response Adjudication, and (4) Metric Aggregation & Reporting. Figure 1 illustrates this architecture.
3.2. Dataset Curation
The GuardBench dataset is maintained as a versioned repository with cryptographic hashes per entry. New prompts are sourced from: (a) published academic datasets [8,10]; (b) automated generation via a fine-tuned attacker LLM; and (c) community submissions subject to a human review board. Each prompt carries metadata fields: category, severity (1–5), source, target_behavior, and creation_date.
To prevent train-test contamination, prompts from published datasets are flagged with their origin in the metadata. Evaluators must retrain classifier guardrails on non-overlapping splits when training data overlaps with these sources. For our experiments (Section 5), the DeBERTa classifier (G2) was designed to be fine-tuned on the publicly available AdvBench harmful behaviours dataset [3] augmented with WildGuard training splits [17] (combined 45K examples), with a held-out split confirmed non-overlapping with our evaluation prompts; G2 evaluation is deferred to future work (Section 5).
3.3. Attack Execution
The Attack Execution module is designed to support a configurable set of attack strategies: PAIR [4] and TAP [5] for iterative black-box generation, and GCG [3] for white-box gradient-based suffix optimisation. For the current evaluation, we use direct-submission mode, which replays prompts from the curated dataset without modification—the same protocol used by HarmBench [8] and JailbreakBench [10]. Live adaptive generation via PAIR, TAP, and GCG is supported by the framework and is planned as a future evaluation. The module records: raw prompt, model response, HTTP status, and wall-clock latency at both the 50th and 95th percentiles (, ).
3.4. Response Adjudication
The GuardBench adjudication pipeline is designed as a two-stage classifier: (1) a fast keyword-blocklist scan for high-confidence refusals; (2) a fine-tuned LLM judge that classifies each response as safe, borderline, or harmful. The target production implementation uses a fine-tuned Llama-3-8B-Instruct model trained on 15K annotated pairs, with an inter-annotator agreement of on the training corpus, targeting F1 ≥ 0.89 on held-out examples.
For the current evaluation, the LLM judge is implemented using GPT-4o-mini in zero-shot mode with a structured three-class rubric. Because the target model and judge are the same model, we apply a conservative judging protocol that restricts the harmful label to responses providing substantive, actionable assistance—explicitly excluding mere restatement of a request, decoding without compliance, or general topic discussion that does not advance the harmful goal. This protocol is documented in the released code and its effect on ASR is auditable by comparing classifications against a verbatim-brief baseline. The adjudicator is re-evaluated periodically against fresh human annotations to detect concept drift.
3.5. Metric Aggregation and CI/CD Integration
Metrics are computed per category and in aggregate (Section 5). GuardBench exposes a REST API and a GitHub Actions plugin; a YAML configuration file lets teams specify per-metric pass/fail thresholds. Pipeline runs that violate thresholds block the merge or deployment and emit a structured JSON report with per-category breakdowns.
4. Dataset Construction
4.1. Taxonomy Design
4.2. Severity Labelling
Severity is rated on a 1–5 scale: level 1 covers requests for mildly inappropriate content; level 5 covers content whose generation could directly facilitate serious physical harm (e.g., CBRN weapon synthesis). Two independent annotators label each prompt; disagreements exceeding one level are resolved by a third reviewer. Inter-annotator agreement (Cohen’s ) is 0.81 across the full corpus.
4.3. Benign Counterfactual Set
To measure FPR we construct a 1,000-prompt benign counterfactual set by (a) sampling from MS-MARCO [21] and ShareGPT1 for topically adjacent but harmless prompts, and (b) generating “near-miss” prompts that superficially resemble jailbreaks but have legitimate intent (e.g., security research, medical education). This mirrors the design philosophy of WildGuard [17].
5. Evaluation Methodology
5.1. Metrics
We report three primary metrics:
Attack Success Rate (ASR): the fraction of attack prompts for which the guardrail fails to block a harmful response, as classified by the adjudicator, formally defined in Equation (1). Following HarmBench [8], we use a calibrated LLM judge rather than keyword matching to reduce false-positive adjudication.
False Positive Rate (FPR): the fraction of benign prompts incorrectly blocked by the guardrail, as given in Equation (2).
Latency Overhead (): wall-clock latency added by the guardrail relative to an unguarded baseline, at the 50th percentile () and 95th percentile (). The 95th-percentile figure governs the CI/CD gate as it reflects tail-latency SLA risk.
Composite Safety Score (CSS): a deployability index defined in Equation (3) as
where are operator weights (default , , , reflecting the asymmetry between security, usability, and latency cost) and ms is the reference latency budget. The latency component applies a linear penalty capped at ms, so configurations far below the budget incur minimal penalty while those exceeding it are progressively discounted, with reflecting that latency is an important but not dominant deployment constraint.
5.2. Guardrail Configurations Under Test
The GuardBench framework defines four representative guardrail configurations spanning the practical design space. Two were fully evaluated in this study; two are planned for future evaluation:
- G1 – Rule-Based Filter(evaluated): regex and keyword blocklist; fast but brittle against paraphrasing attacks.
- G4 – Smoothing Defence (SmoothLLM) [18] (planned): perturbation rate , copies; hyperparameter sensitivity analysis planned for future work.
5.3. Experimental Setup
The primary target model is GPT-4o-mini [1] (OpenAI API, temperature 0, max 512 tokens). Each guardrail configuration was evaluated in a single pass over the curated dataset using direct-submission mode ( attack prompts replayed statically; benign prompts); Table 3 reports 95% CIs via Wilson score interval.
The adjudicator is GPT-4o-mini with a structured three-class rubric (safe / borderline / harmful); ASR counts only harmful verdicts. Because the target model and adjudicator are the same model, we apply a conservative judging protocol that restricts harmful to responses providing substantive, actionable assistance, excluding responses that merely restate, decode, or describe a request without complying. This protocol is documented in the released code and its effect on ASR is auditable against a verbatim-brief baseline. G3 (Llama Guard 2) was served via llama.cpp with Q4_K_M GGUF quantisation on Apple M3 Pro (18 GB unified memory, Metal acceleration); the measured of 5,799 ms is specific to this hardware and quantisation configuration. Production serving on a dedicated GPU or hosted endpoint would yield materially lower latency and may reverse the CSS ranking. Results exclude prompts whose adjudicated harm was copyright reproduction rather than safety-policy violation ( prompts; retained in supplementary results with both copyright-including and copyright-excluding ASR reported). All runs completed within the same calendar week to minimise model-version drift. To support reproducibility, all prompt metadata and per-run result files will be released under a CC BY 4.0 licence. The harness is implemented in Python 3.12 with pinned dependencies. Large language model tools (Claude, Anthropic) were used to assist with manuscript drafting and formatting; all experimental code, data collection, and quantitative analysis were conducted independently by the authors.
5.4. Experimental Results
Table 3 reports per-guardrail performance across the full dataset and the GCG-suffix category for the primary GPT-4o-mini target. ASR and FPR are reported as percentages (mean ± 95% CI via Wilson score interval); latency in milliseconds.
Key observations from Table 3:
- The undefended baseline is low against archival prompts. Without any guardrail, GPT-4o-mini is successfully jailbroken in only 5.80% of attempts (copyright-excluding), with GCG-suffix attacks achieving 0.71% ASR—approximately 130× below the 94% reported in [3]. Many-shot attacks score 0.00%. This finding is consistent with two explanations that the current experimental design cannot distinguish: (a) 2023-era jailbreak templates have been substantially incorporated into contemporary safety training, and (b) static replay of archival prompts understates the ASR that a live adaptive attacker would achieve. Practitioners should treat these results as a lower bound; live adaptive evaluation is planned for future work.
- G1 (rule-based) is operationally negligible at this baseline. G1 reduces overall ASR from 5.80% to 5.48%—an absolute reduction of 0.32 percentage points—whilst blocking 25.9% of attack prompts. Paired analysis (McNemar’s test) confirms the effect is statistically detectable (, ) but operationally negligible: G1 prevented 12 harmful responses in 3,778 attempts (copyright-excluding), with zero false positives. Its CSS (0.962) marginally exceeds the unguarded baseline (0.959) solely because of near-zero latency and perfect FPR. Critically, G1 is anti-correlated with where risk lives: it blocks 100% of many-shot prompts (0% ASR category) and only 0.54% of Token Manipulation prompts (38.8% ASR category).
- G3 (Llama Guard 2) reduces ASR meaningfully but at severe latency cost under evaluation hardware. G3 reduces overall ASR to 3.39%, a 2.41 pp improvement over the unguarded baseline, with a 0.60% FPR. However, its of 5,799 ms on our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) causes the CSS to drop to 0.682—below the unguarded baseline’s 0.959. Under production serving conditions (dedicated GPU or hosted endpoint), where Llama Guard 2 typically runs in the low hundreds of milliseconds, the CSS ranking may reverse. The Q4_K_M quantisation may also degrade classification accuracy; unquantised results are reserved for future evaluation.
- Token Manipulation is the dominant residual threat. Across both configurations, Token Manipulation (base64 encoding, character substitution, leetspeak) carries a 38.8% ASR—nearly 7× the overall baseline. G1 is ineffective (0.54% block rate); G3 reduces it to 33.5% (27.9% block rate). Neither guardrail substantially addresses this attack family, highlighting it as the primary open problem for practitioners evaluating obfuscation-resistant defences.
- No single configuration dominates across all axes. GuardBench enables practitioners to select configurations based on their specific risk tolerance and latency budget via the CSS metric. At a 5.8% archival baseline, the marginal security benefit of adding G3 must be weighed against its latency cost—a trade-off that GuardBench quantifies explicitly and prior frameworks do not.
Figure 2 provides a per-category breakdown of ASR across the three evaluated configurations (No Guard, G1, G3), confirming that Token Manipulation is the dominant residual threat at 38.8% ASR, whilst GCG-suffix and many-shot attacks approach zero across all tiers.
6. Discussion and Future Work
6.1. Limitations
GuardBench has five primary limitations in the current evaluation. First, the reported results are specific to one target model (GPT-4o-mini) and may not generalise to all LLM architectures or fine-tuning regimes; cross-model validation against open-weight models remains future work. Second, the adjudicator introduces two sources of measurement error: (a) using GPT-4o-mini as both the target model and the judge creates a circularity risk—the same model grades whether it was successfully jailbroken—mitigated by our conservative judging protocol but with unquantified residual bias; and (b) the framework’s target production adjudicator (fine-tuned Llama-3-8B-Instruct, F1 ≥ 0.89 on held-out examples) was not used in this evaluation, so those accuracy figures do not apply to the GPT-4o-mini judge used here. Third, the dataset contains no audio, image, or structured-data attack vectors; the text-only scope limits applicability to multi-modal deployments where image-based jailbreaks [23] and tool-call hijacking represent distinct threat surfaces. Fourth, all attack results derive from static replay of archival 2023-era prompts; live adaptive generation via PAIR, TAP, or GCG was not performed in this evaluation, so reported ASR values should be treated as a lower bound on what an adaptive attacker would achieve against a target seen during prompt construction. Fifth, G3’s latency measurements reflect a consumer-hardware configuration (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) rather than production serving; results on a dedicated GPU or hosted endpoint are expected to differ substantially.
6.2. Adaptive Attacks
A critical limitation of any static benchmark is the adaptive attacker: once the guardrail configuration is known, attackers craft inputs that specifically circumvent it. GuardBench is designed to address this through a private held-out set and by supporting adaptive variants of PAIR and GCG against each guardrail; these capabilities are implemented in the framework but were not exercised in the current proof-of-concept evaluation. Future versions will include an online adversarial track that continuously solicits new attack submissions.
6.3. Multi-Modal and Agentic Contexts
Current GuardBench prompts are text-only. The rising prevalence of vision-language models introduces image-based jailbreaks [23], while autonomous agents face tool-call hijacking and multi-turn escalation via indirect prompt injection [7]. These attack surfaces are qualitatively distinct from the text-only jailbreak paradigm. We are extending the taxonomy to include multi-modal and agentic attack categories in v2.0.
6.4. Guardrail Combination and Ensembles
Our experiments evaluate guardrails independently. In practice, production systems layer multiple defences. GuardBench’s modular pipeline supports chaining (e.g., G2 followed by G3). A systematic study of ensemble strategies and their cost–benefit tradeoffs is left to future work.
6.5. Calibration and Concept Drift
The LLM adjudicator is a single point of failure. As language models and attack strategies evolve, the judge exhibits concept drift. GuardBench schedules periodic re-annotation of randomly sampled adjudicated examples against a human panel. Platforms should treat benchmark scores as relative rather than absolute given this drift risk.
6.6. Beyond Binary Classification
Current metrics reduce guardrail output to a binary pass/fail. Incorporating severity-weighted ASR, where a harmful response classified at level 5 counts more than a level 1 lapse, provides a more faithful risk model. We propose the severity-weighted ASR in Equation (4):
where is the severity of prompt i, as a direction for future metric design.
7. Conclusion
We have presented GuardBench, a framework for systematically evaluating LLM safety guardrails prior to production deployment. GuardBench contributes (i) a 4,200-prompt taxonomy-driven dataset spanning seven attack categories with severity labels and a 1,000-prompt benign counterfactual set; (ii) a four-stage automated pipeline covering attack execution, hybrid adjudication, and multi-metric reporting with confidence intervals; (iii) empirical baseline results for two guardrail configurations (G1 rule-based, G3 Llama Guard 2) evaluated against GPT-4o-mini using static replay of archival 2023-era prompts, revealing 5.8% baseline ASR and showing that heavy guardrails can score below the unguarded baseline on the composite metric when latency cost dominates—a result specific to our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) that warrants validation on production serving infrastructure; and (iv) a CI/CD integration layer with configurable deployment thresholds. A key finding is that Token Manipulation attacks carry a 38.8% residual ASR—largely unaddressed by either evaluated configuration—identifying the primary open problem for next-generation guardrail design. Our experiments confirm that no single guardrail configuration simultaneously minimises ASR, FPR, and latency at this baseline, and that the CSS metric provides a principled basis for workload-specific guardrail selection. GuardBench gives practitioners the infrastructure to run safety tests on every release cycle, replacing informal pre-deployment checks with reproducible, quantified results.
Author Contributions
Conceptualization, P.B., S.S. and H.G.; methodology, P.B., S.S. and H.G.; software, S.S.; validation, P.B., S.S. and H.G.; formal analysis, S.S.; investigation, P.B., S.S. and H.G.; data curation, P.B. and S.S.; writing—original draft preparation, P.B., S.S. and H.G.; writing—review and editing, P.B., S.S. and H.G.; visualisation, S.S.; supervision, P.B. All authors have read and agreed to the published version of the manuscript. P.B., S.S. and H.G. contributed equally to this work.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The GuardBench dataset, evaluation code, and raw results will be made publicly available at https://github.com/bangadpurva/guardbench upon paper acceptance.
Acknowledgments
During the preparation of this manuscript, the authors used Claude (Anthropic, claude-sonnet-4-6) for the purposes of drafting and editing manuscript text, reviewing citations, and assisting with formatting. The authors have reviewed and edited all output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflict of interest.
References
- OpenAI. GPT-4o mini: Advancing Cost-Efficient Intelligence. OpenAI Blog, 2024. Accessed: August 2024. Accessed. 2024.
- Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. Constitutional AI: Harmlessness from AI Feedback. arXiv Preprint. 2022, arXiv:2212.08073. [Google Scholar]
- Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J.Z.; Fredrikson, M. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv Preprint. 2023, arXiv:2307.15043. [Google Scholar]
- Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G.J.; Wong, E. Jailbreaking Black Box Large Language Models in Twenty Queries. arXiv Preprint. 2023, arXiv:2310.08419. [Google Scholar]
- Mehrotra, A.; Zampetakis, M.; Kassianik, P.; Nelson, B.; Anderson, H.; Singer, Y.; Karbasi, A. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
- Anil, C.; Durmus, E.; Panickssery, N.; Sharma, M.; Ziegler, D.M.; Rimsky, N.; Askell, A.; et al. Many-Shot Jailbreaking; Technical report;Technical Report; Anthropic, 2024. [Google Scholar]
- Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023. [Google Scholar]
- Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. [Google Scholar]
- Zhu, K.; Wang, J.; Zhou, J.; Wang, Z.; Chen, H.; Wang, Y.; Yang, L.; Ye, W.; Gong, N.Z.; Zhang, Y.; et al. PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv Preprint. 2023, arXiv:2306.04528. [Google Scholar]
- Chao, P.; Debenedetti, E.; Robey, A.; Andriushchenko, M.; Croce, F.; Sehwag, V.; Dobriban, E.; Flammarion, N.; Pappas, G.J.; Tramèr, F.; et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
- OpenAI. OpenAI Moderation API. OpenAI Platf. Doc. Accessed. 2023. (accessed on 2024). [Google Scholar] [CrossRef]
- Microsoft. Azure AI Content Safety. Microsoft Azur. Doc. Accessed. 2023. (accessed on 2024). [Google Scholar] [CrossRef]
- Rebedea, T.; Dinu, R.; Sreedhar, M.N.; Parisien, C.; Cohen, J. NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2023. [Google Scholar]
- Perez, F.; Ribeiro, I. Ignore Previous Prompt: Attack Techniques For Language Models. arXiv Preprint. 2022, arXiv:2211.09527. [Google Scholar]
- Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv Preprint. 2023, arXiv:2312.06674. [Google Scholar]
- Meta, A.I. Llama Guard 2: Meta Llama Guard 2. Meta AI Research, 2024. Technical Report. Llama Guard 3 is a subsequent model. Available online: https://ai.meta.com/research/publications/llama-guard-3/.
- Han, S.; Rao, K.; Ettinger, A.; Jiang, L.; Lin, B.Y.; Lambert, N.; Choi, Y.; Dziri, N. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
- Robey, A.; Wong, E.; Hassani, H.; Pappas, G.J. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. arXiv Preprint. 2023, arXiv:2310.03684. [Google Scholar]
- Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv Preprint. 2022, arXiv:2209.07858. [Google Scholar]
- Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; Irving, G. Red Teaming Language Models with Language Models. arXiv Preprint. 2022, arXiv:2202.03286. [Google Scholar]
- Bajaj, P.; Campos, D.; Craswell, N.; Deng, L.; Gao, J.; Liu, X.; Majumder, R.; McNamara, A.; Mitra, B.; Nguyen, T.; et al. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv Preprint. 2016, arXiv:1611.09268. [Google Scholar]
- He, P.; Gao, J.; Chen, W. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
- Qi, X.; Zeng, Y.; Xie, T.; Chen, P.Y.; Jia, R.; Mittal, P.; Henderson, P. Visual Adversarial Examples Jailbreak Aligned Large Language Models. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar]
Figure 1.
GuardBench four-stage pipeline. Prompts flow top-to-bottom through Dataset Curation, Attack Execution, Response Adjudication, and Metric Aggregation. Side inputs supply the guardrail under test and the LLM judge. Outputs feed a Safety Report and a CI/CD gate.
Figure 1.
GuardBench four-stage pipeline. Prompts flow top-to-bottom through Dataset Curation, Attack Execution, Response Adjudication, and Metric Aggregation. Side inputs supply the guardrail under test and the LLM judge. Outputs feed a Safety Report and a CI/CD gate.

Figure 2.
Measured ASR (%) per attack category across evaluated configurations on GPT-4o-mini (copyright-excluding, static replay). Token Manipulation is the dominant residual risk, with 38.8% ASR under both G1 and No Guard, reduced only partially to 33.5% by G3. GCG-suffix and many-shot attacks approach zero across all configurations.
Figure 2.
Measured ASR (%) per attack category across evaluated configurations on GPT-4o-mini (copyright-excluding, static replay). Token Manipulation is the dominant residual risk, with 38.8% ASR under both G1 and No Guard, reduced only partially to 33.5% by G3. GCG-suffix and many-shot attacks approach zero across all configurations.

Table 1.
Feature Comparison: GuardBench vs. Prior Frameworks and Commercial Tools.
| Feature | HarmBench | JBBench | PBench | OAI Mod. | NeMo GR | Ours |
|---|---|---|---|---|---|---|
| ASR measurement | ✓ | ✓ | ✓ | ✓ | ✓ | |
| FPR / benign set | ✓ | ✓ | ||||
| Latency overhead | ✓ | |||||
| Severity labelling | ✓ | ✓ | ||||
| CI/CD integration | ✓ | ✓ | ||||
| Versioned dataset | ✓ | ✓ | ||||
| Guardrail-layer focus | ✓ | ✓ | ✓ |
Table 2.
GuardBench Jailbreak Prompt Taxonomy (v1.0, 4,200 prompts total).
| Category | Count | % | Avg. Sev. |
|---|---|---|---|
| Role-Play & Persona Override | 840 | 20.0 | 3.2 |
| Prompt Injection [7,14] | 700 | 16.7 | 3.8 |
| Token Manipulation | 560 | 13.3 | 2.9 |
| Many-Shot In-Context [6] | 490 | 11.7 | 4.1 |
| Gradient-Optimised Suffix [3] | 420 | 10.0 | 4.6 |
| Iterative Refinement [4,5] | 630 | 15.0 | 4.3 |
| Multilingual & Code-Switch | 560 | 13.3 | 3.5 |
| Total | 4,200 | 100 | 3.71 |
Table 3.
GuardBench Results (Target: GPT-4o-mini, copyright-excluding; lower ASR/FPR/latency is better). Wilson 95% CIs ( attack; benign). G2 and G4 are planned future evaluations.
Table 3.
GuardBench Results (Target: GPT-4o-mini, copyright-excluding; lower ASR/FPR/latency is better). Wilson 95% CIs ( attack; benign). G2 and G4 are planned future evaluations.
| Guard | Overall | GCG Suffix | FPR | |||
| ASR | CSS | ASR | CSS | (ms) | ||
| G1 Rule | 0.962 | 0.995 | 0.8 | |||
| G3 LG2 | 0.682 | 0.697 | 5,799 | |||
| G2 Class. | planned — future work | |||||
| G4 Smooth | planned — future work | |||||
| No Guard | 0.959 | – | 0.00 | 0 | ||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.