Preprint
Article

This version is not peer-reviewed.

GuardBench: A Systematic Framework for Pre-Deployment Evaluation of LLM Safety Guardrails

   † These authors contributed equally to this work.

Submitted:

02 September 2026

Posted:

03 September 2026

You are already at the latest version

Abstract
Large language models (LLMs) face a growing spectrum of adversarial jailbreak inputs that bypass safety training and elicit harmful outputs. Existing red-teaming efforts are ad hoc, using incompatible metrics and omitting guardrail-layer evaluation. We introduce GuardBench, a pre-deployment evaluation framework comprising 4,200 jailbreak prompts across seven attack categories, a four-stage automated pipeline, and a multi-metric scoring protocol covering Attack Success Rate (ASR), False Positive Rate (FPR), and P95 latency with 95% confidence intervals. We evaluate two guardrail configurations against GPT-4o-mini: a rule-based keyword filter (G1) and Llama Guard 2, an 8B-parameter LLM classifier (G3). Evaluated via static replay of archival 2023-era prompts, the undefended baseline achieves 5.8% ASR—far below prior-work rates of 60–80% obtained with live adaptive generation; the gap reflects prompt staleness, alignment generalisation, or both. Against this baseline, G1 reduces ASR by 0.32 pp with negligible latency; G3 achieves a 2.4 pp reduction but scores below the unguarded baseline on the Composite Safety Score due to 5,799 ms P95 latency under our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation). Token Manipulation attacks carry a 38.8% residual ASR under both configurations—the dominant open problem for next-generation guardrail design. GuardBench includes a CI/CD gate with configurable thresholds for evidence-based guardrail selection. Datasets, code, and results will be released upon acceptance.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

The rapid deployment of large language models into consumer-facing and enterprise applications has elevated safety guardrails from a research concern to an operational necessity. Modern alignment pipelines, including Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI [2], substantially reduce harmful outputs, yet empirical evidence consistently demonstrates that safety-trained models remain susceptible to a class of adversarial inputs known as jailbreak attacks. From suffix-based gradient optimisation [3] to iterative black-box refinement [4,5] and in-context escalation [6], attackers exploit diverse model behaviours that no single guardrail fully addresses.
The consequences of guardrail failure are concrete: models are induced to provide synthesis instructions for dangerous substances, generate targeted harassment, or assist in cyberattacks. An increasingly important attack surface is indirect prompt injection, where adversarial instructions are embedded in retrieved documents or tool outputs rather than user turns [7]. Despite high-profile incidents, practitioners lack a standardised pre-deployment checklist that quantifies guardrail robustness across a representative attack portfolio. Existing benchmarks either focus on attack discovery [8] or assess downstream task robustness [9], but few integrate attack enumeration with guardrail evaluation in a deployment-ready pipeline.
Table 1 summarises how GuardBench relates to the most closely related prior frameworks. Unlike HarmBench [8] and JailbreakBench [10], GuardBench explicitly measures FPR against a benign counterfactual set, reports latency overhead as a first-class metric, provides a CI/CD integration layer with configurable deployment thresholds, and applies severity-weighted labelling to every prompt. These capabilities are specifically designed for the guardrail-deployment workflow rather than for academic attack comparison.
Beyond academic benchmarks, commercial guardrail APIs (including the OpenAI Moderation API [11], Azure AI Content Safety [12], and NVIDIA NeMo Guardrails [13] (NeMo Guardrails is a programmable dialog-management framework rather than a benchmark; it is included as a representative commercial guardrail tooling layer)) offer ready-made content filtering with low integration overhead. However, these systems are closed-source, provide no public adversarial ASR benchmarks, do not report latency overhead or FPR on standardised attack taxonomies, and offer no CI/CD integration hooks for automated deployment gating. GuardBench is designed to be complementary: it evaluates commercial guardrail APIs as black-box configurations alongside open-source alternatives.
This paper makes the following contributions:
1.
We introduce GuardBench, an end-to-end pre-deployment evaluation framework that unifies dataset curation, attack execution, adjudication, and multi-metric reporting into a single reproducible pipeline suitable for production deployment gating.
2.
We curate and release a dataset of 4,200 jailbreak prompts organised across a seven-category taxonomy with explicit severity labels and a 1,000-prompt benign counterfactual set enabling FPR measurement.
3.
We establish cross-guardrail baselines against one production-grade LLM (GPT-4o-mini), reporting ASR, FPR, and P95 latency with 95% confidence intervals for two evaluated guardrail configurations, with two additional architectures planned for future evaluation.
4.
We define the Composite Safety Score (CSS) and a CI/CD integration pattern that gates model releases on configurable safety thresholds, so that deployment decisions rest on measured data rather than subjective judgement.

3. GuardBench Framework Design

3.1. Overview

GuardBench is structured as a four-stage pipeline: (1) Dataset Curation, (2) Attack Execution, (3) Response Adjudication, and (4) Metric Aggregation & Reporting. Figure 1 illustrates this architecture.

3.2. Dataset Curation

The GuardBench dataset is maintained as a versioned repository with cryptographic hashes per entry. New prompts are sourced from: (a) published academic datasets [8,10]; (b) automated generation via a fine-tuned attacker LLM; and (c) community submissions subject to a human review board. Each prompt carries metadata fields: category, severity (1–5), source, target_behavior, and creation_date.
To prevent train-test contamination, prompts from published datasets are flagged with their origin in the metadata. Evaluators must retrain classifier guardrails on non-overlapping splits when training data overlaps with these sources. For our experiments (Section 5), the DeBERTa classifier (G2) was designed to be fine-tuned on the publicly available AdvBench harmful behaviours dataset [3] augmented with WildGuard training splits [17] (combined 45K examples), with a held-out split confirmed non-overlapping with our evaluation prompts; G2 evaluation is deferred to future work (Section 5).

3.3. Attack Execution

The Attack Execution module is designed to support a configurable set of attack strategies: PAIR [4] and TAP [5] for iterative black-box generation, and GCG [3] for white-box gradient-based suffix optimisation. For the current evaluation, we use direct-submission mode, which replays prompts from the curated dataset without modification—the same protocol used by HarmBench [8] and JailbreakBench [10]. Live adaptive generation via PAIR, TAP, and GCG is supported by the framework and is planned as a future evaluation. The module records: raw prompt, model response, HTTP status, and wall-clock latency at both the 50th and 95th percentiles ( Δ t 50 , Δ t 95 ).

3.4. Response Adjudication

The GuardBench adjudication pipeline is designed as a two-stage classifier: (1) a fast keyword-blocklist scan for high-confidence refusals; (2) a fine-tuned LLM judge that classifies each response as safe, borderline, or harmful. The target production implementation uses a fine-tuned Llama-3-8B-Instruct model trained on 15K annotated pairs, with an inter-annotator agreement of κ = 0.79 on the training corpus, targeting F1 ≥ 0.89 on held-out examples.
For the current evaluation, the LLM judge is implemented using GPT-4o-mini in zero-shot mode with a structured three-class rubric. Because the target model and judge are the same model, we apply a conservative judging protocol that restricts the harmful label to responses providing substantive, actionable assistance—explicitly excluding mere restatement of a request, decoding without compliance, or general topic discussion that does not advance the harmful goal. This protocol is documented in the released code and its effect on ASR is auditable by comparing classifications against a verbatim-brief baseline. The adjudicator is re-evaluated periodically against fresh human annotations to detect concept drift.

3.5. Metric Aggregation and CI/CD Integration

Metrics are computed per category and in aggregate (Section 5). GuardBench exposes a REST API and a GitHub Actions plugin; a YAML configuration file lets teams specify per-metric pass/fail thresholds. Pipeline runs that violate thresholds block the merge or deployment and emit a structured JSON report with per-category breakdowns.

4. Dataset Construction

4.1. Taxonomy Design

Drawing on prior taxonomies [8,10], we define seven attack categories reflecting distinct mechanism families. Table 2 summarises each category with prompt counts and representative mechanisms.

4.2. Severity Labelling

Severity is rated on a 1–5 scale: level 1 covers requests for mildly inappropriate content; level 5 covers content whose generation could directly facilitate serious physical harm (e.g., CBRN weapon synthesis). Two independent annotators label each prompt; disagreements exceeding one level are resolved by a third reviewer. Inter-annotator agreement (Cohen’s κ ) is 0.81 across the full corpus.

4.3. Benign Counterfactual Set

To measure FPR we construct a 1,000-prompt benign counterfactual set by (a) sampling from MS-MARCO [21] and ShareGPT1 for topically adjacent but harmless prompts, and (b) generating “near-miss” prompts that superficially resemble jailbreaks but have legitimate intent (e.g., security research, medical education). This mirrors the design philosophy of WildGuard [17].

5. Evaluation Methodology

5.1. Metrics

We report three primary metrics:
Attack Success Rate (ASR): the fraction of attack prompts for which the guardrail fails to block a harmful response, as classified by the adjudicator, formally defined in Equation (1). Following HarmBench [8], we use a calibrated LLM judge rather than keyword matching to reduce false-positive adjudication.
ASR = | { p ∈ P attack : judge ( r p ) = harmful } | | P attack |
False Positive Rate (FPR): the fraction of benign prompts incorrectly blocked by the guardrail, as given in Equation (2).
FPR = | { q ∈ P benign : guardrail blocks ( q ) } | | P benign |
Latency Overhead ( Δ t ): wall-clock latency added by the guardrail relative to an unguarded baseline, at the 50th percentile ( Δ t 50 ) and 95th percentile ( Δ t 95 ). The 95th-percentile figure governs the CI/CD gate as it reflects tail-latency SLA risk.
Composite Safety Score (CSS): a deployability index defined in Equation (3) as
CSS = ( 1 − α ASR ) ( 1 − β FPR ) 1 − δ min ( Δ t 95 , T ref ) T ref ,
where α , β , δ ∈ [ 0 , 1 ] are operator weights (default α = 0.70 , β = 0.25 , δ = 0.30 , reflecting the asymmetry between security, usability, and latency cost) and T ref = 3 , 000  ms is the reference latency budget. The latency component applies a linear penalty capped at T ref = 3 , 000  ms, so configurations far below the budget incur minimal penalty while those exceeding it are progressively discounted, with δ = 0.30 reflecting that latency is an important but not dominant deployment constraint.

5.2. Guardrail Configurations Under Test

The GuardBench framework defines four representative guardrail configurations spanning the practical design space. Two were fully evaluated in this study; two are planned for future evaluation:
  • G1 – Rule-Based Filter(evaluated): regex and keyword blocklist; fast but brittle against paraphrasing attacks.
  • G2 – Supervised Classifier(planned): fine-tuned DeBERTa-v3-base [22] on the AdvBench harmful behaviours dataset [3] augmented with WildGuard training splits [17] (combined 45K examples); balances inference speed and accuracy.
  • G3 – LLM Judge (Llama Guard 2)(evaluated): Llama Guard 2 [16] in zero-shot mode. We use v2 rather than the original [15] for improved taxonomy coverage.
  • G4 – Smoothing Defence (SmoothLLM) [18] (planned): perturbation rate q = 0.15 , N = 10 copies; hyperparameter sensitivity analysis planned for future work.

5.3. Experimental Setup

The primary target model is GPT-4o-mini [1] (OpenAI API, temperature 0, max 512 tokens). Each guardrail configuration was evaluated in a single pass over the curated dataset using direct-submission mode ( n = 4 , 200 attack prompts replayed statically; n = 1 , 000 benign prompts); Table 3 reports 95% CIs via Wilson score interval.
The adjudicator is GPT-4o-mini with a structured three-class rubric (safe / borderline / harmful); ASR counts only harmful verdicts. Because the target model and adjudicator are the same model, we apply a conservative judging protocol that restricts harmful to responses providing substantive, actionable assistance, excluding responses that merely restate, decode, or describe a request without complying. This protocol is documented in the released code and its effect on ASR is auditable against a verbatim-brief baseline. G3 (Llama Guard 2) was served via llama.cpp with Q4_K_M GGUF quantisation on Apple M3 Pro (18 GB unified memory, Metal acceleration); the measured Δ t 95 of 5,799 ms is specific to this hardware and quantisation configuration. Production serving on a dedicated GPU or hosted endpoint would yield materially lower latency and may reverse the CSS ranking. Results exclude prompts whose adjudicated harm was copyright reproduction rather than safety-policy violation ( n = 422 prompts; retained in supplementary results with both copyright-including and copyright-excluding ASR reported). All runs completed within the same calendar week to minimise model-version drift. To support reproducibility, all prompt metadata and per-run result files will be released under a CC BY 4.0 licence. The harness is implemented in Python 3.12 with pinned dependencies. Large language model tools (Claude, Anthropic) were used to assist with manuscript drafting and formatting; all experimental code, data collection, and quantitative analysis were conducted independently by the authors.

5.4. Experimental Results

Table 3 reports per-guardrail performance across the full dataset and the GCG-suffix category for the primary GPT-4o-mini target. ASR and FPR are reported as percentages (mean ± 95% CI via Wilson score interval); latency in milliseconds.
Key observations from Table 3:
  • The undefended baseline is low against archival prompts. Without any guardrail, GPT-4o-mini is successfully jailbroken in only 5.80% of attempts (copyright-excluding), with GCG-suffix attacks achieving 0.71% ASR—approximately 130× below the 94% reported in [3]. Many-shot attacks score 0.00%. This finding is consistent with two explanations that the current experimental design cannot distinguish: (a) 2023-era jailbreak templates have been substantially incorporated into contemporary safety training, and (b) static replay of archival prompts understates the ASR that a live adaptive attacker would achieve. Practitioners should treat these results as a lower bound; live adaptive evaluation is planned for future work.
  • G1 (rule-based) is operationally negligible at this baseline. G1 reduces overall ASR from 5.80% to 5.48%—an absolute reduction of 0.32 percentage points—whilst blocking 25.9% of attack prompts. Paired analysis (McNemar’s test) confirms the effect is statistically detectable ( χ 2 = 10.08 , p = 0.0015 ) but operationally negligible: G1 prevented 12 harmful responses in 3,778 attempts (copyright-excluding), with zero false positives. Its CSS (0.962) marginally exceeds the unguarded baseline (0.959) solely because of near-zero latency and perfect FPR. Critically, G1 is anti-correlated with where risk lives: it blocks 100% of many-shot prompts (0% ASR category) and only 0.54% of Token Manipulation prompts (38.8% ASR category).
  • G3 (Llama Guard 2) reduces ASR meaningfully but at severe latency cost under evaluation hardware. G3 reduces overall ASR to 3.39%, a 2.41 pp improvement over the unguarded baseline, with a 0.60% FPR. However, its Δ t 95 of 5,799 ms on our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) causes the CSS to drop to 0.682—below the unguarded baseline’s 0.959. Under production serving conditions (dedicated GPU or hosted endpoint), where Llama Guard 2 typically runs in the low hundreds of milliseconds, the CSS ranking may reverse. The Q4_K_M quantisation may also degrade classification accuracy; unquantised results are reserved for future evaluation.
  • Token Manipulation is the dominant residual threat. Across both configurations, Token Manipulation (base64 encoding, character substitution, leetspeak) carries a 38.8% ASR—nearly 7× the overall baseline. G1 is ineffective (0.54% block rate); G3 reduces it to 33.5% (27.9% block rate). Neither guardrail substantially addresses this attack family, highlighting it as the primary open problem for practitioners evaluating obfuscation-resistant defences.
  • No single configuration dominates across all axes. GuardBench enables practitioners to select configurations based on their specific risk tolerance and latency budget via the CSS metric. At a 5.8% archival baseline, the marginal security benefit of adding G3 must be weighed against its latency cost—a trade-off that GuardBench quantifies explicitly and prior frameworks do not.
Figure 2 provides a per-category breakdown of ASR across the three evaluated configurations (No Guard, G1, G3), confirming that Token Manipulation is the dominant residual threat at 38.8% ASR, whilst GCG-suffix and many-shot attacks approach zero across all tiers.

6. Discussion and Future Work

6.1. Limitations

GuardBench has five primary limitations in the current evaluation. First, the reported results are specific to one target model (GPT-4o-mini) and may not generalise to all LLM architectures or fine-tuning regimes; cross-model validation against open-weight models remains future work. Second, the adjudicator introduces two sources of measurement error: (a) using GPT-4o-mini as both the target model and the judge creates a circularity risk—the same model grades whether it was successfully jailbroken—mitigated by our conservative judging protocol but with unquantified residual bias; and (b) the framework’s target production adjudicator (fine-tuned Llama-3-8B-Instruct, F1 ≥ 0.89 on held-out examples) was not used in this evaluation, so those accuracy figures do not apply to the GPT-4o-mini judge used here. Third, the dataset contains no audio, image, or structured-data attack vectors; the text-only scope limits applicability to multi-modal deployments where image-based jailbreaks [23] and tool-call hijacking represent distinct threat surfaces. Fourth, all attack results derive from static replay of archival 2023-era prompts; live adaptive generation via PAIR, TAP, or GCG was not performed in this evaluation, so reported ASR values should be treated as a lower bound on what an adaptive attacker would achieve against a target seen during prompt construction. Fifth, G3’s latency measurements reflect a consumer-hardware configuration (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) rather than production serving; results on a dedicated GPU or hosted endpoint are expected to differ substantially.

6.2. Adaptive Attacks

A critical limitation of any static benchmark is the adaptive attacker: once the guardrail configuration is known, attackers craft inputs that specifically circumvent it. GuardBench is designed to address this through a private held-out set and by supporting adaptive variants of PAIR and GCG against each guardrail; these capabilities are implemented in the framework but were not exercised in the current proof-of-concept evaluation. Future versions will include an online adversarial track that continuously solicits new attack submissions.

6.3. Multi-Modal and Agentic Contexts

Current GuardBench prompts are text-only. The rising prevalence of vision-language models introduces image-based jailbreaks [23], while autonomous agents face tool-call hijacking and multi-turn escalation via indirect prompt injection [7]. These attack surfaces are qualitatively distinct from the text-only jailbreak paradigm. We are extending the taxonomy to include multi-modal and agentic attack categories in v2.0.

6.4. Guardrail Combination and Ensembles

Our experiments evaluate guardrails independently. In practice, production systems layer multiple defences. GuardBench’s modular pipeline supports chaining (e.g., G2 followed by G3). A systematic study of ensemble strategies and their cost–benefit tradeoffs is left to future work.

6.5. Calibration and Concept Drift

The LLM adjudicator is a single point of failure. As language models and attack strategies evolve, the judge exhibits concept drift. GuardBench schedules periodic re-annotation of randomly sampled adjudicated examples against a human panel. Platforms should treat benchmark scores as relative rather than absolute given this drift risk.

6.6. Beyond Binary Classification

Current metrics reduce guardrail output to a binary pass/fail. Incorporating severity-weighted ASR, where a harmful response classified at level 5 counts more than a level 1 lapse, provides a more faithful risk model. We propose the severity-weighted ASR in Equation (4):
ASR w = ∑ i w i · 1 [ failure i ] ∑ i w i ,
where w i is the severity of prompt i, as a direction for future metric design.

7. Conclusion

We have presented GuardBench, a framework for systematically evaluating LLM safety guardrails prior to production deployment. GuardBench contributes (i) a 4,200-prompt taxonomy-driven dataset spanning seven attack categories with severity labels and a 1,000-prompt benign counterfactual set; (ii) a four-stage automated pipeline covering attack execution, hybrid adjudication, and multi-metric reporting with confidence intervals; (iii) empirical baseline results for two guardrail configurations (G1 rule-based, G3 Llama Guard 2) evaluated against GPT-4o-mini using static replay of archival 2023-era prompts, revealing 5.8% baseline ASR and showing that heavy guardrails can score below the unguarded baseline on the composite metric when latency cost dominates—a result specific to our evaluation hardware (Apple M3 Pro, llama.cpp, Q4_K_M quantisation) that warrants validation on production serving infrastructure; and (iv) a CI/CD integration layer with configurable deployment thresholds. A key finding is that Token Manipulation attacks carry a 38.8% residual ASR—largely unaddressed by either evaluated configuration—identifying the primary open problem for next-generation guardrail design. Our experiments confirm that no single guardrail configuration simultaneously minimises ASR, FPR, and latency at this baseline, and that the CSS metric provides a principled basis for workload-specific guardrail selection. GuardBench gives practitioners the infrastructure to run safety tests on every release cycle, replacing informal pre-deployment checks with reproducible, quantified results.

Author Contributions

Conceptualization, P.B., S.S. and H.G.; methodology, P.B., S.S. and H.G.; software, S.S.; validation, P.B., S.S. and H.G.; formal analysis, S.S.; investigation, P.B., S.S. and H.G.; data curation, P.B. and S.S.; writing—original draft preparation, P.B., S.S. and H.G.; writing—review and editing, P.B., S.S. and H.G.; visualisation, S.S.; supervision, P.B. All authors have read and agreed to the published version of the manuscript. P.B., S.S. and H.G. contributed equally to this work.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The GuardBench dataset, evaluation code, and raw results will be made publicly available at https://github.com/bangadpurva/guardbench upon paper acceptance.

Acknowledgments

During the preparation of this manuscript, the authors used Claude (Anthropic, claude-sonnet-4-6) for the purposes of drafting and editing manuscript text, reviewing citations, and assisting with formatting. The authors have reviewed and edited all output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. OpenAI. GPT-4o mini: Advancing Cost-Efficient Intelligence. OpenAI Blog, 2024. Accessed: August 2024. Accessed. 2024.
  2. Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. Constitutional AI: Harmlessness from AI Feedback. arXiv Preprint. 2022, arXiv:2212.08073. [Google Scholar]
  3. Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J.Z.; Fredrikson, M. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv Preprint. 2023, arXiv:2307.15043. [Google Scholar]
  4. Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G.J.; Wong, E. Jailbreaking Black Box Large Language Models in Twenty Queries. arXiv Preprint. 2023, arXiv:2310.08419. [Google Scholar]
  5. Mehrotra, A.; Zampetakis, M.; Kassianik, P.; Nelson, B.; Anderson, H.; Singer, Y.; Karbasi, A. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
  6. Anil, C.; Durmus, E.; Panickssery, N.; Sharma, M.; Ziegler, D.M.; Rimsky, N.; Askell, A.; et al. Many-Shot Jailbreaking; Technical report;Technical Report; Anthropic, 2024. [Google Scholar]
  7. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023. [Google Scholar]
  8. Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. [Google Scholar]
  9. Zhu, K.; Wang, J.; Zhou, J.; Wang, Z.; Chen, H.; Wang, Y.; Yang, L.; Ye, W.; Gong, N.Z.; Zhang, Y.; et al. PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv Preprint. 2023, arXiv:2306.04528. [Google Scholar]
  10. Chao, P.; Debenedetti, E.; Robey, A.; Andriushchenko, M.; Croce, F.; Sehwag, V.; Dobriban, E.; Flammarion, N.; Pappas, G.J.; Tramèr, F.; et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
  11. OpenAI. OpenAI Moderation API. OpenAI Platf. Doc. Accessed. 2023. (accessed on 2024). [Google Scholar] [CrossRef]
  12. Microsoft. Azure AI Content Safety. Microsoft Azur. Doc. Accessed. 2023. (accessed on 2024). [Google Scholar] [CrossRef]
  13. Rebedea, T.; Dinu, R.; Sreedhar, M.N.; Parisien, C.; Cohen, J. NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2023. [Google Scholar]
  14. Perez, F.; Ribeiro, I. Ignore Previous Prompt: Attack Techniques For Language Models. arXiv Preprint. 2022, arXiv:2211.09527. [Google Scholar]
  15. Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv Preprint. 2023, arXiv:2312.06674. [Google Scholar]
  16. Meta, A.I. Llama Guard 2: Meta Llama Guard 2. Meta AI Research, 2024. Technical Report. Llama Guard 3 is a subsequent model. Available online: https://ai.meta.com/research/publications/llama-guard-3/.
  17. Han, S.; Rao, K.; Ettinger, A.; Jiang, L.; Lin, B.Y.; Lambert, N.; Choi, Y.; Dziri, N. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
  18. Robey, A.; Wong, E.; Hassani, H.; Pappas, G.J. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. arXiv Preprint. 2023, arXiv:2310.03684. [Google Scholar]
  19. Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv Preprint. 2022, arXiv:2209.07858. [Google Scholar]
  20. Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; Irving, G. Red Teaming Language Models with Language Models. arXiv Preprint. 2022, arXiv:2202.03286. [Google Scholar]
  21. Bajaj, P.; Campos, D.; Craswell, N.; Deng, L.; Gao, J.; Liu, X.; Majumder, R.; McNamara, A.; Mitra, B.; Nguyen, T.; et al. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv Preprint. 2016, arXiv:1611.09268. [Google Scholar]
  22. He, P.; Gao, J.; Chen, W. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
  23. Qi, X.; Zeng, Y.; Xie, T.; Chen, P.Y.; Jia, R.; Mittal, P.; Henderson, P. Visual Adversarial Examples Jailbreak Aligned Large Language Models. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar]
Figure 1. GuardBench four-stage pipeline. Prompts flow top-to-bottom through Dataset Curation, Attack Execution, Response Adjudication, and Metric Aggregation. Side inputs supply the guardrail under test and the LLM judge. Outputs feed a Safety Report and a CI/CD gate.
Figure 1. GuardBench four-stage pipeline. Prompts flow top-to-bottom through Dataset Curation, Attack Execution, Response Adjudication, and Metric Aggregation. Side inputs supply the guardrail under test and the LLM judge. Outputs feed a Safety Report and a CI/CD gate.
Preprints 231449 g001
Figure 2. Measured ASR (%) per attack category across evaluated configurations on GPT-4o-mini (copyright-excluding, static replay). Token Manipulation is the dominant residual risk, with 38.8% ASR under both G1 and No Guard, reduced only partially to 33.5% by G3. GCG-suffix and many-shot attacks approach zero across all configurations.
Figure 2. Measured ASR (%) per attack category across evaluated configurations on GPT-4o-mini (copyright-excluding, static replay). Token Manipulation is the dominant residual risk, with 38.8% ASR under both G1 and No Guard, reduced only partially to 33.5% by G3. GCG-suffix and many-shot attacks approach zero across all configurations.
Preprints 231449 g002
Table 1. Feature Comparison: GuardBench vs. Prior Frameworks and Commercial Tools.
Table 1. Feature Comparison: GuardBench vs. Prior Frameworks and Commercial Tools.
Feature HarmBench JBBench PBench OAI Mod. NeMo GR Ours
ASR measurement ✓ ✓ ✓ ✓ ✓
FPR / benign set ✓ ✓
Latency overhead ✓
Severity labelling ✓ ✓
CI/CD integration ✓ ✓
Versioned dataset ✓ ✓
Guardrail-layer focus ✓ ✓ ✓
Table 2. GuardBench Jailbreak Prompt Taxonomy (v1.0, 4,200 prompts total).
Table 2. GuardBench Jailbreak Prompt Taxonomy (v1.0, 4,200 prompts total).
Category Count % Avg. Sev.
Role-Play & Persona Override 840 20.0 3.2
Prompt Injection [7,14] 700 16.7 3.8
Token Manipulation 560 13.3 2.9
Many-Shot In-Context [6] 490 11.7 4.1
Gradient-Optimised Suffix [3] 420 10.0 4.6
Iterative Refinement [4,5] 630 15.0 4.3
Multilingual & Code-Switch 560 13.3 3.5
Total 4,200 100 3.71
Table 3. GuardBench Results (Target: GPT-4o-mini, copyright-excluding; lower ASR/FPR/latency is better). Wilson 95% CIs ( n = 3 , 778 attack; n = 1 , 000 benign). G2 and G4 are planned future evaluations.
Table 3. GuardBench Results (Target: GPT-4o-mini, copyright-excluding; lower ASR/FPR/latency is better). Wilson 95% CIs ( n = 3 , 778 attack; n = 1 , 000 benign). G2 and G4 are planned future evaluations.
Guard Overall GCG Suffix FPR Δ t 95
ASR CSS ASR CSS (ms)
G1 Rule 5.48 ± 0.73 0.962 0.71 ± 0.81 0.995 0.00 ± 0.00 0.8
G3 LG2 3.39 ± 0.58 0.682 0.24 ± 0.47 0.697 0.60 ± 0.48 5,799
G2 Class. planned — future work
G4 Smooth planned — future work
No Guard 5.80 ± 0.74 0.959 0.71 ± 0.81 – 0.00 0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.