Submitted:
23 January 2025
Posted:
23 January 2025
You are already at the latest version
Abstract
Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that fundamentally reimagines LLM safety by decoupling it from refusal prefixes through the use of humor as an indirect refusal strategy. Rather than explicitly rejecting harmful instructions, HumorReject responds with contextually appropriate humor that naturally defuses potentially dangerous requests while maintaining engaging interactions. Our approach effectively addresses the common "over-defense" issues in existing safety mechanisms, demonstrating superior robustness against various attack vectors while preserving natural and high-quality interactions on legitimate tasks. Our findings suggest that innovations at the data level are even more fundamental than the alignment algorithm itself in achieving effective LLM safety, opening new directions for developing more resilient and user-friendly AI systems. Our code and dataset are available at https://github.com/wooozihui/HumorReject
Keywords:
1. Introduction
- We propose a novel indirect refusal strategy based on humorous responses, which can effectively decouple LLMs’ safety from refusal prefixes (§Section 4.1), significantly lowering the risk of prefix injection attacks;
- We construct and publicly release the HumorReject preference dataset of 400 samples, and demonstrate that using existing alignment algorithm [7] with just 10 epochs of fine-tuning on this dataset can fundamentally enhance the safety of the previously unsafeguarded Mistral-7B-instruct-v0.1 model (§Section 4.2). This effective result indicates that existing alignment algorithms are sufficient for producing highly safe models—innovations at the data level are even more fundamental than the alignment algorithm itself in achieving effective LLM safety;
- Beyond prefix injection attacks, we conduct extensive security evaluations of the HumorReject model through various attack vectors, including mismatched generalization attacks (§Section 4.3) and our novel adaptive attack “HumorDAN” (§Section 4.4). Our experimental results demonstrate the model’s robust resistance against these diverse attack strategies;
- We also perform in-depth analysis of model usability and find that previous defense methods suffer from over-defense issues: 1) models generate refusals even for benign inputs [2,11], and 2) response quality significantly deteriorates under harmful context conditions [12]. HumorReject training effectively avoid these problems (§Section 4.5).
2. Related Work
2.1. LLM Alignment
2.2. Jailbreak Attacks
2.3. LLM with Humor
3. HumorReject Training
3.1. Training Dataset Construction
3.2. Training Settings
4. Empirical Studies
4.1. RQ1: How Effectively Does HumorReject Decouple Safety from Refusal Prefix?

4.2. RQ2: How Effectively Does HumorReject Defend Against Prefix Injection Attacks?

4.3. RQ3: Beyond Prefix Injection, Do Other Types of Attacks Still Pose Threats to Model Safety?

4.4. RQ4: Does the HumorReject Approach Introduce New Security Risks?
- HumorReject Mistral-7B-instruct-v0.1: Safety Rate 99%
- HumorReject Llama3-8B-instruct: Safety Rate 99%


4.5. RQ5: Does HumorReject Affect the Model’s Performance on Benign Inputs?

4.6. RQ6: Why Did Previous Humorous LLM Not Demonstrate Good Safety?

5. Conclusion
Acknowledgments
Appendix A. Case Study
Appendix A.1. Examples of Training Data


Appendix A.2. Prompts
Appendix A.2.1. Prompts for Judge Models


Appendix A.2.2. Prompts for HumorDAN Attack

Appendix A.2.3. HumorReject-like System Prompt in RQ6

Appendix A.3. Defense Cases
Appendix A.3.1. Defense Against GCG Attack

Appendix A.3.2. Defense Against AutoDAN Attack

Appendix A.3.3. Defense Against CodeAttack

Appendix A.3.4. Defense Against ReNeLLM Attack

Appendix A.3.5. Defense Against CodeChameleon Attack



Appendix A.4. Failure Cases


References
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 2022, 35, 27730–27744.
- Qi, X.; Panda, A.; Lyu, K.; Ma, X.; Roy, S.; Beirami, A.; Mittal, P.; Henderson, P. Safety Alignment Should Be Made More Than Just a Few Tokens Deep, 2024, [arXiv:cs.CR/2406.05946].
- Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Panickssery, N.; Gurnee, W.; Nanda, N. Refusal in language models is mediated by a single direction. arXiv preprint arXiv:2406.11717 2024.
- Wei, A.; Haghtalab, N.; Steinhardt, J. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems 2024, 36.
- Zou, A.; Wang, Z.; Kolter, J.Z.; Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043 2023.
- Liu, X.; Xu, N.; Chen, M.; Xiao, C. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451 2023.
- Hong, J.; Lee, N.; Thorne, J. Orpo: Monolithic preference optimization without reference model. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 11170–11189.
- Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D.S.; Casas, D.d.l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. Mistral 7B. arXiv preprint arXiv:2310.06825 2023.
- Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 2024.
- Andriushchenko, M.; Croce, F.; Flammarion, N. Jailbreaking leading safety-aligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151 2024.
- Yuan, Y.; Jiao, W.; Wang, W.; Huang, J.t.; Xu, J.; Liang, T.; He, P.; Tu, Z. Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training. arXiv preprint arXiv:2407.09121 2024.
- Zou, A.; Phan, L.; Wang, J.; Duenas, D.; Lin, M.; Andriushchenko, M.; Kolter, J.Z.; Fredrikson, M.; Hendrycks, D. Improving alignment and robustness with circuit breakers. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
- Christiano, P.F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems 2017, 30.
- Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 2024, 36.
- Ethayarajh, K.; Xu, W.; Muennighoff, N.; Jurafsky, D.; Kiela, D. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306 2024.
- Wu, Z.; Gao, H.; He, J.; Wang, P. The dark side of function calling: Pathways to jailbreaking large language models. arXiv preprint arXiv:2407.17915 2024.
- Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; Zhang, Y. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825 2023.
- Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G.J.; Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419 2023.
- Anil, C.; Durmus, E.; Sharma, M.; Benton, J.; Kundu, S.; Batson, J.; Rimsky, N.; Tong, M.; Mu, J.; Ford, D.; et al. Many-shot Jailbreaking. Anthropic, April 2024.
- Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; Shi, W. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373 2024.
- Lv, H.; Wang, X.; Zhang, Y.; Huang, C.; Dou, S.; Ye, J.; Gui, T.; Zhang, Q.; Huang, X. CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models, 2024, [arXiv:cs.CL/2402.16717].
- Ren, Q.; Gao, C.; Shao, J.; Yan, J.; Tan, X.; Lam, W.; Ma, L. Codeattack: Revealing safety generalization challenges of large language models via code completion. In Proceedings of the Findings of the Association for Computational Linguistics ACL 2024, 2024, pp. 11437–11452.
- Deng, Y.; Zhang, W.; Pan, S.J.; Bing, L. Multilingual jailbreak challenges in large language models. arXiv preprint arXiv:2310.06474 2023.
- Ding, P.; Kuang, J.; Ma, D.; Cao, X.; Xian, Y.; Chen, J.; Huang, S. A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. arXiv preprint arXiv:2311.08268 2023.
- Zhong, S.; Huang, Z.; Gao, S.; Wen, W.; Lin, L.; Zitnik, M.; Zhou, P. Let’s Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13246–13257.
- Tikhonov, A.; Shtykovskiy, P. Humor Mechanics: Advancing Humor Generation with Multistep Reasoning. arXiv preprint arXiv:2405.07280 2024.
- Chen, Y.; Yuan, Y.; Liu, P.; Liu, D.; Guan, Q.; Guo, M.; Peng, H.; Liu, B.; Li, Z.; Xiao, Y. Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor Interpretation. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024, Vol. 38, pp. 17826–17834.
- Vikhorev, D.; Galimzianova, D.; Gorovaia, S.; Zhemchuzhina, E.; Yamshchikov, I.P. CleanComedy: Creating Friendly Humor through Generative Techniques. arXiv preprint arXiv:2412.09203 2024.
- Chen, Y.; Yang, C.; Hu, T.; Chen, X.; Lan, M.; Cai, L.; Zhuang, X.; Lin, X.; Lu, X.; Zhou, A. Are U a Joke Master? Pun Generation via Multi-Stage Curriculum Learning towards a Humor LLM. In Proceedings of the Findings of the Association for Computational Linguistics ACL 2024, 2024, pp. 878–890.
- Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; Hashimoto, T.B. Stanford Alpaca: An Instruction-following LLaMA model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
- Anthropic. Claude 3.5 sonnet. https://www.anthropic.com/news/claude-3-5-sonnet, 2024.
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 2021.
- Zheng, Y.; Zhang, R.; Zhang, J.; Ye, Y.; Luo, Z.; Feng, Z.; Ma, Y. LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Bangkok, Thailand, 2024.
- Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249 2024.
- Tramer, F.; Carlini, N.; Brendel, W.; Madry, A. On adaptive attacks to adversarial example defenses. Advances in neural information processing systems 2020, 33, 1633–1645.
- DAN Template. https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa881516.
- Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300 2020.
- Röttger, P.; Kirk, H.R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; Hovy, D. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv preprint arXiv:2308.01263 2023.
- X.ai. Grok 2. https://x.ai/blog/grok-2, 2024.
- Musk, E. Twitter Status Post. https://x.com/elonmusk/status/1720635518289908042, 2023.
- Adversa.ai. LLM Red Teaming vs. Grok, ChatGPT, Claude, Gemini, Bing, Mistral, LLaMA. https://adversa.ai/blog/llm-red-teaming-vs-grok-chatgpt-claude-gemini-bing-mistral-llama/, 2023.
- Plinius, E. Grok System Prompt Leak. https://github.com/elder-plinius/Grok-System-Prompt-Leak, 2023.
- Team, G.; Riviere, M.; Pathak, S.; Sessa, P.G.; Hardin, C.; Bhupatiraju, S.; Hussenot, L.; Mesnard, T.; Shahriari, B.; Ramé, A.; et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118 2024.
- AI, Q. Qwen 2.5: Advancing AI for Everyone. https://qwen2.org/qwen2-5/, 2024.



| Model: LLaMA3-8B-instruct | Humor Rate (%) | Reject Rate (%) | Safety Rate (%) |
|---|---|---|---|
| Vanilla | 0 | 96 | 97 |
| HumorReject | 95 | 2 | 100 |
| Model | Attack | Vanilla | CB | DeepAug | DeRTa | HumorReject (Ours) |
|---|---|---|---|---|---|---|
| Llama3-8B-instruct | GCG | 88 | 99 | 99 | 97 | 98 |
| AutoDAN | 87 | 98 | 40 | 89 | 99 | |
| Template | 98 | 97 | 100 | 100 | 99 | |
| Prefill | 41 | 95 | 59 | 98 | 100 | |
| Template+Prefill | 2 | 98 | 3 | 32 | 98 | |
| Average | 63.2 | 97.4 | 60.2 | 83.2 | 99.0 | |
| Mistral-7B-instruct-v0.1 | GCG | 4 | 89 | 66 | 61 | 95 |
| AutoDAN | 22 | 86 | 19 | 50 | 97 | |
| Template | 2 | 89 | 8 | 54 | 96 | |
| Prefill | 1 | 99 | 56 | 92 | 98 | |
| Template+Prefill | 4 | 90 | 7 | 53 | 97 | |
| Average | 6.6 | 90.6 | 31.2 | 62.0 | 96.6 |
| Model | Attack | Vanilla | CB | DeepAug | DeRTa | HumorReject (Ours) |
|---|---|---|---|---|---|---|
| Llama3-8B-instruct | ReNeLLM | 44 | 84 | 63 | 86 | 92 |
| CodeAttack | 35 | 89 | 79 | 66 | 77 | |
| CodeChameleon | 44 | 94 | 62 | 68 | 83 | |
| Average | 41.0 | 89.0 | 68.0 | 73.3 | 84.0 | |
| Mistral-7B-instruct-v0.1 | ReNeLLM | 9 | 85 | 19 | 30 | 95 |
| CodeAttack | 7 | 84 | 8 | 26 | 98 | |
| CodeChameleon | 47 | 100 | 70 | 73 | 95 | |
| Average | 21.0 | 89.7 | 32.3 | 56.3 | 96.0 |
| Model | Method | MMLU (%) | MMLU with Harmful Context (%) | XSTEST Compliance Rate (%) |
|---|---|---|---|---|
| Llama3 | Vanilla Model | 58.0 | 54.8 | 95.2 |
| DeRTa | 59.4 | 50.8 | 72.4 | |
| Circuit Breaker | 58.4 | 25.8 | 95.6 | |
| DeepAug | 60.6 | 59.2 | 60.4 | |
| HumorReject (Ours) | 60.8 | 58.2 | 94.8 | |
| Mistral | Vanilla Model | 49.8 | 45.4 | 97.2 |
| DeRTa | 39.6 | 33.6 | 25.6 | |
| Circuit Breaker | 47.4 | 0 | 96.4 | |
| DeepAug | 47.2 | 39.2 | 38 | |
| HumorReject (Ours) | 50.2 | 45.4 | 94.0 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).