Submitted:
15 September 2024
Posted:
17 September 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Methodology
2.1. Reward Model
2.1.1. Choosing the Reward Model
2.1.2. Data Synthesis and Length Bias Mitigation
- 1.
- Model serving
- 2.
- Prompt Space Creation
- 3.
- Response Generation
- 1.
- Reducing Length Bias: By including chunks as short as a single token, this method addresses the tendency of reward models to favour longer responses, as observed in existing studies (Singhal et al., 2023). The inclusion of single-token chunks enables the model to learn fine-grained token-level differences, reducing the inherent bias towards longer sequences.
- 2.
- Balancing Token-Level and Sequence-Level Evaluation: The variable chunk length allows the model to evaluate both short and long segments of text, providing a more balanced training dataset. This ensures that the reward model can assess responses of varying lengths, from individual tokens to nearly complete sequences, fostering a more nuanced understanding of both token-level and sequence-level interactions.
Length Bias in Reward Models
2.1.3. Training
2.2. Selective Assistance

- Acceptance: If the reward score of the candidate token exceeds a predefined threshold, the token is accepted, and the decoding flow continues with the SLM generating the next candidate token.
- Rejection: If the reward score falls below the threshold, the candidate token is rejected, and the prefix used to generate this token is sent to the cloud LLM. The cloud LLM then generates the next token.
3. Experiments
Dataset Creation
Training and Evaluation

- Train/Chosen Reward: Displays the reward scores assigned to tokens generated by the LLM, highlighting their alignment with the probability distribution.
- Train/Reject Reward: Represents the reward scores for tokens produced by the SLM that were rejected, showing the model’s capacity to identify misaligned tokens.
- Train/Preference Loss: Demonstrates the overall preference loss during training. The consistently low loss signifies the model's effectiveness in maintaining a clear distinction between chosen and rejected tokens.
| Dataset | Reward threshold | SLM accuracy (Baseline) |
Hybrid decoding accuracy (Ours) |
LLM accuracy (Baseline) |
Cloud LLM activation ratio (%) |
|---|---|---|---|---|---|
| gsm8k | 1.0 | 50.7 | 66.48 | 77.8 | 56.00 |
| 2.0 | 74.75 | 78.00 | |||
| 4.0 | 77.78 | 87.00 | |||
| mmlu | 1.0 | 52.4 | 69.9 | 70.5 | 92.00 |
| mbpp | 1.0 | 36.6 | 38.8 | 60.00 | 35.00 |
| 1.5 | 52.2 | 68.00 | |||
| 2.0 | 60.0 | 87.00 | |||
| 4.0 | 60.0 | 98.00 | |||
| cnndm | 1.0 | 20.9 | 25.7 | 28.9 | 59.00 |
3.1.1. Reward Threshold v/s Cloud LLM activation ratio

3.1.2. Reward Threshold v/s Accuracy

3.1.3. Throughput & Latency

4. Limitations
5. Future Considerations
6. Conclusions
References
- Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Rühle, V., Lakshmanan, L. V. S., & Awadallah, A. H. (n.d.). Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. OpenReview. https://openreview.net/forum?id=02f3mUtqnM.
- Ong, I., Almahairi, A., Wu, V., Chiang, W., Wu, T., Gonzalez, J. E., Kadous, M. W., & Stoica, I. (2024, June 26). RouteLLM: Learning to Route LLMs with Preference Data. arXiv.org. https://arxiv.org/abs/2406.18665v3.
- Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., . . . Kaplan, J. (2022, December 15). Constitutional AI: Harmlessness from AI Feedback. arXiv.org. http://arxiv.org/abs/2212.08073.
- Kim, S., Bae, S., Shin, J., Kang, S., Kwak, D., Yoo, K. M., & Seo, M. (2023, May 23). Aligning Large Language Models through Synthetic Feedback. arXiv.org. http://arxiv.org/abs/2305.13735.
- Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., Xiong, L., Chen, L., Xi, Z., Xu, N., Lai, W., Zhu, M., Chang, C., Yin, Z., Weng, R., . . . Huang, X. (2023, July 11). Secrets of RLHF in Large Language Models Part I: PPO. arXiv.org. http://arxiv.org/abs/2307.04964.
- Huang, S., Noukhovitch, M., Hosseini, A., Rasul, K., Wang, W., & Tunstall, L. (2024, March 24). The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization. arXiv.org. https://arxiv.org/abs/2403.17031v1.
- Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023, October 5). A long way to go: Investigating length correlations in RLHF. arXiv.org. https://arxiv.org/abs/2310.03716v2.
- Yang, S., Zhang, S., Xia, C., Feng, Y., Xiong, C., & Zhou, M. (2023, June 1). Preference-grounded token-level guidance for language model fine-tuning. arXiv.org. http://arxiv.org/abs/2306.00398.
- Leviathan, Y., Kalman, M., & Matias, Y. (2022, November 30). Fast Inference from Transformers via Speculative Decoding. arXiv.org. https://arxiv.org/abs/2211.17192v2.
- Fu, Y., Bailis, P., Stoica, I., & Zhang, H. (2024, February 3). Break the sequential dependency of LLM inference using lookahead decoding. arXiv.org. https://arxiv.org/abs/2402.02057v1.
- Zhao, W., Huang, Y., Han, X., Xu, W., Xiao, C., Zhang, X., Fang, Y., Zhang, K., Liu, Z., & Sun, M. (2024, February 21). Ouroboros: generating longer drafts phrase by phrase for faster speculative decoding. arXiv.org. https://arxiv.org/abs/2402.13720v2.
- Goel, R., Gagrani, M., Jeon, W., Park, J., Lee, M., & Lott, C. (2024, February 29). Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs. arXiv.org. http://arxiv.org/abs/2403.00858.
- Shen, T., Jin, R., Huang, Y., Liu, C., Dong, W., Guo, Z., Wu, X., Liu, Y., & Xiong, D. (2023, September 26). Large Language Model alignment: a survey. arXiv.org. https://arxiv.org/abs/2309.15025v1.
- Vllm-Project. (n.d.). GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs. GitHub. https://github.com/vllm-project/vllm.
| Dataset | Reward threshold | SLM throughput (Baseline) |
Hybrid decoding throughput (Ours) | LLM throughput (Baseline) |
Cloud LLM activation ratio (%) |
|---|---|---|---|---|---|
| gsm8k | 1.0 | 35.18 | 10.70 | 34.05 | 56.00 |
| 2.0 | 8.71 | 78.00 | |||
| 4.0 | 8.48 | 87.00 | |||
| mbpp | 1.0 | 22.38 | 6.39 | 18.62 | 35.00 |
| 1.5 | 4.82 | 68.00 | |||
| 2.0 | 4.20 | 87.00 | |||
| 4.0 | 4.10 | 98.00 | |||
| cnndm | 1.0 | 23.71 | 4.77 | 19.22 | 59.00 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).