Submitted:
02 January 2026
Posted:
06 January 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Pretrained LLMs, such as Mistral [22], exhibit stable knowledge transfer, whereas smaller models and base models benefit substantially from memory replay.
- Memory replay is especially effective for models without instruction tuning and smaller encoder-decoder models utilized in this work.
- Llama models [23] show a greater tendency to hallucinate undefined relation types than other LLM architectures.
- To the best of our knowledge, we present the first systematic benchmarks evaluating LLM behavior, specifically, knowledge transfer and error patterns, in CRE across multiple architectures and datasets.
- We analyze the effectiveness of memory replay in mitigating forgetting and identify model-specific strengths and limitations.
- We provide a comprehensive analysis of hallucinations, showing how LLMs generate undefined relations in CRE and discussing the implications for KG updates.
2. Related Work
3. Preliminaries
3.1. Continual Relation Extraction
3.2. Pretrained Language Models
- Encoder-only models use a bidirectional transformer encoder to learn the contextual representations. One prominent example is BERT [12] and its variants.
4. Methodology
4.1. Continual Learning with LLMs

4.1.1. Memory Sample Selection
4.1.2. Incremental Instruction Tuning Algorithm
4.2. Instruction Format
5. Evaluation
5.1. Experimental Settings
5.1.1. Dataset
5.1.2. Large Language Models
5.1.3. Parameter Settings
5.1.4. Evaluation Metrics
5.2. Results
5.3. Knowledge Transfer
5.4. Error Analysis
6. Ablation Study
6.1. Memory Size Experiments
| Task Index | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Memory | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
| Flan-T5 Base | No Replay | 96.1 | 96.5 | 95.5 | 95.7 | 95.3 | 94.4 | 93.5 | 94.1 | 94.0 | 92.7 |
| 5 | 96.1 (=) | 96.1 (↓ 0.4) | 95.2 (↓0.3) | 95.2 (↓0.5) | 95.4 (↑ 0.1) | 94.6 (↑ 0.2) | 95.7 (↑ 2.2) | 96.2 (↑ 2.1) | 96.1 (↑ 2.1) | 95.2 (↑ 2.5) | |
| 10 | 96.1 (=) | 96.2 (↑ 0.1) | 95.7 (↑ 0.5) | 96.0 (↑ 0.8) | 95.7 (↑ 0.3) | 95.4 (↑ 0.8) | 96.1 (↑ 0.4) | 96.0 (↓ 0.2) | 96.3 (↑ 0.2) | 95.8 (↑ 0.6) | |
| 15 | 96.1 (=) | 95.8 (↓ 0.4) | 95.7 (=) | 96.0 (=) | 96.3 (↑ 0.6) | 96.3 (↑ 0.9) | 96.9 (↑ 0.8) | 96.7 (↑ 0.7) | 96.7 (↑ 0.4) | 96.3 (↑ 0.5) | |
| Mistral-7B | No Replay | 96.6 | 94.8 | 96.3 | 96.2 | 96.4 | 96.6 | 96.6 | 96.9 | 96.8 | 96.8 |
| 5 | 96.6 (=) | 95.2 (↑ 0.4) | 96.9 (↑ 0.6) | 96.7 (↑ 0.5) | 96.9 (↑ 0.5) | 97.2 (↑ 0.6) | 96.9 (↑ 0.3) | 97.1 (↑ 0.2) | 97.2 (↑ 0.4) | 97.0 (↑ 0.2) | |
| 10 | 96.6 (=) | 95.3 (↑ 0.1) | 96.4 (↓ 0.5) | 95.9 (↓ 0.3) | 96.6 (↓ 0.3) | 97.0 (↓ 0.2) | 96.8 (↓ 0.1) | 96.9 (=) | 95.8 (↓ 1.0) | 96.9 (↓ 0.1) | |
| 15 | 96.6 (=) | 94.8 (↓ 0.5) | 95.2 (↓ 1.2) | 95.7 (↓ 0.2) | 96.5 (↓ 0.1) | 97.1 (↑ 0.1) | 97.2 (↑ 0.4) | 97.3 (↑ 0.4) | 96.8 (↑ 1.0) | 96.7 (↓ 0.2) | |
| Llama2-7B-hf | No Replay | 57.7 | 52.8 | 52.7 | 52.3 | 54.0 | 57.0 | 57.8 | 60.1 | 62.9 | 62.4 |
| 5 | 57.7 (=) | 59.2 (↑ 6.4) | 55.5 (↑ 2.8) | 55.9 (↑ 3.6) | 56.5 (↑ 2.5) | 60.6 (↑ 3.6) | 59.5 (↑ 1.7) | 62.4 (↑ 2.3) | 64.2 (↑ 1.3) | 66.5 (↑ 4.1) | |
| 10 | 57.7 (=) | 57.6 (↓ 1.6) | 54.9 (↓ 0.6) | 55.8 (↓ 0.1) | 57.6 (↑ 1.1) | 62.0 (↑ 1.4) | 62.4 (↑ 2.9) | 65.3 (↑ 2.9) | 67.7 (↑ 3.5) | 70.6 (↑ 4.1) | |
| 15 | 57.7 (=) | 60.3 (↑ 2.7) | 58.4 (↑ 3.5) | 57.6 (↑ 1.8) | 57.2 (↓ 0.4) | 59.2 (↓ 2.8) | 59.9 (↓ 2.5) | 61.3 (↓ 4.0) | 63.1 (↓ 4.6) | 63.9 (↓ 6.7) | |
| Llama3.1-8B | No Replay | 89.1 | 86.4 | 86.9 | 85.4 | 85.7 | 85.2 | 84.7 | 84.5 | 83.8 | 83.4 |
| 5 | 89.1 (=) | 85.6 (↓ 0.7) | 86.3 (↓ 0.5) | 85.4 (↓ 0.1) | 85.2 (↓ 0.4) | 83.7 (↓ 1.6) | 82.9 (↓ 1.8) | 81.4 (↓ 3.0) | 80.2 (↓ 3.6) | 79.1 (↓ 4.3) | |
| 10 | 89.1 (=) | 85.0 (↓ 0.6) | 85.6 (↓ 0.7) | 84.4 (↓ 1.0) | 85.1 (↓ 0.1) | 83.9 (↑ 0.3) | 83.1 (↑ 0.1) | 81.7 (↑ 0.3) | 80.0 (↓ 0.2) | 78.7 (↓ 0.3) | |
| 15 | 89.1 (=) | 85.2 (↑ 0.2) | 86.2 (↑ 0.6) | 84.9 (↑ 0.5) | 84.5 (↓ 0.6) | 83.1 (↓ 0.8) | 82.2 (↓ 0.9) | 81.3 (↓ 0.4) | 79.6 (↓ 0.4) | 78.2 (↓ 0.5) | |
| Qwen2.5-7B | No Replay | 93.0 | 91.9 | 91.3 | 91.2 | 91.4 | 91.7 | 91.7 | 91.8 | 91.8 | 91.2 |
| 5 | 93.0 (=) | 91.8 (↓ 0.2) | 91.1 (↓ 0.2) | 91.2 (=) | 91.6 (↑ 0.1) | 91.9 (↑ 0.2) | 92.0 (↑ 0.3) | 92.0 (↑ 0.2) | 91.9 (↑ 0.2) | 91.4 (↑ 0.2) | |
| 10 | 93.0 (=) | 92.0 (↑ 0.3) | 91.2 (↑ 0.1) | 91.4 (↑ 0.2) | 91.8 (↑ 0.2) | 92.0 (↑ 0.1) | 92.0 (↑ 0.0) | 91.9 (↓ 0.1) | 91.9 (↓ 0.1) | 91.3 (↓ 0.1) | |
| 15 | 93.0 (=) | 91.6 (↓ 0.4) | 91.3 (↑ 0.1) | 91.3 (↓ 0.0) | 91.7 (↓ 0.1) | 91.9 (↓ 0.0) | 92.2 (↑ 0.1) | 92.2 (↑ 0.3) | 92.2 (↑ 0.3) | 91.6 (↑ 0.3) | |
6.2. Hallucination Reduction Approaches
7. Discussion
8. Conclusion
9. Limitations
- Model Coverage. Our benchmark evaluates a set of representative open-source LLMs selected for their architectural diversity and reproducibility. Although our core observations on knowledge transfer, memory replay, and hallucination behavior are expected to generalize across model families, they may not extend to substantially larger models or proprietary systems trained with undisclosed data or objectives.
- Language and Dataset Limitations. All experiments were conducted on the English benchmark datasets (TACRED and FewRel), reflecting the primary training domain of the evaluated LLMs. Although the datasets contain a small number of multilingual examples (Appendix B), our conclusions cannot be generalized to multilingual CRE. Extending the benchmark to multilingual and low-resource settings with multilingual LLMs remains an essential topic for future work.
- Memory Replay on FewRel. We did not perform memory-size ablations on FewRel due to its substantially larger task sequence, which leads to rapid growth in the replay buffer. Nevertheless, the No Replay vs. Memory Size (10) results shown in Figure A12 (Appendix I) clearly support the findings obtained on TACRED with Flan-T5 Base, and memory replay improves the performance of Flan-T5 Base on FewRel as well.
- Impact of Quantization (QLoRA). All models were tuned using a 4-bit QLoRA to ensure feasibility across multiple runs. Although this guarantees fairness, it limits our ability to isolate the effects of quantization on knowledge retention, catastrophic forgetting, and hallucination rates. Prior benchmarks (e.g., [49]) indicate that QLoRA or LoRA can improve the performance of models such as Llama2-7B-hf; however, evaluating these effects directly remains outside the scope of this work.
- Focus on Memory Replay. Our analysis focuses on memory replay as the primary strategy for CL. A diversity-based memory sample selection algorithm, K-means, might miss complex samples. Nonetheless, LLMs depend heavily on sample diversity for robust performance.
- Generality of Findings. The observed trends—e.g., replay benefiting smaller or base models, or instruction-tuned models showing limited improvement—are grounded in our empirical setting.
- Real-World Application. Although experiments were performed locally using open-source models, the ethical risks of deploying CRE systems remain unaddressed. Hallucinated relations can propagate misinformation in real-world settings, particularly in domains such as healthcare and biodiversity. Additional safeguards, such as fact-checking and human oversight, are required but beyond the scope of this work.
- Bias in Models. This work does not assess potential biases in the LLMs or datasets. TACRED and FewRel include entity mentions that may introduce demographic or contextual biases (e.g., gender, age). Since such biases can lead to uneven or skewed relation predictions, the fairness implications of the models remain unexamined.
Funding
Author Contributions: Sefika Efeoglu
Informed Consent Statement
Data Availability Statement
Appendix A. Dataset Statistics
Appendix A.1. FewRel
| Dataset | Train | Validation | Test | # of Relations |
|---|---|---|---|---|
| FewRel | 33600 | 11200 | 11200 | 80 |
Appendix A.2. TACRED
| Dataset | Train | Validation | Test | # of Relations |
|---|---|---|---|---|
| TACRED | 7146 | 1452 | 1223 | 40 |
Appendix B. Languages in Datasets


Appendix C. Parameters
| Model | TACRED | FewRel | ||||||
|---|---|---|---|---|---|---|---|---|
| EP | BS | LR | Trainer | EP | BS | LR | Trainer | |
| Flan-T5 Base | 5 | 8 | 0.001 | Seq2Seq | 5 | 16 | 0.001 | Seq2Seq |
| Mistral-7B | 5 | 4 | 0.0002 | SFT | 5 | 8 | 0.0002 | SFT |
| Llama2-7B | 5 | 4 | 0.0002 | SFT | 5 | 8 | 0.0002 | SFT |
| Llama-3.1-8B | 5 | 4 | 0.0002 | SFT | 5 | 8 | 0.0002 | SFT |
| Qwen2.5-7B | 5 | 4 | 0.0002 | SFT | 5 | 8 | 0.0002 | SFT |
| Model | Rank | Task Type | ||
|---|---|---|---|---|
| Flan-T5 Base | 32 | 4 | 0.01 | Seq2SeqLM |
| Mistral-7B | 16 | 64 | 0.10 | CausalLM |
| Llama2-7B | 16 | 64 | 0.10 | CausalLM |
| Llama-3.1-8B | 16 | 64 | 0.10 | CausalLM |
| Qwen2.5-7B | 16 | 64 | 0.10 | CausalLM |
Appendix D. Permutation Test for Undefined Predictions
| Iteration | TACRED | FewRel | ||
|---|---|---|---|---|
| Diff (%) | p-value | Diff (%) | p-value | |
| 100 | 2.40 | 0.2000 | 433.00 | 0.0300 |
| 1000 | 2.40 | 0.2510 | 433.00 | 0.0250 |
| 10000 | 2.40 | 0.2642 | 433.00 | 0.0305 |
Appendix E. Examples Predictions
Appendix E.1. Llama2 Hallucinations

Appendix E.2. Categories of Undefined Relations
| Category | Similar Context | Part of Ground-Truth | New Relation | Fine-Grained Version of Ground-Truth |
|---|---|---|---|---|
| Ground-Truth | head of government, screenwriter, has part | country, location | – | child, sibling, spouse |
| Prediction | leader, author, component of | country of, located in | Japanese cargoer, Christmas, rail | son or daughter, sister or brother, wife or husband |
| Models | Flan-T5, Llama-3.1, Qwen2.5 | Flan-T5, Llama2, Llama-3.1, Qwen2.5 | all models | Flan-T5, Llama-3.1, Qwen2.5 |
Appendix F. Confusion Matrices on FewRel Dataset




Appendix G. Accuracy Matrices




Appendix H. Permutation Test on Main Results
| Task | Mean (Flan-T5) | Mean (Mistral) | (Mistral-Flan-T5) | 95% CI [low, high] | p-value |
|---|---|---|---|---|---|
| 1 | 0.961 | 0.966 | 0.005 | [0.902, 0.992] | 0.817 |
| 2 | 0.962 | 0.948 | -0.014 | [0.908, 0.976] | 0.595 |
| 3 | 0.957 | 0.964 | 0.007 | [0.951, 0.977] | 0.405 |
| 4 | 0.960 | 0.960 | -0.001 | [0.947, 0.970] | 0.913 |
| 5 | 0.956 | 0.966 | 0.009 | [0.957, 0.974] | 0.468 |
| 6 | 0.954 | 0.970 | 0.016 | [0.960, 0.979] | 0.135 |
| 7 | 0.961 | 0.968 | 0.007 | [0.958, 0.977] | 0.357 |
| 8 | 0.960 | 0.969 | 0.009 | [0.962, 0.977] | 0.278 |
| 9 | 0.963 | 0.958 | -0.005 | [0.931, 0.977] | 0.889 |
| 10 | 0.958 | 0.969 | 0.011 | [0.962, 0.977] | 0.246 |
| Task | Mean (Flan-T5) | Mean (Mistral) | (Mistral-Flan-T5) | 95% CI [low, high] | p-value |
|---|---|---|---|---|---|
| 1 | 0.967 | 0.960 | -0.007 | [0.910, 0.988] | 0.937 |
| 2 | 0.945 | 0.946 | 0.002 | [0.899, 0.983] | 0.952 |
| 3 | 0.946 | 0.947 | 0.002 | [0.919, 0.974] | 0.929 |
| 4 | 0.934 | 0.936 | 0.002 | [0.890, 0.970] | 0.976 |
| 5 | 0.930 | 0.936 | 0.006 | [0.907, 0.959] | 0.722 |
| 6 | 0.925 | 0.923 | -0.003 | [0.905, 0.938] | 0.841 |
| 7 | 0.912 | 0.925 | 0.013 | [0.915, 0.934] | 0.127 |
| 8 | 0.916 | 0.919 | 0.003 | [0.914, 0.923] | 0.714 |
| 9 | 0.907 | 0.923 | 0.016 | [0.906, 0.950] | 0.373 |
| 10 | 0.896 | 0.914 | 0.017 | [0.912, 0.915] | 0.024 |
Appendix I. No Replay vs Memory Replay on FewRel

References
- Sheth, A.; Padhee, S.; Gyrard, A. Knowledge Graphs and Knowledge Networks: The Story in Brief. IEEE Internet Computing 2019, 23, 67–75. [CrossRef]
- Grishman, R. Information Extraction. IEEE Intelligent Systems 2015, 30, 8–15. [CrossRef]
- Biesialska, M.; Biesialska, K.; Costa-jussà, M.R. Continual Lifelong Learning in Natural Language Processing: A Survey. In Proceedings of the Proceedings of the 28th International Conference on Computational Linguistics; Scott, D.; Bel, N.; Zong, C., Eds., Barcelona, Spain (Online), 12 2020; pp. 6523–6541. [CrossRef]
- Chen, Q.; Sun, J.; Palade, V.; Yu, Z. Continual Relation Extraction via Linear Mode Connectivity and Interval Cross Training. Knowledge-Based Systems 2023, 264, 110288. [CrossRef]
- Duan, B.; Liu, X.; Wang, S.; Xu, Y.; Xiao, B. Relational Representation Learning for Zero-Shot Relation Extraction with Instance Prompting and Prototype Rectification. In Proceedings of the ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. [CrossRef]
- Xia, H.; Wang, P.; Liu, T.; Lin, B.; Cao, Y.; Sui, Z. Enhancing Continual Relation Extraction via Classifier Decomposition. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023; Rogers, A.; Boyd-Graber, J.; Okazaki, N., Eds., Toronto, Canada, 7 2023; pp. 10053–10062. [CrossRef]
- Zhao, K.; Xu, H.; Yang, J.; Gao, K. Consistent Representation Learning for Continual Relation Extraction. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2022; Muresan, S.; Nakov, P.; Villavicencio, A., Eds., Dublin, Ireland, 5 2022; pp. 3402–3411. [CrossRef]
- Le, T.T.; Nguyen, M.; Nguyen, T.T.; Ngo Van, L.; Nguyen, T.H. Continual Relation Extraction via Sequential Multi-Task Learning. Proceedings of the AAAI Conference on Artificial Intelligence 2024, 38, 18444–18452. [CrossRef]
- Shen, H.; Ju, S.; Sun, J.; Chen, R.; Liu, Y. Efficient Lifelong Relation Extraction with Dynamic Regularization. In Proceedings of the Natural Language Processing and Chinese Computing; Zhu, X.; Zhang, M.; Hong, Y.; He, R., Eds., Cham, 2020; pp. 181–192. [CrossRef]
- Cui, L.; Yang, D.; Yu, J.; Hu, C.; Cheng, J.; Yi, J.; Xiao, Y. Refining Sample Embeddings with Relation Prototypes to Enhance Continual Relation Extraction. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers); Zong, C.; Xia, F.; Li, W.; Navigli, R., Eds., Online, 8 2021; pp. 232–243. [CrossRef]
- van de Ven, G.M.; Soures, N.; Kudithipudi, D. 1.09 - Continual learning and catastrophic forgetting. In Learning and Memory: A Comprehensive Reference; Wixted, J., Ed.; Academic Press: Oxford, 2025; pp. 153–168. [CrossRef]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Burstein, J.; Doran, C.; Solorio, T., Eds., Minneapolis, Minnesota, 6 2019; pp. 4171–4186. [CrossRef]
- Zhang, H.; Liang, B.; Yang, M.; Wang, H.; Xu, R. Prompt-Based Prototypical Framework for Continual Relation Extraction. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2022, 30, 2801–2813. [CrossRef]
- Chen, X.; Wu, H.; Shi, X. Consistent Prototype Learning for Few-Shot Continual Relation Extraction. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Rogers, A.; Boyd-Graber, J.; Okazaki, N., Eds., Toronto, Canada, 7 2023; pp. 7409–7422. [CrossRef]
- Ye, W.; Zhang, P.; Zhang, J.; Gao, H.; Wang, M. Distilling Causal Effect of Data in Continual Few-Shot Relation Learning. In Proceedings of the Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024); Calzolari, N.; Kan, M.Y.; Hoste, V.; Lenci, A.; Sakti, S.; Xue, N., Eds., Torino, Italia, 5 2024; pp. 5041–5051.
- Shlyk, D.; Groza, T.; Mesiti, M.; Montanelli, S.; Cavalleri, E. REAL: A Retrieval-Augmented Entity Linking Approach for Biomedical Concept Recognition. In Proceedings of the Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, 2024, pp. 380–389. [CrossRef]
- Taffa, T.A.; Usbeck, R. Bridge-Generate: Scholarly Hybrid Question Answering. In Proceedings of the Companion Proceedings of the ACM on Web Conference 2025, New York, NY, USA, 2025; WWW ’25, pp. 1321–1325. [CrossRef]
- Efeoglu, S.; Paschke, A. Retrieval-Augmented Generation-Based Relation Extraction. Semantic Web 2025, 16, 22104968251385519, [. [CrossRef]
- Zhou, D.W.; Sun, H.L.; Ning, J.; Ye, H.J.; Zhan, D.C. Continual Learning with Pre-Trained Models: A Survey. In Proceedings of the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24; Larson, K., Ed. International Joint Conferences on Artificial Intelligence Organization, 8 2024, pp. 8363–8371. [CrossRef]
- Zhang, Y.; Zhong, V.; Chen, D.; Angeli, G.; Manning, C.D. Position-aware Attention and Supervised Data Improve Slot Filling. In Proceedings of the Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing; Palmer, M.; Hwa, R.; Riedel, S., Eds., Copenhagen, Denmark, 9 2017; pp. 35–45. [CrossRef]
- Han, X.; Zhu, H.; Yu, P.; Wang, Z.; Yao, Y.; Liu, Z.; Sun, M. FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation. In Proceedings of the Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing; Riloff, E.; Chiang, D.; Hockenmaier, J.; Tsujii, J., Eds., Brussels, Belgium, 10 2018; pp. 4803–4809. [CrossRef]
- Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D.S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. Mistral 7B, 2023, [arXiv:cs.CL/2310.06825]. [CrossRef]
- Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023, [arXiv:cs.CL/2307.09288]. [CrossRef]
- Wang, H.; Xiong, W.; Yu, M.; Guo, X.; Chang, S.; Wang, W.Y. Sentence Embedding Alignment for Lifelong Relation Extraction. In Proceedings of the Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Burstein, J.; Doran, C.; Solorio, T., Eds., Minneapolis, Minnesota, 6 2019; pp. 796–806. [CrossRef]
- Jialan, L.; Weishan, K.; Lixi, C.; Hua, Y. Improving Continual Relation Extraction with LSTM and Back Forward Projection. In Proceedings of the 2023 20th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), 2023, pp. 1–5. [CrossRef]
- Wu, F.; Zhang, C.; Tan, Z.; Xu, H.; Ge, B. Continual Few-Shot Relation Extraction with Prompt-Based Contrastive Learning. In Proceedings of the Web and Big Data; Song, X.; Feng, R.; Chen, Y.; Li, J.; Min, G., Eds., Singapore, 2024; pp. 312–327. [CrossRef]
- Tirsogoiu, D.M.; Marginean, A. From learned to new relations through generative models combined with relations clustering and few-shot learning. In Proceedings of the 2023 IEEE 19th International Conference on Intelligent Computer Communication and Processing (ICCP). IEEE, 2023, pp. 381–388. [CrossRef]
- Xiong, W.; Song, Y.; Wang, P.; Li, S. Rationale-Enhanced Language Models are Better Continual Relation Learners. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; Bouamor, H.; Pino, J.; Bali, K., Eds., Singapore, 12 2023; pp. 15489–15497. [CrossRef]
- Zhang, L.; Li, Y.; Wang, Q.; Wang, Y.; Yan, H.; Wang, J.; Liu, J. FPrompt-PLM: Flexible-Prompt on Pretrained Language Model for Continual Few-Shot Relation Extraction. IEEE Transactions on Knowledge and Data Engineering 2024, pp. 1–15. [CrossRef]
- Wang, H.; Li, J.; Wu, H.; Hovy, E.; Sun, Y. Pre-Trained Language Models and Their Applications. Engineering 2023, 25, 51–65. [CrossRef]
- Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D.S.; Casas, D.d.l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. Mistral 7B. arXiv preprint arXiv:2310.06825 2023. [CrossRef]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language Models are Unsupervised Multitask Learners. OpenAI 2019. Accessed: 2024-11-15.
- Chung, H.W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. Scaling Instruction-Finetuned Language Models, 2022. [CrossRef]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Jurafsky, D.; Chai, J.; Schluter, N.; Tetreault, J., Eds., Online, 7 2020; pp. 7871–7880. [CrossRef]
- Wang, G.; Hwang, J.N.; Rose, C.; Wallace, F. Uncertainty sampling based active learning with diversity constraint by sparse selection. In Proceedings of the 2017 IEEE 19th International Workshop on Multimedia Signal Processing (MMSP), 2017, pp. 1–6. [CrossRef]
- Li, M.; Yan, Z.; Li, C. Class Incremental Learning with Important and Diverse Memory. In Proceedings of the Image and Graphics; Lu, H.; Ouyang, W.; Huang, H.; Lu, J.; Liu, R.; Dong, J.; Xu, M., Eds., Cham, 2023; pp. 164–175. [CrossRef]
- Tran, Q.; Thanh, N.X.; Anh, N.H.; Hai, N.L.; Le, T.; Ngo, L.V.; Nguyen, T.H. Preserving Generalization of Language Models in Few-Shot Continual Relation Extraction. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 11 2024; pp. 13771–13784. [CrossRef]
- Madaan, A.; Rajagopal, D.; Tandon, N.; Yang, Y.; Bosselut, A. Conditional Set Generation Using Seq2seq Models. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing; Goldberg, Y.; Kozareva, Z.; Zhang, Y., Eds., Abu Dhabi, United Arab Emirates, 12 2022; pp. 4874–4896. [CrossRef]
- Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs, 2023, [arXiv:cs.LG/2305.14314]. [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models, 2021, [arXiv:cs.CL/2106.09685]. [CrossRef]
- Hu, Y.; Cheng, D.; Zhang, D.; Wang, N.; Liu, T.; Gao, X. Task-Aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning; Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; Berkenkamp, F., Eds. PMLR, 7 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 19153–19164. [CrossRef]
- Wu, T.; Li, X.; Li, Y.F.; Haffari, G.; Qi, G.; Zhu, Y.; Xu, G. Curriculum-Meta Learning for Order-Robust Continual Relation Extraction. Proceedings of the AAAI Conference on Artificial Intelligence 2021, 35, 10363–10369. [CrossRef]
- Han, X.; Dai, Y.; Gao, T.; Lin, Y.; Liu, Z.; Li, P.; Sun, M.; Zhou, J. Continual Relation Learning via Episodic Memory Activation and Reconsolidation. In Proceedings of the Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Jurafsky, D.; Chai, J.; Schluter, N.; Tetreault, J., Eds., Online, 7 2020; pp. 6429–6440. [CrossRef]
- Wang, P.; Song, Y.; Liu, T.; Lin, B.; Cao, Y.; Li, S.; Sui, Z. Learning Robust Representations for Continual Relation Extraction via Adversarial Class Augmentation. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing; Goldberg, Y.; Kozareva, Z.; Zhang, Y., Eds., Abu Dhabi, United Arab Emirates, 12 2022; pp. 6264–6278. [CrossRef]
- Zhao, W.; Cui, Y.; Hu, W. Improving Continual Relation Extraction by Distinguishing Analogous Semantics. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Rogers, A.; Boyd-Graber, J.; Okazaki, N., Eds., Toronto, Canada, 7 2023; pp. 1162–1175. [CrossRef]
- Huang, M.; Xiao, M.; Wang, L.; Du, Y. DP-CRE: Continual Relation Extraction via Decoupled Contrastive Learning and Memory Structure Preservation. In Proceedings of the Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024); Calzolari, N.; Kan, M.Y.; Hoste, V.; Lenci, A.; Sakti, S.; Xue, N., Eds., Torino, Italia, 5 2024; pp. 5338–5349.
- Lai, H.; Liu, X.; Gao, J.; Cheng, J.; Qi, Z.; Xu, Y.; Yao, S.; Zhang, D.; Du, J.; Hou, Z.; et al. A Survey of Post-Training Scaling in Large Language Models. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W.; Nabende, J.; Shutova, E.; Pilehvar, M.T., Eds., Vienna, Austria, 2025; pp. 2771–2791. [CrossRef]
- Lavrinovics, E.; Biswas, R.; Bjerva, J.; Hose, K. Knowledge Graphs, Large Language Models, and Hallucinations: An NLP Perspective. Journal of Web Semantics 2025, 85, 100844. [CrossRef]
- Qin, H.; Ma, X.; Zheng, X.; Li, X.; Zhang, Y.; Liu, S.; Luo, J.; Liu, X.; Magno, M. Accurate LoRA-Finetuning Quantization of LLMs via Information Retention. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning; Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; Berkenkamp, F., Eds. PMLR, 21–27 Jul 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 41498–41516. [CrossRef]







| Task Index | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
| TACRED | ||||||||||
| EA-EMR [24] | 47.1 | 40.1 | 38.3 | 29.9 | 28.4 | 27.3 | 26.9 | 25.8 | 22.9 | 19.8 |
| CML [42] | 57.2 | 51.4 | 41.3 | 39.3 | 35.9 | 28.9 | 27.3 | 26.9 | 24.8 | 23.4 |
| EMAR-BERT [43] | 96.6 | 85.7 | 81.0 | 78.6 | 73.9 | 72.3 | 71.7 | 72.2 | 72.6 | 71.0 |
| RP-CRE [10] | 97.6 | 90.6 | 86.1 | 82.4 | 79.8 | 77.2 | 75.1 | 73.7 | 72.4 | 72.4 |
| ACA [44] | 98.2 | 93.8 | 89.9 | 85.9 | 84.2 | 82.7 | 80.5 | 78.4 | 78.6 | 77.5 |
| CRL [7] | 97.7 | 93.2 | 89.8 | 84.7 | 84.1 | 81.3 | 80.2 | 79.1 | 79.0 | 78.0 |
| KIP-Framework [13] | 98.3 | 95.0 | 90.8 | 87.5 | 85.3 | 84.3 | 82.1 | 80.2 | 79.6 | 78.6 |
| CEAR [45] | 97.9 | 93.7 | 90.7 | 86.6 | 84.7 | 84.3 | 81.9 | 80.4 | 80.2 | 79.3 |
| CREST [8] | 97.3 | 91.4 | 82.3 | 82.5 | 79.2 | 75.8 | 78.8 | 77.4 | 78.6 | 79.4 |
| DP-CRE [46] | 97.8 | 93.8 | 91.5 | 87.5 | 85.7 | 84.2 | 82.9 | 81.3 | 81.5 | 80.7 |
| Flan-T5 Base | 96.1 ± 4.2 | 96.2 ± 3.2 | 95.7 ± 1.1 | 96.0 ± 1.0 | 95.7 ± 2.0 | 95.4 ± 1.7 | 96.0 ± 0.9 | 96.0 ± 1.3 | 96.3 ± 0.9 | 95.8 ± 1.6 |
| Mistral-7B | 96.6 ± 6.1 | 95.3 ± 4.4 | 96.4 ± 1.7 | 95.9 ± 1.5 | 96.6 ± 1.1 | 97.0 ± 1.2 | 96.8 ± 1.3 | 96.9 ± 1.0 | 95.8 ± 3.1 | 96.9 ± 1.0 |
| Llama2-7B | 57.7 ± 9.9 | 57.6 ± 7.5 | 54.9 ± 7.6 | 55.8 ± 6.3 | 57.6 ± 4.6 | 62.0 ± 4.5 | 62.4 ± 3.7 | 65.3 ± 2.8 | 67.7 ± 2.3 | 70.6 ± 2.9 |
| Llama-3.1-8B | 89.1 ± 6.9 | 85.0 ± 6.2 | 85.6 ± 4.9 | 84.4 ± 6.8 | 85.1 ± 5.5 | 83.9 ± 6.1 | 83.1 ± 5.2 | 81.7 ± 5.1 | 80.0 ± 7.4 | 78.7 ± 7.3 |
| Qwen2.5-7B | 93.0 ± 5.5 | 92.0 ± 3.4 | 91.2 ± 2.2 | 91.4 ± 2.8 | 91.8 ± 1.7 | 92.0 ± 1.8 | 92.0 ± 1.8 | 91.9 ± 1.6 | 91.9 ± 0.5 | 91.3 ± 0.7 |
| FewRel | ||||||||||
| EA-EMR [24] | 88.5 | 69.0 | 59.1 | 54.2 | 47.8 | 46.1 | 43.1 | 40.7 | 38.6 | 35.1 |
| CML [42] | 91.2 | 74.8 | 68.2 | 58.2 | 53.7 | 50.4 | 47.8 | 44.4 | 43.1 | 39.7 |
| EMAR-BERT [43] | 98.8 | 89.1 | 89.5 | 85.7 | 83.6 | 84.8 | 79.3 | 80.0 | 77.1 | 73.8 |
| RP-CRE [10] | 97.9 | 92.7 | 91.6 | 89.2 | 88.4 | 86.8 | 85.1 | 84.1 | 82.2 | 81.5 |
| KIP-Framework [13] | 98.4 | 93.5 | 92.0 | 91.2 | 90.0 | 88.2 | 86.9 | 85.6 | 84.1 | 82.5 |
| CRL [7] | 98.2 | 94.6 | 92.5 | 90.5 | 89.4 | 87.9 | 86.9 | 85.6 | 84.5 | 83.1 |
| CEAR [45] | 98.3 | 95.6 | 93.5 | 92.0 | 90.8 | 89.3 | 88.0 | 86.8 | 85.6 | 84.0 |
| ACA [44] | 98.4 | 95.1 | 93.0 | 91.5 | 90.5 | 88.9 | 87.9 | 86.7 | 85.8 | 84.4 |
| CREST [8] | 98.7 | 93.6 | 93.8 | 92.3 | 91.0 | 89.9 | 87.6 | 86.7 | 86.0 | 84.8 |
| DP-CRE [46] | 98.5 | 95.4 | 93.7 | 92.1 | 90.9 | 89.4 | 88.5 | 87.4 | 86.3 | 85.1 |
| Flan-T5 Base | 96.7 ± 1.5 | 94.8 ± 1.0 | 95.1 ± 1.9 | 93.5 ± 2.9 | 93.2 ± 2.2 | 92.4 ± 1.4 | 91.4 ± 1.2 | 91.7 ± 1.5 | 91.0 ± 1.3 | 89.6 ± 1.3 |
| Mistral-7B | 96.0 ± 5.4 | 94.6 ± 5.4 | 94.7 ± 1.8 | 93.6 ± 2.4 | 93.6 ± 1.3 | 92.3 ± 1.1 | 92.5 ± 1.2 | 91.9 ± 0.7 | 92.3 ± 2.7 | 91.4 ± 0.7 |
| Llama2-7B | 15.4 ± 3.3 | 27.8 ± 2.9 | 38.9 ± 4.2 | 44.2 ± 3.3 | 52.1 ± 3.5 | 57.4 ± 1.9 | 62.2 ± 1.4 | 67.7 ± 0.4 | 69.4 ± 1.1 | 71.3 ± 1.1 |
| Llama-3.1-8B | 74.8 ± 8.6 | 78.1 ± 3.6 | 81.5 ± 2.5 | 80.2 ± 4.4 | 82.1 ± 4.0 | 82.1 ± 3.7 | 82.1 ± 3.3 | 81.2 ± 2.0 | 81.5 ± 1.8 | 81.1 ± 2.5 |
| Qwen2.5-7B | 83.1 ± 10.2 | 83.9 ± 5.5 | 85.4 ± 4.4 | 86.1 ± 4.1 | 87.8 ± 2.6 | 87.9 ± 2.7 | 88.4 ± 1.7 | 88.5 ± 1.9 | 88.6 ± 1.0 | 88.3 ± 0.6 |
| Method | TACRED | FewRel | Average | |||
|---|---|---|---|---|---|---|
| w (%) | a (%) | w (%) | a (%) | w (%) | a (%) | |
| EA-EMR [24] | 23.0 | 30.0 | 49.0 | 61.2 | 36.0 | 45.6 |
| EMAR-BERT [43] | 31.0 | 36.3 | 53.8 | 68.1 | 42.4 | 52.2 |
| CML [42] | 43.7 | 45.3 | – | – | – | – |
| KIP-Framework [13] | 91.10 | 91.60 | 96.30 | 96.60 | 93.70 | 94.10 |
| Flan-T5 Base | 95.76 ± 1.41 | 95.78 ± 1.35 | 92.68 | 92.70 | ||
| Mistral-7B | 94.93 ± 0.17 | 94.93 ± 0.17 | 95.91 | 95.85 | ||
| Llama2-7B | 71.50 | 71.08 | ||||
| Llama-3.1-8B | ||||||
| Qwen2.5-7B | ||||||
| Model | FewRel (↑) | TACRED (↑) |
|---|---|---|
| Flan-T5 Base (250M) | ||
| Mistral-7B-Instruct-v0.2 | ||
| Qwen2.5-7B-Instruct | ||
| Llama2-7B-hf | ||
| Llama-3.1-8B-Instruct |
| Comparison | Mistral-7B | Flan-T5 Base | Llama2-7B-hf | Llama3.1 | Qwen2.5 |
|---|---|---|---|---|---|
| No Replay vs. 5 | 0.67 | 0.11 | 0.05 | 0.14 | 0.62 |
| No Replay vs. 10 | 0.86 | 0.06 | 0.02 | 0.25 | 0.83 |
| No Replay vs. 15 | 1.00 | 0.03 | 0.17 | 0.31 | 0.43 |
| Test | Flan-T5 Base | Mistral-7B |
|---|---|---|
| same test set | 92.07 | 97.21 |
| full test set | 97.65 | 98.37 |
| Run | Beam | Nucleus (Top-p) | Diff (Beam - Nucleus) | |||
|---|---|---|---|---|---|---|
| Acc (%) | Undefined (%) | Acc (%) | Undefined (%) | Acc (%) | Undefined (%) | |
| 1 | 14.37 | 0.27 | 15.09 | 16.70 | -2.50 | -16.43 |
| 2 | 51.25 | 7.05 | 52.59 | 21.43 | -1.34 | -14.38 |
| 3 | 49.91 | 18.75 | 51.07 | 26.79 | -1.16 | -8.04 |
| 4 | 57.05 | 2.68 | 49.20 | 11.43 | 7.85 | -8.75 |
| 5 | 43.75 | 7.41 | 31.70 | 15.71 | 12.05 | -8.30 |
| Mean | 43.27 | 7.23 | 39.93 | 18.41 | 3.34 | -11.18 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.