Submitted:
04 October 2025
Posted:
06 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Literature Review
3. Methodology
3.1. Symbol and Object Definitions
3.2. Spatiotemporal Lattice Structure and Consistency Criterion Modeling
3.3. Weight Estimation and Monotonic Projection
3.4. Multi-Feature Scoring Model
3.5. Inference, Decision-Making, and Uncertainty
- When and , automatically align to .
- Otherwise, mark it as “requiring human review.”
3.6. LLM-Based Data Augmentation
3.7. Model Training and Implementation Details
- Text normalization (standardizing numbers and units, and normalizing time expressions and place names);
- Feature extraction (temporal anchors and granularity, administrative levels of locations, enumerative field nodes, accident consequence vectors, and sentence semantic similarity);
- Consistency score computation and caching (columnar storage indexed by sample–candidate pairs).
4. Experimental Setup
4.1. Data Sources and Processing
4.2. Baseline Models
4.3. Parameter Settings
4.4. Evaluation Metrics
5. Experimental Results and Analysis
5.1. Overall Performance Analysis

5.2. Stability Analysis Across Different Sources

5.3. Stability Analysis Under Different Decision Thresholds
5.4. Automatic Labeling Efficacy Analysis
6. Results
7. Discussion
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Nothman J, Honnibal M, Hachey B, et al. Event linking: Grounding event reference in a news archive[C]//Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2012: 228-232.
- Humphreys K, Gaizauskas R, Azzam S. Event coreference for information extraction[C]//Operational Factors in Practical, Robust Anaphora Resolution for Unrestricted Texts. 1997.
- Ahn D. The stages of event extraction[C]//Proceedings of the Workshop on Annotating and Reasoning about Time and Events. 2006: 1-8.
- Raghunathan K, Lee H, Rangarajan S, Chambers N, Surdeanu M, Jurafsky D, Manning C D. A multi-pass sieve for coreference resolution[C]//NAACL-HLT 2010. 2010: 492-501.
- Lu J, Ng V. Event coreference resolution with multi-pass sieves[C]//LREC 2016. 2016: 3996-4003.
- Cybulska A, Vossen P. Using a sledgehammer to crack a nut? Lexical diversity and event coreference resolution[C]//LREC 2014. 2014: 4545-4552.
- Cybulska A, Vossen P. “Bag of events” approach to event coreference resolution: supervised classification of event templates[J]. International Journal of Computational Linguistics & Applications, 2015, 6(1): 11-27.
- Cybulska A, Vossen P. Translating granularity of event slots into features for event coreference resolution[C]//Proc. of the 3rd Workshop on EVENTS. 2015: 1-10.
- Hovy E, Mitamura T, Verdejo F, et al. Events are not simple: Identity, non-identity, and quasi-identity[C]//Proc. of the First Workshop on EVENTS. 2013: 21-28.
- Araki J, Hovy E, Mitamura T. Evaluation for partial event coreference[C]//2nd Workshop on EVENTS. 2014: 68-76.
- Chen Z, Ji H. Graph-based event coreference resolution[C]//ACL-IJCNLP Workshop on Graph-based Methods for NLP. 2009: 54-57.
- Chen Z, Ji H, Haralick R. A pairwise event coreference model, feature impact and evaluation for event coreference resolution[C]//Workshop on Events in Emerging Text Types. 2009: 17-22.
- Lee H, Recasens M, Chang A, Surdeanu M, Jurafsky D. Joint entity and event coreference resolution across documents[C]//EMNLP-CoNLL 2012. 2012: 489-500.
- Choubey P K, Huang R. Event coreference resolution by iteratively unfolding inter-dependencies among events[C]//EMNLP 2017. 2017: 2124-2133.
- Yang B, Cardie C, Frazier P. A hierarchical distance-dependent Bayesian model for event coreference resolution[J]. TACL, 2015, 3: 517-528. [CrossRef]
- Kenyon-Dean K, Cheung J C K, Precup D. Resolving event coreference with supervised representation learning and clustering-oriented regularization[C]//*SEM 2018. 2018: 1-10.
- Peng H, Song Y, Roth D. Event detection and co-reference with minimal supervision[C]//EMNLP 2016. 2016: 392-402.
- Barhom S, Shwartz V, Eirew A, Bugert M, Reimers N, Dagan I. Revisiting joint modeling of cross-document entity and event coreference resolution[C]//ACL 2019. 2019: 4179-4189.
- Cremisini A, Finlayson M A. New insights into cross-document event coreference: Systematic comparison and a simplified approach[C]//EMNLP 2020 Workshop on Novel Evaluation Approaches for Text Generation. 2020: 7-16.
- Allaway E, Wang S, Ballesteros M. Sequential cross-document coreference resolution[C]//EMNLP 2021. 2021: 4659-4671.
- Han R, Peng T, Yang C, et al. Is information extraction solved by chatgpt? an analysis of performance, evaluation criteria, robustness and errors[J]. arXiv preprint arXiv:2305.14450, 2023: 48.
- Nath A, Manafi S, Chelle A, et al. Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge Distillation[J]. arXiv preprint arXiv:2404.03196, 2024.
- Wang X, Zhou W, Zu C, et al. Instructuie: Multi-task instruction tuning for unified information extraction[J]. arXiv preprint arXiv:2304.08085, 2023.
- Cybulska A, Vossen P. Guidelines for ECB+ annotation of events and their coreference[J]. Technical Report, 2014.
- Ellis J, Getman J, Fore D, et al. Overview of Linguistic Resources for the TAC KBP 2015 Evaluations: Methodologies and Results[C]//Tac. 2015.
- Chen B, Su J, Pan S J, Tan C L. A unified event coreference resolution by integrating multiple resolvers[C]//IJCNLP 2011. 2011: 102-110.
- Lu J, Venugopal D, Gogate V, Ng V. Joint inference for event coreference resolution[C]//COLING 2016. 2016: 3264-3275.
- Lu J, Ng V. Joint learning for event coreference resolution[C]//ACL 2017. 2017: 90-101.
- Bejan C A, Harabagiu S. Unsupervised event coreference resolution with rich linguistic features[C]//Proc. of ACL 2010. 2010: 1412-1422.
- Bejan C A, Harabagiu S. Unsupervised event coreference resolution[J]. Computational Linguistics, 2014, 40(2): 311-347.
- Araki J, Mitamura T. Joint event trigger identification and event coreference resolution with structured perceptron[C]//EMNLP 2015. 2015: 2074-2080.
- Araki J, Hovy E, Mitamura T. Evaluation for partial event coreference[C]//2nd Workshop on EVENTS. 2014: 68-76.
- Upadhyay S, Gupta N, Christodoulopoulos C, Roth D. Revisiting the evaluation for cross-document event coreference[C]//COLING 2016. 2016: 1949-1960.
- Joshi M, Levy O, Weld D S, Zettlemoyer L. BERT for coreference resolution: Baselines and analysis[C]//EMNLP-IJCNLP 2019. 2019: 5803-5808.
- Joshi M, Chen D, Liu Y, Weld D S, Zettlemoyer L, Levy O. SpanBERT: Improving pre-training by representing and predicting spans[J]. TACL, 2020, 8: 64-77. [CrossRef]
- Cattan A, Eirew A, Stanovsky G, Joshi M, Dagan I. Streamlining cross-document coreference resolution: Evaluation and modeling[OL]. arXiv:2009.11032, 2020.
- Yu X, Yin W, Roth D. Paired representation learning for event and entity coreference[OL]. arXiv:2010.12808, 2020.
- Zeng Y, Jin X, Guan S, Guo J, Cheng X. Event coreference resolution with their paraphrases and argument-aware embeddings[C]//COLING 2020. 2020: 1737-1747.
- Caciularu A, Ravfogel S, Bansal R, et al. CDLM: Cross-Document Language Modeling[OL]. Findings of ACL 2021 / arXiv:2101.00406, 2021.
- Beltagy I, Peters M E, Cohan A. Longformer: The long-document transformer[OL]. arXiv:2004.05150, 2020.
- Ma Y, Cao Y, Hong Y C, et al. Large language model is not a good few-shot information extractor, but a good reranker for hard samples![J]. arXiv preprint arXiv:2303.08559, 2023.
- Li J, Jia Z, Zheng Z. Semi-automatic data enhancement for document-level relation extraction with distant supervision from large language models[J]. arXiv preprint arXiv:2311.07314, 2023.
- Ding B, Min Q, Ma S, et al. A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference Resolution[J]. arXiv preprint arXiv:2404.01921, 2024.
- Min Q, Guo Q, Hu X, et al. Synergetic event understanding: A collaborative approach to cross-document event coreference resolution with large language models[J]. arXiv preprint arXiv:2406.02148, 2024.
- Bugert M, Reimers N, Gurevych I. Generalizing cross-document event coreference resolution across multiple corpora[J]. Computational Linguistics, 2021, 47(3): 575-614. [CrossRef]
- Field D A. Laplacian smoothing and Delaunay triangulations[J]. Communications in applied numerical methods, 1988, 4(6): 709-712. [CrossRef]
- Yuan Q, Cong G, Thalmann N M. Enhancing naive bayes with various smoothing methods for short text classification[C]//Proceedings of the 21st international conference on world wide web. 2012: 645-646.Yuan Q, Cong G, Thalmann N M. Enhancing naive bayes with various smoothing methods for short text classification[C]//Proceedings of the 21st international conference on world wide web. 2012: 645-646.
- Neelon B, Dunson D B. Bayesian isotonic regression and trend analysis[J]. Biometrics, 2004, 60(2): 398-406. [CrossRef]
- De Leeuw J, Hornik K, Mair P. Isotone optimization in R: pool-adjacent-violators algorithm (PAVA) and active set methods[J]. Journal of statistical software, 2010, 32: 1-24.
- Fellegi I P, Sunter A B. A theory for record linkage[J]. Journal of the American statistical association, 1969, 64(328): 1183-1210.
- Eirew A, Caciularu A, Dagan I. Cross-document event coreference search: Task, dataset and modeling[J]. arXiv preprint arXiv:2210.12654, 2022.
- Wu L, Petroni F, Josifoski M, et al. Scalable zero-shot entity linking with dense entity retrieval[J]. arXiv preprint arXiv:1911.03814, 2019.
- Nath A, Manafi S, Chelle A, et al. Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge Distillation[J]. arXiv preprint arXiv:2404.03196, 2024.
- Hsu I, Xue Z, Pochh N, et al. Argument-Aware Approach To Event Linking[J]. arXiv preprint arXiv:2403.15097, 2024.
- Kim T, Oh J, Kim N Y, et al. Comparing kullback-leibler divergence and mean squared error loss in knowledge distillation[J]. arXiv preprint arXiv:2105.08919, 2021.



| Method | Hit@1 | Hit@5 | ROC-AUC | Precision | Recall | F1 | Brier | ECE |
|---|---|---|---|---|---|---|---|---|
| F-S[50] | 33.49 | 69.31 | 83.24 | 74.83 | 46.97 | 57.71 | 0.17 | 3.89 |
| CDECS[51] | 38.58 | 74.71 | 87.56 | 66.71 | 52.43 | 58.71 | 0.15 | 2.79 |
| BLINK[52] | 36.84 | 72.08 | 84.55 | 63.18 | 50.62 | 56.21 | 0.16 | 3.16 |
| AAEL[53] | 42.12 | 75.63 | 87.06 | 69.79 | 53.39 | 60.5 | 0.15 | 2.67 |
| LLM[54] | 40.93 | 77.88 | 86.69 | 67.12 | 55.31 | 60.65 | 0.15 | 2.52 |
| Ours | 41.51 | 77.33 | 87.34 | 73.92 | 54.07 | 62.46 | 0.14 | 1.97 |
| Source | Metric | F-S | CDECS | BLINK | AAEL | LLM | Ours |
|---|---|---|---|---|---|---|---|
| Govt. portals | Hit@1 | 42.21 | 40.33 | 38.17 | 43.56 | 41.82 | 42.97 |
| F1 | 59.87 | 60.25 | 57.89 | 61.34 | 61.02 | 63.11 | |
| Prof’l inst. | Hit@1 | 32.74 | 39.15 | 36.92 | 42.8 | 41.35 | 42.08 |
| F1 | 56.92 | 59.87 | 56.34 | 61.05 | 61.44 | 63.27 | |
| Pub. media | Hit@1 | 31.58 | 35.47 | 34.26 | 39.21 | 38.74 | 39.87 |
| F1 | 55.33 | 56.82 | 54.17 | 58.96 | 59.21 | 60.53 |
| Model | Threshold = 0.3 | Threshold = 0.5 | Threshold = 0.7 | Range (Δ) |
|---|---|---|---|---|
| F-S | 55.2 | 57.7 | 53.8 | 3.9 |
| CDECS | 56.8 | 58.7 | 55.3 | 3.4 |
| BLINK | 53.4 | 56.2 | 51.7 | 4.5 |
| AAEL | 58.9 | 60.5 | 57.2 | 3.3 |
| LLM | 59.1 | 60.7 | 56.9 | 3.8 |
| Ours | 61.2 | 62.5 | 60.8 | 1.7 |
| Model | Auto | Acc (HC auto) | Review | TPR (review set) |
|---|---|---|---|---|
| F-S | 58.8 | 92.3 | 41.2 | 68.5 |
| CDECS | 67.5 | 94.1 | 32.5 | 72.3 |
| BLINK | 62.6 | 93.2 | 37.4 | 70.1 |
| AAEL | 70.2 | 95.4 | 29.8 | 75.6 |
| LLM | 71.4 | 95.8 | 28.6 | 76.2 |
| Ours | 77.7 | 97.5 | 22.3 | 81.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).