Submitted:
26 June 2026
Posted:
29 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Table Retrieval Models Based on Traditional Methods
2.2. Table Retrieval Methods Based on Pre-Trained Language Models
2.3. Table Retrieval Based on Table Matching

3. Problem Formulation
4. Methodology
4.1. Query-Table Relevance
4.1.1. Table-Level Relevance Calculation
4.1.2. Value-Level Relevance Calculation
4.1.3. Comprehensive Relevance Score
4.2. Inter-Table Correlation
4.2.1. Inter-Table Similarity Calculation
4.2.2. Chain Relationship Modeling
4.3. Comprehensive Table Scoring and Target Table Selection
4.3.1. Comprehensive Score Calculation
4.3.2. Target Table Selection
5. Experimental Design and Result Analysis
5.1. Evaluation Metrics
5.2. Datasets
5.2.1. Private Dataset in the Petroleum Domain
5.2.2. Public Datasets
5.3. Experimental Setup
5.4. Baseline Models
5.5. Main Experimental Results
5.6. Ablation Study
5.7. Hyperparameter Sensitivity Analysis
5.7.1. Impact of Weight Parameters ()
5.7.2. Impact of Decay Factor and Thresholds ()
6. Conclusions
References
- Reiter E, Dale R. Building applied natural language generation systems[J]. Natural Language Engineering, 1997, 3(1): 57-87.
- Cafarella M J, Halevy A, Wang D Z, et al. WebTables: Exploring the Power of Tables on the Web[J]. Proceedings of the VLDB Endowment, 2008, 1(1): 538-549. [CrossRef]
- Cafarella M J, Halevy A, Khoussainova N. Data Integration for the Relational Web[J]. Proceedings of the VLDB Endowment, 2009, 2(1): 1090-1101.
- Venetis P, Halevy A Y, Madhavan J, et al. Recovering Semantics of Tables on the Web[J]. Proceedings of the VLDB Endowment, 2011, 4(9): 528-538.
- Zhang S, Balog K. Ad hoc table retrieval using semantic similarity[C]//ACM. Proceedings of the 2018 World Wide Web Conference. Lyon, France: ACM, 2018: 1553-1562.
- Zelle, J. M, Mooney, R. J. Learning to parse database queries using inductive logic programming[C]//AAAI. Proceedings of the Thirteenth National Conference on Artificial Intelligence. Portland, Oregon: AAAI Press, 1996: 1050-1055.
- Trabelsi M, Davison B D, Heflin J. Improved table retrieval using multiple context embeddings for attributes[C]. 2019 IEEE international conference on big data (Big Data). Angeles, California, USA: IEEE, 2019: 1238-1244.
- Chen Z, Jia H, Heflin J, et al. Leveraging schema labels to enhance dataset search[C]//Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020. Lisbon, Portugal: Springer International Publishing, 2020: 267-280.
- Lai S, Xu L, Liu K, et al. Recurrent convolutional neural networks for text classification[C]//AAAI. Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence. Austin, Texas: AAAI Press, 2015: 2267-2273.
- Devlin J, Chang M W, Lee K, et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding[C]//Association for Computational Linguistics. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Minneapolis, Minnesota: Association for Computational Linguistics, 2019: 4171-4186.
- Pengcheng Y, Graham N, Wen-tau Y, Sebastian R, et al. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data[C]//Association for Computational Linguistics. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020: 5116-5123.
- Liu Y H, Ott M, Goyal N, et al. RoBERTa: A robustly optimized BERT pretraining approach[C]//Proceedings of the Eighth International Conference on Learning Representations. Online: OpenReview.net, 2020.
- Herzig J, Nowak P K, Mueller T, et al. TaPas: Weakly Supervised Table Parsing via Pre training[C]//Association for Computational Linguistics. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020: 4320-4333.
- Wang Z, Dong H, Jia R, et al. Tuta: Tree-based transformers for generally structured table pre-training[C]//ACM. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Virtual Event, Singapore: ACM, 2021: 1780-1790.
- Hiroshi I, Dung T, Varun M, Mohit I, et al. TABBIE: Pretrained Representations of Tabular Data[C]//Association for Computational Linguistics. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Association for Computational Linguistics, 2021: 3446-3456.
- Dey R, Salem F M. Gate-variants of gated recurrent unit (GRU) neural networks[C]//IEEE. Proceedings of the 2017 IEEE 60th International Midwest Symposium on Circuits and Systems. Boston, MA, USA: IEEE, 2017: 1597-1603.
- Deng X, Sun H, Lees A, et al. Turl: Table understanding through representation learning[J]. ACM SIGMOD Record, 2022, 51(1): 33-40. [CrossRef]
- Chen D, O’Bray L, Borgwardt K. Structure-aware transformer for graph representation learning[C]//PMLR. Proceedings of the 39th International Conference on Machine Learning. Baltimore, Maryland, USA: PMLR, 2022: 3469-3489.
- Mingyu Z, Xinwei F, Qingyi S, Qiaoqiao S, Zheng L, Wenbin J, Weiping W, et al. Multimodal Table Understanding[C]//Association for Computational Linguistics. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: Association for Computational Linguistics, 2024: 9102-9124.
- Ahmad A, Maik T, Julian E, Wolfgang L, Robert W, et al. Towards a Hybrid Imputation Approach Using Web Tables[C]//Unknown Editor. Proceedings of Big Data Computing. Location Unknown: Publisher Unknown, 2015: 21-30.
- Oliver L, Dominique R, Petar R, Robert M, Heiko P, Christian B, et al. The Mannheim Search Join Engine[J]. Journal of Web Semantics, 2015, 35: 159-166. [CrossRef]
- Anish D S, Lujun F, Nitin G, Alon Y H, Hongrae L, Fei W, Reynold X, Cong Y, et al. Finding Related Tables[C]//ACM. Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data. Scottsdale, Arizona, USA: ACM, 2012: 817-828.
- Chen P B, Zhang Y, Roth D. Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval[C]//Association for Computational Linguistics. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024: 2687-2699.
- Tao Y, Rui Z, Kai Y, Michihiro Y, Dongxu W, Zifan L, James M, Irene L, Qingning Y, Shanelle R, Zilin Z, Dragomir R, et al. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task[C]//Association for Computational Linguistics. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Brussels, Belgium: Association for Computational Linguistics, 2018: 3911-3921.
- Agarwal S. Data mining: Data mining concepts and techniques[C]//IEEE. Proceedings of the 2013 International Conference on Machine Intelligence and Research Advancement. Katra, India: IEEE, 2013: 203-207.
- Lee C H, Polozov O, Richardson M. KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers[J]. ACM Journal of Experimental Algorithmics, 2021, 26: 1-15.
- Gautier I, Mathilde C, Lucas H, Sebastian R, Piotr B, Armand J, Edouard G, et al. Unsupervised Dense Information Retrieval with Contrastive Learning[DB/OL]. [2022]. https://arxiv.org/abs/2112.09131.
- Gautier I, Mathilde C, Lucas H, Sebastian R, Piotr B, Armand J, Edouard G, et al. Unsupervised Dense Information Retrieval with Contrastive Learning[DB/OL]. [2022]. https://arxiv.org/abs/2112.09131.
- Karpukhin V, Oguz B, Min S, et al. Dense Passage Retrieval for Open-Domain Question Answering[C]//Association for Computational Linguistics. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Online: Association for Computational Linguistics, 2023: 6769-6781.
- Jonathan H, Thomas M, Syrine K, Julian M E, et al. Open Domain Question Answering over Tables Via Dense Retrieval[C]//Association for Computational Linguistics. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Association for Computational Linguistics, 2021: 512-519.







| dataset | train | valid |
|---|---|---|
| Bird | 4476 | 639 |
| Spider | 2494 | 356 |
| Dataset | Model | Top-2 | Top-5 | NDCG@2 | NDCG@5 | ||||
|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | ||||
| Spider | Contriever | 76.29 | 72.78 | 74.49 | 39.84 | 93.38 | 55.85 | 76.07 | 71.57 |
| TaBERT | 78.59 | 72.79 | 75.58 | 40.76 | 93.67 | 56.80 | 77.01 | 72.33 | |
| DPR | 81.20 | 77.23 | 79.15 | 40.93 | 94.97 | 57.47 | 82.72 | 72.64 | |
| DTR | 81.45 | 77.47 | 79.41 | 41.05 | 95.88 | 57.49 | 83.88 | 72.84 | |
| Relatab(Ours) | 81.50 | 79.21 | 80.34 | 40.88 | 97.55 | 57.62 | 84.61 | 73.96 | |
| Bird | Contriever | 65.03 | 59.46 | 62.12 | 37.04 | 82.96 | 51.21 | 70.63 | 67.27 |
| TaBERT | 66.23 | 59.17 | 62.50 | 37.96 | 83.21 | 52.14 | 72.25 | 68.08 | |
| DPR | 64.75 | 58.70 | 61.58 | 37.11 | 82.72 | 51.23 | 75.31 | 67.37 | |
| DTR | 64.97 | 58.93 | 61.80 | 37.21 | 82.98 | 51.38 | 75.51 | 67.47 | |
| Relatab(Ours) | 65.43 | 61.26 | 63.28 | 37.86 | 85.47 | 52.48 | 76.12 | 68.40 | |
| Cementing Tables |
Contriever | 67.36 | 73.44 | 70.27 | 39.91 | 94.06 | 56.04 | 72.32 | 68.40 |
| TaBERT | 68.81 | 74.22 | 71.41 | 40.38 | 94.60 | 56.60 | 73.03 | 69.38 | |
| DPR | 71.64 | 76.98 | 73.80 | 41.66 | 96.25 | 58.10 | 76.56 | 71.32 | |
| DTR | 71.76 | 76.03 | 73.83 | 41.72 | 96.32 | 58.22 | 76.66 | 73.96 | |
| Relatab(Ours) | 72.79 | 77.88 | 75.25 | 41.66 | 97.20 | 58.32 | 79.21 | 75.35 | |
| Dataset | F1 | |||
|---|---|---|---|---|
| Spider | 78.41 | 80.34 | 79.82 | 76.34 |
| Bird | 62.85 | 63.28 | 63.04 | 59.11 |
| CementingTables | 72.87 | 75.25 | 73.54 | 70.93 |
| Dataset | Model | P | R | F1 |
|---|---|---|---|---|
| Spider | Relatab (QT) | 79.57 | 76.77 | 78.14 |
| Relatab (QT+TT) | 81.50 | 79.21 | 80.34 | |
| Bird | Relatab (QT) | 66.53 | 58.99 | 62.53 |
| Relatab (QT+TT) | 65.43 | 61.26 | 63.28 | |
| CementingTables | Relatab (QT) | 68.79 | 75.23 | 71.87 |
| Relatab (QT+TT) | 72.79 | 77.88 | 75.25 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).