Submitted:
01 June 2025
Posted:
04 June 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
I want to find proteins that are associated with diseases and have specific functional annotations and UniProt annotation scores
1.1. Our Approach
1.2. Contributions
- We introduce DWCP, a structured and generalizable specification that encodes deep web access workflows in a machine-readable format, enabling automated and reproducible data retrieval across diverse web interfaces.
- We develop FAIRFind, an end-to-end system that autonomously discovers, semantically interprets, and executes deep web access paths associated with FAIRsharing-listed resources, making previously inaccessible data programmatically available.
- We implement a semantic query engine that matches natural language user queries to DWCP-encoded workflows, enabling users to retrieve scientific data without knowing the structure of the underlying database or interface.
- We design and conduct a comprehensive evaluation across multiple language models, query types, and interface complexities and demonstrate high precision and success rates in structured data extraction and access workflow execution.
1.3. Impact and Vision
2. Related Works
3. The DWCP Protocol: Deep Web Communication for FAIR Indexing
3.1. Formal Specification of DWCP
3.2. Message Structure and Data Model
3.3. Protocol Operations
4. System Logic
4.1. Architecture Overview
4.2. Metadata Harvesting
4.3. Form Discovery and Interaction Modeling
4.3.1. Detecting Interactive Inputs
4.3.2. Semantic Filtering and Selection
4.3.3. Simulating Form Interaction
4.3.4. Constructing Access Path Abstractions

4.4. Data Extraction and Access Normalization
4.4.1. Table Detection and Parsing
4.4.2. Access Normalization Across Types
- Data is extracted from the page rendered after simulating a completed form submission. The DWCP output includes the full form object, and the access-type is set to “form”.
- Pages containing static tables are parsed directly without interaction. The resulting output includes headers and sample-rows, and the access-type is set to “html-table”.
- When data is provided via downloadable links (e.g., CSV, XLS), the system downloads and parses the file. The first row is treated as headers, and the next two rows are extracted as samples.
| Algorithm 1: DWCP-Based Deep Web Indexing Pipeline |
![]() |
4.5. Semantic Query Engine
4.5.1. Vectorized Document Indexing
- HTML Content Chunks: Extracted from the rendered page’s visible content.
- Schema Blocks: Representations of DWCP access path metadata.
4.5.2. Stage-One Retrieval: Site Ranking
4.5.3. Stage-Two Retrieval: Access Path Matching
4.5.4. Dynamic Interaction and Data Extraction
4.5.5. Output Rendering
4.6. Prompt Engineering and Output Structuring
5. Evaluation Framework
5.1. Benchmark Construction
- A simple query using exact field names and form inputs.
- A complex query that captures intent abstractly.
5.2. Evaluation Metrics
5.2.1. Exact Match Accuracy
5.2.2. Data Retrieval Success
- LLaMA 3.3 (70B)
- Qwen 3
- DeepSeek-R1
- Gemma
5.3. Execution Pipeline
| Algorithm 2:Evaluation Pipeline for Semantic Query Execution |
|
Input: Natural language query q, ground truth
Output: Evaluation result and response time
|
6. Results and Analysis
6.1. Exact Match Performance
6.2. Data Retrieval Performance
6.3. Retrieval Time and Efficiency
7. Discussion
8. Conclusions
Acknowledgments
System Availability
References
- Consortium, U. UniProt: a worldwide hub of protein knowledge. Nucleic acids research 2019, 47, D506–D515. [Google Scholar] [CrossRef] [PubMed]
- Anderson, N.L.; Anderson, N.G. The human plasma proteome: history, character, and diagnostic prospects. Molecular & cellular proteomics 2002, 1, 845–867. [Google Scholar]
- Keshava Prasad, T.; Goel, R.; Kandasamy, K.; Keerthikumar, S.; Kumar, S.; Mathivanan, S.; Telikicherla, D.; Raju, R.; Shafreen, B.; Venugopal, A.; et al. Human protein reference database—2009 update. Nucleic acids research 2009, 37, D767–D772. [Google Scholar] [CrossRef] [PubMed]
- Piñero, J.; Bravo, À.; Queralt-Rosinach, N.; Gutiérrez-Sacristán, A.; Deu-Pons, J.; Centeno, E.; García-García, J.; Sanz, F.; Furlong, L.I. DisGeNET: a comprehensive platform integrating information on human disease-associated genes and variants. Nucleic acids research 2016, gkw943. [Google Scholar] [CrossRef] [PubMed]
- Davis, A.P.; Wiegers, T.C.; Johnson, R.J.; Sciaky, D.; Wiegers, J.; Mattingly, C.J. Comparative toxicogenomics database (CTD): update 2023. Nucleic acids research 2023, 51, D1257–D1262. [Google Scholar] [CrossRef] [PubMed]
- Sansone, S.A.; McQuilton, P.; Rocca-Serra, P.; Gonzalez-Beltran, A.; Izzo, M.; Lister, A.L.; Thurston, M.; Community, F. FAIRsharing as a community approach to standards, repositories and policies. Nature biotechnology 2019, 37, 358–367. [Google Scholar] [CrossRef] [PubMed]
- Rigden, D.J.; Fernández, X.M. The 27th annual Nucleic Acids Research database issue and molecular biology database collection. Nucleic Acids Research 2020, 48, D1–D8. [Google Scholar] [CrossRef] [PubMed]
- Pampel, H.; Vierkant, P.; Scholze, F.; Bertelmann, R.; Kindling, M.; Klump, J.; Goebelbecker, H.J.; Gundlach, J.; Schirmbacher, P.; Dierolf, U. Making research data repositories visible: the re3data. org registry. PloS one 2013, 8, e78080. [Google Scholar] [CrossRef] [PubMed]
- Robinson-Garcia, N.; Mongeon, P.; Jeng, W.; Costas, R. DataCite as a novel bibliometric source: Coverage, strengths and limitations. Journal of Informetrics 2017, 11, 841–854. [Google Scholar] [CrossRef]
- Zenodo. Zenodo - Research. Shared. https://zenodo.org/, 202. Accessed: 2025-04-05.
- Dryad. Dryad - Publish and Preserve Your Data. https://datadryad.org/stash, 2024. Accessed: 2025-04-05.
- Figshare. Figshare - Share your research. https://figshare.com/, 2024. Accessed: 2025-04-05.
- Rettberg, N.; Schmidt, B. OpenAIRE-Building a collaborative Open Access infrastructure for European researchers. LIBER Quarterly: The Journal of the Association of European research libraries 2012, 22, 160–175. [Google Scholar] [CrossRef]
- Raghavan, S.; Garcia-Molina, H. Crawling the Hidden Web. In Proceedings of the VLDB 2001, Proceedings of 27th International Conference on Very Large Data Bases, September 11-14, 2001, Roma, Italy; Apers, P.M.G.; Atzeni, P.; Ceri, S.; Paraboschi, S.; Ramamohanarao, K.; Snodgrass, R.T., Eds. Morgan Kaufmann, 2001, pp. 129–138.
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 2023. [Google Scholar]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 2023. [Google Scholar]
- Barbosa, L.; Freire, J. Searching for Hidden-Web Databases. In Proceedings of the Proceedings of the Eight International Workshop on the Web & Databases (WebDB 2005), Baltimore, Maryland, USA, Collocated mith ACM SIGMOD/PODS 2005, June 16-17, 2005; Doan, A.; Neven, F.; McCann, R.; Bex, G.J., Eds., 2005, pp. 1–6.
- Barbosa, L.; Freire, J. An adaptive crawler for locating hidden-web entry points. In Proceedings of the Proceedings of the 16th international conference on World Wide Web, 2007, pp. 441–450.
- Madhavan, J.; Ko, D.; Kot, L.; Ganapathy, V.; Rasmussen, A.; Halevy, A.Y. Google’s Deep Web crawl. Proc. VLDB Endow. 2008, 1, 1241–1252. [Google Scholar] [CrossRef]
- Chen, F.K.; Liu, C.H.; You, S.D. Using Large Language Model to Fill in Web Forms to Support Automated Web Application Testing. Information 2025, 16, 102. [Google Scholar] [CrossRef]
- Gur, I.; Nachum, O.; Miao, Y.; Safdari, M.; Huang, A.; Chowdhery, A.; Narang, S.; Fiedel, N.; Faust, A. Understanding html with large language models. arXiv preprint arXiv:2210.03945 2022. [Google Scholar]
- Stafeev, A.; Recktenwald, T.; De Stefano, G.; Khodayari, S.; Pellegrino, G. YURASCANNER: Leveraging LLMs for Task-driven Web App Scanning. In Proceedings of the The Network and Distributed System Security (NDSS) Symposium. CISPA, 2024.
- Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data 2016, 3, 1–9. [Google Scholar] [CrossRef] [PubMed]
- Van de Sompel, H.; Nelson, M.L.; Lagoze, C.; Warner, S. Resource harvesting within the OAI-PMH framework. D-lib magazine 2004, 10. [Google Scholar] [CrossRef]
- Brickley, D.; Burgess, M.; Noy, N. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In Proceedings of the The world wide web conference; 2019; pp. 1365–1375. [Google Scholar]
- Chen, X.; Gururaj, A.E.; Ozyurt, B.; Liu, R.; Soysal, E.; Cohen, T.; Tiryaki, F.; Li, Y.; Zong, N.; Jiang, M.; et al. DataMed–an open source discovery index for finding biomedical datasets. Journal of the American Medical Informatics Association 2018, 25, 300–308. [Google Scholar] [CrossRef] [PubMed]
- Enevoldsen, K.; Chung, I.; Kerboua, I.; Kardos, M.; Mathur, A.; Stap, D.; Gala, J.; Siblini, W.; Krzemiński, D.; Winata, G.I.; et al. MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv preprint arXiv:2502.13595, 2025. [Google Scholar] [CrossRef]
- Wei, J.; Bosma, M.; Zhao, V.Y.; Guu, K.; Yu, A.W.; Lester, B.; Du, N.; Dai, A.M.; Le, Q.V. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021. [Google Scholar]
- Kojima, T.; Gu, S.S.; Reid, M.; Matsuo, Y.; Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems 2022, 35, 22199–22213. [Google Scholar]






| Simple Query | Complex Query |
|---|---|
| Find entries in the PR2 database where Level is “Alveolata” and Class is “Dinoflagellates” with Version 5.0, focusing on taxonomic groups annotated by experts. | I am looking for detailed information on proteins with unique dynamic properties, such as those with chameleon subsequences or Dual Personality Fragments, as analyzed through molecular dynamics simulations. Specifically, I want to explore how these proteins interact with other biological molecules and their domain limits. |
| Find samples with Data Type as “Metagenome” or “Natural Organic Matter” from the National Microbiome Data Collaborative database, focusing on studies related to the Earth Microbiome Project and the 1000 Soils Research Campaign. | I am looking for microbiome data samples that involve metagenomic analysis or natural organic matter studies, particularly those associated with large collaborative projects like the Earth Microbiome Project or soil research initiatives. |
| Model | Query Type | Avg. Time (s) | Site Match | Page Match | Schema Match |
|---|---|---|---|---|---|
| Llama3.3 | Complex | 40.84 | 0.592 | 0.449 | 0.449 |
| Simple | 37.76 | 0.755 | 0.612 | 0.612 | |
| Overall | 39.30 | 0.673 | 0.531 | 0.531 | |
| DeepSeek-R1 | Complex | 13.80 | 0.571 | 0.449 | 0.388 |
| Simple | 13.34 | 0.694 | 0.592 | 0.510 | |
| Overall | 13.57 | 0.633 | 0.520 | 0.449 | |
| Qwen3 | Complex | 24.83 | 0.531 | 0.429 | 0.367 |
| Simple | 34.51 | 0.612 | 0.531 | 0.469 | |
| Overall | 29.67 | 0.571 | 0.480 | 0.418 | |
| Gemma | Complex | 13.81 | 0.531 | 0.408 | 0.367 |
| Simple | 12.64 | 0.612 | 0.490 | 0.449 | |
| Overall | 13.22 | 0.571 | 0.449 | 0.408 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
