Submitted:
08 September 2025
Posted:
10 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Bioarchaeology and Data Complexity
1.2. Assessing FAIRness in Bioarchaeology
1.3. Natural Language Processing and the Role of NER
2. Materials and Methods
2.1. Document Selection
2.2. Annotation Process
- Initial Annotation: Using a word processor, the researcher manually annotated the selected documents. Each osteoarchaeological or palaeopathological term referring to the human body was tagged with a unique identifier derived from the U.S. National Library of Medicine’s Medical Subject Headings (MeSH). MeSH was selected for its extensive vocabulary coverage, though limitations include inconsistencies with British English terminology and a lack of archaeological disease classifications.
- Expert Annotation: A domain expert independently reviewed and re-annotated the same five documents to verify accuracy. Different colours were used to distinguish term types, with a consistent colour key maintained to ensure comparability. A “super-annotator” subsequently reviewed both annotation sets, resolving all discrepancies—44 in total—to create a consistent, verified dataset. Overall, 2,582 annotations covering 252 distinct terms were made, although 28 terms were only observed once, limiting their training utility. These low-frequency terms were excluded from the final model training.
- Structured Annotation in GATE: The final, reconciled annotations were transferred into XML format and imported into the GATE Developer environment. Using MeSH identifiers, these annotations were applied to the documents within GATE to produce the final Gold Standard training set used in model development.
2.3. Model Training and Rationale
- Precision: 0.913;
- Recall: 0.868;
- F1-Score: 0.889;
2.4. Evaluation Framework
- A score of 7 represents the highest level of agreement or satisfaction;
- A score of 1 represents the lowest.
3. Results
3.1. Combined Results
3.2. Students
3.3. Experts
3.4. Public
4. Discussion
4.1. Reliability
4.2. Usefulness
4.3. Time-Saving Potential
4.4. Willingness to Use Again
4.5. Accessibility
4.6. Implications for Heritage AI
5. Conclusions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| OPES | Osteoarchaeological and Palaeopathological Entity Search |
| NLP | Natural Language Processing |
| NER | Named Entity Recognition |
| FAIR | Findable, Accessible, Interoperable, Reusable |
| ADS | Archaeology Data Service |
| LLM | Large Language Model |
| MeSH | Medical Subject Headings |
| aDNA | Ancient DNA |
Appendix A
| Do you think of yourself as an: | Do you think this tool is time saving? | Are the results reliable? | How accessible was this tool to use? | Would you like to use this tool again? | How useful is this tool for research for archaeologists? | |
| Non-archaeologists | 6 | 7 | 7 | 1 | 7 | |
| Non-archaeologists | 7 | 5 | 7 | 1 | 5 | |
| Non-archaeologists | 5 | 6 | 7 | 5 | 6 | |
| Non-archaeologists | 7 | 7 | 7 | 7 | 7 | |
| Student | 6 | 7 | 5 | 6 | 6 | |
| Student | 4 | 5 | 7 | 5 | 6 | |
| Student | 3 | 2 | 6 | 2 | 2 | |
| Non-archaeologists | 5 | 5 | 7 | 6 | 6 | |
| Student | 6 | 6 | 7 | 5 | 6 | |
| Student | 3 | 5 | 6 | 6 | 5 | |
| Student | 7 | 3 | 7 | 6 | 6 | |
| Student | 7 | 6 | 7 | 7 | 7 | |
| Expert | 4 | 5 | 6 | 4 | 6 | |
| Student | 3 | 4 | 6 | 2 | 3 | |
| Non- archaeologists |
4 | 4 | 3 | 1 | 5 | |
| Student | 6 | 5 | 5 | 5 | 6 | |
| Expert | 4 | 4 | 5 | 3 | 5 | |
| Expert | 6 | 5 | 6 | 5 | 5 | |
| Non- archaeologists |
5 | 6 | 7 | 5 | 6 | |
| Student | 5 | 5 | 6 | 6 | 6 | |
| Student | 2 | 5 | 5 | 2 | 3 | |
| Student | 3 | 2 | 2 | 2 | 4 | |
| Student | 5 | 4 | 7 | 6 | 6 | |
| Student | 5 | 5 | 6 | 6 | 5 | |
| Student | 6 | 6 | 6 | 6 | 6 | |
| Student | 5 | 4 | 7 | 7 | 5 | |
| Student | 4 | 6 | 6 | 5 | 4 | |
| Student | 6 | 6 | 6 | 7 | 7 | |
| Student | 5 | 4 | 5 | 4 | 6 | |
| Expert | 1 | 2 | 1 | 1 | 1 | |
| Expert | 4 | 1 | 6 | 3 | 3 | |
| Expert | 3 | 3 | 5 | 3 | 4 | |
| Expert | 5 | 3 | 4 | 4 | 5 | |
| Student | 5 | 5 | 1 | 5 | 6 | |
| Expert | 4 | 3 | 6 | 4 | 4 | |
| Student | 6 | 6 | 7 | 6 | 7 | |
| Student | 5 | 6 | 7 | 6 | 6 | |
| Expert | 7 | 7 | 7 | 7 | 7 | |
| Student | 5 | 4 | 5 | 4 | 4 | |
| Student | 3 | 4 | 5 | 4 | 5 | |
| Student | 7 | 6 | 7 | 7 | 7 | |
| Expert | 3 | 4 | 4 | 3 | 5 | |
| Student | 5 | 5 | 7 | 5 | 5 | |
| Student | 5 | 6 | 7 | 6 | 6 | |
| Student | 5 | 6 | 7 | 6 | 6 | |
| Non- archaeologists |
5 | 5 | 7 | 5 | 7 | |
| Non- archaeologists |
5 | 5 | 7 | 5 | 7 | |
| Non- archaeologists |
5 | 5 | 7 | 5 | 7 | |
| Non- archaeologists |
5 | 6 | 4 | 1 | 7 | |
| Student | 7 | 5 | 4 | 7 | 7 | |
| Student | 7 | 7 | 7 | 7 | 7 | |
| Non- archaeologists |
5 | 7 | 7 | 5 | 7 | |
| Non- archaeologists |
6 | 6 | 7 | 5 | 7 | |
| Expert | 5 | 5 | 7 | 2 | 5 | |
| Student | 3 | 2 | 5 | 3 | 4 | |
| Student | 6 | 6 | 7 | 6 | 7 | |
| Expert | 4 | 4 | 4 | 3 | 4 | |
| Student | 4 | 5 | 3 | 5 | 5 | |
| Expert | 6 | 5 | 6 | 6 | 6 | |
| Expert | 2 | 2 | 6 | 2 | 2 | |
| Student | 4 | 6 | 7 | 5 | 4 | |
| Non- archaeologists |
5 | 5 | 3 | 6 | 4 | |
| Expert | 2 | 4 | 4 | 3 | 2 | |
| Expert | 4 | 4 | 5 | 4 | 5 | |
| Expert | 6 | 5 | 7 | 3 | 5 | |
| Expert | 3 | 2 | 3 | 3 | 3 | |
| Student | 3 | 5 | 6 | 3 | 2 | |
| Expert | 5 | 4 | 6 | 3 | 4 | |
| Expert | 4 | 4 | 6 | 3 | 5 | |
| Expert | 6 | 4 | 4 | 4 | 6 | |
| Student | 3 | 4 | 5 | 4 | 4 | |
| Expert | 5 | 4 | 5 | 1 | 3 | |
| Expert | 1 | 1 | 1 | 1 | 4 | |
| Student | 6 | 5 | 5 | 6 | 5 | |
| Student | 6 | 5 | 7 | 6 | 6 | |
| Student | 4 | 5 | 7 | 5 | 4 | |
| Expert | 5 | 6 | 7 | 5 | 5 | |
| Expert | 3 | 4 | 6 | 4 | 4 | |
| Student | 7 | 6 | 7 | 7 | 7 | |
| Student | 7 | 5 | 7 | 7 | 7 | |
| Expert | 5 | 5 | 6 | 6 | 6 | |
| Student | 3 | 3 | 2 | 2 | 4 | |
| Non- archaeologists |
7 | 7 | 7 | 7 | 7 |
References
- Britton, K.; Richards, M.P. Introducing Archaeological Science. In Archaeological Science: An Introduction; Richards, M.P., Britton, K., Eds.; Cambridge University Press: Cambridge, UK, 2020; pp. 3–10. [Google Scholar] [CrossRef]
- Hendy, J.; Welker, F.; Demarchi, B.; Speller, C.; Warinner, C.; Collins, M.J. A guide to ancient protein studies. Nat. Ecol. Evol. 2018, 2, 791–799. [Google Scholar] [CrossRef] [PubMed]
- Bonino da Silva Santos, L.O.; Bonino, L.; Pariente, T.J.; Kuhn, T. FAIR Data Points Supporting Big Data Interoperability. In Enterprise Interoperability VII; Popplewell, K., Thoben, K.D., Knothe, T., Poler, R., Eds.; STE Press: London, UK, 2016; pp. 270–279. [Google Scholar]
- Lien-Talks, A. How FAIR Is Bioarchaeological Data: With a Particular Emphasis on Making Archaeological Science Data Reusable. J. Comput. Appl. Archaeol. 2024, 7, 246–261. [Google Scholar] [CrossRef] [PubMed]
- Talboom, L. Improving the Discoverability of Zooarchaeological Data with the Help of NLP. Master’s Thesis, University of York, York, UK, 2017. [Google Scholar]
- Jeffrey, S.; Richards, J.D.; Ciravegna, F.; Waller, S.; Chapman, S.; Zhang, Z. The Archaeotools Project. Philos. Trans. R. Soc. A 2009, 367, 2507–2519. [Google Scholar] [CrossRef] [PubMed]
- Tudhope, D.; Binding, C.; Jeffrey, S.; May, K.; Vlachidis, A. A STELLAR Role for Knowledge Organization Systems in Digital Archaeology. Bull. Am. Soc. Inf. Sci. Technol. 2011, 37, 15–18. [Google Scholar] [CrossRef]
- Isaksen, L. Lines of Sight: Modelling straight-line travel in ancient Mediterranean sailing. Internet Archaeol. 2011, 30. [Google Scholar] [CrossRef]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar] [CrossRef]
- Lee, J.; Yoon, W.; Kim, S.; Kim, D.; Kim, S.; So, C.H.; Kang, J. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 2020, 36, 1234–1240. [Google Scholar] [CrossRef] [PubMed]
- Beltagy, I.; Lo, K.; Cohan, A. SciBERT: A Pretrained Language Model for Scientific Text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 3–7 November 2019; pp. 3615–3620. [Google Scholar] [CrossRef]
- OpenAI. GPT-4 Technical Report. arXiv 2023, arXiv:2303.08774. [Google Scholar] [CrossRef]
- Rogers, A.; Kovaleva, O.; Rumshisky, A. A Primer in BERTology: What We Know About How BERT Works. Trans. Assoc. Comput. Linguist. 2020, 8, 842–866. [Google Scholar] [CrossRef]
- Strubell, E.; Ganesh, A.; McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 3645–3650. [Google Scholar] [CrossRef]
- Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
- Bender, E.M.; Gebru, T.; McMillan-Major, A.; Mitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event, Canada, 3–10 March 2021; pp. 610–623. [Google Scholar] [CrossRef]
- D’Ignazio, C.; Klein, L.F. Data Feminism; MIT Press: Cambridge, MA, USA, 2020. [Google Scholar]
- Fiorucci, M.; Khoroshiltseva, M.; Pontil, M.; Traviglia, A.; Del Bue, A.; James, S. Artificial Intelligence for Digital Heritage Innovation: Setting up a R&D Agenda for Europe. Heritage 2024, 7, 1028–1257. [Google Scholar] [CrossRef]
- Albert, W.; Tullis, T. Measuring the User Experience: Collecting, Analyzing, and Presenting Usability Metrics, 2nd ed.; Morgan Kaufmann: Burlington, MA, USA, 2013; pp. 145–170. [Google Scholar]
- van Dis, E.A.M.; Bollen, J.; Zuidema, W.; Bockting, C.L.H.; Schoevers, R.A. ChatGPT: Five Priorities for Research. Nature 2023, 614, 224–226. [Google Scholar] [CrossRef]
- Zhang, Y.; Chen, Q.; Yang, Z.; Lin, H.; Lu, Z. Evolution and emerging trends of named entity recognition: Bibliometric analysis from 2000 to 2023. PLOS ONE 2024, 19, e0297874. [Google Scholar] [CrossRef]
- Donciu, C.M.; Foschini, L.; Martoglia, R. A novel NLP-driven approach for enriching artefact descriptions, provenance, and entities in cultural heritage. Neural Comput. Appl. 2025, 37, 2587–2606. [Google Scholar] [CrossRef]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. arXiv 2020, arXiv:2005.14165. [Google Scholar] [CrossRef]
- Barker, E.; Terras, M.; Nyhan, J. The Heritage Connector: A machine learning framework for building linked open data from museum collections. Digit. Scholarsh. Humanit. 2021, 36, ii23–ii43. [Google Scholar] [CrossRef]
- Meghini, C.; Barker, E.; Binding, C.; Tudhope, D. ARIADNE: A Research Infrastructure for Archaeology. J. Comput. Cult. Herit. 2017, 10, 18. [Google Scholar] [CrossRef]
- Workshop: Big Data and Archaeology: Towards an AI Solution. In Proceedings of the Workshop on Big Data and Archaeology, Mainz, Germany, 6–7 June 2024.




Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).