Preprint
Article

This version is not peer-reviewed.

Preserving Digital Cultural Heritage: AWikidata Approach to Extract Multimodal Collections as Data in Museums

A peer-reviewed version of this preprint was published in:
Heritage 2026, 9(8), 292. https://doi.org/10.3390/heritage9080292

Submitted:

24 June 2026

Posted:

25 June 2026

You are already at the latest version

Abstract
GLAM (Galleries, Libraries, Archives, and Museums) institutions have been making digital collections available for decades. New initiatives to publish, preserve, and reuse data have emerged, with collaborative projects such as Wikidata playing a significant role. This work provides a framework for extracting multimodal collections as data, using Wikidata as the primary source, along with a selection of research scenarios for reusing this data in line with emergent trends in data publication and reuse. Results showed that current trends in data preservation can be achieved through open-source code, cloud services, and collaborative platforms. Future work to be explored includes adopting best practices for provenance documentation and refining reuse scenarios.
Keywords: 
;  ;  

1. Introduction

Preserving digital cultural heritage ensures that cultural and historical information remains accessible, authentic, and usable for future generations. Without active preservation, these materials can become inaccessible over time due to technological change, file corruption, or obsolete software [1]. Simultaneously, there are also increasingly new technologies to ensure computational access and responsible use of collections. Thus Galleries, Libraries, Archives and Museums (GLAMs), as a main provider of cultural heritage data, are presented with challenges of both ensuring long term preservation with shifting technological opportunities and policy instruments for making their digital collections available.
Open Science practices have been defined and implemented to promote open access to scientific publications, research data, models, algorithm collects, software, protocols, notebooks or workflows [2]. In parallel, emergent data sharing initiatives focused on federation, security and sovereignty have emerged during the last years such as Solid and data spaces [3,4]. Community efforts such as GLAM Labs,1 AI4LAM,2 and international initiatives such as Collections as Data,3 FAIR [5] and CARE [6], provide guides for best practices for the publication of machine-actionable collections that support computational access and ethical uses of these data [7,8].
Increasingly institutions have turned to the Semantic Web to describe their digital collections and catalogues [9], with Linked Open Data used as the gold standard for enriching and maximizing the reuse of cultural heritage data. For example, the National Library of Scotland, the Smithsonian American Art Museum or the Rijksmuseum (Amsterdam) published their data using Application Programming Interfaces (API) or as downloadable files [10,11]. They identified existing sources covering different domains with which to enrich their digital collections. Some examples include GeoNames, Virtual International Authority File (VIAF) and Wikidata, a collaborative edition platform owned by Wikimedia Foundation that has played a leading role in the GLAM sector [12,13,14]. Thus, we see Linked Data ontologies and knowledge graph infrastructures in GLAMs are used not only to enrich the information on their collections, but also in making this information publicly accessible as linked open data in a way to efficiently integrate and implement an unprecedented amount of often unstructured, siloed data, at lightning speed.
Despite these affordances, the implementation of linked data technologies in GLAMs remain rare. This is due to a number of issues, largely that the technology and approaches are not ones that traditionally exist in these institutions, nor are taught to collection curators and managers. Despite using catalogue systems with standard ontologies which aid to organize, store and manage collections; the integration of these data with the Semantic Web demands a shift in conceptual data models and data formats. A number of models have been developed for bibliographic records: the IFLA-Library Reference Model (IFLA-LRM), the Resource Description and Access (RDA), and Bibliographic Framework (BIBFRAME); as well as for cultural heritage objects: the Europeana Data Model (EDM). These models are distinct. Consequently the question remains: how can these knowledge bases that use Semantic Web principles to collect structured data be employed for preservation purposes.
To investigate this we take Wikidata as an example, and show this can be extracted and reused in a reproducible manner. We also show the potential (re)uses of the extracted data. The two main contributions of this work are: i) a framework for extracting multimodal collections as data using Wikidata as a primary source; and ii) a selection of research scenarios to reuse this data according to emergent trends in data publication and reuse. With that, this work is relevant to GLAM practitioners, curators, and researchers who wish to reuse and enrich digital collections made available by GLAM institutions.
This article is organized as follows: after a description of the state of the art, we present the framework to extract multimodal collections as data from Wikidata and its application to a selection of museums. The article concludes with the main conclusions and final remarks.

3. Methodology

The proposed methodology is based on a framework in four steps:
1.
identification of sources;
2.
data extraction;
3.
publication; and
4.
reuse.
Figure 1 shows the architecture employed for a tool we named PLOW—Preservation through Linked Open data and Wikidata. The individual steps are described in detail below.

3.1. Identification of Sources

The first step is the identification of a data source, for which several technical options exist. GLAM datasets are commonly published through GLAM Labs—the Data Foundry at the National Library of Scotland being a representative case; and through the dedicated data sections of major libraries, such as the API et jeux de données of the National Library of France or the equivalent position at the Rijksmuseum in Amsterdam.13 These rely chiefly on data dumps that may be downloaded and reused with little friction; others, by contrast, expose their holdings through APIs—SPARQL, OAI-PMH and IIIF among them. Cross-domain hubs occupy a third position: Wikidata, in particular, has become a significant identifier hub in the GLAM sector, enabling the integration of metadata with multimedia content.
Identifying a source, however, is not the same as selecting one, and several criteria bear on that choice. Foremost among them is licensing: while many datasets are released under open licences such as Creative Commons, a substantial number remain copyrighted, or carry no clear statement of reuse at all [19,20]. Considerations of a different order also intrude: among them, the protection of cultural heritage placed at risk by war, as in Ukraine, Gaza and Afghanistan.14 To these may be added the familiar measures of fitness for purpose: data quality, completeness, timeliness, and availability [16], together with thematic relevance, the degree, to which a source speaks to the subject or research question at hand. Such relevance is best made concrete: a study of artists in exile during the Spanish Civil War, of the paintings of a particular movement, or of works looted by the Nazis during the Second World War will each privilege very different collections [42,43].

3.2. Data Extraction

This step involves extracting the data from the source identified. The output may take a range of serialisation formats: JSON, XML or RDF, among the most common, and frequently extends beyond textual metadata to multimedia content such as images, maps or OCR-derived text. Extraction need not rely solely on the originating institution alone: aggregators such as Europeana or collaboratively edited platforms such as Wikidata can serve equally well as points of access.
In some cases, this step can be related to a particular topic, literary movement or event (e.g., World War II). In other cases, the data is extracted based on membership relations such as being part of a particular catalogue (e.g., Rijksmuseum Amsterdam or National Library of France). In this context, Wikidata has created many properties for GLAM institutions in order to link their catalogues [44]. For example, Listing 2 shows a SPARQL query to retrieve artworks from Wikidata linked to the Prado Museum by means of the dedicated property wdt:P8905. Other examples rely on using the CONSTRUCT clause in SPARQL to build a knowledge graph from the triples retrieved by a query.
Preprints 220084 i002
  • Listing 2: SPARQL query to retrieve Art paintings from Wikidata linked to the Prado Museum
Once extracted, metadata, and images alike may also surface in aggregators such as Europeana, where they can be integrated and browsed alongside further digital collections.15

3.3. Publication

The final step concerns the publication of the dataset and its subsequent reuse. Established guidance exists for the former: best-practice frameworks set out the principal steps for publishing machine-actionable collections, within Europeana and beyond [20,33].
A useful organising principle here is Berners-Lee’s five-star deployment scheme (see Figure 2), which ranks open data by degree of accessibility: from material merely placed online under an open licence, through structured and non-proprietary formats, to the use of URIs and, at the highest level, linking to other data to provide context. The scheme is valuable precisely because it reframes publication not as a binary of open versus closed but as a gradient of reusability, against which an institution may locate its own holdings and chart a path of improvement [45].

3.4. Reuse

Reuse, for its part, has been approached from several directions, each building on the content once published. Some work has defined research scenarios grounded in that content [33]. Among the practical instruments, the Jupyter Notebook has gained particular traction, allowing institutions to furnish prototype examples that combine code, charts, documentation and interactive widgets in a single reproducible artefact [34]. Institutions have begun to employ Jupyter Notebooks as a powerful tool to provide examples of use through prototypes. They combine code, charts, documentation, and interactive widgets [31,32]. Further examples and prospective research scenarios drawn from GLAM data are set out in Section 2.
Of particular interest is the integration of published collections with data and research infrastructures. The Common European Data Space for Cultural Heritage and the European Collaborative Cloud for Cultural Heritage (ECCCH) both operate at this level, as does the European Open Science Cloud (EOSC)—whose Jupyter Notebook service, notably, permits reproducible code to be run on dedicated cloud servers. Visibility and reuse are, in turn, served by platforms such as the Social Sciences and Humanities Open Marketplace, Zenodo and OpenAIRE [21].
One further consideration deserves emphasis: versioning. Institutions frequently wish to publish their catalogues as a series of updates rather than a single release, and platforms such as Zenodo accommodate this directly, issuing a distinct DOI and a suggested citation for each version [46].

4. Results

In this section, we describe the application of the framework in Section 3 for several GLAM institutions. We specifically have selected museums, given the acknowledged contributions to Wikidata and thus available data. The museums were selected according to the following criteria:
1.
availability of a dedicated Wikidata property to link artworks;
2.
historical importance, cultural influence, as well as digital transformation;
3.
provision of documentation of the catalogue.
This resulted in the following cases: Musée d’Orsay, National Gallery, National Gallery of Ireland, Prado Museum, and the Royal Museum of Fine Arts Antwerp.
Wikidata was used as the primary repository for extracting multimodal content, including metadata, and images. To extract the metadata, the SPARQL query included in Appendix A was employed. The data was retrieved in May 2026. Note that images are stored on Wikimedia Commons, where Wikidata only preserves a link. Table 1 show a summary of the results following the application of the framework.
A GitHub repository was created including the extracted data, a README file with additional documentation and a collection of Jupyter Notebooks illustrating how to extract and reuse the data.16 Several Python packages, such as SPARQLWrapper, were employed to query the Wikidata SPARQL endpoint.17 The data was extracted as a JSON file. The Jupyter notebooks were hosted and executed using the services provided by EOSC. To do so, a dedicated virtual machine was deployed to run the code provided. This work intends to promote its use amongst the community. Best practices were used to develop and publish the collection of Jupyter Notebooks [30,31].

4.1. Scenarios of Reuse

This section outlines a selection of scenarios in which the extracted data might be reused, each grounded in prior work spanning topics ranging from the integration of data repositories to the application of various AI-based techniques [34,37]. The scenarios are addressed to a mixed readership of practitioners, curators, researchers and students, and their purpose is illustrative rather than prescriptive: they show how data may be extracted and preserved in keeping with the best practices promoted by international initiatives, without pretending to a full specification or a worked implementation. What they offer, in short, is a narrative account of added value rather than a manual.
Federated queries. Institutions tend to publish digital collections as an isolated collections effectively making a data silo and thus hindering collaboration and their reuse. However, advances in technology have created a new context in which methods such as Named Entity Recognition (NER), and Entity Linking are helping enrich digital collections. When linking records to external repositories, new opportunities arise to enrich metadata, improve interoperability, and enable broader access to cultural and research collections. Wikidata federates several SPARQL repositories, enabling access to different repositories from relevant institutions through Wikidata [47].
For example, Listing 3 shows a federated SPARQL query to retrieve bibliographic records from Biblioteca Virtual Miguel de Cervantes (BVMC) in which the painter Diego Velázquez is represented as a subject. Note that the SERVICE command is employed to access the BVMC SPARQL endpoint from Wikidata. The BIND command is employed to filter the bibliographic records using the Picasso’s identifier in BVMC. Other examples could be based on the integration of painting from catalogues such as the Prado Museum and the National Gallery, focused, for instance, on a particular painter such as Diego Velázquez.
Preprints 220084 i003
  • Listing 3: SPARQL query to retrieve bibliographic records from Biblioteca Virtual Miguel de Cervantes in which the painter Diego Velázquez is represented as a subject.
According to these observations, museum data could be enriched with data from libraries and archives, such as the National Library of France, to create integrated research scenarios that combine different types of records. This could help address challenges related to data quality, interoperability, and reuse [16].
AI-based computer vision models. Art works in Wikidata are described using many properties (e.g., title and inception). However, given Wikidata’s collaborative editing approach, properties are not consistently used across all entities. The property wdt:P180 is used to indicate what an item depicts, in particular, to include entities visually depicted in an image (e.g., Spanish Civil War).18 Computer vision models can be employed to enhance the value of this property by leveraging the multimedia content generated by the data extraction process, thereby enabling the refinement and enhancement of image descriptions [48]. Metadata can be automatically generated to describe images provided by GLAM institutions in greater detail, such as paintings, maps, or postcards [34,49].
See, for instance, Figure 3 representing Las Meninas from Diego Velázquez. The Wikidata identifier wd:Q208758 provides metadata about the painting, including the property wdt:P180 with several values such as people included in the painting, the painter himself or a dog.
Federated data storage systems. Solid is an emerging initiative for storing federated data while ensuring security and sovereignty [3]. Previous research has analysed the use of Solid in the GLAM context, combined with Wikidata as a federated storage system.19 Software libraries and tutorials in different programming languages, such as Java and Python, have been developed to facilitate their adoption.20 These new technologies facilitate data portability across applications while enhancing transparency and privacy in data access.
RAG-based search engine. Wikidata may be interrogated by several means—keyword matching or SPARQL, among them—yet each has its limits: a keyword set may fail to capture the intent behind a query, and SPARQL’s steep learning curve places it beyond many users. Wikimedia has accordingly begun exploring semantic search techniques, including vector embeddings, as a more forgiving mode of access.21 Coupled with large language models and natural language processing, such methods can materially improve retrieval [50], and earlier work has already applied Retrieval-Augmented Generation to web-archive content and digitised collections, thereby opening those holdings to conversational access.22 Hence, the artwork descriptions held in Wikidata records might themselves underwrite a RAG search engine—one that affords new routes into the content, curbs hallucination, and injects domain-specific knowledge into the underlying model.
LLM or SLM-based chatbot. A natural specialisation of the retrieval-augmented scenario sketched above is a conversational agent addressed not to the researcher but to the museum visitor: a chatbot, in effect, with which one might interrogate a single art object one is standing before. The premise is straightforward: the multimodal records extracted by the framework: title, creator (wdt:P170), date of inception (wdt:P571), the entities recorded under wdt:P180 and the associated image wdt:P18—furnish a bounded, machine-readable context against which a visitor’s questions may be answered. Rather than draw on its own parametric memory, with all the attendant risk of confabulation, the model retrieves the relevant record and renders it as natural-language prose; and where the visitor asks what a painting depicts, the very property whose enrichment was discussed above supplies the answer. The choice between a large language model (LLM) and a small language model (SLM) is, in this setting, less a technical detail than a design decision with institutional consequences. An LLM offers fluency, breadth, and a capacity to field unanticipated questions, but it is costly to serve continuously to the public and, when cloud-hosted, returns us to the difficulty noted above of data leaving European jurisdiction. An SLM, a model of a few billion parameters capable of running on modest, or even on-premise hardware, inverts these trade-offs. Its narrower internal knowledge is, paradoxically, an advantage rather than a constraint: confined to answering from the retrieved record alone, it is the more readily held to the institution’s own authoritative data, and the less inclined to volunteer plausible invention. Hence, for a bounded domain such as the holdings of a single museum, pairing a modest SLM with strong retrieval grounding may prove not merely the cheaper option but also the more trustworthy one [51].
Wikidata records carry labels in many languages, the mechanism is already exploited by the SERVICE clause in Listing 2—a visitor might both pose questions and receive answers in their own tongue, without additional translation infrastructure. In order to foster transparency, each response may be accompanied by a reference to the Wikidata item from which it was drawn, so that the curious or the sceptical can verify the claim at its source. It should be stressed, however, that such a system inherits the limitations of its underlying data. Where Wikidata’s coverage of the Prado is partial, or where wdt:P180 has been applied unevenly, the agent will be correspondingly silent or incomplete; and no degree of retrieval grounding wholly forecloses error. A visitor-facing use would therefore demand careful evaluation against curatorial ground truth before it could responsibly stand between the public and the collection.
These research scenarios provide a wide range of approaches for reusing the content extracted by the framework proposed in this work.

4.2. Discussion

This work complements existing approaches to the extraction, publication and reuse of digital collections, but it occupies a position none of them does: it takes digital preservation as its governing concern, a collaborative platform—Wikidata—as its means, and the museum collection as its principal content. Its contribution, then, is to enrich the current landscape of worked examples, supplying the community with additional cases that follow established best practices.
Several limitations bound the scope of what has been attempted here. Wikidata does not hold a complete representation of the Prado’s catalogue; a substantial selection of records has nonetheless been linked to it, thanks to the collaborative efforts of the community. It would be valuable, even so, for the institution to offer access by additional means—OAI-PMH, IIIF, or SPARQL, each weighing richness, intelligibility, and complexity differently—and, looking further ahead, through sovereignty-preserving approaches such as data spaces. Using Wikidata entails several challenges that must be addressed. For example, data uploaded to Wikidata is highly reused by AI companies to create commercial products, while GLAM institutions receive no compensation. In exchange, the visibility of the digital collections increases, impacting society. In some cases, the data modelling provided by Wikidata might not be suitable to represent indigenous knowledge. A more inclusive approach is required in these cases in order to facilitate an equitable and non-discriminatory environment.
The EOSC cloud services are constrained in memory and disk storage; for the examples presented here, those resources sufficed, but more demanding AI-based work would require provision with greater GPU and memory capacity [48]. Alternatives such as Google Colab supply additional resources, though at the cost of depositing data outside Europe—a consideration of some weight for European institutions and researchers.
The reuse scenarios have been presented as textual sketches: prototypes in narrative form rather than working systems, and their implementation remains a task for further work; refinement may in any case prove necessary, since several of the data and research infrastructures invoked are themselves still under development.
The work admits of extension in several directions. The SPARQL queries might be refined to draw in additional classes of institutions, libraries, and archives, and data research infrastructures, such as those discussed above, could be pursued in the service of preservation and reuse alike.

5. Conclusions

GLAM institutions have, in recent years, been exploring new ways of making their digital collections available, and the preservation of digital heritage has become a matter of real public consequence—a means of safeguarding cultural heritage for the future. A range of initiatives have emerged to promote the publication and responsible reuse of the data these institutions hold. In this research we showed how Wikidata can be used in these tasks.
We proposed a framework for extracting multimodal collections as data using Wikidata as a primary source. In keeping with Open Science principles, this work has described a reproducible approach for extracting multimodal collections as data from museums via Wikidata. The framework was applied to a selection of museums, with Wikidata serving as the principal repository, and was accompanied by research scenarios demonstrating how the extracted datasets might be reused. Taken together, this demonstrates how data preservation can be met through open-source code, cloud services and collaborative platforms—without recourse to bespoke or proprietary infrastructure.
Many institutions have, to be sure, already turned to Wikidata for a variety of ends; but its adoption, enrichment and completeness remain unevenly realised. The obstacles are partly practical—a good number of institutions lack the resources and the expertise to begin—and it is here that the community’s sustained work in linking and curating entities has proved decisive. Wikibase, moreover, now offers a further avenue: as cloud provisioning matures, it affords full control over data and governance, together with the freedom to tailor the data model in ways that hosted Wikidata does not.
Two directions invite further work. The first is the adoption of best practices to record provenance in a machine-readable form [46]. The second is the refinement of the reuse scenarios themselves—supplying the structured, machine-actionable detail that would allow others to take them up in earnest. Of the two, the latter is likely to do the most to encourage adoption.

Author Contributions

Conceptualisation: GC; Formal Analysis: all; Investigation: all; Methodology: GC; Visualizations: GC; Resources: all; Software: GC; Writing original draft: GC; Writing – review & editing: all.

Acknowledgments

The authors would like to thank the International GLAM Labs Community, the Impact Centre of Competence, Wikimedia Spain and DARIAH-EU. C. Annemieke Romein’s research within the HAICu-project (digital Humanities Artificial Intelligence and Cultural Heritage) is funded by the NWO NWA 1518.22.10 grant. The contributions of Julie M. Birkholz are supported by the FWO-funded Large-Scale Infrastructure project CLARIAH-VL+ (I001525N) and BELSPO FED-tWIN project (Prf-2019-040).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

    The following abbreviations are used in this manuscript:
AI Artificial Intelligence
API Application Programming Interface
BVMC Biblioteca Virtual Miguel de Cervantes
CARE Collective Benefit, Authority to Control, Responsibility, Ethics
CH Cultural Heritage
DOI Digital Object Identifier
ECCCH European Cultural Heritage Cloud
EOSC European Open Science Cloud
FAIR Findable, Accessible, Interoperable, Reusable
GLAM Galleries, Libraries, Archives and Museums
GPU Graphics Processing Unit
IIIF International Image Interoperability Framework
LLM Large Language Model
LOD Linked Open Data
OAI-PMH Open Archives Initiative Protocol for Metadata Harvesting
RAG Retrieval-Augmented Generation
RDF Resource Description Framework
SPARQL SPARQL Protocol and RDF Query Language
VIAF Virtual International Authority File

Appendix A

Listing 4 was employed to extract the metadata from the Wikidata public SPARQL endpoint for the Prado Museum.
Preprints 220084 i004
  • Listing 4: SPARQL query to retrieve Art paintings from Wikidata linked to the Prado museum

References

  1. Terras, M. Opening access to collections: The making and using of open digitised cultural content. Online Information Review 2015, 39, 733–752. [CrossRef]
  2. European Commission. Open Science, 2026. Accessed: 2026-05-15.
  3. Mansour, E.; Sambra, A.V.; Hawke, S.; Zereba, M.; Capadisli, S.; Ghanem, A.; Aboulnaga, A.; Berners-Lee, T. A Demonstration of the Solid Platform for Social Web Applications. In Proceedings of the Proceedings of the 25th International Conference Companion on World Wide Web, Republic and Canton of Geneva, CHE, 2016; WWW ’16 Companion, p. 223–226. [CrossRef]
  4. Trujillo, J.; Candela, G.; Reina-Reina, A. Data Analytics and Artificial Intelligence in the new scenario of Data Spaces. In Proceedings of the Proceedings of the 27th International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP 2025) co-located with the 28th International Conference on Extending Database Technology and the 28th International Conference on Database Theory (EDBT/ICDT 2025), March 25, 2025; Maté, A.; Lissandrini, M., Eds., Barcelona, Spain, 2025; Vol. 3931, CEUR Workshop Proceedings, pp. 91–92.
  5. Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 2016, 3, 160018. [CrossRef]
  6. Carroll, S.R.; Garba, I.; Figueroa-Rodríguez, O.L.; Holbrook, J.; Lovett, R.; Materechera, S.; Parsons, M.A.; Raseroka, K.; Rodriguez-Lonebear, D.; Rowe, R.; et al. The CARE Principles for Indigenous Data Governance. Data Sci. J. 2020, 19, 43. [CrossRef]
  7. Mahey, M.; Al-Abdulla, A.; Ames, S.; Bray, P.; Candela, G.; Derven, C.; Dobreva-McPherson, M.; Gasser, K.; Chambers, S.; Karner, S.; et al. Open a GLAM lab; International GLAM Labs Community, Book Sprint: Doha, Qatar, 2019; p. 164. [CrossRef]
  8. Padilla, T.; Allen, L.; Frost, H.; Potvin, S.; Russey Roke, E.; Varner, S. Final Report — Always Already Computational: Collections as Data, 2019. [CrossRef]
  9. Berners-Lee, T.; Hendler, J.; Lassila, O. The Semantic Web in Scientific American. Scientific American Magazine 2001, 284.
  10. Candela, G. Towards a semantic approach in GLAM Labs: The case of the Data Foundry at the National Library of Scotland. J. Inf. Sci. 2026, 52, 3–21. [CrossRef]
  11. Dijkshoorn, C.; Jongma, L.; Aroyo, L.; van Ossenbruggen, J.; Schreiber, G.; ter Weele, W.; Wielemaker, J. The Rijksmuseum collection as Linked Data. Semantic Web 2018, 9, 221–230. [CrossRef]
  12. Zhao, F. A systematic review of Wikidata in Digital Humanities projects. Digital Scholarship in the Humanities 2022, 38, 852–874, [https://academic.oup.com/dsh/article-pdf/38/2/852/50488385/fqac083.pdf]. [CrossRef]
  13. Freire, N.; Isaac, A. Wikidata’s Linked Data for Cultural Heritage Digital Resources: An Evaluation Based on the Europeana Data Model. International Conference on Dublin Core and Metadata Applications 2020, 2020, 59–68. [CrossRef]
  14. Candela, G.; Cuper, M.; Holownia, O.; Gabriëls, N.; Dobreva, M.; Mahey, M. A Systematic Review of Wikidata in GLAM Institutions: A Labs Approach. In Proceedings of the Linking Theory and Practice of Digital Libraries - 28th International Conference on Theory and Practice of Digital Libraries, TPDL 2024, September 24-27, 2024, Proceedings, Part II; Antonacopoulos, A.; Hinze, A.; Piwowarski, B.; Coustaty, M.; Nunzio, G.M.D.; Gelati, F.; Vanderschantz, N., Eds., Ljubljana, Slovenia, 2024; Vol. 15178, Lecture Notes in Computer Science, pp. 34–50. [CrossRef]
  15. Berners-Lee, T. Linked Data http://www. w3. org/DesignIssues. LinkedData. html 2006.
  16. Candela, G. An automatic data quality approach to assess semantic data from cultural heritage institutions. J. Assoc. Inf. Sci. Technol. 2023, 74, 866–878. [CrossRef]
  17. Shang, C. Quantitative assessment of the value of intangible cultural heritage art supported by multimodal machine learning. Discov. Artif. Intell. 2026, 6, 7. [CrossRef]
  18. Disli, M.; Gabriëls, N.; Chambers, S.; Ames, S.; Knazook, B.; Candela, G. Exploring the adoption of collections as data in the GLAM context. Inf. Res. 2025, 30, 65–77. [CrossRef]
  19. Dişli, M.; Candela, G. Copyright and Licencing for Cultural Heritage Collections As Data. Journal of Open Humanities Data 2025. [CrossRef]
  20. Candela, G.; Gabriëls, N.; Chambers, S.; Dobreva, M.; Ames, S.; Ferriter, M.; Fitzgerald, N.; Harbo, V.; Hofmann, K.; Holownia, O.; et al. A checklist to publish collections as data in GLAM institutions. Global Knowledge, Memory and Communication 2023, 74, 1323–1355. [CrossRef]
  21. Candela, G.; Chambers, S.; Irollo, A. A Collections as Data Workflow on the SSH Open Marketplace. https://marketplace.sshopencloud.eu/workflow/I3JvP6, 2024. SSH Open Marketplace workflow.
  22. Moissinac, J.C.; Rouzé, F.; Wadhera, P.; Germain, B. Toward a Semantic Representation of the Joconde Database. In Proceedings of the Semapro, IARIA, Nice, France, 10 2020.
  23. de la Culture, M. Collections des musées de France: Base Joconde, 2026.
  24. Gonzalez Cifuentes, T.; Torrejon Vazquez, D. Dataset de obras del Museo Nacional del Prado, 2026. [CrossRef]
  25. Sanderson, R. Implementing Linked Art in a Multi-Modal Database for Cross-Collection Discovery. Open Library of Humanities 2024, 10. [CrossRef]
  26. Nguyen, T.; Storås, A.M.; Thambawita, V.; Hicks, S.A.; Halvorsen, P.; Riegler, M.A. Multimedia Datasets: Challenges and Future Possibilities. In Proceedings of the MultiMedia Modeling; Dang-Nguyen, D.T.; Gurrin, C.; Larson, M.; Smeaton, A.F.; Rudinac, S.; Dao, M.S.; Trattner, C.; Chen, P., Eds., Cham, 2023; pp. 711–717.
  27. Candela, G. Browsing Linked Open Data in Cultural Heritage: A Shareable Visual Configuration Approach. ACM Journal on Computing and Cultural Heritage 2025, 18, 9:1–9:15. [CrossRef]
  28. Kelly, P.; Schild, J.; Jafari, A. FolkRAG: A retrieval-augmented generation system for cultural heritage materials. Neural Comput. Appl. 2025, 37, 20281–20297. [CrossRef]
  29. Alkemade, H.; Candela, G.; Claeyssens, S.; Colavizza, G.; Eren, S.; Freire, N.; Irollo, A.; Isaac, A.; Lehmann, J.; Neudecker, C.; et al. Datasheets for Digital Cultural Heritage Datasets, 2025. [CrossRef]
  30. Pimentel, J.F.; Murta, L.; Braganholo, V.; Freire, J. Understanding and improving the quality and reproducibility of Jupyter notebooks. Empir. Softw. Eng. 2021, 26, 65. [CrossRef]
  31. Candela, G.; Chambers, S.; Sherratt, T. An approach to assess the quality of Jupyter projects published by GLAM institutions. J. Assoc. Inf. Sci. Technol. 2023, 74, 1550–1564. [CrossRef]
  32. Sherratt, T. GLAM Workbench, 2025. [CrossRef]
  33. Candela, G.; Rosiński, C.; Margraf, A. A reproducible framework to publish and reuse Collections as data: the case of the European Literary Bibliography. Transformations: A DARIAH Journal 2025, Workflows. [CrossRef]
  34. Candela, G.; Dobreva, M.; Alkemade, H.; Holownia, O.; Mahey, M.; Ames, S.; Renaud, K.; Vodopivec, I.; Lee, B.C.G.; Padilla, T.; et al. A Use Case Lens on Digital Cultural Heritage, 2025, [arXiv:cs.DL/2509.08710].
  35. Dişli, M.; Candela, G.; Gutiérrez, S.; Fontenelle, G. Open Data Practices of Art Museums in Wikidata: A Compliance Assessment. Journal of Open Humanities Data 2025. [CrossRef]
  36. Lindemann, D.; Candela, G.; Marchetti, A.; Pellizzari di San Girolamo, C.C.; Olea, I.; Varvantakis, C.; Assis, T.; Moitinho de Almeida, V.; Obregón Sierra, Á.; Santiago Faria, A.; et al. The Wikibase Ecosystem in DH and GLAM, 2025. [CrossRef]
  37. Thornton, K. Wikidata for Digital Preservationists. Technical report, Digital Preservation Coalition, 2021. [CrossRef]
  38. Thornton, K.; Seals-Nutt, K. Metadata, 2022. In-Person Long Paper. Recording available at https://youtu.be/Xx6-Z4EDxEk. [CrossRef]
  39. Trognitz, M.; Mandell, R.A.; Štuhec, S.; Palacz, J. Wikidata as a Reconciliation Anchor: Curating Data for Long-Term Preservation. Journal of Open Humanities Data 2026. [CrossRef]
  40. Sichani, A.M.; Kono, K.; Winters, J. Building Cultural Heritage Data Infrastructures with Wikidata: The Case of the Congruence Engine Data Register. Journal of Open Humanities Data 2026. [CrossRef]
  41. Smith-Yoshimura, K. Experimentations with Wikidata/Wikibase, 2020. Hanging Together blog. Accessed 2026-05-19.
  42. Salinas, C.G. Las artistas del exilio republicano español: El refugio latinoamericano; Arte Grandes Temas, Ediciones Cátedra: Madrid, 2019.
  43. J. Paul Getty Trust. Getty Provenance Index, n.d. Accessed: 2026-05-19.
  44. Romein, C.A.; Wagner, A.; van Zundert, J.J. Building and Deploying a Classification Schema using Open Standards and Technology. Journal for Digital Legal History 2023, 2. [CrossRef]
  45. ROMEIN, C.A.; KEMMAN, M.; BIRKHOLZ, J.M.; BAKER, J.; DE GRUIJTER, M.; MEROÑO-PEÑUELA, A.; RIES, T.; ROS, R.; SCAGLIOLA, S. State of the Field: Digital History. History 2020, 105, 291–312, [https://onlinelibrary.wiley.com/doi/pdf/10.1111/1468-229X.12969]. [CrossRef]
  46. Romein, C.A.; Hodel, T.; Gordijn, F.; van Zundert, J.J.; Chagué, A.; et al. Exploring Data Provenance in Handwritten Text Recognition Infrastructure. Journal of Data Mining and Digital Humanities 2024. [CrossRef]
  47. Dişli, M.; Osti, G.; Candela, G.; Zijdeman, R.J. Federated LOD Queries as CaD - Notebooks, 2025. [CrossRef]
  48. Candela, G.; Holownia, O.; Odsbjerg, M.; Cuper, M.; Gabriëls, N.; Hofmann, K.; Gray, E.J.; Chambers, S.; Mahey, M. Promoting Computational Access to Digital Collections in the Nordic and Baltic Countries: An Icelandic Use Case. Journal of Open Humanities Data 2025, 11. [CrossRef]
  49. Romein, C.A.; Veldhoen, S.F.; Romein, J.C. Applying (Semi-)Automatic Metadata to Early Modern Normative Texts. Annif and Policeygesetzgebung from the City-state of Bern (1528–1798). Digital Scholarship in the Humanities 2025. [CrossRef]
  50. Kelly, P.; Schild, J.; Jafari, A. FolkRAG: A retrieval-augmented generation system for cultural heritage materials. Neural Computing and Applications 2025, 37, 20281–20297. [CrossRef]
  51. Mahadeshwar, R.; Cranenburgh, A.v.; Caselli, T.; Nissim, M. Evaluating the Impact of Source Diversity for RAG in Historical Research. In Proceedings of the Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026); Piperidis, S.; Bel, N.; van den Heuvel, H.; Ide, N.; Krek, S.; Toral, A., Eds., Palma, Mallorca, Spain, May 2026; pp. 716–734. [CrossRef]
Figure 1. Architecture defined for PLOW—Preservation through Linked Open data and Wikidata.
Figure 1. Architecture defined for PLOW—Preservation through Linked Open data and Wikidata.
Preprints 220084 g001
Figure 2. 5-star Open Data - 5-star Open Data
Figure 2. 5-star Open Data - 5-star Open Data
Preprints 220084 g002
Figure 3. Las Meninas. Source: Wikimedia Commons
Figure 3. Las Meninas. Source: Wikimedia Commons
Preprints 220084 g003
Table 1. Summary of the data retrieved from Wikidata.
Table 1. Summary of the data retrieved from Wikidata.
Institution Property No. records
Musée d’Orsay wdt:P4659 1914
National Gallery wdt:P13325 2472
National Gallery of Ireland wdt:P8906 1052
Prado Museum wdt:P8905 4012
Royal Museum of Fine Arts Antwerp wdt:4905 2065
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.