Submitted:
24 June 2026
Posted:
25 June 2026
You are already at the latest version
Abstract
GLAM (Galleries, Libraries, Archives, and Museums) institutions have been making digital collections available for decades. New initiatives to publish, preserve, and reuse data have emerged, with collaborative projects such as Wikidata playing a significant role. This work provides a framework for extracting multimodal collections as data, using Wikidata as the primary source, along with a selection of research scenarios for reusing this data in line with emergent trends in data publication and reuse. Results showed that current trends in data preservation can be achieved through open-source code, cloud services, and collaborative platforms. Future work to be explored includes adopting best practices for provenance documentation and refining reuse scenarios.
Keywords:
collections as data
; Wikidata
; culture heritage
1. Introduction
Preserving digital cultural heritage ensures that cultural and historical information remains accessible, authentic, and usable for future generations. Without active preservation, these materials can become inaccessible over time due to technological change, file corruption, or obsolete software [1]. Simultaneously, there are also increasingly new technologies to ensure computational access and responsible use of collections. Thus Galleries, Libraries, Archives and Museums (GLAMs), as a main provider of cultural heritage data, are presented with challenges of both ensuring long term preservation with shifting technological opportunities and policy instruments for making their digital collections available.
Open Science practices have been defined and implemented to promote open access to scientific publications, research data, models, algorithm collects, software, protocols, notebooks or workflows [2]. In parallel, emergent data sharing initiatives focused on federation, security and sovereignty have emerged during the last years such as Solid and data spaces [3,4]. Community efforts such as GLAM Labs,1 AI4LAM,2 and international initiatives such as Collections as Data,3 FAIR [5] and CARE [6], provide guides for best practices for the publication of machine-actionable collections that support computational access and ethical uses of these data [7,8].
Increasingly institutions have turned to the Semantic Web to describe their digital collections and catalogues [9], with Linked Open Data used as the gold standard for enriching and maximizing the reuse of cultural heritage data. For example, the National Library of Scotland, the Smithsonian American Art Museum or the Rijksmuseum (Amsterdam) published their data using Application Programming Interfaces (API) or as downloadable files [10,11]. They identified existing sources covering different domains with which to enrich their digital collections. Some examples include GeoNames, Virtual International Authority File (VIAF) and Wikidata, a collaborative edition platform owned by Wikimedia Foundation that has played a leading role in the GLAM sector [12,13,14]. Thus, we see Linked Data ontologies and knowledge graph infrastructures in GLAMs are used not only to enrich the information on their collections, but also in making this information publicly accessible as linked open data in a way to efficiently integrate and implement an unprecedented amount of often unstructured, siloed data, at lightning speed.
Despite these affordances, the implementation of linked data technologies in GLAMs remain rare. This is due to a number of issues, largely that the technology and approaches are not ones that traditionally exist in these institutions, nor are taught to collection curators and managers. Despite using catalogue systems with standard ontologies which aid to organize, store and manage collections; the integration of these data with the Semantic Web demands a shift in conceptual data models and data formats. A number of models have been developed for bibliographic records: the IFLA-Library Reference Model (IFLA-LRM), the Resource Description and Access (RDA), and Bibliographic Framework (BIBFRAME); as well as for cultural heritage objects: the Europeana Data Model (EDM). These models are distinct. Consequently the question remains: how can these knowledge bases that use Semantic Web principles to collect structured data be employed for preservation purposes.
To investigate this we take Wikidata as an example, and show this can be extracted and reused in a reproducible manner. We also show the potential (re)uses of the extracted data. The two main contributions of this work are: i) a framework for extracting multimodal collections as data using Wikidata as a primary source; and ii) a selection of research scenarios to reuse this data according to emergent trends in data publication and reuse. With that, this work is relevant to GLAM practitioners, curators, and researchers who wish to reuse and enrich digital collections made available by GLAM institutions.
This article is organized as follows: after a description of the state of the art, we present the framework to extract multimodal collections as data from Wikidata and its application to a selection of museums. The article concludes with the main conclusions and final remarks.
2. Related Work: The Semantic Web & GLAM
In 2006, Tim Berners-Lee, the inventor of the Web, wrote a memo on the Semantic Web, which outlined “a common framework that allows data to be shared and reused across application, enterprise, and community boundaries”; of which LOD served as a technique to describe knowledge [15]. Thus in contrast to referencing an unstructured description of a place, person or object, for example a textual description in book, linked data through standards such as the Resource Description Framework (RDF) provides a standardised structure to organise, store and link information on these entities. For example, a historical statement referencing the Belgian author of the well-known comic book series Tintin could be stated as such “Georges Prosper Remi wrote The Adventures of Tintin” and expressed as a triplet consisting of: a subject (“:Georges Prosper Remi”), a predicate (“:wrote”), and an object (“:The Adventures of Tintin”). Each of these items are represented with unique identifiers (Uniform Resource Identifiers – URIs) that machines are able to read and retrieve.
Crucially, such a model permits linkage to external repositories, thereby enriching metadata beyond the holding institution’s own records. The product of applying these principles to a digital collection is a knowledge graph, the construction of which hinges on a number of interdependent concerns — data modelling, enrichment and quality foremost among them [10,16].
GLAM institutions hold collections of considerable thematic and material breadth, integrating text, images, maps, metadata, postcards, and manuscripts [17].4 To render this heterogeneity usable, they have adopted a variety of publication strategies, ranging from international standards for serving images as the Library of Congress5 and Europeana.6 A growing body of work has, in turn, taken stock of the state of {collections as data} across several national contexts, formulating best practices and guidelines for the publication of machine-actionable collections [18,19,20,21].
Museums, in particular, have been conspicuously active in publishing, enriching and aggregating their catalogues. At one end stand large national databases such as the French Joconde, which documents over a million works held in French art collections [22,23]. Elsewhere the emphasis falls on machine-actionable release: the Rijksmuseum in Amsterdam offers its catalogue as a data dump expressed in established vocabularies, notably CIDOC-CRM.7 and comparable metadata datasets have been derived from the Prado [24]. A further strand publishes catalogue metadata as Linked Open Data (LOD), as the Smithsonian American Art Museum and the Getty Provenance Index8 have done an approach whose value is far from merely technical, since metadata covering dealer stock books and sales catalogues can support the identification of art looted during the Second World War. Underpinning much of this activity are shared ontologies: CIDOC-CRM,9 widely adopted for describing the physical nature of cultural heritage, and community initiatives such as Linked Art, which promote agreed data models for the domain [25].10 Running alongside these efforts, experiments with multimodal datasets — combining text, images and metadata — have sought to advance knowledge discovery and model development, though not without raising challenges in storage and data modelling [25,26].
Once published, such data lends itself to reuse along several axes. One line of enquiry has applied automated methods to assess the quality of the LOD repositories released by cultural heritage (CH) institutions [16]; another has supplied a reproducible, machine-readable framework for visualising those repositories within the GLAM sector [27]. More recently, the coupling of large language models with external information retrieval — in the form of Retrieval-Augmented Generation (RAG) systems — has been deployed to widen access to archival materials [28]. None of this proceeds without adequate description, and detailed documentation has accordingly become indispensable, both to encourage reuse and to render the structure and content of CH datasets intelligible [29].
Computational notebooks, and Jupyter Notebooks in particular, have meanwhile acquired a notable place in the GLAM toolkit, furnishing a varied repertoire of reproducible and inventive examples [30,31,32], and an array of use cases and research scenarios has been articulated across an equally wide thematic range [33]. Their uptake, however, remains uneven. Adoption within (Western) GLAM and across data and research infrastructures remains emergent, marked by considerable disparities in technical capacity, workflow design, and integration with cloud services [34].
2.1. Wikidata
Wikimedia projects serve a range of functions in the GLAM context, from linking and enriching catalogue records to publishing multimedia content [12]. Many institutions have created properties in Wikidata to connect their records to external sources [35]. Listing 1 shows a SPARQL query for retrieving Wikidata properties used to link to this institutions. Through its connection to Wikimedia Commons, Wikidata can also surface the multimedia content held there, such as photographs of artworks in museum collections.11 Properties such as wdt:P18 and wdt:P6802, accordingly, serve to record the links to these images.12 Underpinning these capabilities is Wikibase, the software on which Wikidata runs and which institutions increasingly deploy in their own right to produce machine-readable, FAIR data through cloud services [5,36].

- Listing 1: SPARQL query to retrieve Wikidata properties to link museums
Wikidata’s adoption in the GLAM sector has been examined from several angles, and the recurring verdict is favourable. The infrastructure sustained by the Wikimedia Foundation and maintained by an active community has emerged as a viable option for cultural heritage institutions [37,38], not least because it addresses needs that are otherwise costly to meet: within the Digital Humanities it serves both to recover missing identifiers across all categories of named entity and as a source for controlled-vocabulary integration [39], while affording a transparent, collaborative environment in which cultural heritage data can be described and linked [40]. The balance is not wholly one-sided—the benefits and the attendant challenges of working in the sector have both been documented [41]—yet the barrier to entry is falling, as recent cloud services make Wikibase more accessible to newcomers in the DH and GLAM communities [36].
The body of work surveyed above is, taken together, substantial, but it leaves one issue unaddressed. Existing efforts have, for the most part, reused established datasets: whether to generate visualisations or to evaluate (new) AI-based services—and where frameworks for data extraction have been proposed, complete with reproducible code, they have not taken the collection as their principal object. What is missing, in short, is an approach that brings these three concerns together at once: the GLAM institution as primary content, a governing purpose, and using reproducible code as the means.
3. Methodology
The proposed methodology is based on a framework in four steps:
- 1.
- identification of sources;
- 2.
- data extraction;
- 3.
- publication; and
- 4.
- reuse.
Figure 1 shows the architecture employed for a tool we named PLOW—Preservation through Linked Open data and Wikidata. The individual steps are described in detail below.
3.1. Identification of Sources
The first step is the identification of a data source, for which several technical options exist. GLAM datasets are commonly published through GLAM Labs—the Data Foundry at the National Library of Scotland being a representative case; and through the dedicated data sections of major libraries, such as the API et jeux de données of the National Library of France or the equivalent position at the Rijksmuseum in Amsterdam.13 These rely chiefly on data dumps that may be downloaded and reused with little friction; others, by contrast, expose their holdings through APIs—SPARQL, OAI-PMH and IIIF among them. Cross-domain hubs occupy a third position: Wikidata, in particular, has become a significant identifier hub in the GLAM sector, enabling the integration of metadata with multimedia content.
Identifying a source, however, is not the same as selecting one, and several criteria bear on that choice. Foremost among them is licensing: while many datasets are released under open licences such as Creative Commons, a substantial number remain copyrighted, or carry no clear statement of reuse at all [19,20]. Considerations of a different order also intrude: among them, the protection of cultural heritage placed at risk by war, as in Ukraine, Gaza and Afghanistan.14 To these may be added the familiar measures of fitness for purpose: data quality, completeness, timeliness, and availability [16], together with thematic relevance, the degree, to which a source speaks to the subject or research question at hand. Such relevance is best made concrete: a study of artists in exile during the Spanish Civil War, of the paintings of a particular movement, or of works looted by the Nazis during the Second World War will each privilege very different collections [42,43].
3.2. Data Extraction
This step involves extracting the data from the source identified. The output may take a range of serialisation formats: JSON, XML or RDF, among the most common, and frequently extends beyond textual metadata to multimedia content such as images, maps or OCR-derived text. Extraction need not rely solely on the originating institution alone: aggregators such as Europeana or collaboratively edited platforms such as Wikidata can serve equally well as points of access.
In some cases, this step can be related to a particular topic, literary movement or event (e.g., World War II). In other cases, the data is extracted based on membership relations such as being part of a particular catalogue (e.g., Rijksmuseum Amsterdam or National Library of France). In this context, Wikidata has created many properties for GLAM institutions in order to link their catalogues [44]. For example, Listing 2 shows a SPARQL query to retrieve artworks from Wikidata linked to the Prado Museum by means of the dedicated property wdt:P8905. Other examples rely on using the CONSTRUCT clause in SPARQL to build a knowledge graph from the triples retrieved by a query.

- Listing 2: SPARQL query to retrieve Art paintings from Wikidata linked to the Prado Museum
Once extracted, metadata, and images alike may also surface in aggregators such as Europeana, where they can be integrated and browsed alongside further digital collections.15
3.3. Publication
The final step concerns the publication of the dataset and its subsequent reuse. Established guidance exists for the former: best-practice frameworks set out the principal steps for publishing machine-actionable collections, within Europeana and beyond [20,33].
A useful organising principle here is Berners-Lee’s five-star deployment scheme (see Figure 2), which ranks open data by degree of accessibility: from material merely placed online under an open licence, through structured and non-proprietary formats, to the use of URIs and, at the highest level, linking to other data to provide context. The scheme is valuable precisely because it reframes publication not as a binary of open versus closed but as a gradient of reusability, against which an institution may locate its own holdings and chart a path of improvement [45].
3.4. Reuse
Reuse, for its part, has been approached from several directions, each building on the content once published. Some work has defined research scenarios grounded in that content [33]. Among the practical instruments, the Jupyter Notebook has gained particular traction, allowing institutions to furnish prototype examples that combine code, charts, documentation and interactive widgets in a single reproducible artefact [34]. Institutions have begun to employ Jupyter Notebooks as a powerful tool to provide examples of use through prototypes. They combine code, charts, documentation, and interactive widgets [31,32]. Further examples and prospective research scenarios drawn from GLAM data are set out in Section 2.
Of particular interest is the integration of published collections with data and research infrastructures. The Common European Data Space for Cultural Heritage and the European Collaborative Cloud for Cultural Heritage (ECCCH) both operate at this level, as does the European Open Science Cloud (EOSC)—whose Jupyter Notebook service, notably, permits reproducible code to be run on dedicated cloud servers. Visibility and reuse are, in turn, served by platforms such as the Social Sciences and Humanities Open Marketplace, Zenodo and OpenAIRE [21].
One further consideration deserves emphasis: versioning. Institutions frequently wish to publish their catalogues as a series of updates rather than a single release, and platforms such as Zenodo accommodate this directly, issuing a distinct DOI and a suggested citation for each version [46].
4. Results
In this section, we describe the application of the framework in Section 3 for several GLAM institutions. We specifically have selected museums, given the acknowledged contributions to Wikidata and thus available data. The museums were selected according to the following criteria:
- 1.
- availability of a dedicated Wikidata property to link artworks;
- 2.
- historical importance, cultural influence, as well as digital transformation;
- 3.
- provision of documentation of the catalogue.
This resulted in the following cases: Musée d’Orsay, National Gallery, National Gallery of Ireland, Prado Museum, and the Royal Museum of Fine Arts Antwerp.
Wikidata was used as the primary repository for extracting multimodal content, including metadata, and images. To extract the metadata, the SPARQL query included in Appendix A was employed. The data was retrieved in May 2026. Note that images are stored on Wikimedia Commons, where Wikidata only preserves a link. Table 1 show a summary of the results following the application of the framework.
A GitHub repository was created including the extracted data, a README file with additional documentation and a collection of Jupyter Notebooks illustrating how to extract and reuse the data.16 Several Python packages, such as SPARQLWrapper, were employed to query the Wikidata SPARQL endpoint.17 The data was extracted as a JSON file. The Jupyter notebooks were hosted and executed using the services provided by EOSC. To do so, a dedicated virtual machine was deployed to run the code provided. This work intends to promote its use amongst the community. Best practices were used to develop and publish the collection of Jupyter Notebooks [30,31].
4.1. Scenarios of Reuse
This section outlines a selection of scenarios in which the extracted data might be reused, each grounded in prior work spanning topics ranging from the integration of data repositories to the application of various AI-based techniques [34,37]. The scenarios are addressed to a mixed readership of practitioners, curators, researchers and students, and their purpose is illustrative rather than prescriptive: they show how data may be extracted and preserved in keeping with the best practices promoted by international initiatives, without pretending to a full specification or a worked implementation. What they offer, in short, is a narrative account of added value rather than a manual.
Federated queries. Institutions tend to publish digital collections as an isolated collections effectively making a data silo and thus hindering collaboration and their reuse. However, advances in technology have created a new context in which methods such as Named Entity Recognition (NER), and Entity Linking are helping enrich digital collections. When linking records to external repositories, new opportunities arise to enrich metadata, improve interoperability, and enable broader access to cultural and research collections. Wikidata federates several SPARQL repositories, enabling access to different repositories from relevant institutions through Wikidata [47].
For example, Listing 3 shows a federated SPARQL query to retrieve bibliographic records from Biblioteca Virtual Miguel de Cervantes (BVMC) in which the painter Diego Velázquez is represented as a subject. Note that the SERVICE command is employed to access the BVMC SPARQL endpoint from Wikidata. The BIND command is employed to filter the bibliographic records using the Picasso’s identifier in BVMC. Other examples could be based on the integration of painting from catalogues such as the Prado Museum and the National Gallery, focused, for instance, on a particular painter such as Diego Velázquez.

- Listing 3: SPARQL query to retrieve bibliographic records from Biblioteca Virtual Miguel de Cervantes in which the painter Diego Velázquez is represented as a subject.
According to these observations, museum data could be enriched with data from libraries and archives, such as the National Library of France, to create integrated research scenarios that combine different types of records. This could help address challenges related to data quality, interoperability, and reuse [16].
AI-based computer vision models. Art works in Wikidata are described using many properties (e.g., title and inception). However, given Wikidata’s collaborative editing approach, properties are not consistently used across all entities. The property wdt:P180 is used to indicate what an item depicts, in particular, to include entities visually depicted in an image (e.g., Spanish Civil War).18 Computer vision models can be employed to enhance the value of this property by leveraging the multimedia content generated by the data extraction process, thereby enabling the refinement and enhancement of image descriptions [48]. Metadata can be automatically generated to describe images provided by GLAM institutions in greater detail, such as paintings, maps, or postcards [34,49].
See, for instance, Figure 3 representing Las Meninas from Diego Velázquez. The Wikidata identifier wd:Q208758 provides metadata about the painting, including the property wdt:P180 with several values such as people included in the painting, the painter himself or a dog.
Federated data storage systems. Solid is an emerging initiative for storing federated data while ensuring security and sovereignty [3]. Previous research has analysed the use of Solid in the GLAM context, combined with Wikidata as a federated storage system.19 Software libraries and tutorials in different programming languages, such as Java and Python, have been developed to facilitate their adoption.20 These new technologies facilitate data portability across applications while enhancing transparency and privacy in data access.
RAG-based search engine. Wikidata may be interrogated by several means—keyword matching or SPARQL, among them—yet each has its limits: a keyword set may fail to capture the intent behind a query, and SPARQL’s steep learning curve places it beyond many users. Wikimedia has accordingly begun exploring semantic search techniques, including vector embeddings, as a more forgiving mode of access.21 Coupled with large language models and natural language processing, such methods can materially improve retrieval [50], and earlier work has already applied Retrieval-Augmented Generation to web-archive content and digitised collections, thereby opening those holdings to conversational access.22 Hence, the artwork descriptions held in Wikidata records might themselves underwrite a RAG search engine—one that affords new routes into the content, curbs hallucination, and injects domain-specific knowledge into the underlying model.
LLM or SLM-based chatbot. A natural specialisation of the retrieval-augmented scenario sketched above is a conversational agent addressed not to the researcher but to the museum visitor: a chatbot, in effect, with which one might interrogate a single art object one is standing before. The premise is straightforward: the multimodal records extracted by the framework: title, creator (wdt:P170), date of inception (wdt:P571), the entities recorded under wdt:P180 and the associated image wdt:P18—furnish a bounded, machine-readable context against which a visitor’s questions may be answered. Rather than draw on its own parametric memory, with all the attendant risk of confabulation, the model retrieves the relevant record and renders it as natural-language prose; and where the visitor asks what a painting depicts, the very property whose enrichment was discussed above supplies the answer. The choice between a large language model (LLM) and a small language model (SLM) is, in this setting, less a technical detail than a design decision with institutional consequences. An LLM offers fluency, breadth, and a capacity to field unanticipated questions, but it is costly to serve continuously to the public and, when cloud-hosted, returns us to the difficulty noted above of data leaving European jurisdiction. An SLM, a model of a few billion parameters capable of running on modest, or even on-premise hardware, inverts these trade-offs. Its narrower internal knowledge is, paradoxically, an advantage rather than a constraint: confined to answering from the retrieved record alone, it is the more readily held to the institution’s own authoritative data, and the less inclined to volunteer plausible invention. Hence, for a bounded domain such as the holdings of a single museum, pairing a modest SLM with strong retrieval grounding may prove not merely the cheaper option but also the more trustworthy one [51].
Wikidata records carry labels in many languages, the mechanism is already exploited by the SERVICE clause in Listing 2—a visitor might both pose questions and receive answers in their own tongue, without additional translation infrastructure. In order to foster transparency, each response may be accompanied by a reference to the Wikidata item from which it was drawn, so that the curious or the sceptical can verify the claim at its source. It should be stressed, however, that such a system inherits the limitations of its underlying data. Where Wikidata’s coverage of the Prado is partial, or where wdt:P180 has been applied unevenly, the agent will be correspondingly silent or incomplete; and no degree of retrieval grounding wholly forecloses error. A visitor-facing use would therefore demand careful evaluation against curatorial ground truth before it could responsibly stand between the public and the collection.
These research scenarios provide a wide range of approaches for reusing the content extracted by the framework proposed in this work.
4.2. Discussion
This work complements existing approaches to the extraction, publication and reuse of digital collections, but it occupies a position none of them does: it takes digital preservation as its governing concern, a collaborative platform—Wikidata—as its means, and the museum collection as its principal content. Its contribution, then, is to enrich the current landscape of worked examples, supplying the community with additional cases that follow established best practices.
Several limitations bound the scope of what has been attempted here. Wikidata does not hold a complete representation of the Prado’s catalogue; a substantial selection of records has nonetheless been linked to it, thanks to the collaborative efforts of the community. It would be valuable, even so, for the institution to offer access by additional means—OAI-PMH, IIIF, or SPARQL, each weighing richness, intelligibility, and complexity differently—and, looking further ahead, through sovereignty-preserving approaches such as data spaces. Using Wikidata entails several challenges that must be addressed. For example, data uploaded to Wikidata is highly reused by AI companies to create commercial products, while GLAM institutions receive no compensation. In exchange, the visibility of the digital collections increases, impacting society. In some cases, the data modelling provided by Wikidata might not be suitable to represent indigenous knowledge. A more inclusive approach is required in these cases in order to facilitate an equitable and non-discriminatory environment.
The EOSC cloud services are constrained in memory and disk storage; for the examples presented here, those resources sufficed, but more demanding AI-based work would require provision with greater GPU and memory capacity [48]. Alternatives such as Google Colab supply additional resources, though at the cost of depositing data outside Europe—a consideration of some weight for European institutions and researchers.
The reuse scenarios have been presented as textual sketches: prototypes in narrative form rather than working systems, and their implementation remains a task for further work; refinement may in any case prove necessary, since several of the data and research infrastructures invoked are themselves still under development.
The work admits of extension in several directions. The SPARQL queries might be refined to draw in additional classes of institutions, libraries, and archives, and data research infrastructures, such as those discussed above, could be pursued in the service of preservation and reuse alike.
5. Conclusions
GLAM institutions have, in recent years, been exploring new ways of making their digital collections available, and the preservation of digital heritage has become a matter of real public consequence—a means of safeguarding cultural heritage for the future. A range of initiatives have emerged to promote the publication and responsible reuse of the data these institutions hold. In this research we showed how Wikidata can be used in these tasks.
We proposed a framework for extracting multimodal collections as data using Wikidata as a primary source. In keeping with Open Science principles, this work has described a reproducible approach for extracting multimodal collections as data from museums via Wikidata. The framework was applied to a selection of museums, with Wikidata serving as the principal repository, and was accompanied by research scenarios demonstrating how the extracted datasets might be reused. Taken together, this demonstrates how data preservation can be met through open-source code, cloud services and collaborative platforms—without recourse to bespoke or proprietary infrastructure.
Many institutions have, to be sure, already turned to Wikidata for a variety of ends; but its adoption, enrichment and completeness remain unevenly realised. The obstacles are partly practical—a good number of institutions lack the resources and the expertise to begin—and it is here that the community’s sustained work in linking and curating entities has proved decisive. Wikibase, moreover, now offers a further avenue: as cloud provisioning matures, it affords full control over data and governance, together with the freedom to tailor the data model in ways that hosted Wikidata does not.
Two directions invite further work. The first is the adoption of best practices to record provenance in a machine-readable form [46]. The second is the refinement of the reuse scenarios themselves—supplying the structured, machine-actionable detail that would allow others to take them up in earnest. Of the two, the latter is likely to do the most to encourage adoption.
Author Contributions
Conceptualisation: GC; Formal Analysis: all; Investigation: all; Methodology: GC; Visualizations: GC; Resources: all; Software: GC; Writing original draft: GC; Writing – review & editing: all.
Acknowledgments
The authors would like to thank the International GLAM Labs Community, the Impact Centre of Competence, Wikimedia Spain and DARIAH-EU. C. Annemieke Romein’s research within the HAICu-project (digital Humanities Artificial Intelligence and Cultural Heritage) is funded by the NWO NWA 1518.22.10 grant. The contributions of Julie M. Birkholz are supported by the FWO-funded Large-Scale Infrastructure project CLARIAH-VL+ (I001525N) and BELSPO FED-tWIN project (Prf-2019-040).
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| AI | Artificial Intelligence |
| API | Application Programming Interface |
| BVMC | Biblioteca Virtual Miguel de Cervantes |
| CARE | Collective Benefit, Authority to Control, Responsibility, Ethics |
| CH | Cultural Heritage |
| DOI | Digital Object Identifier |
| ECCCH | European Cultural Heritage Cloud |
| EOSC | European Open Science Cloud |
| FAIR | Findable, Accessible, Interoperable, Reusable |
| GLAM | Galleries, Libraries, Archives and Museums |
| GPU | Graphics Processing Unit |
| IIIF | International Image Interoperability Framework |
| LLM | Large Language Model |
| LOD | Linked Open Data |
| OAI-PMH | Open Archives Initiative Protocol for Metadata Harvesting |
| RAG | Retrieval-Augmented Generation |
| RDF | Resource Description Framework |
| SPARQL | SPARQL Protocol and RDF Query Language |
| VIAF | Virtual International Authority File |
Appendix A
Listing 4 was employed to extract the metadata from the Wikidata public SPARQL endpoint for the Prado Museum.

- Listing 4: SPARQL query to retrieve Art paintings from Wikidata linked to the Prado museum
References
- Terras, M. Opening access to collections: The making and using of open digitised cultural content. Online Information Review 2015, 39, 733–752. [CrossRef]
- European Commission. Open Science, 2026. Accessed: 2026-05-15.
- Mansour, E.; Sambra, A.V.; Hawke, S.; Zereba, M.; Capadisli, S.; Ghanem, A.; Aboulnaga, A.; Berners-Lee, T. A Demonstration of the Solid Platform for Social Web Applications. In Proceedings of the Proceedings of the 25th International Conference Companion on World Wide Web, Republic and Canton of Geneva, CHE, 2016; WWW ’16 Companion, p. 223–226. [CrossRef]
- Trujillo, J.; Candela, G.; Reina-Reina, A. Data Analytics and Artificial Intelligence in the new scenario of Data Spaces. In Proceedings of the Proceedings of the 27th International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP 2025) co-located with the 28th International Conference on Extending Database Technology and the 28th International Conference on Database Theory (EDBT/ICDT 2025), March 25, 2025; Maté, A.; Lissandrini, M., Eds., Barcelona, Spain, 2025; Vol. 3931, CEUR Workshop Proceedings, pp. 91–92.
- Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 2016, 3, 160018. [CrossRef]
- Carroll, S.R.; Garba, I.; Figueroa-Rodríguez, O.L.; Holbrook, J.; Lovett, R.; Materechera, S.; Parsons, M.A.; Raseroka, K.; Rodriguez-Lonebear, D.; Rowe, R.; et al. The CARE Principles for Indigenous Data Governance. Data Sci. J. 2020, 19, 43. [CrossRef]
- Mahey, M.; Al-Abdulla, A.; Ames, S.; Bray, P.; Candela, G.; Derven, C.; Dobreva-McPherson, M.; Gasser, K.; Chambers, S.; Karner, S.; et al. Open a GLAM lab; International GLAM Labs Community, Book Sprint: Doha, Qatar, 2019; p. 164. [CrossRef]
- Padilla, T.; Allen, L.; Frost, H.; Potvin, S.; Russey Roke, E.; Varner, S. Final Report — Always Already Computational: Collections as Data, 2019. [CrossRef]
- Berners-Lee, T.; Hendler, J.; Lassila, O. The Semantic Web in Scientific American. Scientific American Magazine 2001, 284.
- Candela, G. Towards a semantic approach in GLAM Labs: The case of the Data Foundry at the National Library of Scotland. J. Inf. Sci. 2026, 52, 3–21. [CrossRef]
- Dijkshoorn, C.; Jongma, L.; Aroyo, L.; van Ossenbruggen, J.; Schreiber, G.; ter Weele, W.; Wielemaker, J. The Rijksmuseum collection as Linked Data. Semantic Web 2018, 9, 221–230. [CrossRef]
- Zhao, F. A systematic review of Wikidata in Digital Humanities projects. Digital Scholarship in the Humanities 2022, 38, 852–874, [https://academic.oup.com/dsh/article-pdf/38/2/852/50488385/fqac083.pdf]. [CrossRef]
- Freire, N.; Isaac, A. Wikidata’s Linked Data for Cultural Heritage Digital Resources: An Evaluation Based on the Europeana Data Model. International Conference on Dublin Core and Metadata Applications 2020, 2020, 59–68. [CrossRef]
- Candela, G.; Cuper, M.; Holownia, O.; Gabriëls, N.; Dobreva, M.; Mahey, M. A Systematic Review of Wikidata in GLAM Institutions: A Labs Approach. In Proceedings of the Linking Theory and Practice of Digital Libraries - 28th International Conference on Theory and Practice of Digital Libraries, TPDL 2024, September 24-27, 2024, Proceedings, Part II; Antonacopoulos, A.; Hinze, A.; Piwowarski, B.; Coustaty, M.; Nunzio, G.M.D.; Gelati, F.; Vanderschantz, N., Eds., Ljubljana, Slovenia, 2024; Vol. 15178, Lecture Notes in Computer Science, pp. 34–50. [CrossRef]
- Berners-Lee, T. Linked Data http://www. w3. org/DesignIssues. LinkedData. html 2006.
- Candela, G. An automatic data quality approach to assess semantic data from cultural heritage institutions. J. Assoc. Inf. Sci. Technol. 2023, 74, 866–878. [CrossRef]
- Shang, C. Quantitative assessment of the value of intangible cultural heritage art supported by multimodal machine learning. Discov. Artif. Intell. 2026, 6, 7. [CrossRef]
- Disli, M.; Gabriëls, N.; Chambers, S.; Ames, S.; Knazook, B.; Candela, G. Exploring the adoption of collections as data in the GLAM context. Inf. Res. 2025, 30, 65–77. [CrossRef]
- Dişli, M.; Candela, G. Copyright and Licencing for Cultural Heritage Collections As Data. Journal of Open Humanities Data 2025. [CrossRef]
- Candela, G.; Gabriëls, N.; Chambers, S.; Dobreva, M.; Ames, S.; Ferriter, M.; Fitzgerald, N.; Harbo, V.; Hofmann, K.; Holownia, O.; et al. A checklist to publish collections as data in GLAM institutions. Global Knowledge, Memory and Communication 2023, 74, 1323–1355. [CrossRef]
- Candela, G.; Chambers, S.; Irollo, A. A Collections as Data Workflow on the SSH Open Marketplace. https://marketplace.sshopencloud.eu/workflow/I3JvP6, 2024. SSH Open Marketplace workflow.
- Moissinac, J.C.; Rouzé, F.; Wadhera, P.; Germain, B. Toward a Semantic Representation of the Joconde Database. In Proceedings of the Semapro, IARIA, Nice, France, 10 2020.
- de la Culture, M. Collections des musées de France: Base Joconde, 2026.
- Gonzalez Cifuentes, T.; Torrejon Vazquez, D. Dataset de obras del Museo Nacional del Prado, 2026. [CrossRef]
- Sanderson, R. Implementing Linked Art in a Multi-Modal Database for Cross-Collection Discovery. Open Library of Humanities 2024, 10. [CrossRef]
- Nguyen, T.; Storås, A.M.; Thambawita, V.; Hicks, S.A.; Halvorsen, P.; Riegler, M.A. Multimedia Datasets: Challenges and Future Possibilities. In Proceedings of the MultiMedia Modeling; Dang-Nguyen, D.T.; Gurrin, C.; Larson, M.; Smeaton, A.F.; Rudinac, S.; Dao, M.S.; Trattner, C.; Chen, P., Eds., Cham, 2023; pp. 711–717.
- Candela, G. Browsing Linked Open Data in Cultural Heritage: A Shareable Visual Configuration Approach. ACM Journal on Computing and Cultural Heritage 2025, 18, 9:1–9:15. [CrossRef]
- Kelly, P.; Schild, J.; Jafari, A. FolkRAG: A retrieval-augmented generation system for cultural heritage materials. Neural Comput. Appl. 2025, 37, 20281–20297. [CrossRef]
- Alkemade, H.; Candela, G.; Claeyssens, S.; Colavizza, G.; Eren, S.; Freire, N.; Irollo, A.; Isaac, A.; Lehmann, J.; Neudecker, C.; et al. Datasheets for Digital Cultural Heritage Datasets, 2025. [CrossRef]
- Pimentel, J.F.; Murta, L.; Braganholo, V.; Freire, J. Understanding and improving the quality and reproducibility of Jupyter notebooks. Empir. Softw. Eng. 2021, 26, 65. [CrossRef]
- Candela, G.; Chambers, S.; Sherratt, T. An approach to assess the quality of Jupyter projects published by GLAM institutions. J. Assoc. Inf. Sci. Technol. 2023, 74, 1550–1564. [CrossRef]
- Sherratt, T. GLAM Workbench, 2025. [CrossRef]
- Candela, G.; Rosiński, C.; Margraf, A. A reproducible framework to publish and reuse Collections as data: the case of the European Literary Bibliography. Transformations: A DARIAH Journal 2025, Workflows. [CrossRef]
- Candela, G.; Dobreva, M.; Alkemade, H.; Holownia, O.; Mahey, M.; Ames, S.; Renaud, K.; Vodopivec, I.; Lee, B.C.G.; Padilla, T.; et al. A Use Case Lens on Digital Cultural Heritage, 2025, [arXiv:cs.DL/2509.08710].
- Dişli, M.; Candela, G.; Gutiérrez, S.; Fontenelle, G. Open Data Practices of Art Museums in Wikidata: A Compliance Assessment. Journal of Open Humanities Data 2025. [CrossRef]
- Lindemann, D.; Candela, G.; Marchetti, A.; Pellizzari di San Girolamo, C.C.; Olea, I.; Varvantakis, C.; Assis, T.; Moitinho de Almeida, V.; Obregón Sierra, Á.; Santiago Faria, A.; et al. The Wikibase Ecosystem in DH and GLAM, 2025. [CrossRef]
- Thornton, K. Wikidata for Digital Preservationists. Technical report, Digital Preservation Coalition, 2021. [CrossRef]
- Thornton, K.; Seals-Nutt, K. Metadata, 2022. In-Person Long Paper. Recording available at https://youtu.be/Xx6-Z4EDxEk. [CrossRef]
- Trognitz, M.; Mandell, R.A.; Štuhec, S.; Palacz, J. Wikidata as a Reconciliation Anchor: Curating Data for Long-Term Preservation. Journal of Open Humanities Data 2026. [CrossRef]
- Sichani, A.M.; Kono, K.; Winters, J. Building Cultural Heritage Data Infrastructures with Wikidata: The Case of the Congruence Engine Data Register. Journal of Open Humanities Data 2026. [CrossRef]
- Smith-Yoshimura, K. Experimentations with Wikidata/Wikibase, 2020. Hanging Together blog. Accessed 2026-05-19.
- Salinas, C.G. Las artistas del exilio republicano español: El refugio latinoamericano; Arte Grandes Temas, Ediciones Cátedra: Madrid, 2019.
- J. Paul Getty Trust. Getty Provenance Index, n.d. Accessed: 2026-05-19.
- Romein, C.A.; Wagner, A.; van Zundert, J.J. Building and Deploying a Classification Schema using Open Standards and Technology. Journal for Digital Legal History 2023, 2. [CrossRef]
- ROMEIN, C.A.; KEMMAN, M.; BIRKHOLZ, J.M.; BAKER, J.; DE GRUIJTER, M.; MEROÑO-PEÑUELA, A.; RIES, T.; ROS, R.; SCAGLIOLA, S. State of the Field: Digital History. History 2020, 105, 291–312, [https://onlinelibrary.wiley.com/doi/pdf/10.1111/1468-229X.12969]. [CrossRef]
- Romein, C.A.; Hodel, T.; Gordijn, F.; van Zundert, J.J.; Chagué, A.; et al. Exploring Data Provenance in Handwritten Text Recognition Infrastructure. Journal of Data Mining and Digital Humanities 2024. [CrossRef]
- Dişli, M.; Osti, G.; Candela, G.; Zijdeman, R.J. Federated LOD Queries as CaD - Notebooks, 2025. [CrossRef]
- Candela, G.; Holownia, O.; Odsbjerg, M.; Cuper, M.; Gabriëls, N.; Hofmann, K.; Gray, E.J.; Chambers, S.; Mahey, M. Promoting Computational Access to Digital Collections in the Nordic and Baltic Countries: An Icelandic Use Case. Journal of Open Humanities Data 2025, 11. [CrossRef]
- Romein, C.A.; Veldhoen, S.F.; Romein, J.C. Applying (Semi-)Automatic Metadata to Early Modern Normative Texts. Annif and Policeygesetzgebung from the City-state of Bern (1528–1798). Digital Scholarship in the Humanities 2025. [CrossRef]
- Kelly, P.; Schild, J.; Jafari, A. FolkRAG: A retrieval-augmented generation system for cultural heritage materials. Neural Computing and Applications 2025, 37, 20281–20297. [CrossRef]
- Mahadeshwar, R.; Cranenburgh, A.v.; Caselli, T.; Nissim, M. Evaluating the Impact of Source Diversity for RAG in Historical Research. In Proceedings of the Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026); Piperidis, S.; Bel, N.; van den Heuvel, H.; Ide, N.; Krek, S.; Toral, A., Eds., Palma, Mallorca, Spain, May 2026; pp. 716–734. [CrossRef]
| 1 | |
| 2 | |
| 3 | |
| 4 | See, for example, https://www.glamlabs.io/projects. |
| 5 | |
| 6 | |
| 7 | |
| 8 | |
| 9 | |
| 10 | |
| 11 | See, for example, https://commons.wikimedia.org/wiki/Museo_del_Prado
|
| 12 | See, for example, Picasso’s painting Guernica in Wikidata available at https://www.wikidata.org/wiki/Q175036. |
| 13 | See, for example, https://api.bnf.fr/. |
| 14 | See, for example, the AISTER project available at https://aister.uni.lu/. |
| 15 | |
| 16 | |
| 17 | |
| 18 | See, for example, the Guernica entity in Wikidata available at https://www.wikidata.org/wiki/Q175036
|
| 19 | |
| 20 | |
| 21 | |
| 22 |
Figure 1.
Architecture defined for PLOW—Preservation through Linked Open data and Wikidata.

Figure 2.
5-star Open Data - 5-star Open Data
Figure 2.
5-star Open Data - 5-star Open Data

Figure 3.
Las Meninas. Source: Wikimedia Commons

Table 1.
Summary of the data retrieved from Wikidata.
| Institution | Property | No. records |
|---|---|---|
| Musée d’Orsay | wdt:P4659 | 1914 |
| National Gallery | wdt:P13325 | 2472 |
| National Gallery of Ireland | wdt:P8906 | 1052 |
| Prado Museum | wdt:P8905 | 4012 |
| Royal Museum of Fine Arts Antwerp | wdt:4905 | 2065 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.