Submitted:
04 September 2026
Posted:
07 September 2026
You are already at the latest version
Abstract
This paper argues that today’s mainstream generative AI models do not yet provide the level of culturalization required in the creative and cultural industries (CCIs). Debates on generative AI in cultural domains have focused primarily on productivity, innovation, intellectual property, labor disruption and bias. While these concerns are important, they do not fully capture a more basic issue: cultural production depends on forms of meaning that are historically situated, socially mediated and unevenly distributed across languages, communities and markets. In such contexts, fluency, safety and technical usefulness are not sufficient measures of system quality. The paper discusses how an existing line of thought from localization, cross-cultural design and culturalization research should be carried into the era of generative AI as explicit attention to cultural context, symbolic meaning and interpretive plurality. The CCIs provide a particularly revealing testbed for this claim because they expose the limits of generic AI systems in domains where value depends on tone, memory, symbolism and representation. To ground the argument, the article revisits earlier work on targeted narratives, personalization and adaptive cultural mediation across museums, heritage sites and pilgrimage routes, and uses recent exploratory attempts with widely used generative AI models to show what remains missing. On that basis, the article proposes a framework for culturalized generative AI and outlines an integrated research and governance agenda spanning AI, the humanities and cultural-sector practice.
Keywords:
generative AI
; culturalization
; creative and cultural industries
; adaptive cultural mediation
; cultural heritage
; personalization
; AI governance
1. Introduction
Generative AI is rapidly transforming the creative and cultural industries (CCIs). Across writing, audiovisual production, design, music and interactive media, generative systems are increasingly being used to support ideation, prototyping, editing and content generation. Recent scholarship suggests that these technologies may lower some barriers to creative production, accelerate workflows and expand experimentation, while also reshaping the organization of labor and the conditions under which creative work is performed [1,2,3,4]. In this sense, generative AI is not simply another production tool: it is becoming part of the socio-technical infrastructure through which cultural goods are imagined, produced, circulated and valorized.
At the same time, the growing use of generative AI in the CCIs has uncovered a major limitation. These systems may produce fluent, polished and commercially usable outputs, yet remain poorly equipped to engage with cultural nuance. The challenge is not reducible to technical quality alone. In cultural production, meaning is rarely universal or context-free; it is shaped by language, symbolism, memory, genre, social norms and audience interpretation. What counts as compelling, respectful, humorous, authentic or emotionally resonant in one culture may be inappropriate, offensive or banal in another. In other words, cultural adequacy is not a single standard; it varies across communities and situations. For this reason, the adoption of generative AI in the CCIs raises not only questions of efficiency and creativity, but also questions of cultural adequacy and ethics. This concern is increasingly echoed in current research and policy discussions on AI and culture, which warn that AI systems are being embedded in cultural ecosystems faster than corresponding frameworks for cultural governance, diversity and rights are being developed [5,6].
The problem becomes especially visible when generative systems are used in concrete cultural workflows. In film and television, for example, cultural adequacy involves more than translation: it can concern the representation of historical antagonists, the framing of political violence or the symbolic weight of particular settings and visual tropes [7,8]. In games, generative systems may reshape worlds, dialogue and visual assets that draw on mythology, religion or political identities [9,10]. In music, adaptation often depends on rhythm, instrumentation, affect and style as much as on lyrics [11,12]. In museum and heritage settings, AI-mediated storytelling can influence how visitors encounter the past, interpret contested histories and position themselves in relation to cultural memory [13]. Across these and other contexts, the same basic tension appears: systems that increase speed and accessibility can also encourage standardization, flatten difference and steer creators toward outputs that are more predictable, more culturally generic and more similar to one another. Experimental evidence already suggests that while generative AI can enhance individual creative performance, especially for less experienced creators, it may simultaneously reduce collective diversity and cultural novelty in outputs [14,15].
The empirical studies developed in this article are drawn primarily from museums, cultural heritage and pilgrimage contexts. These domains are treated as demanding testbeds for culturalized generative AI because they make historically grounded knowledge, source provenance, symbolic meaning, interpretive plurality and audience-sensitive mediation especially visible. Although these characteristics are especially explicit in heritage, they also arise in other creative and cultural domains whenever outputs depend on historically situated knowledge, symbolic interpretation or culturally differentiated audiences, including some work in games, publishing, audiovisual media, advertising and music. The cases therefore provide a focused empirical lens from which to derive broader design principles for the CCIs.
This article argues that generative AI for the CCIs should therefore be culturalized. Culturalization means more than translation, localization, personalization or generic bias mitigation. It refers instead to the design, deployment and evaluation of AI systems in ways that take culture seriously as dynamic, relational, contested and interpretively plural. This position draws on a longer tradition in localization, cross-cultural design and prior work on culturalization in digital media, where scholars have shown that successful adaptation across contexts requires more than surface-level linguistic transfer; it requires engagement with customs, beliefs, symbolic repertoires and situated meanings [9,16]. Culturalization, in this stronger sense, is not an optional market layer added after technical development. It is a design orientation that recognizes that cultural products do not travel unchanged across contexts and that “universal” outputs often reflect the norms of already dominant groups.
Recent work on large language models reinforces this concern. Studies of cultural bias and cultural alignment show that widely used models tend, by default, to reflect some value patterns and symbolic assumptions more strongly than others rather than behaving as culturally neutral systems [17]. At the same time, survey work on “culture” in LLM research notes that the field still lacks conceptual clarity: many papers operationalize culture through narrow proxies without explicitly defining what culture is or how it should be represented in computational systems [18]. These findings suggest that the problem is not only that current generative systems are insufficiently diverse in their outputs; it is also that the conceptual apparatus for building and evaluating culturally aware systems remains underdeveloped.
Against this background, this article advances a conceptual, diagnostic and normative argument. It draws on previously published work in adaptive cultural mediation, where targeted narratives, personalization, cross-venue cultural connections and route-based multimedia storytelling were used to construct situated interpretive pathways for different publics and contexts [19,20,21,22,23,24]. For the present article, cultural heritage experts familiar with those earlier projects participated in exploratory tests with recent generative AI models, using current prompting practices and analyzing the resulting outputs for factual validity, curatorial adequacy, cultural meaning and audience-sensitive adaptation.
The article makes four contributions:
- It clarifies culturalization for generative AI in the creative and cultural industries.
- It distinguishes culturalization from localization, personalization and fairness.
- It proposes a pipeline framework for culturalized generative AI.
- It identifies a research agenda for future work across AI, social sciences, humanities and cultural-sector practice.
The remainder of the article develops these contributions in stages. Section 2 clarifies what culturalization means, distinguishes it from adjacent notions and explains why the creative and cultural industries are a critical testbed for the problem. Section 3 examines why current production-grade generative AI foundation models tend to fall short in culturally sensitive domains. Section 4 turns to earlier work in adaptive cultural mediation and recent generative AI explorations in order to show what such mediation had already achieved before current generative AI and why present systems still fail to match that level of culturalization in practice. Section 5 proposes an integrated framework for what is currently missing, combining pipeline requirements with design and deployment principles. Section 6 outlines a research agenda for developing, evaluating and governing culturalized generative AI before the conclusion.
2. From Localization to Culturalization
The central claim of this article is that the challenges posed by generative AI in the creative and cultural industries cannot be addressed adequately through localization alone. In many practical discussions, localization is considered the obvious answer to cross-cultural adaptation. But in the context of generative systems and especially in the context of cultural production, that answer is too narrow.
2.1. What “Culturalization” Means
In its narrowest sense, localization usually refers to adapting a product to a target market through changes in language, formatting, interface conventions and selected references. In digital media industries, localization often includes translation, subtitling, interface adaptation and some market-specific editing [25]. Culturalization goes further. As work in cross-cultural design and game studies has argued, successful adaptation requires attention not only to linguistic transfer but also to local customs, values, beliefs, symbolic repertoires and contextual expectations. Pyae’s review of digital game localization explicitly describes culturalization as a layer that addresses the cultural suitability of game context and content for local audiences, beyond translation alone [9]. Sun’s work on cross-cultural technology design similarly argues that technologies must be designed as culture-sensitive artefacts rather than merely exported and superficially adjusted for local users [16].
For the purposes of this article, culturalization can therefore be defined as the design, development, deployment and evaluation of generative AI systems in ways that take culture seriously as a dynamic, relational and interpretive dimension of meaning-making. This definition matters because the kinds of outputs produced in the CCIs—scripts, songs, images, characters, exhibitions, game worlds, promotional materials, voiceovers and narrative structures—are not simply containers of semantic information. They are culturally coded artefacts whose significance depends on context. A line of dialogue may be grammatically correct but socially inappropriate; a visual motif may be aesthetically attractive but historically insensitive; a narrative adaptation may be structurally coherent but symbolically tone-deaf. In this setting, cultural adequacy is not a thread woven in after the fact, but part of the fabric of technical performance. It is an integral component to guarantee CCI content success.
2.2. Culturalization Versus Localization
The distinction between localization and culturalization is especially important because, in practice, the two are often conflated. Localization tends to assume that there is a relatively stable original product that can be transferred into another market by adjusting its language and selected surface features [25]. Culturalization questions that assumption. It starts from the premise that products do not travel unchanged across contexts because meaning is co-produced by audiences, histories, institutions and repertoires of interpretation. In this stronger sense, culturalization is not simply a downstream adaptation phase. It may require rethinking the content itself, the assumptions embedded in it and the way the system generating it encodes what counts as normal, plausible, tasteful or safe.
This distinction becomes sharper in the case of generative AI. Localization is often compatible with a pipeline in which a model first generates an allegedly generic output and the output is then adapted for specific markets. Culturalization, by contrast, implies that the model, the prompting process, the training data, the human review process and the evaluation criteria may all need to be rethought from the outset and continuously. The problem is not merely that today’s mainstream systems sometimes mistranslate or miss culturally specific references. It is that they often generate from latent baselines shaped by uneven training data and dominant norms. Recent survey work on culture in large language model research shows that the field still relies heavily on partial proxies for culture and rarely defines the concept explicitly [18]. In parallel, empirical work on cultural alignment indicates that widely used models tend to default toward particular dominant value patterns rather than culturally plural baselines [17]. These findings support the argument that culturalization cannot be reduced to post hoc adaptation; it must be seen as a design problem at the level of the system itself.
2.3. Culturalization Versus Personalization and Bias Mitigation
Culturalization should also be distinguished from two adjacent frameworks that are often invoked in AI debates: personalization and bias mitigation. Personalization aims to tailor outputs to individual users based on preferences, behavior or inferred traits. In some settings, personalization can support culturally responsive experiences, especially where users can articulate their own preferences or identities [26]. Yet personalization is not equivalent to culturalization. First, culture is not reducible to individual preference profiles. It involves shared histories, symbolic systems, social norms and contested meanings that exceed the individual user. Second, personalization often optimizes for convenience, engagement or predicted relevance [27], whereas culturalization asks whether a representation is contextually appropriate, socially situated and interpretively responsible.
Nor is culturalization exhausted by bias mitigation. Bias mitigation typically aims to reduce unfairness, stereotyping or disparate performance across groups. That is necessary, but it is still insufficient for the kinds of generative tasks addressed here. A system may avoid obviously harmful stereotypes and still fail to produce outputs that are culturally resonant, historically literate or aesthetically grounded in a particular setting. Put differently, bias mitigation is often defensive: it asks how to avoid distortion or harm. Culturalization is also constructive: it asks how systems might support plural, situated and meaningful forms of expression.
2.4. Why Culturalization Matters Especially in Creative and Cultural Domains
The need for culturalization is especially pronounced in creative and cultural domains because cultural artefacts circulate through layers of interpretation that are thicker than those usually involved in transactional or informational systems [28]. A banking app or navigation interface can often be localized successfully while preserving most of its functional logic. Creative artefacts are different. Their value frequently depends on tone, symbolism, narrative voice, emotional texture, humor, genre convention, memory and recognition. What appears universal at the level of plot or image may carry very different connotations in practice.
Generative AI intensifies this issue because it does not merely distribute finished cultural objects; it participates in producing them [29]. When such systems are used to ideate, rewrite, visualize, remix or narrate, they influence not only how content is circulated but what kinds of content are likely to be created in the first place. This is why the move from localization to culturalization is more than a terminological refinement. It marks a shift from thinking of culture as an external market variable to treating it as constitutive of the creative process itself.
The creative and cultural industries are therefore a particularly compelling domain for generative AI. They make explicit something that is often obscured in broader AI debates: the quality of an output cannot be assessed independently of culture, interpretation and context. In cultural production, technical polish and commercial usability are not enough. A generated image, script, soundtrack, subtitle, exhibition narrative or game scenario may still be culturally thin, stereotyped or symbolically misplaced. The CCIs thus expose the limits of evaluation frameworks centered only on fluency, usability or safety [1,30].
More fundamentally, cultural products are rooted in narratives, aesthetic codes, symbols and histories that do not automatically travel well across contexts. Humor, irony, beauty, authenticity, historical sensitivity and emotional resonance cannot be derived from formal correctness alone. The CCIs are therefore a demanding environment for generative AI because what matters is not only what is said or shown, but how, in relation to whom and against what background of meaning.
3. Why Current Generative AI Falls Short
Current generative AI systems can speed up, scale, democratize and facilitate creative processes. Yet widely used foundation models remain weakly equipped for culturally sensitive contexts. This shortfall is not accidental or limited to a few edge cases. It follows from structural characteristics of today’s dominant generative AI architectures: asymmetries in training data, weak representations of situated cultural meaning, inherited stereotyping and value defaults and optimization regimes that privilege general fluency and safety over contextual adequacy.
3.1. Training Data Asymmetries and Dominant Cultural Baselines
A first reason contemporary commercial LLMs fall short is that their training corpora are profoundly uneven. Contemporary foundation models are typically trained on vast internet-scale datasets whose linguistic, geographic and cultural distributions are highly imbalanced. In practice, this means that some languages, genres, aesthetic conventions and value-laden discourses are massively overrepresented, while others appear only sparsely or in mediated form. Models trained under these conditions do not merely learn language; they also absorb patterned assumptions about what is typical, natural, desirable, humorous, prestigious, safe or coherent. The result is that systems often generate from culturally dominant baselines while presenting their outputs as generic or universal. Survey work on multilingual and low-resource LLMs repeatedly notes that model capabilities remain significantly weaker outside high-resource languages and contexts, largely because representative data and evaluation resources are lacking [31]. This asymmetry is likely to extend beyond linguistic performance: when low-resource languages are thinly represented, the cultural concepts, local knowledge and communicative norms carried by those languages are also less available to models, increasing the risk that dominant cultural frames substitute for situated ones [32].
3.2. Multimodal Generation and the Loss of Situated Meaning
A second limitation concerns the nature of multimodal generation itself. In the creative and cultural industries, generative systems increasingly operate across text, image, audio and video. Yet cultural meaning is rarely stored in isolated tokens or objects; it emerges from relations among style, context, history, symbolism, embodiment and audience expectation. A system may generate an image that is visually convincing, a script that is grammatically smooth or a melody that is structurally plausible while still failing to grasp what those outputs signify in a particular cultural situation. General surveys of multimodal large language models emphasize their growing capacity for cross-modal generation and understanding, but they also underline persistent limitations in grounding, reasoning and reliable contextual interpretation across modalities [33].
This is especially problematic in creative domains, where a small symbolic shift can transform the social meaning of an artefact. In museum storytelling, music or audiovisual production, an output may preserve factual or denotative content while losing affect, irony, memory or intertextual resonance. The shortfall is therefore epistemic as well as technical: the model lacks a sufficiently grounded account of what cultural elements mean in practice.
3.3. Stereotypes, Flattening and Symbolic Misreadings
A third reason current production-grade generative AI foundation models fall short is that they inherit and reproduce stereotypes embedded in both data and design processes. The problem is not limited to offensive outputs. More often, stereotyping appears as flattening: cultures are represented through a narrow set of recognizable signs, social roles are mapped onto familiar cliches and ambiguity is reduced to legible categories. Such outputs may appear harmless or even useful, yet they still restrict the range of imaginable representations. Design processes matter here because many everyday decisions about datasets, filters, evaluation criteria, safety policies, interface defaults and acceptable outputs are made in engineering and product environments where abstraction, scalability and deployment constraints may have greater institutional weight than historically informed cultural interpretation, unless social-science, humanities and cultural-sector expertise is given real authority in the process [34].
Recent empirical work reinforces this concern. Research on cultural bias and cultural alignment has found that prominent language models often reproduce a narrow subset of value orientations more reliably than a genuinely plural range of cultural settings [17]. Related work on multilingual models shows that stereotypes can leak across languages, meaning that biases learned in one linguistic context may shape model behavior in another [35]. For creative applications, this is especially consequential because open-ended generation is precisely where subtle stereotypes can be normalized under the guise of plausibility, inspiration or taste. A model helping to brainstorm characters, scenes or narratives may not produce overtly hateful content, yet it can still repeatedly associate prestige with some bodies, voices, geographies or lifestyles and marginality with others.
3.4. Broad Acceptability and Cultural Flattening
A fourth limitation lies in how contemporary generative systems are optimized. Commercial models are commonly tuned to be broadly helpful, safe and acceptable across very large user bases. This becomes problematic in cultural production because many artistic practices depend on confrontation, contradiction, provocation, irony and situated forms of critique. A model that has been heavily optimized to avoid controversy may generate outputs that are smooth, cautious and legible, but also aesthetically flattened and culturally timid. In this sense, safety alignment can drift toward genericness, often mistaken for universality. A “safe” output may avoid explicit offense while erasing the tensions, histories or symbols that make a work meaningful in context.
3.5. Why the Shortfall Is Structural Rather Than Incidental
These issues suggest that the limitations of today’s dominant generative AI pipelines in culturally sensitive creative contexts are structural. They stem from deeper assumptions in present-day AI pipelines: that scale can substitute for situated knowledge, that culture can be proxied crudely without major loss, that multilingual coverage is equivalent to cultural competence and that general-purpose helpfulness can serve as a universal standard of quality. Recent survey literature supports this diagnosis by showing that the field still lacks shared conceptual and methodological foundations for representing and evaluating culture in generative systems [18]. Treating culturalization as a first-class design requirement would reorganize data selection, model adaptation, prompting, interface design, human review and evaluation around culturally situated creativity.
4. Adaptive Cultural Mediation as a Blueprint to Culturalized Generative AI
An instructive precursor to the argument developed in the previous sections can be found in earlier work on adaptive cultural mediation across museums, heritage sites, city-scale cultural settings and pilgrimage routes [19,20,21,22,23,24]. The applications differed, but the underlying move was similar: cultural objects, places and stories were organized into meaningful pathways for different publics and situations. Some systems reorganized museum collections around reflective topics, visitor profiles and temporal cues; others connected local heritage with broader historical themes, other sites or route-based narratives unfolding over several days.
In all these cases, the aim was not simply to deliver more information, but to construct situated pathways through cultural material. Objects, paintings, sites and stories were not treated as self-explanatory units whose value could be reduced to factual correctness. Their significance depended on narrative framing, thematic emphasis, comparison, symbolic resonance and audience position. Human experts selected the materials, defined the relations among them, shaped the interpretive lenses and evaluated whether the resulting experiences provoked curiosity, empathy, comparison and reflection rather than mere consumption of information [19,21,36]. In that sense, the earlier work already embodied several principles that this article groups under culturalization: curated cultural data, situated interpretation, profile-sensitive adaptation, expert mediation and evaluation beyond usability alone.
Once generative AI became widely available, it was reasonable to expect it to amplify this earlier approach. The promise was not to replace curators, historians, educators or writers, but to multiply their work: to produce many more targeted itineraries, narrative framings, reflective prompts and audience-sensitive explanations at lower cost and with less manual authoring effort, while remaining grounded in curatorial intent, collection-specific knowledge and situated interpretation. This expectation motivated a recent comparison across different AI models, including ChatGPT, Perplexity, Copilot, Gemini, Deepseek and Claude [29]. Based on those comparisons and on the popularity of AI tools measured by StatCounter1, ChatGPT was selected for the exploratory studies presented in the following subsections.
The cases should be read as diagnostic case studies, not as attempts to establish statistical generality. Their purpose is to make visible, through concrete materials and controlled exploratory tests, the kinds of cultural scaffolding that curated mediation can preserve and that generic generation can easily miss. They are useful here because museums, heritage interpretation and pilgrimage routes make several dimensions of culturalization simultaneously observable: provenance, interpretive scaffolding, symbolic meaning, audience adaptation, expert mediation, spatial and temporal context and plural readings. To examine whether current generative AI can genuinely extend adaptive cultural mediation, this article reports three complementary exploratory studies. The first study establishes a generic-generation baseline using museum narratives. The second investigates whether disciplined prompting can preserve cultural fidelity while enabling personalization. The third, finally, asks whether cultural mediation remains adequate when adaptation must also account for location, sequence, motivation and time, as in pilgrimage routes. The studies, therefore, progressively increase the degree of control over cultural generation to support a nuanced assessment of what is currently missing in mainstream generative AI pipelines.
4.1. Initial Exploratory Comparison with Generic Generation
The source narratives previously designed for the Archaeological Museum of Tripolis offer a useful baseline. In [19,20], the museum was not treated as an isolated local collection, but part of a wider network of European cultural venues whose objects could be connected to other collections, artworks and educational materials. This cross-venue logic made the Tripolis narratives especially relevant for culturalization, because the interpretive task was to preserve local specificity while building meaningful links to broader European cultural repertoires.
The sample shown in Table 1 illustrates the complexity of curated mediation, as it built a layered pathway around a funerary stele with the inscription “ARKADIA CHAIRE”, in connection with a well-known painting by Titian, Diana and Callisto, kept at the National Gallery London. It first identified the tombstone historically as a first- or second-century inscription for a deceased woman; then moved through the myth of Callisto, Lycaon, Zeus, Artemis, Hera and Arcas; then linked the myth to the Greek terms for bear and to the constellations Ursa Major and Ursa Minor; then used Titian’s painting as an art-historical bridge to the National Gallery; and finally connected the episode to educational material for contemporary students and to a reflective question about the relevance of myths and art today. The accompanying analysis explicitly marked the layers involved: history, mythology, history of art, education and psychology, as well as the fact that the narrative had passed institutional review. In other words, the curated text was rich not because it was more ornate, but because it held several validated interpretive registers together in one visitor-facing explanation.
A small set of exploratory prompts on material from the Archaeological Museum of Tripolis then illustrated the gap between such curated mediation and generic generation. In one prompt, ChatGPT was asked to produce a short visitor narrative connecting the funerary inscription “” with Titian’s Diana and Callisto. The generated entry in Table 1 produced a fluent and emotionally coherent bridge between farewell, Arcadian myth, female loss and memory. As a first draft, it was not useless. The output sounded curatorial, yet its connection remained thin and it lacked a full sequence of mediating moves, which made the original narrative pedagogically and culturally meaningful.
A second prompt asked for a short route through several exhibits kept at the Archaeological Museum of Tripolis focused on women in antiquity, including Gortsouli figurines, a headless Athena statue, a Roman portrait of a girl from Mantinea, a tondo of Hercules and Auge and the “” tombstone. The model identified a plausible overall theme: women as sacred figures, civic ideals, daughters, mythic heroines and remembered individuals. At the same time, it introduced unsupported claims about the Gortsouli figurines, apparently confusing the requested material with clay figurines from a different context. This is a typical culturalization failure rather than a simple stylistic weakness. The narrative thread was attractive, but the grounding of the thread in the actual collection was unstable.
A third prompt asked for three short narratives about the tombstone for (i) a family with children, (ii) a visitor interested in gender history and (iii) a visitor with limited prior knowledge of antiquity. Here the model did show some profile-sensitive adaptation. The family version used simple language and a contemporary analogy with photographs and memory; the gender-history version foregrounded women’s social visibility and remembrance; the introductory version explained the function of the tombstone in accessible terms. However, the adaptation mostly changed tone and emphasis around a stable object description. It did not demonstrate a deeper reconfiguration of the interpretive pathway, nor did it show how such visitor profiles should affect object selection, sequencing, comparison or reflective prompts. The pilgrimage examples of Section 4.3 show the same pattern in a longitudinal setting.
Finally, when asked to find connections between Tripolis exhibits and works at the National Gallery, the model heavily reused the earlier Titian framing. It proposed several apparently coherent pairings, but the search space had narrowed around the immediately preceding conversation. This illustrates another practical limitation of generic chat-based cultural generation: conversational memory can create a false sense of situatedness. The system appears to build on context, but it may instead become anchored to an earlier association and under-explore alternative curatorial possibilities.
These examples should be read cautiously. They are not a benchmark and they do not establish population-level claims about all models or all museum tasks. Their value is diagnostic. They show the difference between a fluent cultural draft and a culturally disciplined interpretive layer. In each case, the model could produce usable language, but the human expert still had to determine whether the output was grounded, whether its symbolic associations were warranted, whether its profile adaptation was meaningful and whether the generated route preserved the intended interpretive frame.
4.2. Structured Prompting for Controlled Cultural Adaptation
A related pilot study provides a more structured version of the same diagnosis in a different museum context, through two experiments that used a matched factorial design involving:
- Two Egyptian heritage artefacts: the Golden Burial Mask of Tutankhamun, kept at the Grand Egyptian Museum, and the Golden Armchair of Tutankhamun, kept at the Museum of Egyptian Antiquities.
- Two visitor profiles: an 11-year-old primary-school visitor interested in sports and games; and a 34-year-old visitor with a PhD in technology and an interest in fashion.
- Two conditions: provide a baseline general-audience explanation, or a personalized visitor-specific explanation.
- Two independent runs per condition, which produced 16 evaluated outputs per experiment.
Experiment 1 used minimally-constrained prompts that asked the model to personalize explanations while using only the supplied artefact data. Experiment 2 used a constraint-aware prompt architecture that added explicit safeguards: source-only rules, limits on unsupported historical, religious, symbolic, emotional or functional interpretation, requirements to preserve cultural meaning, boundaries around sports, game, fashion and design analogies and anti-stereotype instructions.
The evaluation deliberately separated two questions that are often conflated: whether the response was well adapted to the visitor, and whether it remained accurate, complete and faithful to the supplied cultural-heritage data. It was carried out in two passes: a content-integrity pass assessing factual error, insight preservation, symbolism control and length compliance, and an adaptation-quality pass assessing personalization quality and, in Experiment 2, tone appropriateness, narrative framing and stereotype risk. This design is important because a fluent and well-tailored explanation can otherwise mask a loss of curatorial fidelity.
Table 2 and Table 3 show two outputs, out of the 16 produced in each experiment, that turned out to contain factual errors. The first comes from Experiment 1 and illustrates an invented functional claim, namely that the mask served a “protective” role. The second comes from Experiment 2 and shows that even the constraint-aware prompt could still allow minor unsupported interpretive language, in this case the phrase “court identity”.
In Experiment 1, all personalized responses received the maximum personalization score, but personalization did not guarantee cultural or curatorial fidelity. Across the 16 evaluated runs, factual issues appeared in 11 outputs, insight issues in 4 and symbolism-control issues in 6. More importantly, the error profile differed by visitor type. Child-oriented explanations tended toward compression loss: sports and game analogies made the responses accessible, but simplification sometimes removed key curatorial insights, such as the role of symbolic materials in reinforcing divinity or the Armchair’s early-reign context and Amarna influence. Adult-oriented explanations tended toward elaborative drift: the fashion and design framing made the responses rhetorically sophisticated, but encouraged unsupported claims about status, power, transformation, innovation, legitimacy or symbolic meaning. The result is not simply that personalization increases hallucination. Rather, personalization redistributes risk: child adaptation can omit, while adult adaptation can over-interpret.
In Experiment 2, the constraint-aware prompt architecture preserved personalization while reducing several forms of drift. Across 16 evaluated runs, factual issues fell from 11 to 3, insight issues from 4 to 1 and symbolism-control issues from 6 to 1; length compliance rose from 13 to all 16 outputs. Personalization quality remained at the maximum score for all personalized responses, tone appropriateness averaged 2/2, narrative framing averaged 1.88/2 and no stereotype risk was detected in the evaluated sample. The most useful mechanism was not the removal of analogy, but its epistemic containment: responses marked sports, game, fashion or design comparisons as modern explanatory analogies rather than as historical facts or ancient intentions. In other words, analogy could remain pedagogically useful without becoming an unauthorized interpretation of the artefact.
It cannot be asserted that systematic prompting solves culturalization. The sample is small, the constraint-aware prompt architecture bundles several safeguards, the outputs are model- and time-specific and the evaluation was expert-coded rather than based on independent raters or visitor outcomes. The more defensible conclusion is that culturalized generation requires a separation between a stable cultural core and an adaptive communicative layer. The former preserves approved facts, provenance, interpretive boundaries, symbolism and curatorial intent; the latter varies tone, analogy, narrative address and explanatory route for different publics. Cultural adequacy is not equivalent to either factual correctness or personalization quality, but depends on controlling how adaptation interacts with source-defined meaning.
4.3. Extending cultural mediation to longitudinal experiences
Prior route-based work on multimedia narratives for pilgrims provides a complementary example because it moves cultural mediation outside the bounded space of the museum [22]. Pilgrimage is a slow, mobile and longitudinal cultural experience. Interpretation is not delivered in one visit or around one object. It unfolds over successive days, in changing physical conditions, across landscapes, villages, accommodation stops, detours, pauses and encounters with other walkers.
A further experiment used curated narratives created to accompany pilgrims along the Way of St. James in the Galician province of Ourense [23]. These were not generic Camino descriptions. They were written through situated literary voices, with different stretches associated with different writers or historical figures, including Fefa Vila, Miguel de Cervantes, Hydatius of Limia, Inês de Castro, Vicente Risco, Otero Pedrayo, Graham Greene, Gildeberta of Flanders and Rosalía de Castro. This device did more than add style. It allowed local heritage to be interpreted from different temporal, literary, gendered and cultural standpoints, connecting concrete places on the route with archival traces, local history, rural memory, literary imagination, religious practice or contemporary reflection.
The curated narratives were organized as sequences of geolocated units: each entry combined coordinates, a route location, narrative text, media resources and source references. They became seeds to test profile-sensitive generation for four targets: a young international pilgrim, an older Spanish pilgrim with strong cultural background, a religious pilgrim and a secular cultural tourist. Table 4 and Table 5 show two examples of narrative texts produced by AI from the original material.
The outputs were generally fluent and profile-sensitive, and the factual accuracy remained as high as with disciplined prompting in the Egyptian pilot of Section 4.2. However, a small exploratory feedback exercise with eight participants (two representatives of each of the aforementioned profiles), provided evidence that sharpens a point about interpretive plurality:
- Younger international participants responded better when the narrative opened the literary cues toward discovery and personal journeys, whereas culturally knowledgeable older participants valued more specific point, like the one in Table 4 that Tirso de Molina’s positive representation of Galicia was unusual within much Golden Age Castilian literature.
- Prior knowledge affected tolerance for simplification. Participants with less cultural background appreciated familiar entry points and clearer scaffolding, while more expert participants were less satisfied by versions that merely made the text accessible. For them, cultural adequacy required a sharper literary or historical payoff.
- Motivation changed what counted as resonance. Religious pilgrims found value in versions that connected death, hope, beauty and reflection, while secular cultural tourists responded more strongly to concrete social or historical interpretation, such as the luctuosa as evidence of rural obligations, inheritance and local authority in Table 5.
- Fatigue, available time and walking situation mattered. Participants did not treat shorter versions as acceptable simply because they were shorter. They valued them when the reduced text preserved the core interpretive point and helped decide whether to pause, continue walking or save attention for a later stop.
- Cultural provenance made cross-cultural bridging especially delicate. Participants appreciated connections to their own cultural repertoires only when those connections were grounded in the supplied material. Ungrounded parallels were perceived as artificial, trivial or wrong, confirming that adaptation goes beyond guessing national analogies.
This experiment reveals that pilgrimage routes pose a demanding test case for culturalization, as it cannot be equated with generating more variants. A generic model could easily produce a plausible story about the Way, a church, a rural legend or a famous Galician writer. The harder task is to know which story belongs to each place, why that narrative voice fits that stretch, which local facts and sources should be trusted, when the story should be offered, how long it should be, and how it should connect to what the pilgrim has already heard on previous days. A culturalized system would therefore need to preserve a stable, curated account of place-based heritage while adapting its presentation to route segment, location, timing, language, fatigue, available time and visitor motivation.
5. Toward a Framework for Culturalized Generative AI
The three studies presented in the preceding section point in the same direction. Popular generative AI foundation models can produce fluent and profile-sensitive outputs, and disciplined prompting can reduce some forms of factual, symbolic and interpretive drift. However, they do not by themselves fully realize the expected role that motivated the exploratory work. Reliable culturalization still depends on conditions outside generic generation: curated source material, explicit interpretive constraints, profile-aware but bounded adaptation, situated delivery across objects or routes, expert mediation and evaluation criteria that treat cultural respect as epistemic fidelity rather than mere output fluency. That diagnosis motivates the pipeline framework developed in this section.
A culturalized approach cannot be reduced to outputs alone, nor can it be treated only as a layer of prompting, retrieval or human review around an otherwise generic model. It must also inform the internal mechanics and design choices through which LLMs tokenize languages, encode representations, retrieve and rank knowledge, learn cultural associations and align outputs with preferred norms. Culturalization therefore has to be pursued across the full pipeline, from data collection and model adaptation to interfaces, human mediation and evaluation.
Cultural adaptation can take at least two forms:
- The first is cross-context robustness: creating a product that can circulate across settings without alienation, distortion or obvious offense.
- The second is plural generation: producing culturally tailored versions of a concept that resonate with specific audiences while preserving coherence and avoiding caricature.
The second mode is especially important for generative AI because rapid variation without cultural grounding can simply produce shallow permutations of the same dominant baseline. Culture is therefore not a downstream adjustment but a cross-cutting design concern that should inform every stage of the pipeline—from data curation and knowledge grounding to generation, interfaces, human mediation, and evaluation. Figure 1 summarizes this view, consistent with work showing that cultural competence in LLMs depends not only on scale, but also on cultural resources, benchmarks and deployment conditions [37,38].
The pipeline in Figure 1 can be read as six mutually dependent layers:
- Training data. Any framework for culturalized generative AI has to begin with data. Contemporary models inherit many limitations from training data in which cultural worlds are unevenly and hierarchically represented. A culturalized approach therefore requires attention not only to language coverage, but to who is represented, how, under what conditions and with what degree of internal plurality. Data curation should seek underrepresented cultural materials, preserve context wherever possible and avoid treating culture as a fixed national label. Recent benchmark work supports this point by showing both that LLM performance varies substantially across cultures and languages and that progress depends on more diverse, human-verified cultural resources rather than scale alone [39,40].
- Knowledge grounding. Today’s dominant generative AI architectures tend to absorb culture implicitly through statistical regularities in data, which can support surface fluency but not deeper cultural adequacy. Culturalization therefore requires retrieval, grounding and knowledge representation practices that preserve provenance, historical context, symbolic repertoires, regional conventions and contested meanings. The point is not to make culture fully formalizable, but to structure systems so that they can better recognize when outputs depend on situated cultural knowledge.
- Generation. The generation stage is where grounded materials, model behavior and communicative goals are composed into concrete outputs. Culturalization at this layer concerns the selection of narrative frame, register, examples, analogies, omissions and degrees of uncertainty. It should support more than one culturally credible rendering of a prompt, while avoiding both generic globalized prose and caricatured adaptation.
- Interface. Even when the underlying model is fixed, interfaces and prompting practices shape what cultural outputs are likely to be produced. Many failures attributed to the model are mediated by generic prompting environments that reward speed, convenience and broad legibility rather than contextual sensitivity. Interfaces should make cultural assumptions more visible, invite users to specify audience, historical frame, symbolic references or interpretive stance and provide ways of comparing alternative culturally inflected outputs rather than offering a single default completion.
- Human mediation. Cultural adequacy is interpretive and socially situated, so human-in-the-loop design is not merely a quality-control stage. Cultural practitioners, translators, curators, historians, designers and community stakeholders are needed not only to spot errors, but to continuously shape what counts as adequacy in the first place. Participatory approaches in cultural heritage research point in a similar direction: evaluation and design are more robust when they include affected communities and domain experts rather than relying solely on technical metrics [41].
- Evaluation. Systems must be assessed not only for fluency, relevance and safety, but for cultural adequacy: contextual sensitivity, symbolic meaning, historical awareness and support for plural interpretation. Alongside conventional AI metrics, evaluation should consider cultural adequacy, cultural resonance, plurality and reflexivity: whether outputs respect relevant contexts, connect with intended audiences, sustain more than one culturally credible rendering and surface uncertainty or sensitivity rather than presenting culturally loaded outputs as neutral. Recent benchmark research is beginning to make cultural evaluation more systematic, but it also confirms that current methods remain narrow and incomplete [40].
These layers make culturalization a pipeline orientation rather than a single technical feature. Training data, knowledge grounding, generation, interface design, human mediation and evaluation all need to be organized around cultural context, plurality and accountable interpretation. The framework is not heritage-specific: the same pipeline questions arise when systems support culturally grounded world-building in games, historically situated narrative generation in film or publishing, local symbolic references in advertising or culturally meaningful adaptation in music and audiovisual translation. This pipeline orientation also has a sustainability dimension. One reason generative AI is attractive in cultural settings is that it could potentially exploit and recombine reusable narrative structures, semantic standards and durable cultural associations at much greater scale than earlier handcrafted systems allowed. But that potential depends on keeping such structures explicit enough to support revision, comparison, institutional reuse and accountable adaptation across venues and audiences [20,21].
If culturalization is to function as more than a descriptive label, it also needs design and deployment principles. These concern how systems should be built, who should shape them, what values they should protect and how trade-offs should be handled. Existing AI governance frameworks already emphasize diversity, human oversight, accountability, transparency and context-sensitive risk management [42,43,44]. For culturalized generative AI, the main principles are:
- Culture as a dynamic, relational and contested setting. Culture should not be modeled as fixed, homogeneous or reducible to stable group traits. A system that merely tags users or outputs with broad cultural labels may appear sensitive while still reproducing essentialism. Culturalization should instead begin from the recognition that meanings are relational, historically situated and often internally contested, in line with broader commitments to diversity and inclusiveness in AI governance [42].
- Human-centered co-creation rather than replacement. In culturally sensitive domains, human involvement is constitutive of quality: curators, artists, translators, community representatives, educators and domain experts are needed to determine resonance, appropriateness and legitimacy. Culturalized systems should preserve human creative agency and attribution, especially when generative tools shape ideation, drafting, remixing or adaptation at scale and legal ownership questions remain unsettled [45]. For a culturalized approach, the goal is not human in the loop in the narrow procedural sense, but human-led interpretation around the loop [42,43].
- Transparency and accountability in cultural adaptation. When systems adapt content across audiences, they should communicate relevant cultural assumptions, allow scrutiny of sensitive adaptations and create lines of responsibility when representational harms occur [42,44]. This includes documenting why a particular cultural framing, source, omission, analogy or level of simplification was used, because opacity in audience adaptation can normalize dominant symbolic frameworks while making representational decisions hard to contest.
- Diversity-preserving design (when relevant) rather than market-driven flattening. Generative AI can intensify the tendency toward dominant frameworks because systems often reward what is familiar, statistically central and easy to optimize across large markets. Culturalized design should therefore be able to sustain diversity in datasets, outputs, interface options and evaluation criteria [6]. This is not only a matter of avoiding offensive content: in creative and heritage domains, repeated flattening can affect the symbolic environment in which communities recognize themselves and others, so cultural rights and public value depend on preserving multiple histories, languages, styles and interpretive communities [46].
- Participation by creators, communities and domain experts. Communities affected by AI-mediated cultural production should have opportunities to shape the criteria by which these systems are designed and evaluated, especially where memory, representation, heritage and legitimacy are at stake [42,43]. Participation is also a way to address cultural ownership: system designers should ask not only whether cultural materials, styles, archives or community traditions can be used, but under what conditions their use is legitimate, reciprocal and properly recognized [6,45].
- Proportionality and context-sensitive deployment. A lightweight promotional copy tool, an internal brainstorming assistant, a heritage interpretation system and a public-facing narrative generator do not require identical thresholds of review. The level of scrutiny should depend on stakes, sensitivity and likely consequences of misrepresentation [43,44]. Deployment choices should also consider labor and institutional effects: generative AI can complement human expertise and improve work quality, but cost-cutting adoption can reduce autonomy, increase monitoring or shift value away from cultural workers [47,48].
- Freedom of expression, critique and cultural friction. Culturalized design should not eliminate friction, conflict or critique. Creative work often depends on tension, provocation, irony and disagreement. The design aim is to distinguish meaningful cultural friction from careless representational harm, supporting creativity without defaulting either to reckless insensitivity or sterile neutrality [42].
6. A Research Agenda
If culturalization is to become more than a normative aspiration, it has to be translated into a research program: concepts, resources, methods and institutional practices for studying and building culturally aware generative systems [6,18]. The main priorities are:
- Clarifying the concept of culture in AI research. Research on culture in LLMs often relies on partial proxies such as nationality, language, values or demographic markers [18]. Future work should specify whether it is studying language, symbolic repertoire, historical memory, aesthetic convention, value orientation, community practice, audience interpretation or some combination of these.
- Building culturally diverse and context-rich resources. Culturalized generative AI requires better resources for underrepresented languages, contexts, genres and symbolic traditions, not just larger generic corpora [6,31]. These resources should preserve provenance, contextual metadata and interpretive annotations, and should include multimodal cultural artefacts such as scripts, exhibition narratives, musical descriptions, subtitling variants and visual motifs.
- Developing stronger benchmarks and evaluation protocols. Benchmark work such as CDEval makes visible that models do not behave as culturally neutral systems [40], but creative and cultural production also requires assessment of tone, symbolism, resonance and interpretive plurality. Building on the exploratory studies reported here, future work should implement larger-scale evaluation protocols in which multiple generative AI systems are assessed across curated cultural scenarios. Such protocols should not ask only which model gives the best answer. They should also examine whether retrieval is grounded in culturally appropriate sources, whether uncertainty and contested interpretations are preserved, whether interfaces expose cultural assumptions and whether human review improves cultural adequacy.
- Advancing interdisciplinary methods. Computational work is needed to build datasets, adapt models and measure variation, but ethnographic, qualitative and design-based methods are also needed to understand how creators, curators, translators and audiences actually interact with generative systems [49].
- Participatory validation with creators, communities and audiences. Because cultural adequacy is interpretive, evaluation should combine complementary perspectives rather than collapse them into a single score. Domain experts such as curators, historians and translators can assess historical accuracy, symbolism, provenance and interpretive discipline. Creative practitioners such as writers, artists, designers and localizers can assess expressive quality, genre logic and cultural authenticity. Intended audiences can assess resonance, accessibility, engagement and perceived appropriateness, while community representatives can identify harms or reductions that might be invisible to model developers.
- Studying concrete sectoral use cases. Future work should validate the framework beyond heritage through comparative studies in games, publishing, audiovisual production, advertising, music and other CCI settings. Cultural adequacy in a museum explanation is not identical to cultural adequacy in a game world, song adaptation or advertisement, but each may involve grounding, symbolic fidelity, plurality, audience resonance and accountable human mediation. Sectoral work can keep culture from becoming too abstract to operationalize.
- Connecting technical research to governance and public policy. Technical work should be linked to governance questions from the outset: which applications deserve heightened scrutiny, what documentation should accompany culturally adaptive systems and what forms of auditing are appropriate when outputs vary across audiences. Future work should study lifecycle mechanisms for accountability, including source documentation, records of cultural assumptions, audit trails for adaptations, review responsibilities and institutional policies for deployment in museums, archives, media, education and other CCI settings [42,43]. It should also examine how cultural rights, public value and democratic participation can be protected when AI systems influence what is produced, translated, recommended, archived or narrated [46].
These priorities suggest three commitments for future work: conceptual clarity about culture, socio-technical attention to the full generative pipeline and sustained engagement with creative and cultural sectors as sites where symbolic meaning and public value are central. A benchmark for culturalization should therefore evaluate the pipeline represented in Figure 1, not only final outputs. It should test how data selection, knowledge grounding, generation, interface design, human mediation and evaluation interact across specific cultural scenarios, and how documentation, review and accountability mechanisms shape those interactions. This would provide a more rigorous empirical basis for validating the framework proposed in this article, comparing alternative technical approaches and eventually translating cultural adequacy into measurable, revisable and accountable research practices without reducing culture to a conventional NLP target or treating governance as separate from system design.
7. Conclusion
Generative AI is already reshaping the creative and cultural industries, but its growing presence in these domains has revealed a conceptual and practical gap. Much of the current debate focuses on productivity, innovation, copyright, labor disruption and bias. These issues matter, but they understate a basic point: creative and cultural production depends on meanings that are historically situated, symbolically dense and interpretively plural. Fluency, efficiency and broad safety are therefore insufficient measures of system quality.
The concept of culturalization helps name what existing categories do not fully capture. Localization is too narrow when the challenge is adaptation of meaning across contexts. Personalization is too individualizing when the relevant issues concern shared histories and collective interpretation. Bias mitigation is necessary but insufficient when the goal is also to support richer, more plural and more context-sensitive expression. The examples discussed here show the difference between culturally plausible output and culturally disciplined mediation: current production-grade generative AI foundation models can generate fluent drafts, but they still struggle to preserve the interpretive scaffolding, source grounding, situated timing and human judgment that adaptive cultural mediation requires.
From that diagnosis, the article has proposed a framework spanning training data, knowledge grounding, generation, interface design, human mediation and evaluation. The framework does not assume that culture can be fully encoded, that AI systems can perfectly represent cultures or that they can produce culturally authentic outputs in any final sense. Nor does every application require the same depth of culturalization: the relevant level will vary by domain, audience and stakes [6]. The framework argues instead that the complexity of culture is a reason for explicit design and governance, not a reason to ignore cultural meaning behind generic fluency or universalizing assumptions.
The evidence developed here comes primarily from museums, cultural heritage and pilgrimage because these domains make cultural mediation unusually explicit. They are demanding testbeds, not the limits of the argument. The framework is proposed for the broader CCIs, while its application in games, publishing, audiovisual media, advertising and music requires further sector-specific validation.
The research agenda that follows from this framework is therefore empirical, institutional and conceptual. It requires clearer definitions of culture in AI research, richer context-preserving resources, interdisciplinary methods, governance mechanisms and evaluation protocols that test the full pipeline rather than isolated model outputs. Such work should involve domain experts, creative practitioners, communities and intended audiences, so that cultural adequacy can be assessed as a matter of accuracy, expressive quality, resonance, legitimacy and public accountability rather than as a conventional benchmark score alone.
The broader implication is that culturalized generative AI is not merely a way to refine products. It is at once a technical challenge, an institutional challenge, a governance challenge and a cultural challenge. It asks whose meanings are privileged, whose forms of expression are normalized and whose participation becomes easier or harder when AI enters cultural production. To take culturalization seriously is to treat generative AI as a pipeline-wide design orientation accountable to the plurality of worlds in which it now operates.
Author Contributions
Conceptualization, M.L.N., A.A. and I.L.; methodology, M.L.N., A.A., M.M.K., A.D. and I.L.; investigation, M.L.N., A.A., M.M.K., A.D. and I.L.; formal analysis, M.L.N., A.A., M.M.K., A.D. and I.L.; writing—original draft preparation, M.L.N.; writing—review and editing, M.L.N., A.A., M.M.K., A.D. and I.L.; supervision, M.L.N., A.A. and I.L. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
No new datasets were created or analyzed in this study.
Acknowledgments
The authors thank the cultural heritage experts and participants who contributed to the exploratory activities discussed in this manuscript.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Amankwah-Amoah, J.; Khan, Z.; Wood, G.; Knight, G. The impending disruption of creative industries by generative AI: Opportunities, challenges, and research agenda. Technol. Forecast. Soc. Change 2024, 200, 123379. [Google Scholar] [CrossRef]
- Erickson, K. AI and work in the creative industries: digital continuity or discontinuity? Creat. Ind. J. 2024, 1–21. [Google Scholar] [CrossRef]
- Wahid, R.; Mero, J.; Ritala, P. Technology-enabled democratization: Impact of generative AI on content marketing agencies. Ind. Mark. Manag. 2025, 131, 1–16. [Google Scholar] [CrossRef]
- Parra Pennefather, P. AI and the Future of Creative Work. In Creative Prototyping with Generative AI: Augmenting Creative Workflows with Generative AI; Apress: Berkeley, CA, 2023; pp. 387–410. [Google Scholar]
- Tiribelli, S.; Pansoni, S.; Frontoni, E.; Giovanola, B. Ethics of artificial intelligence for cultural heritage: Opportunities and challenges. IEEE Trans. Technol. Soc. 2024, 5, 293–305. [Google Scholar] [CrossRef]
- UNESCO. Artificial Intelligence and Culture: Report of the Independent Expert Group on Artificial Intelligence and Culture (CULTAI); Technical report; UNESCO, 2025. [Google Scholar]
- Guillot, M.N. The pragmatics of audiovisual translation: Voices from within in film subtitling. J. Pragmat. 2020, 170, 317–330. [Google Scholar] [CrossRef]
- Villegas-Simón, I.; Soto-Sanfiel, M.T. Similarities in the adaptation of scripted television formats: The global and the local in transnational television culture. Poetics 2021, 87, 101549. [Google Scholar] [CrossRef]
- Pyae, A. Understanding the role of culture and cultural attributes in digital game localization. Entertain. Comput. 2018, 26, 105–115. [Google Scholar] [CrossRef]
- Bosman, F.G. The Sacred and the Digital: Critical Depictions of Religions in Video Games. Religions 2019, 10, 130. [Google Scholar] [CrossRef]
- Franzon, J. Choices in song translation: Singability in print, subtitles and sung performance. The Translator 2008, 14, 373–399. [Google Scholar] [CrossRef]
- Bosseaux, C. The Translation of Song. In The Oxford Handbook of Translation Studies; Malmkjaer, K., Windle, K., Eds.; Oxford University Press, 2012; pp. 183–197. [Google Scholar] [CrossRef]
- Podara, A.; Giomelakis, D.; Nicolaou, C.; Kotsakis, R.; Botsou, K.; Veglis, A. Digital Storytelling in Cultural Heritage: Audience Engagement in the Interactive Documentary New Life. Sustainability 2021, 13, 1193. [Google Scholar] [CrossRef]
- Doshi, A.R.; Hauser, O.P. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 2024, 10, eadn5290. [Google Scholar] [CrossRef]
- Moon, K.; Green, A.; Kushlev, K. Homogenizing effect of large language models (LLMs) on creative diversity: An empirical comparison of human and ChatGPT writing. Comput. Hum. Behav. Artif. Hum. 2025, 6, 100207. [Google Scholar] [CrossRef]
- Sun, H. Cross-Cultural Technology Design: Creating Culture-Sensitive Technology for Local Users; Oxford University Press, 2012. [Google Scholar] [CrossRef]
- Tao, Y.; Viberg, O.; Baker, R.S.; Kizilcec, R.F. Cultural bias and cultural alignment of large language models. PNAS Nexus 2024, 3, pgae346. [Google Scholar] [CrossRef]
- Adilazuarda, M.F.; Mukherjee, S.; Lavania, P.; Singh, S.; Aji, A.F.; O’Neill, J.; Modi, A.; Choudhury, M. Towards Measuring and Modeling “Culture” in LLMs: A Survey. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024; pp. 15763–15784. [Google Scholar] [CrossRef]
- Antoniou, A.; Reboreda Morillo, S.; Lepouras, G.; Diakoumakos, J.; Vassilakis, C.; Lopez Nores, M.; Jones, C.E. Bringing a peripheral, traditional venue to the digital era with targeted narratives. Digit. Appl. Archaeol. Cult. Herit. 2019, 14, e00111. [Google Scholar] [CrossRef]
- Lopez-Nores, M.; Bravo-Quezada, O.G.; Bassani, M.; Antoniou, A.; Lykourentzou, I.; Jones, C.E.; Kontiza, K.; González-Soutelo, S.; Reboreda-Morillo, S.; Naudet, Y.; et al. Technology-Powered Strategies to Rethink the Pedagogy of History and Cultural Heritage through Symmetries and Narratives. Symmetry 2019, 11, 367. [Google Scholar] [CrossRef]
- Kontiza, K.; Antoniou, A.; Daif, A.; Reboreda-Morillo, S.; Bassani, M.; González-Soutelo, S.; Lykourentzou, I.; Jones, C.E.; Padfield, J.; Lopez-Nores, M. On How Technology-Powered Storytelling Can Contribute to Cultural Heritage Sustainability across Multiple Venues—Evidence from the CrossCult H2020 Project. Sustainability 2020, 12, 1666. [Google Scholar] [CrossRef]
- López Salas, E. A collection of narrative practices on cultural heritage with innovative technologies and creative strategies. Open Res. Eur. 2021, 1, 130. [Google Scholar] [CrossRef]
- López-Nores, M.; Pazos-Arias, J.J.; Reboreda-Morillo, S.; Penín-Romero, Ó. The Horizon 2020 project rurAllure. Studying pilgrimage as slow tourism, territorial development, social cohesion. Engramma 2023, 204. [Google Scholar] [CrossRef]
- Dahroug, A.; Vlachidis, A.; Liapis, A.; Bikakis, A.; López-Nores, M.; Sacco, O.; Pazos-Arias, J.J. Using Dates as Contextual Information for Personalized Cultural Heritage Experiences; Preprint submitted to; Elsevier, 10 October 2019. [Google Scholar]
- Jiménez-Crespo, M.A. Localization in Translation; Routledge, 2024. [Google Scholar]
- Ober, T.M.; Lehman, B.A.; Gooch, R.; Oluwalana, O.; Solyst, J.; Phelps, G.; Hamilton, L.S. Technical Report 2023(1); Culturally responsive personalized learning: Recommendations for a working definition and framework. ETS Research Report Series, 2023.
- Lemon, K.N.; Verhoef, P.C. Understanding customer experience throughout the customer journey. J. Mark. 2016, 80, 69–96. [Google Scholar] [CrossRef]
- Geertz, C. The Interpretation of Cultures; Basic Books, 2017. [Google Scholar]
- Antoniou, A.; Theodoropoulos, A.; Chaleplioglou, A.; Roinioti, E.; Dafiotis, P.; Lepouras, G.; Sousa, C.; et al. AI-Enabled Cultural Experiences: A Comparison of Narrative Creation Across Different AI Models. Electronics 2025, 14, 4043. [Google Scholar] [CrossRef]
- Belfiore, E. Whose cultural value? Representation, power and creative industries. Int. J. Cult. Policy 2020, 26, 383–397. [Google Scholar] [CrossRef]
- Alam, F.; Choudhury, M.; Di Nunzio, G.M.; et al. LLMs for Low Resource Languages in Multilingual Settings: Overview, Challenges and Opportunities. In Proceedings of the Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts, 2024; pp. 43–49. [Google Scholar] [CrossRef]
- Helm, P.; Bella, G.; Koch, G.; Giunchiglia, F. Diversity and language technology: how language modeling bias causes epistemic injustice. Ethics Inf. Technol. 2024, 26. [Google Scholar] [CrossRef]
- Khanal, S.; et al. A Comprehensive Survey and Guide to Multimodal Large Language Models in Generative AI. arXiv 2024, arXiv:2411.06284. [Google Scholar] [CrossRef]
- Selbst, A.D.; Boyd, D.; Friedler, S.A.; Venkatasubramanian, S.; Vertesi, J. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Proceedings of the Conference on Fairness, Accountability, and Transparency, New York, NY, USA, 2019; pp. 59–68. [Google Scholar] [CrossRef]
- Cao, Y.T.; Sotnikova, A.; Zhao, J.; Zou, L.X.; Rudinger, R.; Daumé, H., III. Multilingual Large Language Models Leak Human Stereotypes across Language Boundaries. In Proceedings of the Proceedings of the Fourth Workshop on NLP for Positive Impact (NLP4PI), Vienna, Austria, 2025; pp. 175–188. [Google Scholar] [CrossRef]
- Not, E.; Petrelli, D. Empowering cultural heritage professionals with tools for authoring and deploying personalised visitor experiences: the meSch ecosystem. User Model. User-Adapt. Interact. 2019, 29, 67–120. [Google Scholar] [CrossRef]
- Li, C.; Chen, M.; Wang, J.; Sitaram, S.; Xie, X. CultureLLM: Incorporating Cultural Differences into Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems, 2024; 37. [Google Scholar]
- Sarraj, G.; et al. CulturalBench: A Robust, Diverse and Challenging Cultural Knowledge Benchmark for Large Language Models. arXiv 2024, arXiv:2410.02677. [Google Scholar] [CrossRef]
- Myung, J.; Lee, N.; Zhou, Y.; Jin, J.; Putri, R.A.; Antypas, D.; Borkakoty, H.; Kim, E.; Perez-Almendros, C.; Ayele, A.A.; et al. BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages. In Proceedings of the Advances in Neural Information Processing Systems 37 (NeurIPS 2024), 2024; Datasets and Benchmarks Track. [Google Scholar] [CrossRef]
- Wang, Y.; et al. A Benchmark for Measuring the Cultural Dimensions of Large Language Models. In Proceedings of the Proceedings of the 4th Workshop on Cross-Cultural Considerations in NLP (C3NLP), 2024; pp. 1–13. [Google Scholar] [CrossRef]
- Gravagnuolo, A.; Angrisano, M.; Bosone, M.; Buglione, F.; De Toro, P.; Fusco Girard, L. Participatory evaluation of cultural heritage adaptive reuse interventions in the circular economy perspective: A case study of historic buildings in Salerno (Italy). J. Urban Manag. 2024, 13, 107–139. [Google Scholar] [CrossRef]
- UNESCO. Recommendation on the Ethics of Artificial Intelligence; Technical report; UNESCO, 2021. [Google Scholar]
- Tabassi, E.; et al. Artificial Intelligence Risk Management Framework (AI RMF 1.0); Technical Report NIST AI 100-1; National Institute of Standards and Technology, 2023. [Google Scholar] [CrossRef]
- High-Level Expert Group on Artificial Intelligence. Ethics Guidelines for Trustworthy AI. Technical report, 2019; European Commission.
- World Intellectual Property Organization. Generative AI: Navigating Intellectual Property. Tech. Rep. WIPO 2024. [Google Scholar] [CrossRef]
- UNESCO. Artificial intelligence and culture. UNESCO Themat. Page AI Cult. 2025. [Google Scholar] [CrossRef]
- Gmyrek, P.; Berg, J.; Bescond, D. Generative AI and Jobs: A Global Analysis of Potential Effects on Job Quantity and Quality. In Technical Report Working Paper 96, International Labour Organization; 2023. [Google Scholar] [CrossRef]
- International Labour Organization. Generative AI at work: What it means for jobs in Europe and beyond. In ILO article; 2025. [Google Scholar]
- Heigl, R. Generative artificial intelligence in creative contexts: a systematic review and future research agenda. Manag. Rev. Q. 2026, 76, 955–992. [Google Scholar] [CrossRef]
Figure 1.
Pipeline view of culturalized generative AI.

Table 1.
The curated narrative for the Tripolis tombstone and the newly generated narrative.
| Curated narrative |
|---|
| Look at this simple tombstone, between the first and the second century AD. Despite its simplicity, this particular item is of great importance for Arcadia. It bears the inscription “Hail Arcadia”, in honor of a deceased woman who has the same name as this area. But the name Arcadia has its root in mythology. According to the myth, once there was the beautiful daughter of Lycaon, king of this area, and she was called Callisto, meaning the most beautiful. Callisto was Artemis’ nymph and had taken a vow to remain a virgin. But Zeus seduced Callisto and she got pregnant. Artemis soon realized and became furious, and so did Hera, the wife of Zeus. The myth is not clear, but one of the two goddesses transformed the beautiful girl into a bear. As a bear, Callisto gave birth to a son called Arcas. Remember here that in Greek bear is called arctos or arkouda, and the names Arcas and Arcadia come from that. Arcas was separated from his mother, since he could not grow up with an animal. One day he went hunting and nearly killed his own mother. Zeus intervened and prevented the murder by placing mother and son amongst the stars, as Ursa Major and Ursa Minor, constellations known as Great and Little Bear. The myth goes further to explain why these two constellations are always visible in the northern hemisphere and never set: Hera cursed them never to rest. On your screen you can see a famous painting by Titian, housed in the National Gallery in London, picturing the scene in which Artemis discovers Callisto’s pregnancy. Imagine how she must have felt. Titian beautifully shows her emotions. Would you like to look further into the issue of art, myths and hidden pregnancies, or sexual education more broadly? |
| Newly generated narrative |
| At the entrance of the Archaeological Museum of Tripolis, a simple funerary stele greets us with the words “ARKADIA CHAIRE”—“Farewell, Arcadia”. Carved in stone nearly two thousand years ago, it is both a goodbye to a woman and an invocation of a place whose name is inseparable from myth. That same Arcadia lives again in Diana and Callisto by Titian at the National Gallery. Here, the nymph Callisto —-daughter of the Arcadian king Lycaon in myth-— stands at the tragic moment when her secret is revealed. Surrounded by Diana and her companions, she is caught between innocence and exile, between belonging and loss. The stone in Tripolis speaks with only two words; Titian’s canvas speaks with bodies, gestures, and color. Yet both tell the same ancient Arcadian story: how beauty, fate, and sorrow become memory. One preserves a human farewell, the other a mythic fall. Together, they remind us that Arcadia is not only a place on the map, but a landscape of remembrance, where history and myth continue to echo one another across centuries. |
Table 2.
Excerpt from an Experiment 1 output, aimed at a general audience, coded as containing a factual error.
Table 2.
Excerpt from an Experiment 1 output, aimed at a general audience, coded as containing a factual error.
| Artefact_Text | Artefact Name: Golden Burial Mask of Tutankhamun. Period: New Kingdom. Material: Gold; Glass; Lapis Lazuli; Obsidian; Carnelian; Faience; Quartzite. Object Type: Funerary Mask. Description: Solid gold funerary mask placed over the head and shoulders of Tutankhamun’s mummy; combines realistic facial features with divine representation. Symbolism: Osiris (underworld); Re (sun god); immortality (gold); divine identity. Key Insights: Religious syncretism; dual identity (human + god); symbolic materials reinforce divinity. |
| Visitor_Text | Age: 11. Education Level: Primary School. Field: 5th Grade Egyptian Curriculum. Primary Interest: Sports and Games. Interest Category: Entertainment. |
| Prompt | You are a museum guide. Explain the following Egyptian artefact for a general audience. Requirements: be accurate; use only the provided information; do not invent facts; keep the explanation between 100 and 120 words. Artefact data: Golden Burial Mask of Tutankhamun; New Kingdom; gold, glass, lapis lazuli, obsidian, carnelian, faience and quartzite; funerary mask; solid gold mask placed over the head and shoulders of Tutankhamun’s mummy; combines realistic facial features with divine representation; symbolism of Osiris, Re, immortality and divine identity; key insights of religious syncretism, dual human–god identity and symbolic materials reinforcing divinity. |
| GPT Response | This Golden Burial Mask of Tutankhamun dates to the New Kingdom and is a funerary mask made from gold, glass, lapis lazuli, obsidian, carnelian, faience, and quartzite. It was placed over the head and shoulders of the king’s mummy, serving both a protective and symbolic role. The mask combines realistic facial features with elements that represent divinity, showing the ruler as both human and god. Its symbolism connects to Osiris, the god of the underworld, and Re, the sun god, while the use of gold reflects immortality. Overall, the mask illustrates religious syncretism and expresses a dual identity, where materials and design reinforce the king’s divine nature. |
Table 3.
Excerpt from an Experiment 2 output, aimed at the adult profile, coded as containing a factual error.
Table 3.
Excerpt from an Experiment 2 output, aimed at the adult profile, coded as containing a factual error.
| Artefact_Text | Artefact Name: Golden Armchair of Tutankhamun Period: New Kingdom. Material: Wood; Gold leaf; Glass; Faience; Semi-precious stones. Object Type: Furniture/Throne. Description: Gold-covered ceremonial armchair with inlays; depicts king and queen under Aten sun rays; includes winged uraeus and royal cartouches. Symbolism: Royal authority; Aten symbolism; protection (uraeus); royal intimacy. Key Insights: Early reign artefact; Amarna influence; combines political and personal representation. |
| Visitor_Text | Age: 34. Education Level: PhD. Field: Technology. Primary Interest: Fashion. Interest Category: Lifestyle. |
| Prompt | You are a museum guide. Explain the following Egyptian artefact for the visitor described below. Requirements: Adapt the tone and narrative framing to the visitor’s education level, field, and stated interests. Use an intellectually engaging style appropriate for an adult visitor with a PhD in technology and an interest in fashion. You may highlight craftsmanship, materials, visual design, construction, or representation only when these aspects are supported by the artefact data. You may use modern design or fashion comparisons only as clearly marked explanatory analogies. Do not present modern concepts such as branding, innovation, identity design, luxury culture, or technology as historical intentions unless explicitly supported by the artefact data. Do not stereotype the visitor based on education, profession, age, or interest. Be culturally respectful. Be accurate. Use only the information provided in the artefact data. Do not invent facts or add unsupported historical, religious, symbolic, emotional, or functional interpretations. Preserve the cultural meaning stated in the artefact data. Keep the explanation between 100 and 120 words. Visitor profile: Age: 34; Education Level: PhD; Field: Technology; Primary Interest: Fashion; Interest Category: Lifestyle. Artefact data: Artefact Name: Golden Armchair of Tutankhamun; Period: New Kingdom; Material: Wood; Gold leaf; Glass; Faience; Semi-precious stones; Object Type: Furniture/Throne. Description: Gold-covered ceremonial armchair with inlays; depicts king and queen under Aten sun rays; includes winged uraeus and royal cartouches. Symbolism: Royal authority; Aten symbolism; protection (uraeus); royal intimacy Key Insights: Early reign artefact; Amarna influence; combines political and personal representation |
| GPT Response | The Golden Armchair of Tutankhamun, from the New Kingdom, is a ceremonial throne made of wood covered with gold leaf and enriched with glass, faience, and semi-precious stone inlays. Its decoration depicts the king and queen beneath the rays of the Aten, while a winged uraeus and royal cartouches reinforce themes of royal authority, Aten symbolism, protection, and royal intimacy. As an artefact from Tutankhamun’s early reign, it reflects the continuing influence of the Amarna period and combines political and personal representation. As a modern analogy only, you might compare its carefully coordinated materials and visual composition to a couture ensemble, where every element contributes to a unified, deliberate design. |
Table 4.
Profile-sensitive generated outputs from curated materials for A Limia.
| Location | A Limia plains. |
| Original cues | Although death is a very present theme, life and its passions also have their place. The Baroque author Tirso de Molina defines A Limia in La gallega María-Hernández as a flowery and fertile valley. The play narrates the journey of the Portuguese nobleman Don Álvaro, who falls in love with María in Galicia. Through these verses, Tirso expresses an exercise in love for Galicia that contrasts with much Golden Age Castilian literature. |
| Profile A | As you cross the wide plain of A Limia, you might expect a route devoted to death to speak only of endings. But this landscape also makes room for life, love and new beginnings. The Baroque writer Tirso de Molina imagined A Limia as a flowering, fertile valley. In his play La gallega María-Hernández, the Portuguese nobleman Don Álvaro arrives here and falls in love with María after entering Galicia. It is a reminder that the same road can carry very different stories. One traveller reflects on mortality. Another discovers love. As you continue walking, perhaps your own journey will become part of that landscape too. |
| Profile B | The broad landscape of A Limia offers a welcome counterpoint to the route’s recurring reflections on death. Here, Tirso de Molina presents the valley as fertile and flourishing in La gallega María-Hernández. The journey of the Portuguese nobleman Don Álvaro into Galicia culminates not in conflict or loss, but in his love for María. What makes these verses particularly striking is their generous appreciation of Galicia. Instead of treating the region as a distant setting, Tirso allows the landscape itself to become worthy of admiration, with A Limia occupying a central place in that literary vision. As you walk through the plain, it is worth remembering that landscapes are preserved not only in memory, but also in literature. |
| Profile C | The theme of death accompanies much of this route, yet this open plain offers another perspective. Tirso de Molina describes A Limia as a fertile and beautiful valley. Here, the journey of Don Álvaro into Galicia becomes a story of love rather than loss. Pilgrimage often brings together both realities. We carry thoughts about mortality, yet we also remain open to encounters, beauty and hope. This landscape quietly reminds us that life and death are not opposites, but companions along the same road. |
| Profile D | This stretch of A Limia invites you to pause and look across one of the great open landscapes of the route. The Baroque writer Tirso de Molina described this valley as fertile and flourishing in La gallega María-Hernández. His story follows the Portuguese nobleman Don Álvaro, whose journey into Galicia leads him to fall in love with María. Even after centuries, the landscape still lends itself to the same quiet observation. Wide plains encourage travellers to slow their pace, lift their eyes from the path and allow the place itself to become part of the experience. |
Table 5.
Profile-sensitive generated outputs from curated materials cues for Cimadevila.
| Location | Cimadevila. |
| Original cues | Death was also a source of significant economic activity. When a head of household died, the heirs were responsible for paying the luctuosa, a tribute to which nobles were entitled upon the death of their vassal. Vassalage even existed in the afterlife. |
| Profile A | Today, death is usually something deeply personal. In the past, it also carried financial consequences. When the head of a household died, the heirs were expected to pay a tax known as the luctuosa. Local nobles were entitled to receive it after the death of one of their vassals. It is a surprising thought. Even after someone had died, obligations remained behind for those who continued living. Here, death was not only about grief. It also shaped everyday life, inheritance and responsibility. |
| Profile B | Cimadevila preserves the memory of an aspect of death that is easy to overlook today: its place within the feudal economy. The death of the head of a household triggered the payment of the luctuosa, a tribute owed by the heirs because the deceased had been a vassal of the local lord. The expression that survives is memorable: “There was even vassalage in the afterlife.” Whether taken literally or rhetorically, it captures the remarkable extent to which feudal obligations reached into every stage of life, and beyond death itself. |
| Profile C | Death has always invited spiritual reflection, but it has also shaped the practical life of communities. Here, the death of the head of a household required the heirs to pay the luctuosa, a tribute claimed by the local noble after the death of a vassal. As you continue your pilgrimage, this serves as a quiet reminder that every age has understood death through more than one lens. Alongside faith and remembrance came responsibilities that those left behind had to carry. |
| Profile D | This stop reveals how closely death and everyday society could be connected. When the head of a household died, the heirs had to pay the luctuosa, a tribute that local nobles were entitled to receive after the death of one of their vassals. It is a small detail, but one that opens a window onto the organisation of rural life. Death was not only a personal event; it also affected property, obligations and the relationship between families and local authority. Even a short stop like this reminds us that history often survives in the customs that once seemed entirely ordinary. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.