Submitted:
28 July 2026
Posted:
29 July 2026
You are already at the latest version
Abstract
The Heritage-Aligned Reconstruction Framework (HARF) was developed through the Western Buddha of Bamiyan to structure AI-mediated heritage reconstruction through evidentially constrained prompt design, prompt-sufficiency assessment and expert-in-the-loop evaluation. Its transferability beyond a figural-sculpture case has not been empirically examined. This article tests that question through a bounded cross-site evaluation of the Temple of Bel at Palmyra, substantially destroyed in 2015. The case is not treated as a second full reconstruction study, but as a test of whether HARF’s schema and evaluative logic can be instantiated for a monumental architectural site governed by different material, proportional and iconographic conventions. The Dynamic Prompt Blueprint was adapted by replacing Bamiyan-specific descriptors with Palmyra-specific architectural evidence, including limestone construction, fluted Corinthian columns, inscriptions, cella dimensions, pseudoperipteral colonnade, thalamoi and zodiac ceiling reliefs. A fourteen-item Palmyra block within the Phase I expert questionnaire was completed by 32 respondents. The findings show that HARF’s schema transfers without redesign, while dominant failure points shift towards architectural and decorative details, especially relief, surface, inscription and roofline form. HARF is therefore best understood as a diagnostic evaluation framework, not as a reconstruction engine.
Keywords:
cultural heritage reconstruction
; generative artificial intelligence
; HARF
; Temple of Bel
; Palmyra
; bounded cross-site evaluation
; expert-in-the-loop evaluation
; digital heritage
1. Introduction
The deliberate destruction of cultural heritage during armed conflict has placed sustained pressure on digital reconstruction to develop methods that are both technically reliable and ethically accountable (Brosché et al., 2017; Cunliffe et al., 2016). Among the most consequential losses of the twenty-first century are the sixth-century Buddhas of Bamiyan, destroyed in March 2001 (Flood, 2002), and the Temple of Bel at Palmyra, intentionally destroyed with explosives in August 2015 (Schmidt-Colinet, 2019). Digital reconstruction methods have been applied to Palmyra through visual reconstruction, photogrammetry, crowdsourced imagery, and immersive 3D environments (Rihani, 2023; Silver et al., 2018). Generative-AI reconstruction raises a different evaluation problem: its outputs may appear visually coherent while remaining archaeologically uncertain. The methodological issue is therefore not only how to generate plausible images, but how to distinguish visual plausibility from evidential plausibility [7,8,9].
The Heritage-Aligned Reconstruction Framework (HARF) was introduced to address that gap (Arzomand et al., 2026). HARF integrates a Dynamic Prompt Blueprint, a Prompt Sufficiency Index, computational pre-evaluation, and expert-in-the-loop assessment within an evidentially constrained workflow. A companion empirical study evaluated HARF on the Western Buddha of Bamiyan and reported a structural divergence between algorithmic similarity scoring and expert archaeological judgement, with the most unstable reconstruction domains concentrated in drapery articulation, cranial morphology, and hand gesture (Arzomand & Kalganova, 2026b). The Bamiyan validation, however, leaves one substantive question open. A framework calibrated on a figural Gandhāran sculpture may behave differently on a heritage site of an entirely different typological order, and this has not yet been examined empirically.
This paper reports on a bounded cross-site evaluation of HARF on the Temple of Bel at Palmyra. The contrast with Bamiyan is deliberate. Bamiyan is a figural sculpture in mud-straw plaster on a sandstone cliff, governed by Gandhāran iconographic conventions; Palmyra is a Graeco-Roman temple in locally quarried limestone, governed by classical orders and an architrave with relief frieze (Arzomand et al., 2024; Seyrig et al., 1975). Each typology stresses different aspects of a reconstruction framework. The exercise reported here is questionnaire-based: it tests whether the framework’s schema and evaluative logic transfer, not whether AI outputs constitute defensible reconstructions of the temple itself. Specialist Palmyrene archaeology and debates on post-destruction preservation and reconstruction are addressed elsewhere (Abdulkarim & Seigne, 2024; Munawar, 2017; Schmidt-Colinet, 2019). Syrian and Palmyrene community perspectives are not examined empirically in this paper and are treated as a limitation.
This paper makes three contributions. First, it provides framework-level generalisation evidence by showing that HARF’s top-level schema and evaluative logic can be instantiated for a heritage typology fundamentally different from its original Bamiyan validation case without redesigning the framework architecture. Secondly, it identifies a typology-dependent instability pattern in generative reconstruction: when the framework is held constant and the heritage typology changes, the dominant failure domains shift from figural micro-grammar to architectural and decorative micro-grammar. Thirdly, it clarifies what HARF adds beyond existing evaluation approaches. Photogrammetric validation, visual plausibility assessment, and computational similarity pipelines can assess geometry, appearance, or structural correspondence, but they do not by themselves explain whether an evidence-constrained generative-AI evaluation framework remains stable across heritage typologies, or where expert review is likely to expose evidential instability. What is newly shown here is that HARF transfers at the level of schema and evaluative logic, while historical plausibility remains site-dependent. The contribution is therefore not a new reconstruction of the Temple of Bel, but a set of principles for evaluating when, how, and why generative-AI reconstruction frameworks transfer across heritage typologies, and where their outputs require expert intervention.
2. Materials and Methods
2.1. Adapting the Dynamic Prompt Blueprint to Palmyra
The Dynamic Prompt Blueprint is structured as a tiered slot schema with three levels: core structural anchors, auxiliary contextual descriptors, and technical control parameters (Arzomand et al., 2026). The schema is intended to remain stable across heritage contexts; only its slot contents are expected to change with the case. To test that property, the slot contents were re-populated for the Temple of Bel using UNESCO records, the principal Temple of Bel architectural study, and relevant digital-reconstruction literature [10,13,14,15]. Bamiyan-specific descriptors such as mud-straw plaster, Gandhāran wet-fold drapery, and Abhaya mudrā were replaced by limestone construction, fluted Corinthian columns of recorded height, a pseudoperipteral colonnade, a cella of recorded dimensions, an architrave frieze depicting Palmyrene deities, two interior thalamoi with zodiac ceiling reliefs, and Palmyrene–Aramaic inscription content. Negative-prompt content was reformulated to exclude Petra-like façades and undocumented decorative elements, and to mark uncertain roofline treatment as interpretive in line with the recent specialist reconsideration of the temple’s crenellations (Schmidt-Colinet & Seigne, 2022). The transferability test is therefore not whether a fixed prompt can be reused, but whether a fixed evaluative schema can be re-instantiated when the evidential content changes. The DPB supports transferability because it separates the stable structure of evaluation from the variable content of the case. Core anchors, contextual descriptors, negative constraints, and paradata remain constant as categories, while their contents change according to the monument or artefact under examination (Arzomand et al., 2026; Denard, 2009). In the Palmyra case, transferability is assessed by whether those categories remain usable when the object changes from a figural sculpture to an architectural monument. Table 1 reports the resulting cross-site instantiation. The schema, the tier structure, and the paradata-logging conventions were not altered.
2.2. Evaluation Instrument
The evaluation instrument was designed in response to a gap in existing digital-heritage evaluation practice. Photogrammetric and image-based reconstruction workflows can assess spatial correspondence and geometric reconstruction quality, but they do not address the specific problem posed by generative AI: visually coherent outputs may remain archaeologically or culturally unsupported (Rihani, 2023; Wahbeh et al., 2016). Digital-heritage charters emphasise transparency, paradata, and the distinction between documented evidence and interpretive reconstruction, but they do not provide a site-transfer test for generative-AI outputs (Denard, 2009; López-Menchero Bendicho & Grande, 2011). Recent benchmark work on historical and cultural artefacts likewise confirms the need for structured, expert-verified evaluation rather than reliance on output plausibility alone (Ghaboura et al., 2025). The Palmyra instrument therefore assesses framework behaviour rather than final reconstruction validity: whether HARF’s prompt architecture, prompt-sufficiency procedure, paradata record, and expert-review process can identify the domains in which generative outputs remain evidentially unstable.
A fourteen-item Palmyra block was embedded within the Phase I expert questionnaire used for the Bamiyan study (Arzomand et al., 2026; Arzomand & Kalganova, 2026b). The block was preceded by a structured monument briefing reproducing the temple’s principal dimensions, its architectural and artistic style, and its cultural and religious significance, so that respondents could rate against a defined evidentiary baseline rather than against personal knowledge alone (UNESCO, 2026). Six items elicited five-point Likert ratings (5 = Highly Accurate/Highly Sufficient; 4 = Mostly; 3 = Moderately; 2 = Slightly; 1 = Not Accurate/Insufficient) of structural fidelity, proportional accuracy of podium, staircase and colonnade, Corinthian-column representation, fidelity of remaining columns, walls and roofline against the photographic record, realism of restored relief, inscription and texture detail, and overall keyword sufficiency.
Four multi-select items elicited the architectural features judged missing or inaccurate, the aspects perceived to improve across iterations, the elements appearing most divergent from the documented ruins, and the descriptive domains where additional keyword coverage was required. Three single-choice items recorded the most accurate iteration for relief iconography, the most convincing iteration for the symbolic objects held by figures within the relief, and the preferred candidate among the ten reconstructions assessed. One open-ended item invited free-text suggestions for prompt refinement. Respondents evaluated the candidates against the image plates reproduced in Figure 1, Figure 2 and Figure 3.
2.3. Methodological Justification for Framework-Level Evaluation
The Palmyra evaluation tests framework transferability rather than the historical truth of the generated reconstructions. Keyword sufficiency is therefore treated as a scalar indicator of evidential coverage: it records whether the prompt contained enough site-specific information to constrain generation, while the accompanying keyword-domain item identifies where that coverage remained weak. It is not a full recalculation of the Bamiyan PSI.
Likert aggregation is used descriptively rather than psychometrically. The items are not treated as a single validity scale; each captures a distinct evidential domain, including structural fidelity, proportional coherence, columnar representation, ruin-state comparison, restored detail, and keyword sufficiency. Reporting means, medians, interquartile ranges, and threshold proportions provides a transparent account of how the panel evaluated each domain.
HARF therefore evaluates whether the framework can be re-instantiated, whether prompt sufficiency can be assessed, and whether expert review can locate instability domains. It does not, by itself, establish the historical validity of a reconstruction. Historical plausibility remains dependent on documentary evidence, specialist archaeological authority, and, where appropriate, community-grounded interpretation.
2.4. Respondents and Analytical Strategy
The Palmyra block was completed by thirty-two respondents drawn from the Phase I interdisciplinary expert panel. Familiarity with the Temple of Bel was elicited by a dedicated four-point item earlier in the instrument. Of the thirty-two Palmyra completers, ten reported no familiarity with the Temple of Bel, ten slight familiarity, six moderate familiarity, and six very high familiarity. The findings are therefore interpreted as interdisciplinary framework-level evidence rather than as Palmyrene specialist judgement.
Throughout this paper, n denotes the number of valid respondents for the relevant item, not the number of images. Unless otherwise stated, n = 32 for the Palmyra block. For multi-select items, percentages use the valid respondent count as the denominator and totals may exceed n because respondents could select more than one option.
Likert items are reported as means and standard deviations (SD), medians and interquartile ranges (IQR), and the proportion rated four or five. Multi-select and single-choice categorical items are reported as counts and percentages. Wilson 95% confidence intervals (CI) are reported for headline proportions (Wilson, 1927).
No computational similarity pipeline (SSIM, ORB+RANSAC, proportional similarity) was conducted on the Palmyra candidates. The Bamiyan composite-score procedure (Arzomand & Kalganova, 2026b) was not extended here, and the cross-site evaluation is restricted to questionnaire-based framework transferability. The implications of this restriction are addressed in the limitations.
3. Results
3.1. Architectural Fidelity
Ratings of architectural and structural accuracy clustered in the moderate band (Table 2). The structural-fidelity rating returned a mean of 3.28 (SD 1.20), with 46.9% of respondents (95% CI 30.9–63.6%) endorsing the candidates as Mostly or Highly Accurate; the proportional-accuracy rating returned 3.22 (SD 1.13) with the same proportion above threshold; the Corinthian-column rating was the highest of the macro-architectural items at 3.47 (SD 1.02), with half the panel above threshold (95% CI 33.6–66.4%). The modal response on the structural-fidelity and proportional-accuracy items was Mostly Accurate; the Corinthian-column item was split evenly between Mostly Accurate and Moderately Accurate. The pattern indicates cautious endorsement of macro-architectural plausibility rather than strong expert consensus.
Half of the panel selected the same single candidate (16/32) as the most accurate of the ten reconstructions presented in Figure 1, with a long tail across the remainder. The selected candidate corresponds to the largest of the candidates in Figure 1, an outcome consistent with the macro-architectural orientation of the Likert ratings. Architectural features identified as missing or inaccurate concentrated on column placement and height (40.6%, 95% CI 25.5–57.7%), interior thalamoi (37.5%), roofline merlons (34.4%), zodiac ceiling reliefs (31.2%), and architrave relief carvings (28.1%). Corinthian-column representation was therefore rated more favourably than restored relief and texture detail, but column placement and height remained a non-trivial source of expert concern.
This result has a practical implication for future applications of HARF. Macro-structural recognisability should not be treated as sufficient evidence of reconstruction adequacy. A generative output may preserve the broad silhouette or architectural massing of a site while failing in features that carry archaeological or cultural specificity. When HARF is transferred to another artefact, monument, or evidence-rich reconstruction context, the prompt-design and expert-review stages should first identify the culturally diagnostic features through which the object’s site-specific identity is expressed. In the Palmyra case, these features include relief carving, surface treatment, roofline form, inscriptions, and interior cult-architectural elements; in another case, the corresponding high-risk features would need to be defined from the documentary record before evaluation begins.
3.2. Iterative Refinement
Iteration did not resolve expert-panel concerns about relief iconography. As shown in Figure 4, 56.2% of respondents (18/32; 95% CI 39.3–71.8%) selected “none of the iterations are historically accurate” when asked to identify the iteration most historically accurate in its depiction of the original relief. The remaining responses were distributed across the four iterations, with no single iteration commanding more than 15.6% of the panel.
The result translates into a design constraint for generative reconstruction workflows. Relief iconography should be treated as a high-risk evidential domain rather than as a detail expected to improve through general iterative refinement. Iteration may improve surface coherence, but it does not necessarily improve historical or iconographic adequacy. In future HARF applications, iconographic, inscriptional, and culturally specific surface features should be isolated as explicit evaluation targets from the outset, rather than left to late-stage visual refinement.
On the iteration that most convincingly reconstructed the symbolic objects held by figures within the relief, the panel split evenly between “none” and the most refined iteration (10/32 each, 31.2%). Across iterations, respondents recognised gains in background architectural elements (46.9%) and decorative jewellery (43.8%) more readily than in inscriptions and iconography (34.4%) or facial expressions and proportions (31.2%). Iterative refinement, on this evidence, did not consistently translate visual gains into iconographic plausibility. The implication is that refinement cycles can be retained within HARF, but only as controlled diagnostic stages: each iteration should be assessed against predefined evidential criteria rather than against general visual improvement.
3.3. Comparison with the Photographic Record
When the candidates were judged against pre-2015 photographs of the temple, ratings were systematically lower than against the briefing description alone. The proportional-fidelity rating against the remaining columns, walls and roofline returned a mean of 2.88 (SD 1.18); the realism rating for restored relief, inscription and surface detail was the lowest of any item in the block at 2.47 (SD 1.11), with under 20% of respondents above threshold (Table 2). The categories most frequently identified as divergent from the photographic record are reported in Figure 5. Relief sculptures and wall carvings were selected by 62.5% of respondents (95% CI 45.3–77.1%), followed by surface textures and weathering (59.4%), roofline and decorative merlons (53.1%), entrance proportions (37.5%), and column height and fluting (31.2%). Surface, relief, and roofline detail therefore dominate the divergence profile; columnar geometry, while present in the distribution, sits below them.
3.4. Prompt-Keyword Sufficiency
The keyword-sufficiency rating returned a mean of 3.41 (SD 1.10), with half of the panel rating the keyword set as Mostly or Highly Sufficient and 6.2% as Insufficient. This is broadly consistent with the Bamiyan PSI value of 3.6 reported in the original HARF validation study, although the Palmyra measure is methodologically narrower because it is based on a single keyword-sufficiency item rather than a seven-domain weighted index (Arzomand et al., 2026). Domains where additional keyword coverage was requested concentrated on environmental and surrounding structures (53.1%, 95% CI 36.4–69.1%), architectural proportions and layout (46.9%), artistic elements such as friezes, frescoes and relief carvings (37.5%), cultural and religious symbolism (37.5%), and material and texture detail (31.2%). Four respondents (12.5%) considered the keyword set fully sufficient. Nine respondents provided substantive free-text suggestions; these recommended explicit specification of the cella footprint, removal of evocative or impressionistic phrasing in favour of concrete descriptors, the addition of weathering and patina vocabulary, the explicit inclusion of stepped merlons, and the grouping of keywords for use across iterative refinement steps. The pattern is consistent with the principle, articulated for Bamiyan, that prompt completeness is improved through site-specific evidential anchoring rather than through stylistic embellishment (Arzomand et al., 2026).
4. Discussion
The Palmyra evaluation supports a framework-level interpretation of HARF’s transferability. The principal finding is not that the Temple of Bel outputs were historically accurate; they were not consistently judged as such. Rather, HARF’s schema remained usable when transferred from a figural Gandhāran sculpture to a Graeco-Roman architectural monument: the DPB could be re-instantiated, the questionnaire could be extended to a new site, and the expert panel could identify domains of evidential instability. Macro-architectural features were assessed more favourably than relief, inscription, surface, and roofline detail; 56.2% of respondents judged that none of the relief iterations was historically accurate. The Palmyra case therefore does not validate the Temple of Bel reconstructions as historically adequate. It supports a narrower claim: HARF can structure the evaluation of a different heritage typology and expose where generative outputs fail.
This finding supports a typology-dependent instability principle. The term refers to the layer of features most likely to fail under generative reconstruction because they depend on culturally specific surface grammar rather than on generic macro-form. At Bamiyan, this instability concentrated in figural micro-grammar: drapery articulation, cranial morphology, and hand gesture (Arzomand & Kalganova, 2026b). At Palmyra, it shifted to architectural and decorative micro-grammar: relief carving, surface texture, inscription, roofline treatment, and interior cult-architectural features. The schema remained stable, but the failure layer changed. This matters because framework transferability does not imply output reliability. A transferable evaluation framework should be judged by whether it can locate the new failure domains produced by a different object type, not by whether its outputs appear visually coherent. The cross-site evidence can be represented as a three-layer model of transferability.
This model keeps the contribution of the Palmyra case bounded. The study does not show that HARF produces historically reliable reconstructions across heritage sites. It shows that, when applied to a new typology, HARF can preserve its evaluative architecture while revealing a different pattern of failure. A future application should therefore not begin from a blank methodological position, but neither should it assume that a successful Bamiyan prompt structure will automatically produce a valid reconstruction elsewhere. The schema can be retained; the descriptors and high-risk evidential domains must be redefined from the documentary record.
The Palmyra results also qualify the role of iterative refinement. The iteration sequence did not resolve relief-iconographic concerns, despite some perceived gains in background architecture and decorative detail. Iteration should therefore be treated as a controlled diagnostic stage, not as an automatic route to historical improvement. Each refinement cycle should be assessed against predefined evidential criteria, particularly where iconography, inscription, surface weathering, or roofline form is archaeologically significant. This is consistent with the wider problem identified in generative-AI heritage work: visual coherence can improve while evidential adequacy remains weak, especially where outputs depend on culturally specific surface features that may be under-represented or distorted in training data (Foka & Griffin, 2024).
The study therefore establishes Phase I transferability rather than complete validation. Phase I shows that HARF’s schema and evaluative logic can be re-instantiated across one major typological transition, from figural sculpture to architectural monument. Phase II remains necessary where site-specific archaeological authority, community grounding, or computational verification is required. HARF is consequently best understood as a diagnostic instrument for responsible generative-AI evaluation, not as a reconstruction engine.
The cross-site comparison can therefore be stated precisely. HARF transfers at the level of schema and evaluative logic; historical plausibility remains dependent on site-specific evidence, specialist judgement, and, where relevant, community interpretation. Table 3 summarises this distinction between framework-level transferability and site-specific evidential validity.
5. Limitations
Five limitations qualify these findings. First, the panel is interdisciplinary rather than composed of specialists in Palmyrene archaeology. As reported in Section 2.4, twenty of the thirty-two Palmyra completers reported no or only slight familiarity with the Temple of Bel. The structured monument briefing mitigated this constraint but cannot substitute for a Palmyrene specialist panel. Secondly, no Phase II Palmyra specialist evaluation accompanies this paper, so the depth of site-specific authority achieved at the second phase of the Bamiyan study is not reproduced here (Arzomand & Kalganova, 2026b). Thirdly, no community-grounded Syrian or Palmyrene respondent layer accompanies the paper, of the kind developed for Bamiyan in the polyvocal community study extending the wider research programme (Arzomand & Kalganova, 2026a). This is a substantive limitation, particularly because recent work on the Temple of Bel after its destruction stresses the importance of the Palmyrene population and collective memory in the site’s future (Abdulkarim & Seigne, 2024). Fourthly, no Palmyra-specific computational similarity pipeline was conducted; the cross-site evaluation is therefore restricted to questionnaire-based Phase I transferability, and full computational evaluation remains future work. Fifthly, reproducibility constraints arising from the volatility of generative platforms apply equally here. Paradata logging mitigates the constraint but does not eliminate it (Huvila, 2022).
6. Conclusions
A bounded cross-site evaluation indicates that HARF’s prompt blueprint and evaluative logic can be transferred from a figural Gandhāran sculpture to a Graeco-Roman architectural monument without redesigning the framework’s top-level schema. The transfer is evaluative rather than reconstructive: the Palmyra outputs were not consistently judged historically accurate. The error profile nevertheless changes with heritage typology, shifting from drapery, cranial morphology, and hand gesture at Bamiyan to relief, surface, inscription, roofline, and interior cult-architectural detail at Palmyra. This establishes a typology-dependent instability principle for generative heritage reconstruction.
The broader implication is that frameworks such as HARF are best understood as diagnostic instruments rather than reconstruction engines. Their value lies in structuring evidence, recording prompt logic, assessing prompt sufficiency, and identifying where generative outputs remain unstable. They do not remove the need for specialist judgement, nor do they convert visual plausibility into historical plausibility. A Phase II evaluation by specialists in Palmyrene archaeology, a Syrian and Palmyrene community layer, and a Palmyra-specific computational similarity pipeline remain necessary extensions.
Author Contributions
Author Contributions: Conceptualization, K.A.; methodology, K.A.; investigation, K.A.; formal analysis, K.A.; data curation, K.A.; visualization, K.A.; project administration, K.A.; writing—original draft preparation, K.A.; writing—review and editing, T.K.; supervision, T.K. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by the British Council Warm Welcome Scholarships scheme.
Institutional Review Board Statement
This study was approved by the Brunel University of London Research Ethics Committee under reference 50187-LR-Jan/2025-53624-2.
Data Availability Statement
Data Availability Statement: The aggregated data supporting the findings of this study are reported in the article. Raw respondent-level data are not publicly available because of ethical restrictions, respondent confidentiality, and the risk of deductive identification within a small expert panel. Further details may be made available from the corresponding author upon reasonable request, subject to ethical approval and data-protection requirements.
Acknowledgments
The author thanks the Phase I expert respondents and the British Council Warm Welcome Scholarships scheme.
Abbreviations
The following abbreviations are used in this manuscript:
| AI | Artificial intelligence |
| CI | Confidence interval |
| DPB | Dynamic Prompt Blueprint |
| GenAI | Generative artificial intelligence |
| HARF | Heritage-Aligned Reconstruction Framework |
| IQR | Interquartile range |
| ORB | Oriented FAST and Rotated BRIEF |
| PSI | Prompt Sufficiency Index |
| RANSAC | Random Sample Consensus |
| SD | Standard deviation |
| SSIM | Structural Similarity Index Measure |
| UNESCO | United Nations Educational, Scientific and Cultural Organization |
References
- Abdulkarim, M.; Seigne, J. The Future of the Temple of Bel in Palmyra after Its Destruction. Bull. Am. Soc. Overseas Res. 2024, 391, 93–106. [Google Scholar] [CrossRef] [PubMed]
- Arzomand, K.; Kalganova, T. Community-Grounded Authenticity at Bamiyan: Afghan Perspectives on AI-mediated Heritage Reconstruction [Preprint]. In Zenodo; 2026a. [Google Scholar] [CrossRef]
- Arzomand, K.; Kalganova, T. Evaluating Generative AI Reconstructions of the Bamiyan Buddhas: Computational Similarity and Expert Archaeological Assessment [Preprint]. In Zenodo; 2026b. [Google Scholar] [CrossRef]
- Arzomand, K.; Kalganova, T.; Rustell, M. HARF: A Human–AI Collaborative Framework for Cultural Heritage Reconstruction with Expert-Guided Multi-Platform Generative AI and Systematic Prompt Engineering [Preprint]. In Zenodo; 2026. [Google Scholar] [CrossRef]
- Arzomand, K.; Rustell, M.; Kalganova, T. From Ruins to Reconstruction: Harnessing Text-to-Image AI for Restoring Historical Architectures. Chall. J. Struct. Mech. 2024, 10(2), 69. [Google Scholar] [CrossRef]
- Brosché, J.; Legnér, M.; Kreutz, J.; Ijla, A. Heritage under attack: motives for targeting cultural property during armed conflict. Int. J. Herit. Stud. 2017, 23(3), 248–260. [Google Scholar] [CrossRef]
- Cunliffe, E.; Muhesen, N.; Lostal, M. The Destruction of Cultural Property in the Syrian Conflict: Legal Implications and Obligations. Int. J. Cult. Prop. 2016, 23(1), 1–31. [Google Scholar] [CrossRef]
- Denard, H. The London Charter for the Computer-Based Visualisation of Cultural Heritage; King’s College London, 2009; Available online: https://www.londoncharter.org/.
- Flood, F. B. Between Cult and Culture: Bamiyan, Islamic Iconoclasm, and the Museum. Art. Bull. 2002, 84(4), 641–659. [Google Scholar] [CrossRef]
- Foka, A.; Griffin, G. AI, Cultural Heritage, and Bias: Some Key Queries That Arise from the Use of GenAI. Heritage 2024, 7(11), 6125–6136. [Google Scholar] [CrossRef]
- Ghaboura, S.; More, K.; Thawkar, R.; Alghallabi, W.; Thawakar, O.; Khan, F. S.; Cholakkal, H.; Khan, S.; Anwer, R. M. Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts 2025. [CrossRef]
- Huvila, I. Improving the usefulness of research data with better paradata. Open Inf. Sci. 2022, 6(1), 28–48. [Google Scholar] [CrossRef]
- López-Menchero Bendicho, V. M.; Grande, A. The Seville Principles: International Principles of Virtual Archaeology (2011); International Forum of Virtual Archaeology (IFVA), 2011; Available online: https://smartheritage.com/seville-principles/.
- Munawar, N. A. Reconstructing Cultural Heritage in Conflict Zones: Should Palmyra be Rebuilt? Ex. Novo J. Archaeol. 2017, 2, 33–48. [Google Scholar] [CrossRef]
- Rihani, N. Interactive Immersive Experience: Digital Technologies for Reconstruction and Experiencing Temple of Bel Using Crowdsourced Images and 3D Photogrammetric Processes. Int. J. Archit. Comput. 2023, 147807712311682. [Google Scholar] [CrossRef]
- Schmidt-Colinet, A. No temple in Palmyra! Opposing the reconstruction of the Temple of Bel. Syr. Stud. 2019, 11(2), 63–85. Available online: https://ojs.st-andrews.ac.uk/index.php/syria/article/view/2015.
- Schmidt-Colinet, A.; Seigne, J. The battlements of temples in Hellenistic-Roman Syria. The crenellations of the Sanctuary of Bel in Palmyra reconsidered. Rev. Archéologique 2022, n° 73(1), 57–77. [Google Scholar] [CrossRef]
- Seyrig, H.; Amy, R.; Will, E. Le temple de Bel à Palmyre. Vol. I: Texte et planches; Vol. II: Album. In Revue Archéologique; (Number 1). Librairie Orientaliste Paul Geuthner., 1975; Available online: https://www.google.co.uk/books/edition/Le_Temple_de_B%C3%AAl_a_Palmyre/0JtbtgEACAAJ?hl=en.
- Silver, M.; Fangi, G.; Denker, A. Reviving Palmyra in Multiple Dimensions: Images, Ruins and Cultural Memory; Whittles Publishing Ltd, 2018. [Google Scholar]
- UNESCO. Site of Palmyra. UNESCO World Heritage Centre. 2026. Available online: https://whc.unesco.org/en/list/23.
- Wahbeh, W.; Nebiker, S.; Fangi, G. Combining public domain and professional panoramic imagery for the accurate and dense 3D reconstruction of the destroyed Bel Temple in Palmyra. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2016, III–5, 81–88. [Google Scholar] [CrossRef]
- Wilson, E. B. Probable Inference, the Law of Succession, and Statistical Inference. J. Am. Stat. Assoc. 1927, 22(158), 209–212. [Google Scholar] [CrossRef]
Figure 1.
Architectural and structural-accuracy stimuli: the ten AI-generated Temple of Bel reconstruction candidates evaluated by the panel. AI-generated evaluation stimuli, not presented as final reconstructions; platform metadata, model identifiers, prompt version, seed values, and refinement-cycle records are recorded in the project paradata in line with HARF and the London Charter (Arzomand et al., 2026; Denard, 2009).
Figure 1.
Architectural and structural-accuracy stimuli: the ten AI-generated Temple of Bel reconstruction candidates evaluated by the panel. AI-generated evaluation stimuli, not presented as final reconstructions; platform metadata, model identifiers, prompt version, seed values, and refinement-cycle records are recorded in the project paradata in line with HARF and the London Charter (Arzomand et al., 2026; Denard, 2009).

Figure 2.
Iterative-refinement stimuli: the original Palmyra relief (top left) and four AI-generated iterations. AI-generated evaluation stimuli, not presented as final reconstructions.
Figure 2.
Iterative-refinement stimuli: the original Palmyra relief (top left) and four AI-generated iterations. AI-generated evaluation stimuli, not presented as final reconstructions.

Figure 3.
Comparison of the original ruins with AI-generated reconstructions. Images 11 and 12 show pre-2015 views of the Temple of Bel ruins; Images 13 and 14 are AI-generated evaluation stimuli. The images were used to assess proportional fidelity of the remaining columns, walls, and roofline; the realism of restored relief carvings, inscriptions, and wall textures; and the elements appearing most divergent from the documented ruin state. The AI-generated images are evaluation stimuli only and are not presented as final reconstructions.
Figure 3.
Comparison of the original ruins with AI-generated reconstructions. Images 11 and 12 show pre-2015 views of the Temple of Bel ruins; Images 13 and 14 are AI-generated evaluation stimuli. The images were used to assess proportional fidelity of the remaining columns, walls, and roofline; the realism of restored relief carvings, inscriptions, and wall textures; and the elements appearing most divergent from the documented ruin state. The AI-generated images are evaluation stimuli only and are not presented as final reconstructions.

Figure 4.
Iteration judged most historically accurate for Temple of Bel relief iconography. The modal response was “None of the iterations are historically accurate” (56.2%; 95% CI 39.3–71.8%), indicating that iterative refinement did not resolve relief-iconographic concerns.
Figure 4.
Iteration judged most historically accurate for Temple of Bel relief iconography. The modal response was “None of the iterations are historically accurate” (56.2%; 95% CI 39.3–71.8%), indicating that iterative refinement did not resolve relief-iconographic concerns.

Figure 5.
Categories identified as most divergent from the documented Temple of Bel ruins. Bars denote the proportion of respondents flagging each category; error bars denote Wilson 95% confidence intervals. Surface, relief, and roofline detail concentrate at the top of the distribution.
Figure 5.
Categories identified as most divergent from the documented Temple of Bel ruins. Bars denote the proportion of respondents flagging each category; error bars denote Wilson 95% confidence intervals. Surface, relief, and roofline detail concentrate at the top of the distribution.

Table 1.
Cross-site instantiation of the Dynamic Prompt Blueprint for the Temple of Bel, with Bamiyan referents shown for comparability. The schema is preserved; the slot contents are replaced with site-specific descriptors drawn from documentary sources.
Table 1.
Cross-site instantiation of the Dynamic Prompt Blueprint for the Temple of Bel, with Bamiyan referents shown for comparability. The schema is preserved; the slot contents are replaced with site-specific descriptors drawn from documentary sources.
| Tier | Slot | Bamiyan referent | Temple of Bel instantiation |
| Core anchor | Monument identity | Western Buddha (Salsal) | Temple of Bel, Palmyra |
| Core anchor | Period | 6th century AD | 1st century AD; consecrated AD 32 |
| Core anchor | Geometry | Statue ~55 m, niche-aligned | Cella 39.45 × 13.86 × 14 m; 41 Corinthian columns at 15.81 m, 1.33 m base |
| Core anchor | Material | Mud-straw plaster on sandstone | Locally quarried limestone; fluted Corinthian columns with gilded capital treatment where documented |
| Auxiliary | Figural / architectural context | Standing Buddha, Abhaya mudrā, wet-fold drapery | Pseudoperipteral colonnade, fluted shafts, architrave frieze, monumental staircase, elevated podium |
| Auxiliary | Environmental setting | Sandstone cliff, monastic caves, Bamiyan Valley | Palmyrene urban sanctuary in oasis landscape; Great Colonnade and civic context |
| Auxiliary | Symbolic content | Indo-Hellenistic Buddhist iconography, halo, lotus | Palmyrene–Aramaic inscriptions; relief friezes of Palmyrene deities; ceiling reliefs of seven planets and twelve zodiac signs |
| Technical | Negative prompts | Exclude Petra-like façades, modern materials, anachronistic vegetation | Render roofline only where documented; mark uncertain roofline treatment as interpretive; exclude Roman-imperial conflations |
| Technical | Paradata | Seed, model, prompt version, refinement cycle | Seed, model, prompt version, refinement cycle |
Table 2.
Likert ratings of the AI reconstructions of the Temple of Bel (n = 32). Counts denote the number of respondents in each scale category. Wilson 95% confidence intervals are reported for the proportion rated ≥ 4.
Table 2.
Likert ratings of the AI reconstructions of the Temple of Bel (n = 32). Counts denote the number of respondents in each scale category. Wilson 95% confidence intervals are reported for the proportion rated ≥ 4.
| Rating dimension | Mean (SD) | Median (IQR) | 5 | 4 | 3 | 2 | 1 | % ≥ 4 [95% CI] |
| Structural fidelity | 3.28 (1.20) | 3.0 (3–4) | 5 | 10 | 9 | 5 | 3 | 46.9% [30.9–63.6%] |
| Proportions of podium, staircase, colonnade | 3.22 (1.13) | 3.0 (3–4) | 3 | 12 | 9 | 5 | 3 | 46.9% [30.9–63.6%] |
| Corinthian columns | 3.47 (1.02) | 3.5 (3–4) | 5 | 11 | 11 | 4 | 1 | 50.0% [33.6–66.4%] |
| Remaining columns, walls, roofline | 2.88 (1.18) | 3.0 (2–4) | 2 | 9 | 9 | 7 | 5 | 34.4% [20.4–51.7%] |
| Restored relief, inscriptions, textures | 2.47 (1.11) | 2.0 (2–3) | 1 | 5 | 9 | 10 | 7 | 18.8% [8.9–35.3%] |
| Keyword sufficiency | 3.41 (1.10) | 3.5 (3–4) | 5 | 11 | 10 | 4 | 2 | 50.0% [33.6–66.4%] |
Table 3.
Three-layer model of HARF transferability and typology-specific instability across Bamiyan and Palmyra.
Table 3.
Three-layer model of HARF transferability and typology-specific instability across Bamiyan and Palmyra.
| Layer | Role in the evaluation | Bamiyan | Palmyra | Implication |
| Stable framework layer | Prompt structure, paradata, sufficiency assessment, expert review | Preserved | Preserved | Can transfer without schema redesign |
| Site-specific evidence layer | Material, geometry, iconography, environmental context | Gandhāran sculpture, mud-straw plaster, drapery, mudrā | Limestone temple, Corinthian colonnade, thalamoi, inscriptions | Must be re-specified for each case |
| Typology-specific instability layer | Features most likely to fail under generation | Drapery, cranial morphology, hand gesture | Relief, surface, inscription, roofline | Must be identified early and evaluated explicitly |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.