Submitted:
15 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Activities related to the maintenance and conservation of architectural heritage require tools that provide reliable, spatially referenced geometric, diagnostic, and maintenance information on site. This article presents a critical review of the research carried out for Milan Cathedral and outlines a methodological workflow for using Mixed Reality in inspection and maintenance activities. The approach integrates four main components: the semantic and multi-resolution structuring of point clouds through classification and segmentation; the hierarchical organization of information according to the operational requirements of the Veneranda Fabbrica del Duomo; the localization of the device and the spatial registration of digital content within the real environment; and the design of a multimodal interface for consulting and entering data directly on site. The prototype developed demonstrates the feasibility of using classified point clouds directly as queryable information models in MR systems and enables maintenance information to be linked to specific architectural elements within their physical context. The results show that the main contribution of this approach lies not in a specific hardware platform, but in defining a structured, potentially transferable workflow that links digital surveying, information organization, and operational conservation activities.
Keywords:
mixed reality
; cultural heritage
; heritage conservation
; 3D survey
; point cloud classification
; immersive maintenance
1. Introduction
In recent decades, digital technologies have come to play an increasingly central role in the preservation and dissemination of cultural heritage [1,2]. In particular, the rapid development of 3D survey techniques has transformed how cultural heritage is documented and analyzed [3,4,5]. At the same time, Extended Reality (XR) has emerged as a transformative tool for cultural heritage interaction [4,5,6,7,8,9,10,11]. XR is an umbrella term encompassing Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), a set of immersive technologies spanning the reality-virtuality continuum [12,13]. VR refers to systems in which users engage with a fully virtual environment. AR overlays virtual content onto the real environment. MR blends virtual and real environments by enabling real-time interaction among the user, digital elements, and real-world objects.
Over the past decade, advances in head-mounted hardware have moved MR from a largely experimental concept to deployable spatial information platforms. Optical see-through devices like Microsoft HoloLens 2 and Magic Leap 2 provide improved display resolution, multimodal interaction, and upgraded spatial mapping capabilities [14,15].
The recent maturation of MR technologies is particularly relevant for domains in which decision-making is intrinsically linked to the physical environment. Cultural heritage conservation is one such domain, as inspection, documentation, and maintenance activities are carried out directly on site and require the continuous interpretation of spatially referenced information [16,17,18]. Decisions and conservation measures for a cultural asset must be based on an in-depth understanding of its historical, structural, material, and geometric characteristics, supplemented by ongoing inspections and periodic surveys that keep the picture of its current state of conservation up to date. In addition, it is a multidisciplinary practice that requires geometric and non-geometric data collected by specialists from different domains [19].
Managing and accessing this diverse, spatially referenced information in operational contexts remains a major challenge. More specifically, implementing MR in routine conservation practice entails four key methodological issues: preparing and structuring complex datasets, ensuring accurate user localization and digital content registration, integrating the system into existing maintenance workflows, and achieving adequate usability and user acceptance.
First, high-resolution digital surveys produce extremely large datasets that are difficult to manage, visualize, and query in real time. This challenge is compounded by the dense, heterogeneous nature of the content, which combines geometries, point clouds, images, descriptive attributes, diagnostic data, and interpretative information. To make this data truly usable in MR/XR environments, a semantic structure is required to link geometric data with non-geometric information, enabling selective consultation and interaction [20,21].
A second critical issue concerns user localization and registering digital content in the real world. In controlled environments or at the scale of a single room, MR headsets can provide relatively stable registration and effective re-localization. However, in historic buildings, geometric complexity, a spatial extent exceeding the normal operational range of the sensors, and the presence of repetitive architectural elements reduce the reliability of tracking and increase the risk of drift, recognition ambiguity and misalignment between the digital model and the physical context [22,23,24].
Building on this context, the 3D Survey Group at Politecnico di Milano has conducted a series of studies from 2021 that explored several important technical aspects of developing MR applications to support maintenance activities in large, monumental settings [25,26,27]. This research series aimed to provide practical solutions to address the limitations in content preparation, digital content registration, and user localization.
This work is part of a longer-running digitalization programme for the Milan Cathedral, initiated in 2008 by the 3D Survey Group in collaboration with the Veneranda Fabbrica del Duomo di Milano [28,29]. The latter is the historic institution responsible for the conservation, preservation, and enhancement of the Cathedral [30]. The programme includes an extensive 3D survey; the production of detailed 3D point cloud models and 2D representations from the survey data; the establishment of a master data management (MDM) database; and the implementation of a Web-BIM platform in a VR environment [31,32,33,34]. The Cathedral’s scale, architectural complexity, and ongoing preservation demands make it a representative case for developing an expert-facing MR system.
The Cathedral comprises heterogeneous spaces that vary in size, lighting conditions, function, and degree of decoration, requiring a digital surveying, data management, and data querying approach that is flexible, scalable, and reliable (Figure 1). The Veneranda Fabbrica requires technical documentation (up to 1:20 scale) that is readily available and accessible anywhere, along with a complete block-by-block representation of the marble components. At present, expert operators conduct on-site visual inspections of marble blocks, statues, and structural elements, assess their condition, and identify required interventions. At the same time, technical information is recorded in paper tables, and photographs are collected, while degraded marble blocks are marked directly on site for partial or total replacement.
The MR application was conceived to make maintenance information directly accessible and spatially anchored on site, thereby reducing potential data loss, improving information continuity, and streamlining the overall maintenance workflow. In this way, the system establishes a direct link between physical inspection and digital data.
This paper presents a critical synthesis of the research conducted to date and outlines its future development. Rather than focusing on a specific device or technological platform, the study aims to formalize a robust, structured, and transferable methodological workflow that can remain adaptable as hardware and software systems evolve.
The proposed workflow comprises four main phases. The first involves transforming digital survey data into a structured, semantically organized model suitable for use within the MR application. The second consists of defining an information structure consistent with the operational requirements of the Veneranda Fabbrica del Duomo, translating maintenance procedures into classification criteria for architectural elements and into forms for recording interventions. The third phase involves developing the MR application and the procedures for locating the device and spatially registering the digital content within the physical environment. Finally, the fourth phase concerns the design of the interface and interaction methods, intended to be accessible even to operators with limited familiarity with digital tools, and to support hands-free interaction compatible with on-site working conditions.
2. Materials and Methods
Due to the integrated digital survey campaigns conducted by the 3D Survey Group from 2008 to the present, which have evolved with technological advancements, a comprehensive and accurate point cloud model of the Milan Cathedral has been generated. The laser scanning survey consists of approximately 1,300 scans, whereas the photogrammetric survey includes 16,000 photos (Figure 2). The comprehensive model comprises 60 billion points, with an average resolution of around 5 mm, adequate for generating technical drawings at scales of 1:150 or 1:20.
The starting point for the research was to understand how to make detailed digital data not only accessible but also usable in maintenance processes. Survey data is traditionally used to produce two-dimensional technical drawings, which remain indispensable tools for documenting and planning maintenance work. However, the research explored the possibility of extending their use by employing them to organize, contextualize, and record maintenance activities, both past and future. The central question was therefore how digital survey data could provide operational support for a complex, permanent, and constantly evolving construction site such as Milan Cathedral, beyond the mere production of technical drawings. From this perspective, MR was identified as an interface particularly well-suited to linking the physical space with the maintenance information associated with the digital model.
The processed point cloud model, produced through the ongoing digital survey campaign, therefore serves as both the geometric basis for developing the MR application and the spatial reference for the in-situ entry, visualization, and retrieval of information relating to maintenance work. This concept represents a central aspect of the research project: the proposed method begins by digitizing the physical environment to enrich and support an augmented experience of real space through immersive and interactive means, making technical reference data directly accessible and available for consultation within the operational context (Figure 2).
Implementing this strategy necessitated converting the survey model into a systematic and functional MR environment while tackling four interconnected methodological challenges connected to the four workflow phases: (i) creation of a 3D model suitable as base structure in the MR application; (ii) spatial localization and aligment of the digital content within the physical environment, referencing of data for 1:1 superimposition; (iii) design of a useful, customized user interface; (iv) and future integration with an external database for information storage.
2.1. Management, classification and segmentation of point cloud dataset
A primary challenge was handling the vast 3D survey dataset associated with the Milan Cathedral and classifying it to convert the point cloud into an informative point model. The objective is to bypass the laborious, time-consuming manual modelling process, regardless of the adopted paradigm—CAD, BIM, 3D NURBS surfaces, or mesh-based modelling—as each requires substantial manual interpretation, reconstruction, and editing of survey data. Furthermore, Veneranda Fabbrica requires models with sufficient detail to visualize and identify individual blocks. The advantages of the point cloud include a less time-consuming elaboration process and an accurate and detailed three-dimensional representation. Nonetheless, it lacks the requisite semantic structure and is challenging to handle due to its large size (Figure 3).
A semi-automatic classification strategy based on Machine Learning methods was employed to convert the raw point cloud dataset into a structured, semantically enriched model. The extensive and complex architectural environment, coupled with significant computational requirements, renders one-shot classification of the full-resolution model unfeasible. Consequently, a Multi-Level-Multi-Resolution (MLMR) classification approach was used, which allows 3D data classification at multiple geometric resolutions, each corresponding to a different classification level. As the geometric detail increases, the single data class was split into sub-classes. The methodological approach is described in detail in previous publications [26,35].
The semi-automatic classification was organized into three progressively detailed levels, each associated with a different spatial resolution (Figure 4 and Figure 5). At Level 1, a point cloud with approximately 5 cm resolution was used to classify macro-architectural elements. At Level 2, these macro-elements were further subdivided into their main architectural components; for example, columns were split into bases, shafts, and capitals, using a higher-resolution point cloud with approximately 2 cm resolution. At Level 3, the full-resolution dataset was used to subdivide each component into individual blocks or discrete elements, including monolithic units, sculptures, and decorative features. This hierarchical organization allows the same dataset to support both an overall reading of the monument and detailed, block-level maintenance activities.
Finally, instance segmentation identified and assigned a unique index to individual architectural components within each semantic class, enabling unambiguous linking to maintenance records.
The semantic classes and classification levels have been defined in accordance with the operational standards and maintenance requirements of the Veneranda Fabbrica. On this basis, a hierarchical structure for the monument has been drawn up, organized into the following levels: 1) Zone: portion exhibiting analogous intervention approaches; 2) Area: location within the Zone based on architectural characteristics; 3) Sector: a designated location inside the area, such as a bay for the nave; 4) Architectural element: a recognizable object has distinct architectural functions and characteristics; 5) Block: the ultimate component with no potential for additional subdivision [36].
The hierarchical data structure interfaces seamlessly with the Unity Scene system used to design the MR application (Figure 6). In this environment, each scene represents a distinct Sector of the Cathedral. Simultaneously, the data for the final two categories (Architectural element and Block) are nested within each scene and arranged by classification levels (Level 1 - macro architectural elements, Level 2 - individual architectural elements, Level 3 - individual components). The Scene system architecture directly mirrors the outcomes of the MLMR classification structure and the Milan Cathedral database. The identical structure is employed to store intervention data and images locally (on the HoloLens 2) (Figure 6).
This approach optimizes loading and rendering: high-resolution models are used only for close-up, detailed analysis, while lower-resolution versions are used for overviews, maintaining a good balance between visual quality and performance. Scalable point-cloud management at varying resolutions also enhances the user experience by dynamically adapting detail based on distance and user requirements. In this way, the system effectively balances performance and computational sustainability.
In conclusion, the defined workflow – from point clouds to MLMR classification, from the hierarchical structuring of scenes to their visualization in MR – and the designed structure enable maintenance work to be documented and tracked by linking each operation to a specific architectural element, its precise location and the most appropriate scale of representation.
2.2. Holographic Content Alignment and Positioning System
A crucial aspect of the MR application’s operation concerns the device’s ability to localize itself in space and maintain a stable alignment between the digital content and the real environment. Before the model of the Cathedral can be superimposed onto the physical structures, a headset device must construct a local geometric representation of the environment and continuously update its position. The accuracy of this process depends on both the characteristics of the tracking system and environmental conditions.
The HoloLens 2, used during the research and testing phases, generates a spatial map of the surrounding environment, consisting of a 3D mesh, which enables the device to estimate and dynamically update its position relative to the real-world space. Environmental conditions influence the accuracy of this map. Localization is based on four visible-light tracking cameras, used for real-time visual-inertial SLAM, whilst the long-range, low-frame-rate depth sensor acquires the data required to construct the spatial map [37]. Tracking is affected by lighting conditions, although infrared frame acquisition is less sensitive to variations in ambient visible light.
From the experimental tests carried out, the device creates accurate local spatial maps for small, complex spaces [38,39], whilst it is less effective in environments – such as the naves of the Cathedral – where the dimensions and heights exceed the depth sensor’s range. When the device is in use, the virtual content, also known as holograms, is anchored to the acquired spatial maps and remains stable and well-aligned within a radius of approximately 6–7 metres from the user. However, in very large spaces, accuracy decreases with distance, leading to increasingly noticeable deviations, especially for content located at high or distant points (for example, information at the highest point of the vaults, 24 metres above ground level).
As described in previous research, device positioning and hologram alignment were addressed by combining the HoloLens spatial mapping capabilities with Microsoft’s World Locking Tools (WLTs) [40], topographically referenced QR codes, and a simplified ICP (Iterative Closest Point) algorithm that operators can activate during use to correct both large- and small-scale misalignments and limit drift. Under the tested conditions, deviations measured at the QR control points remained within approximately 15 cm [23,24].
Overall, the strategy adopted in the HoloLens 2-based prototype enabled, under the experimental conditions considered, a sufficiently stable and consistent alignment between the digital content and the physical context (Figure 7). Unfortunately, this result cannot be interpreted as a definitive solution or as one that can be automatically transferred to other devices, primarily due to the absence of a standardized methodology for how these devices manage the map or experience drift over time. [41].
Localizing the headset and spatially registering the content are, in fact, fundamental methodological requirements for reliably associating maintenance information with the actual elements of the monument. Its technical implementation depends on sensor and tracking-system characteristics, as well as the functionalities available on each platform. The transition to different devices, such as the currently trialled Magic Leap 2, will therefore require a new performance assessment and may entail adopting different alignment and correction strategies. The methodological contribution therefore lies not in the specific technological solution adopted, but in recognizing spatial registration as an indispensable stage of the workflow and establishing the need to validate it against operating conditions and the required level of accuracy.
2.3. User Experience and Application Usability
The primary objective of the MR application is to support the maintenance of the Cathedral by facilitating the planning of maintenance interventions and improving the management and accessibility of technical information. However, the main challenge lies in enhancing the user experience (UX) by developing bespoke solutions for user–application interaction and integrating them into operational workflows. The key issues concern the system’s ability to adapt effectively to the requirements of on-site maintenance activities, as well as the operators’ propensity to adopt innovative solutions based on emerging technologies.
Users can interact with nearby content through hologram direct manipulation (translation, rotation and resizing) and through virtual hand-ray selection and mid-air gestures (air-tap) (Figure 8). These functions simulate direct physical manipulation, allowing users to select or move models that are beyond their physical reach.
The MR application is defined by adhering to a systematic sequence of steps. A typical scenario and sequence of actions might be as follows.
2.3.1. Step 1: Selecting the Work Area (Mandatory)
In the virtual environment of the Cathedral, the user must choose a working area, beginning with a low-resolution model. The selection process occurs step-by-step: the entire building, interiors/exteriors, and specific sections (e.g., naves, transept, apse) (Figure 9).
2.3.2. Step 2: Digital Model – Real Space Alignment (Mandatory)
After selecting the work area, the user aligns the 1:1-scale point cloud model within the physical space. HoloLens 2 first maps the surrounding environment, whilst the WLT system – supported by the topographical references provided by the QR codes – stabilizes the spatial relationship between the device, the hologram and the Cathedral’s reference system. Once registered, the digital model can be viewed at actual size, allowing maintenance information to be linked to individual marble blocks in situ.
2.3.3. Step 3: Visualise High-Res Models (Optional)
The user has access to detailed models of architectural elements that can be interacted with (rotated, moved, zoomed). The aim is to enable in-depth analysis and inspection.
2.3.4. Step 4: Creating a New Intervention (Optional)
The user can select a specific point—and therefore the corresponding architectural element—and create a new maintenance record by completing the associated technical form. Structured according to the operational requirements of the Veneranda Fabbrica del Duomo, the form includes a unique identifier, block type, intervention type, date, operator, and notes. Images can also be captured using the HoloLens camera and stored locally. This procedure supports the standardized documentation and traceability of maintenance interventions.
2.3.5. Step 5: Visualise Existing Interventions (Optional)
Users can access records of previous works via hotspots (red dots) linked to the high-resolution point cloud model. In this way, historical technical data (chronology of works) can be contextualized within the physical location (Figure 10).
2.3.6. Step 6: Visualize a Close-up of a Table (Optional)
The user can open a maintenance form to review, edit, or update its contents, attach photographs, and revise the intervention status. Records can be accessed either by selecting the corresponding physical element in situ or through the associated 3D model.
2.3.7. Step 7: Reset local alignment (optional)
The user can optimize hologram alignment in physical space and overlay accuracy. Indeed, during the MR section, the point model accumulates misalignment over time, but the drift errors can be corrected using the ICP algorithm, optimized for use within the app.
In summary, the operational workflow design focused on three main aspects. The first concerns multimodal interaction, combining natural gestures and voice commands, enabling the application to be used in hands-free mode. The second concerns the use of 3D icons and hotspots to make critical points and previously recorded interventions immediately recognizable on the structure. The third aspect concerns voice data entry, which enables technicians to complete maintenance forms by dictation, thereby reducing the need to interrupt manual tasks.
3. Results
The tangible outcome of this research project is a prototype MR application developed for use within the Milan Cathedral. The application enables digital content and information related to inspection and maintenance activities to be displayed directly in the physical space. Through an immersive interface, operators can view, query, and analyze digital models on site, linking maintenance information to the architectural elements to which it relates.
Beyond the prototype, the main methodological achievement lies in having demonstrated the possibility of using classified and segmented point clouds directly as semantic models within MR applications. This approach reduces the need for intermediate manual geometric modelling and makes the transition from digital surveying to operational data use more straightforward. The semantic and multi-scale structuring of the dataset also enables complex three-dimensional models to be managed and queried at different levels of detail, depending on requirements for visualization, analysis, and maintenance.
Overall, the research demonstrated the feasibility of directly using point clouds, on-site spatial alignment, and employing multi-scale interaction within a complex monumental context. The results obtained therefore constitute an initial validation of the proposed workflow and a basis for subsequent phases of experimentation under real-world operating conditions.
4. Discussion
The next phase, therefore, shifts the focus to long-term usability in real-world scenarios. In fact, once its technical feasibility has been established under test conditions, the next question is whether it can remain reliable in real site conditions and be adaptable to other heritage contexts.
This shift in focus is also shaped by the broader instability of the commercial MR headset market, where multiple platforms have been discontinued or significantly restructured within short development cycles. The previous stage of the research was developed around HoloLens 2, which provided a coherent experimental platform supported by a mature software development ecosystem, including MRTK, Azure-based services, and World Locking Tools. That environment reduced the custom engineering required during the prototype phase. The discontinuation of HoloLens 2 made future development on this platform difficult. Nevertheless, the problem is not specific to one device. For a research line aimed at operational deployment, long-term usability cannot be grounded in any single headset. It must be grounded in a development approach that remains transferable as hardware evolves. The emergence of platforms such as Apple Vision Pro further illustrates the rapid evolution and diversification of MR technologies. Unlike optical see-through devices, Apple Vision Pro combines virtual content with a camera-mediated view of the physical environment through high-resolution video passthrough. While this approach enables effective integration of physical and digital content, optical see-through perception remains particularly relevant in this application, where operators must directly inspect the material appearance and surface condition of heritage elements.
The technical part of the research therefore moves toward using Magic Leap 2. The device satisfies the hard requirement of optical see-through perception, which is essential for on-site inspection of material condition and surface decay. At the same time, it offers a larger field of view, a lighter headset, and improved computing power. More importantly, its Android-based environment opens a development path that is not tied to a single proprietary ecosystem, making application-level work more likely to transfer to other devices as they become available.
Against this background, the next stage of the project can be framed as a move towards scalable and long-term usability. One priority concerns the visualization of dense survey data. Earlier studies have shown the value of a multi-level approach, in which classified point clouds can support interaction at different scales within the monument. The next step is to make this logic more flexible and more closely aligned with conservation practice. The system should allow the user to move smoothly between architectural-element and marble-block-level interactions.
A second line of work concerns localization and alignment in a broader range of spatial conditions. The previous experiments were conducted primarily in indoor settings and yielded a benchmark focused on horizontal deviation. This has set up a concrete foundation. Nevertheless, real conservation work does not occur in controlled spaces. It takes place on scaffolds, on lifting platforms, in exterior areas, in narrow passages, and under changing illumination. These conditions can alter the headset’s tracking behaviour, reduce stability, and expose weaknesses that remain less visible in regular environments. The next stage of the research should therefore test alignment methods across a wider range of situations, including exterior settings, elevated work positions, lighting variations, and temporary obstructions. This is where robustness will be decided.
Another unresolved issue is the relationship between the MR interface and the Cathedral’s broader information infrastructure. The long-term objective is to connect on-site inspection with structured, updatable, and spatially referenced information. At present, the prototype does not provide round-trip access to the Milan Cathedral MDM database. The establishment of this link can move the MR system from a visualization and consultation device to an active working node within the conservation workflow. The next phase must therefore address database connectivity, data insertion, local storage, and later synchronization as part of a single operational logic. This is especially important because internet access cannot be guaranteed throughout the Cathedral. Offline use will be a practical requirement. Data generated during inspection should be stored locally, preserved without interruption, and synchronized to the database when the connection becomes available again. Only then can the application support the continuity of records rather than isolated consultation sessions.
Usability also needs to be tested in real working conditions. What matters is not only whether the interface is understandable in isolation. It is whether it fits the rhythm of actual inspection work. The next phase should therefore involve structured user studies with operators from the Veneranda Fabbrica del Duomo di Milano, who are directly responsible for the Cathedral’s day-to-day maintenance. Their involvement will allow evaluation to move beyond technical performance metrics. That gap can only be closed by testing in the situations for which the application is intended.
At the same time, the system is expected to be extended as a modular framework rather than as a site-specific solution tied to the Milan Cathedral. The Cathedral remains an especially demanding and productive test bed because of its size, complexity, and maintenance intensity. However, the broader relevance of the research depends on whether the core components of the application can be transferred to other heritage sites with different characteristics and operational routines. Alignment logic, point cloud handling, interaction methods, and data exchange should therefore be structured to allow recalibration rather than complete redevelopment. This is a technical issue, but also a conceptual one. The research should aim to define a method adaptable across heritage contexts, not a closed tool optimized for a single monument.
5. Conclusions
The research highlights that a crucial barrier to adopting MR in maintenance activities and cultural heritage conservation practices lies less in the availability of hardware technologies than in defining robust, structured methodological workflows that are sufficiently independent of the evolution of individual platforms. The contribution of this work therefore does not lie in the development of a solution tied to a specific headset – which is inevitably subject to rapid obsolescence – but in the formalization of a method that is transferable and adaptable to changing devices and software ecosystems.
The proposed approach is divided into four main stages as shown in Figure 10: (i) structuring and semantically organizing digital survey data for the MR application; (ii) creating an information structure that meets Veneranda Fabbrica operational needs which is reflected in the criteria for classifying architectural elements and in the forms used to record restoration work; (iii) develops the application and processes for spatial device localization and digital content registration in the physical world; (iv) designs the interface and interaction mechanisms, approachable to operators with limited digital tool experience and suitable for on-site operational settings.
However, one crucial issue remains unresolved, one that cannot be addressed solely through technological development: the effective integration of the application into day-to-day maintenance practices. The next stages of the research will therefore need to assess not only the system’s technical reliability, but also its practical usefulness, ease of use, and ability to fit into established timetables, routines, and procedures on site. It will also be necessary to assess the extent to which operators can accept the application and help overcome the natural organizational inertia associated with introducing new tools.
Table 1.
This is a table. Tables should be placed in the main text near the first time they are cited.
Table 1.
This is a table. Tables should be placed in the main text near the first time they are cited.
|
Methodological issue |
Solution developed |
Evidence obtained |
Current limitation |
| Large point clouds | MLMR classification and segmentation | multi-scale visualization | point-level interaction (the prototype needs mesh colliders) and point clouds with different LoD visualisation management |
| Spatial registration | WLT + QR + ICP | deviations at control points | session drift |
| Maintenance information | hierarchical records | association with individual elements | database synchronization |
| User interaction | multimodal interface | working prototype | user validation |
Author Contributions
Conceptualisation, F. Fassi; methodology, F. Fassi and S. Teruggi; software, S. Teruggi; validation, S. Teruggi and F. Fassi; investigation, S. Teruggi; resources, F. Fassi; data curation, S. Teruggi; writing—original draft preparation, F. Fiorillo and Y. Lei; writing—review and editing, F. Fiorillo, Y. Lei, F. Fassi and S. Teruggi; visualization, S. Teruggi, F. Fiorillo and Y. Lei; supervision, F. Fassi; project administration, F. Fassi. All authors have read and agreed to the published version of the manuscript.” Please turn to the CRediT taxonomy for the term explanation. Authorship must be limited to those who have contributed substantially to the work reported
Data Availability Statement
The point cloud datasets of the Milan Cathedral used in this study are proprietary and cannot be made publicly available due to restrictions imposed by the data owner. Data are available from the corresponding author upon reasonable request and subject to the approval of the data owner.
Acknowledgments
In this section, you can acknowledge any support given that is not covered by the author contribution or funding sections. This may include administrative and technical support, or donations in kind (e.g., materials used for experiments). Where GenAI has been used for purposes such as generating text, data, or graphics, or for study design, data collection, analysis, or interpretation of data, please add “During the preparation of this manuscript/study, the author(s) used [tool name, version information] for the purposes of [description of use]. The authors have reviewed and edited the output and take full responsibility for the content of this publication.”
Abbreviations
The following abbreviations are used in this manuscript:
| ICP | Iterative Closest Point |
| MLMR | Multi-Level-Multi-Resolution |
| MR | Mixed Reality |
| XR | Extended Reality |
References
- Yu, Y.; Abu Raed, A.; Peng, Y.; Pottgiesser, U.; Verbree, E.; van Oosterom, P. How Digital Technologies Have Been Applied for Architectural Heritage Risk Management: A Systemic Literature Review from 2014 to 2024. npj Herit. Sci. 2025, 13, 45. [Google Scholar] [CrossRef]
- Mendoza, M.A.D.; De La Hoz Franco, E.; Gómez, J.E.G. Technologies for the Preservation of Cultural Heritage—A Systematic Review of the Literature. Sustainability 2023, 15, 1059. [Google Scholar] [CrossRef]
- Zhang, J.; Wang, G.; Chen, H.; Huang, H.; Shi, Y.; Wang, Q. Internet of Things and Extended Reality in Cultural Heritage: A Review on Reconstruction and Restoration, Intelligent Guided Tour, and Immersive Experiences. IEEE Internet Things J. 2025, 12, 19018–19042. [Google Scholar] [CrossRef]
- Chen, F.; Ma, P.; Chen, S.; Hu, Q.; Guo, H. Remote Sensing for Cultural Heritage: A Systematic Review. Int. J. Appl. Earth Obs. Geoinf. 2026, 146, 105039. [Google Scholar] [CrossRef]
- Xing, Y.; Yang, S.; Fahy, C.; Harwood, T.; Shell, J. Capturing the Past, Shaping the Future: A Scoping Review of Photogrammetry in Cultural Building Heritage. Electronics 2025, 14, 3666. [Google Scholar] [CrossRef]
- Chatsiopoulou, A.; Michailidis, P.D. Augmented Reality in Cultural Heritage: A Narrative Review of Design, Development and Evaluation Approaches. Heritage 2025, 8, 421. [Google Scholar] [CrossRef]
- Omran, W.; Ramos, R.F.; Casais, B. Virtual Reality and Augmented Reality Applications and Their Effect on Tourist Engagement: A Hybrid Review. J. Hosp. Tour. Technol. 2024, 15, 497–518. [Google Scholar] [CrossRef]
- Dordio, A.; Lancho, E.; Merchán, M.J.; Merchán, P. Cultural Heritage as a Didactic Resource through Extended Reality: A Systematic Review of the Literature. Multimodal Technol. Interact. 2024, 8, 58. [Google Scholar] [CrossRef]
- Comes, R.; Buna, Z.L. Virtual Reality in Cultural Heritage: A Scientometric Analysis and Review of Long-Term Use and Usability Trends. Appl. Sci. 2026, 16, 1013. [Google Scholar] [CrossRef]
- Zhang, J.; Wan Yahaya, W.A.J.; Sanmugam, M. The Impact of Immersive Technologies on Cultural Heritage: A Bibliometric Study of VR, AR, and MR Applications. Sustainability 2024, 16, 6446. [Google Scholar] [CrossRef]
- Innocente, C.; Ulrich, L.; Moos, S.; Vezzetti, E. A Framework Study on the Use of Immersive XR Technologies in the Cultural Heritage Domain. J. Cult. Herit. 2023, 62, 268–283. [Google Scholar] [CrossRef]
- Milgram, P.; Takemura, H.; Utsumi, A.; Kishino, F. Augmented Reality: A Class of Displays on the Reality–Virtuality Continuum. Proceedings of SPIE—The International Society for Optical Engineering, 1995; Available online: https://www.scopus.com/pages/publications/0029211386?origin=resultslist (accessed on 10 April 2026).
- Rauschnabel, P.A.; Felix, R.; Hinsch, C.; Shahab, H.; Alt, F. What Is XR? Towards a Framework for Augmented and Virtual Reality. Comput. Hum. Behav. 2022, 133, 107289. [Google Scholar] [CrossRef]
- Microsoft. HoloLens 2 Hardware. Available online: https://learn.microsoft.com/en-us/hololens/ (accessed on 10 April 2026).
- Magic Leap. Welcome to Magic Leap 2. Available online: https://developer-docs.magicleap.cloud/docs/guides/ml2-overview/ (accessed on 10 April 2026).
- Maia Avelino, R.; Yang, W.; Weichbrodt, A.; Ochsendorf, J.; Flatt, R.J. Augmented Reality for Structural Inspection of Historic Monuments: The Case of Lausanne Cathedral. Int. J. Archit. Herit. 2025, 1–16. [Google Scholar] [CrossRef]
- Patankar, Y.; Tennenini, C.; Bischof, R.; Khatri, I.; Maia Avelino, R.; Yang, W.; Mahamaliyev, N.; Scotto, F.; Mitterberger, D.; Bickel, B.; Girardet, F.; Amsler, C.; Bomou, B.; Flatt, R.J. Heritage++: A Spatial Computing Approach to Heritage Conservation. RILEM Tech. Lett. 2025, 9, 50–60. [Google Scholar] [CrossRef]
- Savini, F.; Castiglia, M.; Gargaro, D.; Trizio, I.; Fabbrocino, G. Mixed Reality Procedures for the Maintenance of Existing Bridges and Retaining Walls. ce/papers 2023, 6, 1382–1390. [Google Scholar] [CrossRef]
- Parente, M.; Bruno, N.; Ottoni, F. HBIM and Information Management for Knowledge and Conservation of Architectural Heritage: A Review. Heritage 2025, 8, 306. [Google Scholar] [CrossRef]
- El-Alailyi, A.; Mazzacca, G.; Alami, A.; Padkan, N.; Takhtkeshha, N.; Fassi, F.; Remondino, F. 2D and 3D Semantic Segmentation for Interpreting and Understanding 3D Heritage Spaces. In Digital Heritage 2025; Campana, S., Ferdani, D., Graf, H., Guidi, G., Hegarty, Z., Pescarin, S., Remondino, F., Eds.; The Eurographics Association, 2025. [Google Scholar] [CrossRef]
- Sanz-Honrado, P.; Santamaría-Maestro, R.; Sánchez-Aparicio, L. J. Seg4D: An Open-Source Solution for Supporting the Diagnosis of Historic Constructions Using 3D Point Clouds — A Case Study Application. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2026, XLVIII-2/W12-2026, 439–446. [Google Scholar] [CrossRef]
- Liu, Z.; Blut, C.; Blankenbach, J. Synergizing Natural Visual Features and 3D Building Models for Robust Indoor Localisation in Mixed Reality Environments. Geo-Spat. Inf. Sci. 2025, 28, 3152–3177. [Google Scholar] [CrossRef]
- Teruggi, S.; Fassi, F. HoloLens 2 Spatial Mapping Capabilities in Vast Monumental Heritage Environments. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2022, XLVI-2/W1-2022, 489–496. [Google Scholar] [CrossRef]
- Teruggi, S.; Fassi, F. Mixed Reality Content Alignment in Monumental Environments. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2022, XLIII-B2-2022, 901–908. [Google Scholar] [CrossRef]
- Teruggi, S.; Fassi, F. Machines Learning for Mixed Reality: The Milan Cathedral from Survey to Holograms. In Augmented Reality, Virtual Reality, and Computer Graphics; Springer: Cham, Switzerland, 2021. [Google Scholar] [CrossRef]
- Teruggi, S.; Grilli, E.; Fassi, F.; Remondino, F. 3D Surveying, Semantic Enrichment and Virtual Access of Large Cultural Heritage. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, VIII-M-1-2021, 155–162. [Google Scholar] [CrossRef]
- Teruggi, S. Mixed Reality Applications as Support for the Maintenance of Monumental Architectures . Ph.D. Thesis, Politecnico di Milano, Milan, Italy, 30 September 2024. Available online: https://hdl.handle.net/10589/225352 (accessed on 13 July 2026).
- Fassi, F.; Achille, C.; Fregonese, L. Surveying and Modelling the Main Spire of Milan Cathedral Using Multiple Data Sources. Photogramm. Rec. 2011, 26, 462–487. [Google Scholar] [CrossRef]
- Achille, C.; Fassi, F.; Fregonese, L. 4 Years History: From 2D to BIM for CH—The Main Spire on Milan Cathedral. In Proceedings of the 18th International Conference on Virtual Systems and Multimedia (VSMM 2012), Milan, Italy, 2–5 September 2012; pp. 377–382. [Google Scholar] [CrossRef]
- Veneranda Fabbrica del Duomo di Milano. Available online: https://www.duomomilano.it/en/about-us/veneranda-fabbrica-duomo/ (accessed on 13 July 2026).
- Fassi, F.; Parri, S. Complex Architecture in 3D: From Survey to Web. Int. J. Herit. Digit. Era 2012, 1, 379–398. [Google Scholar] [CrossRef]
- Fassi, F.; Achille, C.; Mandelli, A.; Rechichi, F.; Parri, S. A New Idea of BIM System for Visualisation, Web Sharing and Using Huge Complex 3D Models for Facility Management. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2015, XL-5/W4, 359–366. [Google Scholar] [CrossRef]
- Fassi, F.; Mandelli, A.; Teruggi, S.; Rechichi, F.; Fiorillo, F.; Achille, C. VR for Cultural Heritage: A VR-Web-BIM for the Future Maintenance of Milan’s Cathedral. In Digital Heritage. Progress in Cultural Heritage: Documentation, Preservation, and Protection; Springer: Cham, Switzerland, 2016; pp. 139–157. [Google Scholar] [CrossRef]
- Achille, C.; Fassi, F.; Mandelli, A.; Perfetti, L.; Rechichi, F.; Teruggi, S. From a Traditional to a Digital Site: 2008–2019, the History of Milan Cathedral Surveys. In Research for Development; Springer: Cham, Switzerland, 2020; pp. 331–341. [Google Scholar] [CrossRef]
- Teruggi, S.; Grilli, E.; Russo, M.; Fassi, F.; Remondino, F. A Hierarchical Machine Learning Approach for Multi-Level and Multi-Resolution 3D Point Cloud Classification. Remote Sens. 2020, 12, 2598. [Google Scholar] [CrossRef]
- Spettu, F.; Teruggi, S.; Canali, F.; Achille, C.; Fassi, F. A Hybrid Model for the Reverse Engineering of the Milan Cathedral: Challenges and Lessons Learnt. Proceedings of ARQUEOLÓGICA 2.0—9th International Congress & 3rd GEORES—GEOmatics and pREServation, Valencia, Spain, 26–28 April 2021; pp. 281–291. [Google Scholar]
- Ungureanu, D.; Bogo, F.; Galliani, S.; Sama, P.; Duan, X.; Meekhof, C.; Stühmer, J.; Cashman, T.J.; Tekin, B.; Schönberger, J.L.; Olszta, P.; Pollefeys, M. HoloLens 2 Research Mode as a Tool for Computer Vision Research. arXiv 2020, arXiv:2008.11239. [Google Scholar]
- Hübner, P.; Landgraf, S.; Weinmann, M.; Wursthorn, S. Evaluation of the Microsoft HoloLens for the Mapping of Indoor Building Environments. In Proceedings of the Dreiländertagung der DGPF, der OVG und der SGPF; Vienna, Austria, 2019; pp. 44–53. [Google Scholar]
- Navares-Vázquez, J.C.; Qiu, Z.; Arias, P.; Balado, J. HoloLens 2 Performance Analysis for Indoor/Outdoor 3D Mapping. J. Build. Eng. 2025, 108, 112826. [Google Scholar] [CrossRef]
- Microsoft. WorldLockingTools-Unity Releases. GitHub. 2022. Available online: https://github.com/microsoft/MixedReality-WorldLockingTools-Unity/releases (accessed on 30 June 2025).
- Hu, T.; Du, T.; Qu, Z.; Gorlatova, M. XR Reality Check: What Commercial Devices Deliver for Spatial Tracking. In Proceedings of the 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), Daejeon, Republic of Korea, 8–12 October 2025; pp. 932–942. [Google Scholar] [CrossRef]
Figure 1.
Spaces of varying sizes and exposure conditions at Milan Cathedral.

Figure 2.
Methodological approach to the research project: from the digital survey of physical space to its interaction in an augmented reality setting.
Figure 2.
Methodological approach to the research project: from the digital survey of physical space to its interaction in an augmented reality setting.

Figure 3.
Comparison between manually modelled NURBS representations and reality-based point-cloud models (not semantic) in terms of processing effort, metric content, semantics, and data size.
Figure 3.
Comparison between manually modelled NURBS representations and reality-based point-cloud models (not semantic) in terms of processing effort, metric content, semantics, and data size.

Figure 4.
Relation between semantic classes and classification levels. Relationship among the hierarchical subdivision of Milan Cathedral, the MLMR classification levels, point-cloud resolution, and representation scale. Adapted and further developed from Teruggi [27].
Figure 4.
Relation between semantic classes and classification levels. Relationship among the hierarchical subdivision of Milan Cathedral, the MLMR classification levels, point-cloud resolution, and representation scale. Adapted and further developed from Teruggi [27].

Figure 5.
. Progressive selection and multi-scale exploration of the semantically structured point-cloud model within the MR application.
Figure 5.
. Progressive selection and multi-scale exploration of the semantically structured point-cloud model within the MR application.

Figure 6.
. Integration of MLMR classification with the hierarchical data organization adopted in the MR application: (a) progressive semantic classification of the Milan Cathedral point cloud, adapted from Teruggi et al. [35]; (b) translation of this hierarchy into the Unity scene structure, adapted from Teruggi [27].
Figure 6.
. Integration of MLMR classification with the hierarchical data organization adopted in the MR application: (a) progressive semantic classification of the Milan Cathedral point cloud, adapted from Teruggi et al. [35]; (b) translation of this hierarchy into the Unity scene structure, adapted from Teruggi [27].

Figure 7.
On-site visualization of the Milan Cathedral point cloud model spatially registered to the physical environment via the MR application.
Figure 7.
On-site visualization of the Milan Cathedral point cloud model spatially registered to the physical environment via the MR application.

Figure 8.
Near interaction paradigms with holographic content: (a) near touch; (b) button press.

Figure 9.
Area selection process: (a) whole Cathedral, (b) split in “area”, (c) split in “zone”, (d) split in “sector”.
Figure 9.
Area selection process: (a) whole Cathedral, (b) split in “area”, (c) split in “zone”, (d) split in “sector”.

Figure 10.
New intervention creation: (a) entering maintenance information through dictation; (b) intervention visualization: data can be visualized referenced in the real
Figure 10.
New intervention creation: (a) entering maintenance information through dictation; (b) intervention visualization: data can be visualized referenced in the real

Figure 10.
The proposed methodological approach.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.