Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Computer Vision and Graphics

Yu Jiao

,

Tingting Shi

,

Bing Zhao

,

Ao Wang

Abstract: Assessing scale mismatch, visual obstruction, and circulation conflict from two-dimensional drawings remains difficult in public art design. Finally, a VR simulation framework is developed, which integrates real scale calibration, space drawing, synchronous behavior trajectory, and dynamic scene analysis. Continuous interaction records were converted into behavioral, semantic, and pedestrian envelope. In a 10-week controlled trial, 180 students were enrolled in the study, with class level stratification, standardized familiarisation, matching tasks, and blind scores. About 4.3 million interaction records supported comparisons with single-view, multi-view, and distance-aware baselines. The proposed approach achieves 4.4% visual area error, 0.903 occlusion-recognition F1, and 0.872 passage-collision F1, and 28.6 ms for each key point. Ablative analysis has demonstrated the contribution of trajectory correction, adaptive weighting, semantic occlusion, and passage-envelope modeling. Compared with the conventional design, the VR group had lower elevation, better visual field visibility, fewer circulation conflicts, and higher site suitability scores, while the viewpoint coverage mediated 26.1% of the overall effect.

Review
Computer Science and Mathematics
Computer Vision and Graphics

Xi Jiang

,

Bingzhang Hu

,

Yunkang Cao

,

Jinbao Wang

,

Feng Zheng

Abstract: Industrial anomaly detection supports automated quality inspection. Traditionally, each product has required its own detector, trained on images without defects and evaluated on benchmarks with fixed classes. Recent studies use vision encoders pretrained on web data, multimodal large language models, expert mixtures, and generative models to relax this setup. They transfer visual features, add reasoning through language, share models across categories and modalities, or generate abnormal training data. Existing taxonomies centered on reconstruction, embedding, distillation, and memory do not describe these developments well. We therefore organize the literature published over the past three years according to five sources of information that replace training on normal samples for each category, namely visual, reasoning, geometric and multimodal, universal, and synthesis priors. Reported results show progress on curated benchmarks, but evidence remains limited under domain shift, protocols for open world settings, and operating points with low false positive rates. We also analyze recent benchmarks and recurring deployment bottlenecks. Industrial video grounded in physical dynamics provides an emerging direction beyond inspection of single frames. The evidence points to a change in research focus, although current methods do not yet provide a general replacement for separate training on normal data for each product in deployed inspection.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Dong Zhou

,

Zhe Huang

,

Haiyang Li

,

Mengyun Cao

,

Chih-Cheng Chen

Abstract: Vision-based 3D sensing provides an effective approach for constructing human-centric digital twins in applications such as ergonomic assessment and human–robot collaboration. However, reconstructing reliable dynamic human digital twins from visual sensor observations remains challenging due to large articulated deformation, which often causes surface discontinuities in existing 3D Gaussian Splatting (3DGS)-based human avatars. Although current methods achieve high image-level rendering quality, their Gaussian primitives are not explicitly adapted to pose-induced surface deformation, resulting in cracks and holes during motion. In this work, we propose ElasticGS, a geometry-consistent human digital twin reconstruction framework for vision-based 3D sensing. ElasticGS introduces a training-free Pose-Aware Dynamic Covariance Adaptation strategy, which dynamically adjusts the rotation and anisotropic scale of Gaussian primitives according to local human surface deformation. Specifically, localized Principal Component Analysis (PCA) is performed on deformed SMPL-X mesh neighborhoods to estimate surface stretching directions and magnitudes, enabling Gaussian representations to maintain continuous surface coverage under complex poses. Furthermore, we propose the Avatar Surface Integrity Score (ASIS), a reference-free metric that evaluates structural completeness and surface compactness of reconstructed human avatars without requiring pixel-aligned ground truth. Experiments on the X-Humans and AvatarReX datasets demonstrate that ElasticGS significantly improves surface integrity while maintaining competitive image-level sensing consistency. Compared with ExAvatar, ElasticGS improves ASIS from 0.642 to 0.846 on X-Humans and from 0.603 to 0.688 on AvatarReX. These results demonstrate the effectiveness of pose-aware covariance adaptation for robust human digital twin reconstruction in vision-based 3D sensing systems.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Giorgio Nordo

,

Ulderico Wanderlingh

,

Lorenzo Affé

,

Saeid Jafari

Abstract: Fluorescence time–lapse microscopy is a fundamental tool for investigating dynamic cellular processes, yet automatic cell detection and temporal signal extraction remain strongly dependent on intensity–driven thresholds and rigid decision rules. Such approaches may become unstable in the presence of noise, photobleaching, overlapping structures, and intrinsic biological variability. In this work, we introduce an uncertainty–aware computational framework that preserves the classical analysis workflow—temporal projection, peak detection, and region–of–interest (ROI) signal extraction—while extending it through neutrosophic morphological enhancement. The proposed implementation emphasizes practical reproducibility and flexibility: the analysis pipeline can automatically operate on heterogeneous inputs, including single images, temporal projections, full 3D stacks, or image sequences, while maintaining consistent detection parameters across classical and neutrosophic processing. By spatially flattening the image stack and enhancing the resulting activity landscape through neutrosophic morphology, the framework explicitly models reliable information, ambiguity, and background contributions, allowing weak but spatially coherent structures to emerge as stable candidates for detection. Beyond detection, the method extracts temporal ROI signals and provides detailed visualization and computational reporting, enabling a deeper interpretation of dynamic cellular behaviour. The framework remains fully compatible with standard fluorescence analysis pipelines while improving robustness, interpretability, and reproducibility under challenging imaging conditions, offering a principled extension for analysing uncertain or borderline cellular signals in fluorescence time–lapse microscopy.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Parsa Hassani Shariat Panahi

,

Amir Hossein Jalilvand

,

M. Hassan Najafi

Abstract: Line-segment detection is fundamental to robotics, autonomous navigation, and industrial inspection. While transformer-based detectors achieve the highest accuracy, their deployment on microcontrollers remains impractical. The STM32N6, with its Neural-ART NPU, promises to enable deep vision at the extreme edge. However, existing detectors rely on attention, grid-sampling, and normalization, operators unsupported by the convolution-oriented NPU. This architectural mismatch is characterized operator by operator: these operators lack accelerator primitives, and the decoder's self-attention alone materializes a 39 MB tensor exceeding on-chip memory. To address this, NPLSD is introduced as a pair of NPU-compatible line-segment detectors. NPLSD-H retains the HGNetv2 backbone of LINEA and replaces the transformer head with a fully-convolutional design. NPLSD-M adapts the M-LSD-tiny trunk to the supported operator set. Warm-started from ImageNet and trained on Wireframe, NPLSD-H reaches sAP10 = 37.9 (35.9 int8); NPLSD-M reaches 41.9 (41.1 int8) with 0.62M parameters. A controlled ablation isolates the trunk as the only variable, and initialization alone accounts for 4.6 points.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Yingcai Wan

,

Baoyu Wang

,

Jiqian Xu

,

Gao Yue

,

Huaizhen Wang

Abstract: Monocular Gaussian SLAM must recover camera motion, surface structure, and appearance from an RGB sequence without metric depth input or benchmark geometry during reconstruction. This setting is challenged by scale-ambiguous predictions, spatially varying reliability, and the tendency of an unconstrained Gaussian map to absorb pose and depth errors into its geometric and appearance parameters. We present Mono3DGS-SLAM, a reliability-guided Gaussian–TSDF framework with a sequence-precalibrated prediction stage and an incremental mapping backend. Deterministic reliability gates select admissible observations, while one frozen scale jointly transports predicted depth, camera translation, and length-valued mapping parameters into an internally consistent canonical coordinate. A confidence-weighted colorized TSDF stores the persistent surface, base color, and visibility support. A capacity-bounded Gaussian layer is restricted to reliable residual regions and primarily restores appearance detail over the fused surface. On Replica, Mono3DGS-SLAM achieves 0.249 cm mean ATE and is most accurate in five of eight scenes. With identical poses, depths, and views, the hybrid output reaches 30.678 dB PSNR, exceeding either component alone. ScanNet and TUM-RGBD results further assess transfer to real monocular sequences.

Review
Computer Science and Mathematics
Computer Vision and Graphics

Beizhen Zhao

,

Boyi Fu

,

Sicheng Yu

,

Zijian Wang

,

Kaiyong Zhao

,

Pengcheng Wu

,

Hao Wang

,

Steven Hoi

,

Hui Xiong

Abstract: 3D scene reconstruction and novel view synthesis have undergone a paradigm shift with the advent of 3D Gaussian Splatting (3DGS). Unlike computationally intensive volumetric rendering, 3DGS leverages a point-based representation with a differentiable tile-based rasterizer. This elegant synthesis achieves state-of-the-art visual fidelity at real-time rendering frame rates. As the 3DGS literature grows rapidly, existing surveys have primarily organized the field around downstream applications. In contrast, this survey provides a rasterization-centric analysis, treating the differentiable rasterization pipeline as the core computational engine of 3DGS. We begin by delineating the mathematical foundations of Gaussian splatting and tracing its lineage from classical volume rendering. Subsequently, we propose a fine-grained taxonomy that categorizes the literature across three hierarchical dimensions: (1) representation and optimization, (2) rasterization pipeline innovations, and (3) scenario-driven rasterization extensions. By deconstructing these algorithmic advances in rasterization, we offer quantitative insights into hardware-algorithm co-design and outline critical trajectories for future research. To support the community, we maintain a continually updated repository of relevant literature and open-source implementations at https://github.com/3DAgentWorld/Advanced3DGS.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Yingcai Wan

,

Huaizhen Wang

,

Jiqian Xu

,

Yue Gao

,

Baoyu Wang

Abstract: Monocular reconstruction under severe low light is ill-posed because underexposure removes repeatable texture, view-dependent camera processing violates photometric consistency, and pose or depth errors can be absorbed as blurred or distorted Gaussian primitives. We propose LL-3DGS Reconstruction, a coupled geometric–photometric framework that estimates a canonical normally illuminated Gaussian scene while retaining raw low-light observations as explicit evidence. Its geometry front end preserves view identity and order, selects a raw or fixed enhanced stream for learned multi-view inference, and converts the estimated cameras, depth, and points into a density-aware Gaussian seed with optional Sim(3) co-registration. The Gaussian back end constrains one shared 3D field through a bounded luminance restoration path and a low-light re-degradation path, preventing independent per-view corrections from becoming inconsistent scene content. Dark-region weighting, reliability-aware supervision, staged activation, bounded densification, and scale–anisotropy constraints further regularize flexible Gaussian primitives against noise, illumination errors, and weak geometry. Finally, co-registered metric depth supplies surface geometry while the canonical Gaussian rendering supplies restored color for colored-point and TSDF fusion. With paired normal-light supervision, the method obtains 25.44 dB PSNR, 0.898 SSIM, and 0.247 LPIPS on five LOM scenes and averages 21.195 dB, 0.449 SSIM, and 0.523 LPIPS across eight LLRS scenes. A controlled LOM-sofa ablation shows that low-light consistency supports observation explainability, while four-scene TSDF outputs document the explicit reconstruction path without claiming geometric accuracy in the absence of 3D ground truth.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Pengyue Jia

,

Song Gao

,

Sharon Li

,

Xiangyu Zhao

Abstract: Worldwide image geolocalization aims to predict the geographic location of an image taken anywhere on Earth, expressed as GPS coordinates, a geographic cell, or an administrative region. The task is open-world by nature: no reference collection provides complete imagery coverage of the planet, so a model must generalize to locations it has never observed. Classical methods pursue this generalization statistically, learning visual-to-geographic mappings with classification heads, retrieval embeddings, or continuous probabilistic models. More recently, foundation models have opened a second route based on knowledge-driven reasoning, in which the location answer is generated from world knowledge internalized during pretraining, ranging from retrieval-augmented generation (RAG) to tool-using agents. This paradigm changes the mechanism by which location answers are produced. This survey provides a systematic review of worldwide image geolocalization with a focus on the foundation-model era. We introduce a two-level taxonomy that organizes classical paradigms by their output mechanism and foundation-model-era methods by the role the foundation model plays, covering RAG, reasoning, agentic, and hybrid designs. We further present a unified review of datasets and benchmarks, a cross-method comparison of reported results, and a dedicated discussion of privacy, fairness, and ethics. Finally, we outline open challenges and future directions. A continuously updated paper list is also available at https://github.com/Jia-py/Awesome-Worldwide-Image-Geolocalization.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Yufeng Dong

,

Minqing Zhang

,

Chao Jiang

,

Shen Peng

Abstract: To address the limited channel robustness of invertible neural networks against Gaussian noise, JPEG compression, and scaling/resampling in image information hiding, while maintaining their embedding capacity, we propose an invertible steganography algorithm that integrates a dual-branch attention mechanism in the frequency domain. First, a dual-branch attention module is designed to adaptively generate embedding weights from both global subband energy and local texture pixel dimensions after wavelet multi-frequency decomposition, thereby constructing a four-channel extended secret payload. Second, a four-channel extension strategy is developed to enhance embedding capacity without introducing random noise distortion. Additionally, dual enhancement mechanisms are employed to preprocess the original carrier before embedding, optimize frequency-domain feature distribution, and restore images contaminated by noise, compression, or scaling attacks during reconstruction. Experimental results demonstrate that under a 1:1.34 embedding ratio, the proposed algorithm achieves a clean-channel C-PSNR of 36.27 dB; under Gaussian noise (σ=10) and JPEG(QF=50) attacks, PSNR values for all channels reach 27.96 dB and 29.43 dB respectively. Compared to comparable models without attention mechanisms, this approach improves carrier quality and secret recovery metrics by over 2 dB on average, combining high transmission capacity with excellent channel robustness and visual imperceptibility.

Brief Report
Computer Science and Mathematics
Computer Vision and Graphics

Enio Ibrišagić

,

Dražen Domijan

Abstract: We tested the ability of multimodal LLMs to reason about spatial relations using a maze-solving task. The task required the model to find and draw, on a supplied image of a maze, a path from the starting point to the exit without crossing the walls. We tested proprietary models Grok, Gemini-3, ChatGPT-4o, and ChatGPT-5.2 on 20 mazes (10 with a path from the start to the exit and 10 without one). Each problem was presented five times in random order, resulting in a total of 100 trials per model. We also tested the effect of language by presenting separate prompts in English and Croatian. We found that only ChatGPT-5.2 was able to solve the task and draw the correct path when a solution existed. The other models made errors such as visual hallucinations (inventing a new, unrelated maze), failing to connect the starting point to the exit, crossing the walls, and walking on the walls. We also asked the LLMs to provide metacognitive judgments about their performance. Interestingly, ChatGPT-5.2 typically gave lower confidence estimates (around 95%) than the others, which were 100% confident they had correctly solved the task.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Chetanpal Singh

,

Santoso Wibowo

,

Srimannarayana Grandhi

,

Satria Mandala

Abstract: Early and accurate cancer detection from medical imaging remains challenging because clinically relevant evidence is distributed across local image appearance, structural relationships between suspicious regions, and ordered imaging context. This study proposes a graph-aware sequence-aware deep learning framework for cancer analysis from medical imaging, combining a convolutional neural network (CNN) backbone for spatial feature extraction, a graph attention network (GAT) for lesion-structure modelling, and a bidirectional long short-term memory (BiLSTM) module for ordered-view or slice-sequence representation learning. Multimodal fusion is evaluated in a leakage-controlled setting only for the RSNA mammography task, where the metadata branch is restricted to patient age and implant status, both available before diagnosis. In contrast, the primary LIDC-IDRI experiment is conducted as an image-only analysis because radiologist malignancy scores and semantic nodule attributes are annotation-derived variables and are not treated as independent clinical predictors. The framework is evaluated on the RSNA Breast Cancer Detection dataset and the LIDC-IDRI lung CT dataset using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC). Additional ablation experiments assess the contribution of graph learning, sequence-aware modelling, and leakage-safe metadata fusion, while SHAP analysis is used to quantify the influence of the included RSNA metadata variables on the final multimodal predictions. The proposed framework provides an interpretable and leakage-aware architecture for integrating spatial, relational, and ordered imaging information, and demonstrates how limited pre-diagnostic metadata can be incorporated without overstating multimodal novelty or clinical realism.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Reka Sandaruwan Gallena Watthage

,

Anil Fernando

Abstract: Holographic video is widely regarded as the next frontier of three-dimensional visual communication, promising glasses-free, full-parallax imagery for telepresence, medical imaging, immersive education, and next-generation AR/VR. Its deployment, however, is presently gated by an unresolved compression bottleneck: a single complex-valued computer-generated hologram (CGH) can exceed several gigabytes, and existing image and video codecs encode statistical assumptions that are systematically violated by interferometric wavefront data.(1) Background: standardised block-based codecs (H.265/HEVC, H.266/VVC) rely on local motion compensation and piecewise-smooth residual priors that fail on holograms, whose fringe patterns exhibit aperture-spanning correlations, wrapped-phase periodicity, and non-local motion coupling. Existing learned hologram codecs (NHVC, HiFiHC, HoloZip) address a subset of these issues but do not simultaneously exploit temporal redundancy in the fringe domain, respect the algebraic structure of C, capture aperture-scale correlations, or optimise a holography-specific perceptual metric.(2) Methods: we propose HoloWave-NVC, an end-to-end learned neural video codec composed of four novel modules: a Dual-Branch Complex Analysis Encoder with amplitude-phase cross-attention fusion built on complex-valued convolutions and a unit-modulus phase embedding; a Fringe-Domain Wavefront Motion Predictor that operates via a bank of learnable angular-spectrum propagation kernels, deformable complex convolutions, and a complex ConvGRU; a Sparse Swin-Attention Entropy Model with global-token shortcuts and frequency-position embeddings for aperture-spanning context; and a Complex-Valued Refinement Network for decoder-side speckle suppression trained jointly under a rate-distortion-perception Lagrangian regularised by a physics-informed angular-spectrum reconstruction loss and a wrapping-aware phase-consistency term. Perceptual quality is quantified by HoloPQN, the first differentiable holography-specific metric, calibrated to mean-opinion-score data from paired-comparison experiments on a custom near-eye holographic display prototype.(3) Results: we report a comprehensive experimental protocol on the JPEG Pleno Holography Common Test Conditions v3.0 dataset against HEVC, VVC, JPEG Pleno Holography (ISO/IEC 21794-5), NHVC, HiFiHC, HoloZip, and a specifically constructed DCVC-Complex baseline, together with an eight-configuration ablation matrix isolating each novel module. Projected BD-rate savings against HEVC are 58.4% on reconstruction-plane PSNR and 68.2% on HoloPQN; the codec is projected to reach subjective “good” quality (MOS 4) at approximately 0.35 bpp, a 4.6× perceptual bit-rate saving over HiFiHC, and to achieve a decoder latency of 13.9 ms, meeting the 16 ms real-time threshold for 60-fps interactive streaming. The largest single ablation gain, approximately 35 percentage points of BD-rate, comes from the fringe-domain motion predictor, falsifying the received wisdom that hologram video is a temporal-redundancy-free medium.(4) Conclusions: HoloWave-NVC directly addresses six research gaps identified in the current holographic-coding literature, establishing a wave-optically principled path toward practical real-time holographic video streaming over 5G and 6G networks. The projected performance targets constitute design goals for the 36-month research programme this article initiates rather than measured outcomes; the release of models, synthetic training corpora, and subjective study data will accompany the primary follow-up publication.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Zhe Geng

,

Linyi Wu

,

Minjie Sun

,

Yu Zhang

,

Yuan Meng

,

Lujia Yao

,

Daiyin Zhu

Abstract: In slow-time colorized subaperture image (CSI), anisotropic targets that reflect strongly when viewed from specific angles appear in vivid colors, which makes them stand out against isotropic background that reflects energy uniformly across all angles. It leads to more accurate annotation labels for ships, vehicles, and airplanes in SAR images and better SAR automatic target detection (ATD) performance. Unfortunately, although many port-related CSI products collected by satellite-borne SAR systems are released for free public-access and could be leveraged for ship detection research, those could support vehicle and airplane detection are rare. To investigate performance improvement in deep-learning based SAR ATD that could be brought by colored SAR images, three novel SAR-ATD frameworks are proposed for ship, vehicle, and aircraft detection, respectively. 1) Context-Guided Ensemble Learning (CGEL) is proposed for ship detection, where state-of-the-art high-resolution colorized spotlight SAR images are exploited to enhance the visual features of ships and reduce false alarms, while the potential ship berthing/docking areas are delimited with adaptive intensity shading (AIS). 2) Context-Driven SAR image Recoloring and Enhancement Mechanism (CD-SAR-REM) is proposed to generate a context-driven color-enhanced version of the original SAR image based on AIS so that potential parking regions are highlighted. 3) Color-feature-aided aircraft detection. In case that CSI products are unavailable, pseudo-color SAR images are generated based on phase congruency and the contextual information extracted by the segmentation module is used to refine the initial predictions generated by the core detection network. Experimental results show that the performance of the proposed context-driven ship, vehicle, and aircraft detection methods based on colored SAR images are superior to many state-of-the-art SAR ATD models.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Hongbo Lu

,

Liang Yao

,

Chenghao He

,

Hao Han

,

Fan Liu

,

Wenlong Liao

,

Tao He

,

Pai Peng

Abstract: World Action Models (WAMs) improve planning by incorporating future world evolution into action generation, yet existing methods allocate a fixed imagination budget to every scene. We propose RISE (Refining Imagination through SElective Rollout), a system-level adaptive imagination framework that makes sequential ROLL/STOP decisions according to the expected planning benefit of continued rollout. At each step, a Latent Evaluator estimates the risk revealed by the current prefix and how much planning could improve if imagination continues, while a Rollout Gate weighs this expected benefit against additional computation cost. Since factual driving logs expose only one realized future, we further construct CounterDrive, a counterfactual dataset with diverse outcomes and risk levels, to enrich future dynamics and provide localized risk supervision. Each retained sample undergoes expert verification and annotation of trajectory validity, incident onset, and causal category, providing a reusable resource for safety-critical world-modeling research. Experiments on NAVSIM and nuScenes show that RISE achieves the best overall planning performance while reducing unnecessary rollout, with additional transfer results supporting its plug-in generality across WAM architectures.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Lingxin Xu

,

Long Wang

,

Qingyun Zuo

,

Jinzhi Zhang

,

Xiaomeng Cui

,

Haisu Zhang

Abstract: Accurate road graph extraction from satellite imagery is essential for large-scale mapping and geospatial analysis. Recent one-shot graph extraction frameworks based on foundation models have achieved promising performance, but their effectiveness decreases in complex environments where road structures are partially obscured by vegetation, buildings, shadows, and other surface conditions. These occlusion-induced disturbances lead to incomplete connectivity and degraded topology reconstruction, particularly under out-of-domain scenarios. This study proposes an occlusion-aware refinement framework to improve the robustness of satellite image road graph extraction while maintaining the original backbone architecture. The proposed framework introduces three complementary strategies: synthetic occlusion augmentation for explicit occlusion-aware representation learning, an occlusion-adaptive extended-line strategy with hard-mining topology optimization for improved connectivity reasoning, and an occlusion-adaptive node-guided resampling mechanism for reliable graph node localization. Experiments conducted on the Global-Scale road graph extraction benchmark demonstrate that the proposed method consistently improves topology reconstruction performance. Compared with the reproduced SAM-Road++ baseline, the proposed framework improves TOPO F1 from 61.81 to 62.56 on the in-domain split and from 46.93 to 51.51 on the out-of-domain split. Furthermore, the ID-OOD performance gap is reduced from 14.88 to 11.05, indicating enhanced robustness under unseen geographic conditions. The results demonstrate that explicitly modeling occlusion as a structured factor can effectively improve the generalization capability of satellite road graph extraction systems.

Review
Computer Science and Mathematics
Computer Vision and Graphics

Ridwan Salahudeen

,

Mathias Fonkam

,

Narasimha Rao Vajjhala

,

Sahalu Balarabe Junaidu

,

Aliyu Garba

Abstract: Talking face generation (TFG), the synthesis of photorealistic speaking video from a portrait and an audio signal, has moved through five architectural generations since 2016: Long Short-Term Memory (LSTM) lip-sync systems, Generative Adversarial Networks (GANs), Neural Radiance Fields (NeRFs), Diffusion models, and now Diffusion Transformers (DiT) and Gaussian Splatting. With real-time photorealism, TFG is entering healthcare, education, identity management, and public communication. This systematic literature review synthesises 92 primary studies (January 2016 to December 2025), selected from 18,343 records across six repositories under Kitchenham’s SLR guidelines and PRISMA 2020, to answer six research questions on architecture evolution, dataset bias, evaluation metrics, generative paradigm shifts, ethical safeguards, and socio-technical deployment. DiT and Gaussian Splatting together account for roughly 36% of the most recent studies retrieved and GAN usage has fallen below 3%. Training datasets are demographically skewed, and the dominant metrics (PSNR, SSIM, SyncNet) do not measure what deployment requires. Most seriously, only one of the 92 systems embeds any technical safeguard (a non-compliant post-hoc watermark), even though the EU AI Act (Regulation 2024/1689), the US TAKE IT DOWN Act (2025), and Coalition for Content Provenance and Authenticity (C2PA) v2.0 entered enforcement during the review period. None integrates C2PA-compliant watermarking, consent, or provenance mechanisms. We propose three deployment-fitness metrics: Perceived Trustworthiness Score (PTS), Cross-Cultural Authenticity Index (CCAI), and Deepfake Detectability Rate (DDR); and the Responsible Talking Face Generation in Socio-Technical Systems (RTFG-STS) Framework, grounded in AI4People principles and socio-technical systems theory, with five operational mechanisms aligned to EU AI Act Article 14.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Yu Jiao

,

Tingting Shi

,

Bing Zhao

,

Ao Wang

Abstract: The evaluation of 3D forms in VR requires continuous process level computation, whereas traditional art training systems retain the final art and subjective scores. This paper presents a VR spatial-sketching platform, which integrates 60 Hz head-mount display posture, dual controller trajectory, stroke event, coordinate registration, state awareness filtering, multi-view coverage estimation, and geometric feature extraction. The trajectory of the algorithm is based on a three sample adaptive window, a low speed threshold of 0.03 m/s, a break point of 120 ms, and a break point of 80 mm. The feedback is updated every 2 s based on viewpoint-coverage entropy, the content ratio error, the center offset, and the interruption cost. A 12-week quasi-experiment involved 216 students, 1 296 works, 7 776 orthographic images, and approximately 7.2 million interaction records. The results of the ablation were 4.9 mm, the turning retention rate was 95.1%, the invalid trajectory was 93.6%, the error was 8.9%, and the delay was 42.6 ms. The results of the experiment were 85.4 ± 5.8 and 79.7 ± 6.4 in the control group, which confirmed the accuracy, continuity, interpretation, and adaptability of VR assessment.

Article
Computer Science and Mathematics
Computer Vision and Graphics

Jian Deng

,

Tao Wang

Abstract: Facial aging is not merely a linear degeneration of the epidermis, but rather a complex, multilevel biomechanical cascade involving skeletal remodeling, dynamic redistribution of fat compartment volumes, and degradation of the skin matrix [1]. Modern anatomical studies have confirmed that subcutaneous fat is precisely divided into superficial and deep fat compartments with distinct boundaries and different aging trajectories, the atrophy of the deep fat compartments, combined with the displacement of the superficial fat compartments, leads to a significant reduction in facial three-dimensional volume [3]. At the same time, the functional decline of facial supporting ligaments is closely related to changes in the extracellular matrix's molecular structure, this reduction in tissue stiffness and tensile strength provides the pathophysiological basis for the transformation of dynamic wrinkles into static wrinkles [17]. In the field of computer vision, facial aging analysis primarily focuses on either inferring an individual’s physiological age from image sequences or enabling cross-age identity recognition, these technologies have broad application value in public security surveillance and digital forensics [7]. With the introduction of deep learning technologies, convolutional neural networks (CNNs) have become the mainstream analytical method in this field due to their exceptional feature extraction capabilities [8]. Addressing the significant variability in individual aging patterns, research has proposed models that incorporate attention mechanisms to achieve precise modeling of aging trajectories [11]. Generative adversarial networks (GANs), through age-conditional generation techniques, have successfully achieved age transformation while preserving identity features, effectively resolving the challenge of maintaining gender and racial consistency that traditional methods struggle with [14]. In age estimation tasks, ordinal regression and label distribution learning have further improved prediction accuracy by capturing the temporal order of age labels [13]. The core paradigm of cross-age face recognition is feature disentanglement, which aims to decompose facial representations into identity-inherent and age-related components, thereby minimizing the interference of age-related changes on identity recognition performance [12]. Despite significant progress, the field still faces severe challenges, such as insufficient data quality and diversity, as well as vast differences in individual aging patterns [7]. The “black-box” nature of deep neural networks results in a lack of sufficient theoretical explanation for the relationship between the learned aging representations and actual biological mechanisms, limiting their in-depth application in the biomedical field [15]. The scarcity of longitudinal paired data severely limits the models’ generalization ability, while the insufficient scale and diversity of mainstream datasets make overfitting a common issue [8]. Although the emerging line-scan confocal optical coherence tomography (OCT) technology can provide three-dimensional microscopic insights into the dermal fiber network, it still faces technical bottlenecks in aligning and standardizing multi-source data during clinical translation [16]. Future research must move beyond a mere race for performance and shift toward an in-depth exploration of the intrinsic structure and causal interpretability of aging characteristics. By integrating multi-source data, such as genomics, and establishing a new evaluation paradigm to assess the biological plausibility of models, thereby bridging the gap between deep network features and actual biological mechanisms of aging [37].

Review
Computer Science and Mathematics
Computer Vision and Graphics

Alexandros Mitsou

,

Evaggelos Spyrou

Abstract: Egocentric vision has become a central paradigm for studying human behaviour, interaction, and intent from a first-person perspective, supporting a growing range of tasks such as action recognition, temporal segmentation, and action anticipation. Progress in this area has been driven largely by the availability of public datasets, which differ substantially in sensing configurations, annotation strategies, application domains, and temporal structure. Despite the rapid expansion of available resources, the egocentric dataset landscape remains fragmented, with limited cross-dataset interoperability and uneven coverage across domains, modalities, and task formulations. This survey presents a comprehensive, dataset-centric analysis of the current egocentric and action-related video ecosystem. We curate and systematically analyse 75 publicly documented datasets spanning more than a decade of research, covering both egocentric and selected exocentric benchmarks that are widely used for action-centric and anticipatory modelling. To organise this landscape, we introduce a unified taxonomy structured around five complementary design dimensions: perspective, domain, sensing modality, annotation structure, and task support. This taxonomy enables consistent comparison across heterogeneous datasets and provides a principled framework for analysing how design choices influence supported tasks and evaluation practices. Beyond cataloguing datasets, we examine temporal trends, distributional patterns, and co-occurrence statistics across the curated corpus, revealing recurring biases toward specific domains, limited multimodal coverage, and a widespread reliance on post hoc constructions for action anticipation. We further discuss annotation heterogeneity, scale–diversity trade-offs, and the ethical and legal constraints that shape egocentric data collection and release. By consolidating dispersed resources into a coherent taxonomy and identifying systematic gaps in current benchmarks, this survey aims to support informed dataset selection, facilitate cross-dataset analysis, and guide future efforts in egocentric dataset design and benchmarking.

of 39