Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Signal Processing

Ioannis A. Bartsiokas

,

Maria-Lamprini A. Bartsioka

,

Panagiotis K. Gkonis

,

George Vardoulias

,

Anastasios Papazafeiropoulos

Abstract: Sixth-generation (6G) wireless systems target ubiquitous connectivity by integrating terrestrial (TNs) and non-terrestrial networks (NTN) cooperation, including low-Earth orbit (LEO) satellites, high-altitude platform stations (HAPS), and unmanned aerial vehicles (UAVs). In such a vertically stratified architecture characterized by massive multi connectivity, the selection of the most suitable tier or tier-combination for each mobile user is a complicated task, as the decision depends jointly on the propagation environment, user requirements, and the instantaneous characteristics of the topology. This paper proposes an altitude-aware Deep Learning (DL) framework that casts multi-tier network selection as a seven-class classification problem spanning standalone TN, LEO, HAPS, and UAV access as well as their TN-assisted multi-connectivity combinations. A fourteen-feature physics-based dataset is generated through extensive simulations based on standardized channel and geometry models, including 3GPP TR 38.901 terrestrial pathloss and line-of-sight probability, satellite-constellation elevation angles, and air-to-ground link geometry for HAPS and UAV. A deep neural network (DNN) is trained on the aforementioned dataset and assessed using stratified five-fold cross-validation. The proposed model achieves an overall classification accuracy of 88.1% with a macro-averaged F1-score of 0.88 and a single-sample inference latency of approximately 2 ms, well within typical 6G handover budgets. The results demonstrate that requirement- and geometry-aware tier selection can be performed accurately, enabling real-time decision-making in ambiguous multi-connectivity configurations for 6G networks.

Article
Computer Science and Mathematics
Signal Processing

Hiwa Asadpour

Abstract: Automatic speech recognition (ASR) evaluation for Garrusi, a Kurdish variety written in a Latin-based orthography, is complicated when an ASR system outputs Arabic script because direct scoring can confuse recognition errors with writing-system differences. This study evaluates MMS-1B-all with the Central Kurdish adapter on 1,722 speech segments from five Garrusi speakers (9,763 reference tokens; 117.9 minutes) without adaptation. A common-reference scoring design was used: the reference was folded once and kept fixed, while the hypothesis representation was changed. The unmodified Arabic-script output produced 111.70% WER and 100.92% CER. Transliteration to Latin reduced these to 102.36% and 57.89%, while additional orthographic folding reduced them to 97.45% and 51.13%. Thus, the full transformation reduced measured WER by 14.25 percentage points and CER by 49.79 points, but substantial error remained. A round-trip control showed that part of the remaining error comes from the transliteration and folding process itself, with 98.5% of its substitutions linked to two orthographic patterns. A separate comparison with a Southern Kurdish fine-tuned system showed a 10.69-point WER difference, which fell to 2.13 points when short i was treated as non-contrastive. These results show that orthographic representation can strongly affect ASR evaluation and that recognition performance should be assessed separately from differences introduced by the scoring pipeline.

Article
Computer Science and Mathematics
Signal Processing

Sunwoo Yeon

,

Jaeyong Kim

,

Hyeonjung Kim

,

Jihwan Won

,

Hyeonwoo Kim

,

Donggyu Sim

,

Cheolsoo Park

Abstract: Intelligent systems deployed in smart cities, smart grids, environmental monitoring, and other data-driven applications increasingly depend on reliable multivariate time-series forecasting. Recent deep forecasting models have achieved strong performance on various benchmarks, but their final predictions are often generated through a fixed forecast-generation process. This one-size-fits-all approach may be suboptimal because input windows exhibit different values, trends, and periodic patterns. We propose gap-aware residual selection (GARS), a plug-in refinement module for time-series forecasting that improves the prediction ability of an existing forecasting model rather than replacing it. GARS constructs deterministic reference forecasts from the observed input window and uses the differences between these forecasts and the initial forecast of the model as gap-aware residual information. Using only the observed input window and initial forecast, GARS constructs three component forecasts: the initial forecast, a gap-aware residual component that adjusts the initial forecast using these reference differences, and a direct residual component learned from the input window. The final forecast is obtained by assigning soft weights to the components across different segments of the forecast horizon rather than using fixed combinations. Future target values are used exclusively for training and evaluation and not for constructing reference forecasts, residuals, or soft weights. Experiments on five multivariate benchmark datasets related to energy, environment, electricity, traffic, and finance, using four representative forecasting models, show that GARS consistently improves average forecasting performance, reducing the mean squared error by 7.95% on average compared with the original models.

Article
Computer Science and Mathematics
Signal Processing

Samir Brahim Belhaouari

,

Yunis Carreon Kahalan

,

Ismael Belhaouari

,

Ejmen Al-Ubejdij

Abstract: Encoding a one-dimensional time series as a two-dimensional image and classifying it with a convolutional neural network is a popular and accurate paradigm, yet the standard encoders, Gramian Angular Fields, Markov Transition Fields, recurrence plots, spectrograms, scalograms, and polar or spiral embeddings, are validated almost exclusively on downstream accuracy. A second, practically essential property is almost never measured: whether a human can look at the produced image and tell the classes apart. We propose the Generalized Iterative Polar Transform (GIPT), a readable time-series imaging framework, and present a significant and novel study that takes a polar-spiral encoder reporting strong benchmark numbers yet producing humanly unreadable images, and asks whether classifier accuracy and visual legibility can be achieved together. We make five contributions. First, we give an explicit geometric diagnosis proving why the existing encoder is opaque: the time term dominates the radius and collapses the angle, so the trajectory occupies a vanishing wedge and the number of spiral turns is provably inert. Second, we derive GIPT, a principled readable encoding that swaps the polar axes, renders a filled silhouette, and uses an amplitude-aware normalization, and we prove a lobe-counting property that renders a frequency difference countable by eye. Third, we isolate a genuine readability-accuracy trade-off and show it is governed entirely by the normalization. Fourth, we show a striking and useful consequence: because GIPT enforces cross-sample consistency, a parameter-free one-shot nearest-neighbor classifier on the encoded images matches or exceeds the convolutional network at a fraction of the cost, eliminating the need for a trained model. Fifth, we propose a paired evaluation protocol that couples a stabilized multi-seed accuracy benchmark with a vision zero-shot legibility test and a human test-set study, turning legibility into a number on the same images. Across a broad suite of twenty-eight univariate UCR archive datasets the GIPT multi-channel encoder is competitive with the opaque baseline (76.0% vs 73.8% CNN accuracy), the training-free one-shot classifier on GIPT images reaches 83.1%, on par with the classical DTW nearest-neighbor standard while running about 1,200 times faster than training the network, and the human-facing rendering is genuinely eye-classifiable where the class difference is morphological, a property we verify with ablation, robustness, and statistical significance analyses. We conclude that readability and accuracy are not in conflict; they are reconciled by rendering the human-facing and network-facing images under different normalizations.

Article
Computer Science and Mathematics
Signal Processing

Jorge Ortiz Ceballos

,

Itzel Maria Abundez Barrera

,

Eréndira Rendón Lara

Abstract: Surface electromyography (sEMG) signals acquired with low-cost sensors tend to exhibit variable degradation that can compromise the reliability of myoelectric control systems outside controlled conditions. This work presents an adaptive processing pipeline for sEMG signals composed of three chained stages: a motor-intent classifier based on dilated temporal convolutional networks, whose function is to determine whether a signal window contains muscle activity associated with a voluntary gesture; a dual-output quality assessor that estimates a continuous score and a binary acceptability label, aimed at deciding whether the signal can be used directly or requires intervention; and a convolutional autoencoder that recovers the morphology of degraded windows before they are used in prosthetic control. The decision policy for reconstruction operates on two independent thresholds and classifies each window into one of three states: signal discard, direct acceptance, or active restoration. The models within the proposed architecture are trained on the public NinaPro DB1, DB3, and DB10 datasets and validated without recalibration on signals recorded from two participants with transtibial amputation over four weekly sessions. The results show that the pipeline correctly handles signal profiles with opposing characteristics. Inference latency remained below 5 ms at the 50th percentile across all scenarios, consistent with real-time operation. An integrated prototype using an Arduino UNO R4 WiFi was used to evaluate the results, and validate the feasibility of the system on physical hardware.

Review
Computer Science and Mathematics
Signal Processing

Arosh S. Perera Molligoda Arachchige

,

Fatemeh Darvizeh

Abstract: Photon-counting computed tomography (PCCT) represents a detector-level transformation in CT imaging. Unlike conventional energy-integrating detectors, photon-counting detectors directly convert individual X-ray interactions into electrical pulses and classify them according to energy. This architecture enables electronic-noise rejection, smaller detector pixels, improved geometric dose efficiency, and intrinsic spectral acquisition. However, the images available to radiologists are not produced directly by the detector; energy-resolved photon counts must first undergo calibration, correction, projection formation, reconstruction, and material decomposition. This narrative review provides an educational framework linking X-ray attenuation physics, detector materials and architectures, energy thresholds, and detector nonidealities to the resulting PCCT images. It describes conventional polyenergetic and ultra-high-resolution images, virtual monoenergetic imaging, iodine maps, virtual non-contrast imaging, calcium and bone subtraction, virtual non-calcium imaging, effective atomic number maps, electron-density maps, and emerging K-edge techniques. Particular emphasis is placed on the clinical purpose and limitations of each reconstruction, including noise, artifacts, partial-volume effects, misregistration, incomplete subtraction, calibration dependence, and limited cross-platform comparability. Practical considerations for protocol design, image selection, interpretation workflow, and spectral-data archiving are also discussed. Understanding the pathway from photon detection to image formation is essential for selecting the appropriate reconstruction, avoiding misinterpretation, and integrating PCCT effectively into clinical radiology.

Concept Paper
Computer Science and Mathematics
Signal Processing

Ramakrishna Sen

,

Dhruv Singh

,

Arun Sourie

,

M Kiran Reddy

Abstract: This paper proposes a switching mode beamforming (SMVB) technique for a two-microphone array setup in a multi-source acoustic environment with one dominant target and two interferers. The proposed method adaptively switches between a robust minimum variance beamformer (RMVB) and a linearly constrained minimum variance (LCMV) beamformer based on the eigenvalue characteristics of the spatial covariance matrix. The eigenvalue distribution is used to infer the signal environment and select the appropriate beamforming strategy. Simulation results show that the proposed SMVB method outperforms conventional MVDR, RMVB, and LCMV beamformers in terms of interference suppression and target signal preservation, achieving improved output signal-to-interference-plus-noise ratio (SINR) and overall signal quality.

Article
Computer Science and Mathematics
Signal Processing

Jakub Sobczyk

,

Krzysztof Fonał

Abstract: Always-on keyword spotting (KWS) on microcontrollers must reconcile high recognition accuracy with severe constraints on memory, computation, and energy. Most small-footprint KWS research targets English, and its conclusions remain largely unverified for typologically different, consonant-rich languages such as Polish. In this work, we present a complete, hardware-validated KWS solution for Polish on the RP2350 microcontroller (Raspberry Pi Pico 2). We design three compact architectures from the convolutional (CNN), convolutional-recurrent (CRNN), and depthwise-separable (DS-CNN) families and benchmark them against BC-ResNet, a state-of-the-art reference, on a 25-keyword Polish vocabulary using MFCC features. All models are evaluated in full precision (Float32) and after 8-bit integer (INT8) post-training quantization, with inference latency measured directly on the target hardware. BC-ResNet attains the highest full-precision accuracy (97.81%) but is the most fragile under quantization, whereas the proposed CRNN becomes the most accurate quantized model (94.28%), the DS-CNN the smallest (46.15 KB), and the CNN the fastest (107.5 ms); all quantized models meet a one-second real-time budget. We further show that the accuracy ranking inverts after quantization, that memory savings are highly architecture-dependent, and that phonetically similar Polish words are the dominant source of error. These results offer practical guidance for deploying small-footprint KWS in Polish and other under-represented languages.

Article
Computer Science and Mathematics
Signal Processing

Jianjun Zhong

,

Wei Zhang

,

Linzhao Hao

Abstract: Fine-grained and open-set fault diagnosis of analog circuits poses two challenges that generic deep classifiers fail to address. First, under a closed-set assumption, models assign unseen fault types to known classes with unwarranted confidence, lacking any mechanism to reject out-of-distribution inputs. Second, component manufacturing tolerances induce structured intra-class variation that causes the frequency responses of distinct fault modes to overlap, severely degrading fine-grained discrimination. This paper proposes a Tolerance-Aware Contrastive Siamese Network (TCSN), a metric-learning framework that constructs a discriminative embedding space jointly addressing both issues. The core contribution is the Tolerance-Aware Contrastive Loss (TACL), comprising a tolerance term \( (\mathcal{L}_{\mathrm{tol}}) \) that suppresses tolerance-induced intra-class scatter across Monte Carlo realizations, and a fine-grained term \( (\mathcal{L}_{\mathrm{fg}}) \) that enforces separation of near-identical overlapping classes. For open-set rejection, an energy-based scoring mechanism maps prototype distances into in-/out-of-distribution scores, identifying unknown faults without retraining. The framework is validated on a Sallen-Key second-order Butterworth low-pass filter benchmark (10 known and 2 unknown fault classes). It attains an open-set AUROC of 0.9309 and an FPR@95%TPR of 0.2500; for closed-set classification it reaches 90% accuracy (macro-F1: 0.88), and it retains 83.0% accuracy under one-shot conditions (n = 1). Under corrupted inputs it degrades gracefully, outperforming all baselines at 10 dB SNR. Ablations confirm the necessity of each component: removing \( \mathcal{L}_{\mathrm{tol}} \) drops the open-set AUROC by 0.3793, removing \( \mathcal{L}_{\mathrm{fg}} \) reduces fine-grained accuracy by 11 percentage points, and replacing energy scoring with distance thresholding lowers the AUROC by 0.2683.

Article
Computer Science and Mathematics
Signal Processing

Rongfeng Li

,

Junchen Liu

,

Zijin Li

,

Ya Li

,

Linfeng Fan

,

Pei Huang

Abstract: The digitized models of Chinese traditional musical instruments have fallen behind those of Western musical instruments, therefore they have restricted the preservation work and the research acquiring. This article deals with score-to-audio generation for the Chinese bamboo flute (dizi), it produces expressive audio from MIDI scores which has human performance detail. The current existing methods have two obstacles that they face. First, the large-scale paired training data (MIDI-audio) for traditional Chinese musical instruments is not much. Second, the already existing systems are not able to control the techniques that are specific to instruments—vibrato, pitch bends, ornaments—which the standard MIDI has no ability to encode. We put forward a expandable working flow that produces infinite artificial MIDI-audio matching groups through rule-based music score random arrangement, automatic DAW drawing output by using business sample storage banks, and systemized character pick-up. The information of technique is directly coded in the MIDI stream through the use of note numbers that lie outside the normal range, which thus proves to be more effective than the method of external conditioning. This system is constructed on the basis of MIDI-DDSP. Pre-training makes use of large-scale man-made data to obtain timbre consistency; The fine adjustment carries out the use of actual recording materials (URMP flute corpus) to realize expressive dynamic changes. Only the synthetic data can generate output that is stable but has no life. The real data which exist alone can bring about the problem of overfitting. This combined result can give realistic tone, which has controllable dynamic changes and playing skills.

Article
Computer Science and Mathematics
Signal Processing

Osama Al Maaini

,

Khizar Hayat

,

Khalil Al Ruqeishi

,

Baptiste Magnier

Abstract: The accuracy of spectro-temporal features for Makhaarij al-Huroof and Sifaat distinguishes between the ten canonical Qiraát recitation styles of the Holy Quran. However, real-world room reverberations blur formant contours and corrupt inter-word energies, thus making Qiraat discrimination difficult. The current dereverberation methods were designed to work under ordinary speech conditions and are not capable of preserving phonetic qualities for domain-specific purposes. This paper introduces a four-step, phase-consistent signal processing approach prioritizing phonetic preservation over direct reverberation suppression. The four steps are: (1) adaptive noise floor attenuation; (2) soft voice activity detection using power-law boundary decay; (3) application-specific spectral contour adjustment from clean Quranic reference audio; and (4) Griffin-Lim algorithm-based phase correction. A total of 48 real-world room recordings were utilized for the evaluation of this approach based on Energy Ratio (ER), Spectral Contrast (SC), and Formant Clarity (FC) – measures specific to the Quran audio domain – alongside conventional speech quality metrics. The proposed approach yielded the highest scores in four out of seven metrics, namely ER (+19.58 dB), SC (+40.11), FC (+822.94), and PESQ (+1.251), while being superior to Spectral Subtraction, Wiener Filtering and WPE Dereverberation approaches. Moreover, the perceptual enhancement was verified in a synthetic controlled experiment where the proposed approach scored an improved PESQ metric (+2.495; SNR −1.874dB). The results illustrate the fact that an optimization for general purpose metrics does not necessarily ensure phonetic preservation required for specific classification.

Article
Computer Science and Mathematics
Signal Processing

Jinhwan Kwon

Abstract: Interpersonal synchrony is a time-dependent coordination pattern in which interacting partners' body movements become temporally aligned. This study frames interpersonal synchrony as a human motion analysis problem and presents Synchrony Vision, an RGB-D sensor-based system for real-time monitoring and event-level analysis of interpersonal motion synchrony in free dialog. The system transforms Kinect-derived skeletal positions into joint acceleration signals, applies sensor-specific conditioning, detects movement peaks, and estimates event-level phase differences between two participants within a ±1.0 s window. The operator-facing interface supports live RGB-D monitoring, acceleration visualization, joint selection, millisecond-scale phase-difference histograms, four synchrony metrics (Frequency, Direction, Width, and Strength), and exportable acceleration, timestamp, peak-pairing, and summary artifacts. We evaluated the deployed pipeline on 25 Kinect-tracked dyads engaged in unconstrained conversation. Across 200.6 min comprising 245,835 frames, the system detected 2,681 synchrony events, exceeded circular-surrogate baselines, and preserved between-dyad ordering across reasonable parameter settings. Motion Energy Analysis-style cross-correlation on the same acceleration signals also confirmed above-chance synchrony but produced different dyad rankings. These findings show that RGB-D skeletal sensing can extend human motion analysis from individual movement capture to transparent, event-level quantification of interpersonal coordination.

Article
Computer Science and Mathematics
Signal Processing

Yutaka Yoshida

,

Kiyoko Yokoyama

Abstract: Accurate R-peak detection in electrocardiogram (ECG) signals is fundamental for cardiovascular analysis. However, most existing methods address differences in sampling frequency (fs) through signal resampling or transfer learning, which may alter the temporal definition of annotated events. In this study, we propose a fs consistent framework for ECG R-peak detection that avoids both resampling and retraining. The proposed method is based on low-sampling morphological learning combined with physiological temporal constraints (PTC). A lightweight classifier (Extreme Gradient Boosting) is trained on 128 Hz ECG data (MIT-BIH Normal Sinus Rhythm Database, XGB) to learn local morphological structures, and feature extraction is defined in milliseconds with time-normalized derivatives to ensure consistency across fs. The trained model is directly applied to higher- fs datasets (360 Hz, 500 Hz, and 1000 Hz) without modification. Final peak locations are determined through deterministic processing, including PTC and local snap processing. Experimental results demonstrated that the proposed method achieved stable detection performance across multiple sampling frequencies. When evaluated in a sample-wise manner, the proposed method achieved mean F1-scores of 0.885 on MIT-BIH Arrhythmia Database (360 Hz), 0.848 on Lobachevsky University Electrocardiography Database (LUDB, 500 Hz, sinus rhythm), 0.837 on LUDB (500 Hz, arrhythmia), and 0.953 on PTB Diagnostic ECG Database (1000 Hz), without any resampling or retraining. The integration of probabilistic candidate detection and deterministic temporal alignment enables consistent peak localization under cross-frequency conditions. These findings demonstrate that augmenting machine learning with deterministic decision mechanisms provides a principled framework for fs -consistent ECG peak detection.

Article
Computer Science and Mathematics
Signal Processing

Nahed H. Solouma

,

Michael R. Gardner

,

Noura Negm

,

Sadeq S. AlSharfi

Abstract: Optical imaging is among the safest and most highly impactful biomedical imaging modalities. Aberration in the optical imaging systems leads to distorted images. This distortion is almost nonlinear and hence affects the relative size, intensity and appearance of image details. Image aberration has many types with some or all of them can be imposed on the image based on the quality of the imaging system and/or surrounding conditions. Many approaches have been introduced to remove or minimize aberration from optical images. If the transfer function of an imaging system and the function of the noise added during the imaging process are known, then an ideal image can be obtained from the image produced by this system. The point spread function (PSF) of the optical imaging system is the image it produces for a point object. PSF is the observable form of the transfer function. The transfer function itself is the exit pupil function or typically the system aberration. The nonlinearity and multiplicity of the aberration imposed on the image together with the added noise makes it difficult to obtain the transfer function from the degraded images. In this work, optimization and global search techniques are utilized in an iterative image restoration algorithm. The proposed technique updates an initially suggested solution of transfer function by optimizing the aberration coefficients. The final solution of the transfer function and hence the PSF is reached when the optimum restored image is obtained. The proposed algorithm is validated by a testing image and then its performance is assessed by a set of aberrated images with different degradation. The results obtained in this work showed 100% success rate to obtain the PSF.

Article
Computer Science and Mathematics
Signal Processing

Siyuan Liu

,

Hangcheng Wu

,

Cheng Sun

,

Yuanbin Qiu

,

Haoliang Wu

,

Yucong Wei

,

Yang Lv

,

Zheng Yang

Abstract: This paper proposes a novel Topological Data Analysis (TDA) pipeline to extract robust structural features from functional near-infrared spectroscopy (fNIRS) signals for the classification of Alzheimer's Disease (AD) stages. Alzheimer's disease is increasingly understood as a disconnection syndrome, where the disruption of functional brain net-works precedes gross anatomical atrophy. However, traditional graph-theoretic ap-proaches rely on arbitrary connectivity thresholds, which can obscure critical multi-scale topological information and are sensitive to noise. To address this, our framework lev-erages Persistent Homology (PH) to analyse the topological evolution of brain networks across a continuous range of scales. By modeling 48-channel hemoglobin concentration time-series as high-dimensional point clouds via Granger causality metrics, we construct filtration sequences of Vietoris-Rips complexes. The resulting topological invari-ants—specifically 0-dimensional connected components, 1-dimensional loops, and 2-dimensional voids—are captured in Persistence Diagrams and subsequently vectorized into Persistence Images (PIs) using Gaussian kernel smoothing. This transformation enables the integration of complex topological features into standard machine learning workflows. Our experimental results on 284 recordings demonstrate that this topolo-gy-driven feature extraction method yields high discriminative power, achieving 77% accuracy in multi-class diagnosis (NC vs. MCI vs. AD). This study validates the efficacy of TDA as a sophisticated signal processing tool for revealing intrinsic neurodegenerative patterns in hemodynamic data, offering a potential non-invasive biomarker for early detection.

Article
Computer Science and Mathematics
Signal Processing

Rongyan Zhou

,

Weijie Tan

,

Meng Li

,

Baosheng Wang

Abstract: This article investigates sensor placement strategies for three-dimensional (3-D) time-of arrival (TOA)-based target localization in Underwater Acoustic Sensor Networks (UASNs). To mitigate underestimation of localization accuracy in complex marine environments, the actual acoustic ray propagation time is derived, and the TOA measurement variance is estimated using the ray acoustic propagation model. These formulations enable a novel 3-D TOA measurement model, and the trace of Cramér-Rao lower bound (CRLB) for this model serves as the optimization criterion for sensor placement. The MinMax k-Means algorithm is proposed to determine the optimal sensor placement by minimizing the average of the trace of CRLB. Extensive numerical simulations are conducted to demonstrate the effectiveness of the placement strategies.

Article
Computer Science and Mathematics
Signal Processing

Wei Li

,

Jiazhu Li

,

Shuyu Wang

,

Yan Chen

,

Jian Chen

Abstract: The health status of rolling bearings is critical to the normal operation of rotating machinery. To effectively extract vibration signal features and accurately identify different fault types, a novel method based on enhanced composite multi-scale slope entropy (ECMSE) and a honey badger algorithm-optimized kernel extreme learning machine (HBA-KELM) is proposed. Specifically, ECMSE integrates high-order differences into the composite multi-scale framework to capture high-frequency information while preserving low-frequency characteristics, thereby enhancing the discriminability of time-series representations. Meanwhile, an average coarse-graining strategy is incorporated to achieve a more comprehensive characterization of the signals. The extracted features are then input into the HBA-KELM classifier for fault identification. Experiments conducted on two public and private rolling bearing datasets demonstrate that our method achieves superior performance in distinguishing different fault types and damage levels compared with several existing approaches.

Article
Computer Science and Mathematics
Signal Processing

Fernando Martín-Rodríguez

,

Mónica Fernández-Barciela

,

Ainhoa Morales-Fernández

,

María Marante-Boado

Abstract: This paper evaluates the practical use of the Three-Dimensional Discrete Cosine Transform (3D-DCT) for video and volumetric image compression. While one- and two-dimensional DCT transforms are widely used in modern multimedia standards, their three-dimensional extension has received limited attention in real coding systems. In this work, a complete 3D-DCT–based encoder is developed by extending a JPEG-like pipeline to operate on three-dimensional data blocks. The proposed approach processes groups of video frames as 3D cubes and applies a separable 3D-DCT followed by quantization, coefficient serialization, and entropy coding. Unlike conventional video codecs that rely on motion estimation and compensation, the proposed system exploits temporal redundancy directly through the transform domain, resulting in a simpler coding structure with reduced algorithmic complexity. Different configurations are evaluated, including alternative methods for computing the 3D quantization matrix, serialization schemes, and variable block depth along the temporal dimension.Experimental results obtained from multiple test videos and volumetric medical datasets, including CT and MRI studies, demonstrate that the proposed method achieves competitive compression ratios while maintaining good reconstruction quality measured in terms of Peak Signal-to-Noise Ratio (PSNR). The results indicate that 3D-DCT provides a flexible and computationally simple solution for both video compression and three-dimensional medical image coding, particularly in applications where implementation simplicity and frame-level accessibility are important.

Article
Computer Science and Mathematics
Signal Processing

Xuchao Gao

,

Mingqiang Li

,

Kai Guan

,

Jianjun Ge

Abstract: To address the high computational complexity and insufficient real-time performance of traditional multi-radar trajectory planning methods in complex electromagnetic interference environments, this paper proposes an imitation learning-based trajectory planning method for multi-radar systems. This method designs a trajectory policy neural network architecture based on multiple semantic information. It proposes a training data construction method with coverage rate as the optimization objective. Then the trajectory policy neural network is trained by using an imitation learning algorithm with an auxiliary target. Simulation results show that the proposed method achieves an average coverage rate of 93.95%, and improves the single-step decision efficiency by a factor of 6.7 compared with heuristic-based trajectory optimization methods.

Article
Computer Science and Mathematics
Signal Processing

Lin Li

,

Jichun Zhu

,

Mingxing Jiang

,

Jingli Fang

Abstract: With the increasing demand for high-quality imaging in consumer electronics, image aesthetic assessment (IAA) has been widely applied to electronic cameras and display devices. Although the deformable attention mechanism has been introduced into IAA due to its perceptual capabilities, enabling models to refine attention regions by learning interest points and their corresponding offsets, existing methods often lack guidance from aesthetic composition features during the offset generation process, which limits their performance in aesthetic evaluation tasks. To address this issue, we propose a Graph Neural Network (GNN)-guided deformable attention module that incorporates composition information into the generation of interest points by modeling image features as graphs and applying GNN to guide interest point selection. In addition, we design an improved Transformer model that employs neighborhood attention to further enhance IAA performance. We evaluate the proposed model on two aesthetic datasets, AVA and TAD66K, and the experimental results demonstrate its effectiveness in improving overall model performance.

of 8