Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Computer Science

Fatih Şahin

,

Necibe Sare Mert

Abstract: Automated alert triage can reduce Security Operations Center (SOC) workload, yet the validation-tuned thresholds deployed systems rely on carry no finite-sample control of their operational error rates and degrade unpredictably under distribution shift. We present a model-agnostic dual-threshold conformal deferral architecture: high-score alerts are auto-escalated under finite-sample marginal class-conditional control of the benign-escalation probability (budget α), low-score alerts are auto-closed under matching control of the threat-miss probability (budget β), and the rest are deferred to an analyst. It needs no retraining, and closes an automatic zone rather than certifying what the calibration data cannot support. We evaluate it on a reinforcement-learning investigation agent in a simulated SOC and on four classifiers trained on CIC-IDS2017 and tested on CSE-CIC-IDS2018, using stratified 25,000-flow calibration and evaluation samples, with attack-type recall computed over the full 16.2-million-flow corpus. Pooling episodes from ten trained policies across two evaluation datasets, the architecture automated 73.7% of decisions at α = β = 0.01, realizing benign auto-escalation and threat auto-close rates of 0.0099 and 0.0101, and deferring the hardest ~26% of alerts (≈49% threat prevalence). After recalibration on labeled target-domain data, severe cross-dataset degradation appears not as silent error but as sharply reduced certifiable automation, with deferral rising to 79–99% for the most affected classifiers.

Review
Computer Science and Mathematics
Computer Science

Janaka Ishan Senarathna

,

Hirushi Dilpriya

Abstract: This review paper discusses the critical role of cryptography in blockchain technology, focusing on how various cryptographic methods ensure the security, integrity, and functionality of blockchain systems. The paper provides a comprehensive analysis of the fundamental principles of blockchain and cryptography, exploring their synergistic relationship and the core features they enable. It examines the application of hashing algorithms, asymmetric cryptography, digital signatures, and symmetric cryptography in the context of blockchain, highlighting their contributions to data integrity, user authentication, and transaction non‐repudiation. The review also investigates the strengths and limitations of cryptography in blockchain, considering potential vulnerabilities and future challenges, such as the threat of quantum computing. Real‐world examples and use cases, including supply chain management and healthcare, demonstrate the practical implementation of cryptographic techniques in blockchain applications. The paper further explores future directions and potential advancements, including post‐quantum cryptography, enhanced privacy‐preserving techniques, and the integration of blockchain with other emerging technologies. The review concludes by emphasizing the indispensable nature of cryptography in blockchain technology, underlining its role in underpinning the core value propositions of security, transparency, and trust. The findings provide valuable insights for researchers, practitioners, and stakeholders interested in understanding the critical importance of cryptography in the rapidly evolving field of blockchain technology.

Article
Computer Science and Mathematics
Computer Science

Shuyi Wang

,

Baoping Wang

Abstract: Visual Internet-of-Things (IoT) sensors are increasingly used to collect artistic images in museums, galleries, cultural heritage sites, and public spaces. Centralizing these images for analysis, however, can expose sensitive information concerning artwork ownership, exhibition layouts, visitor activities, and institutional collections. Federated learning offers a decentralized alternative, but its application is challenged by non-independent and identically distributed image data, resource-constrained sensor nodes, communication overhead, and privacy leakage from model updates. This paper proposes FedArtSense, a privacy-preserving federated learning framework for artistic image analytics in visual IoT sensor networks. FedArtSense introduces prototype-guided representation alignment to reduce client drift caused by heterogeneous artistic styles and collection distributions. An adaptive privacy mechanism dynamically determines gradient-clipping thresholds and noise levels according to update sensitivity, while a Rényi differential privacy accountant provides quantifiable privacy guarantees. In addition, importance-aware sparse aggregation reduces communication costs by transmitting only informative model updates. Experiments on the WikiArt, ArtBench-10, and Behance Artistic Media datasets under realistic non-IID and resource-constrained IoT settings demonstrate that FedArtSense consistently improves classification performance and convergence stability compared with representative federated learning and privacy-preserving baselines. It also achieves a favorable balance among analytical accuracy, privacy protection, and communication efficiency. These results indicate that FedArtSense provides an effective solution for secure and scalable artistic image analysis across distributed visual IoT infrastructures.

Article
Computer Science and Mathematics
Computer Science

Zhipeng Hong

,

Tianyi Xu

,

Huangyin Chen

Abstract: Cloud service continuous delivery involves computing, storage, permission, scheduling, and monitoring modules. Because complex service dependencies may hide cross-service anomalies under insufficient test coverage, this study proposes a quality assessment method combining test coverage mapping and release risk prediction. A directed dependency graph is built for interfaces, resource creation, volume mounting, permission verification, read/write performance, exception recovery, cross-version compatibility, and rollback paths. GraphSAGE learns associations among service nodes, test cases, and historical failures, while CatBoost predicts release failure probability. The dataset contains 31 pipelines, 126 modules, 460 integration cases, 9 release-change types, 2,800 release records, and 72 million execution data points. The coverage graph identifies 37 high-risk uncovered nodes, with storage mounting, permission propagation, resource initialization, and cross-version compatibility accounting for 72.9%. CatBoost achieves 90.8% accuracy and 0.934 AUC. Fault injection shows that dependency timeouts, permission anomalies, and storage latency raise failure probability by 31.5%, 24.7%, and 19.2%. After adding critical-path tests, core coverage increases from 68.4% to 91.2%, and monthly rollbacks fall from 26 to 14. This method supports risk control for banking, insurance, medical, education, and SaaS cloud services.

Article
Computer Science and Mathematics
Computer Science

Yongjian Wang

,

Aibo Song

Abstract: Trusted Data Spaces (TDS) have emerged as the core infrastructure for secure, privacy-preserving data circulation across industries and jurisdictions. However, state-of-the-art TDS implementations suffer from centralized platform monopoly, rigid cross-border governance failure, unfair value distribution, and poor scalability for global-scale collaboration. This paper proposes DAO-TDS, a novel decentralized autonomous trusted data space paradigm that enables centerless, cryptography-governed, and value-closed-loop data circulation. We make three core contributions: (1) We formalize the first anti-monopoly, incentive-compatible game-theoretic model for distributed TDS governance, with rigorous provable security guarantees; (2) We design an original Proof of Data Contribution (PoDC) consensus mechanism and a post-quantum secure Crypto-DAO governance protocol, with formal security proofs under the Universal Composability (UC) framework; (3) We implement a full prototype of DAO-TDS and conduct comprehensive, reproducible evaluations, showing that it supports 10,000+ distributed nodes with >12,000 TPS and <2s 99th-percentile confirmation latency, while delivering >80% of generated value to data contributors (vs. <50% in centralized platforms). While the proposed paradigm demonstrates strong performance and security guarantees, it still faces challenges in adaptive cross-jurisdictional compliance and lightweight edge node deployment, which require further investigation.

Article
Computer Science and Mathematics
Computer Science

Fatma Yasmine Loumachi

,

Karim Ouazzane

,

Anthony Phipps

Abstract: Payment infrastructures are security-critical networked systems whose control status may change as attacks, remediation, and recovery unfold. In such environments, PCI DSS requirement applicability and control satisfaction are not fixed properties of a final assessment point, but may vary across intermediate operational states. Endpoint-oriented standards assessment can consequently fail to retain violations that arise during malware infection, lateral movement, segmentation failure, service disruption, or subsequent restoration. This paper introduces SCF-PCI, a formal and executable framework for path-sensitive PCI DSS valuation over adversarially evolving payment-system states. The framework defines a PCI valuation domain, separates requirement-check applicability from satisfaction, and assigns state-indexed valuations to each reached state. These valuations are then lifted to execution paths, whilst endpoint loss is characterised through restoration, applicability-window closure, and domain exit. SCF-PCI is implemented in OPA Rego using OSCAL-represented PCI DSS control artefacts and evaluated across representative payment-network attack scenarios involving malware infection, remote compromise, DDoS disruption, spyware-assisted fraud, and recovery paths. The results show that state-indexed valuation preserves compliance-relevant security violations that endpoint-only assessment fails to retain. The framework contributes a standards-based assurance method for network and system security analysis in adversarial payment environments.

Review
Computer Science and Mathematics
Computer Science

Ratana Soth

,

Tharoeun Thap

,

Sa Math

Abstract: Kubernetes has become the dominant orchestration platform for cloud-native applications, where autoscaling plays a critical role in maintaining performance, availability, and infrastructure efficiency. Traditional Kubernetes autoscaling mechanisms, including the Horizontal Pod Autoscaler (HPA), Vertical Pod Autoscaler (VPA), Cluster Autoscaler (CA), and Kubernetes Event-Driven Autoscaling (KEDA), primarily rely on reactive threshold-based scaling policies. Although these approaches are effective for relatively stable workloads, they often struggle to handle highly dynamic and bursty traffic patterns commonly observed in modern microservices, edge systems, and artificial intelligence (AI)-driven applications.Recent advances in machine learning (ML) and AI have significantly influenced Kubernetes autoscaling research. Researchers increasingly explored predictive forecasting, reinforcement learning, graph neural networks, and hybrid optimization frameworks to improve scaling responsiveness, reduce latency, minimize infrastructure cost, and optimize service-level objective (SLO) compliance. This paper presents a comprehensive review of ML-based Kubernetes autoscaling techniques published recently. The review is organized into four major autoscaling categories: HPA, VPA, CA, and event-driven autoscaling through KEDA. The paper further analyzes emerging trends, including transformer-based forecasting, multi-agent reinforcement learning, graph-enhanced orchestration, GPU-aware autoscaling, and in-place vertical scaling. Finally, open research challenges such as cross-workload generalization, explainability, scaling conflicts, and edge-cloud deployment constraints are discussed.

Article
Computer Science and Mathematics
Computer Science

Porter E. Coggins III

Abstract: This paper introduces Hill-Enigma-SPN (HESPN), a 128-bit byte-oriented substitution–permutation network (SPN) research construction that uses Hill-cipher matrix diffusion as a foundation while addressing the Hill cipher’s classical linearity and lack of nonlinear substitution. HESPN integrates rotor-scheduled admissible GF(2) byte-matrix diffusion based on the Enigma Encoding Machine, AES S-box substitution, a round-dependent inter-byte permutation, and a specified Argon2id profile for password-based key derivation. Across four independent experimental sessions, HESPN exhibits behavior consistent with a well-diffused SPN on the evaluated metrics. At 16 rounds, plaintext avalanche averages 64.0 bits, and final-round counter-mode keystream testing passes the implemented NIST SP 800-22 core battery, with a 300-sequence confirmation run supporting borderline 100-sequence runs/serial results. Differential probes show no sampled differential collisions from rounds 8–12 and, in a 16-round confirmation across twelve chosen input differences, none at 16 rounds either, at the tested sample size (resolution 2×10⁻⁵); and random-mask linear-bias screening at 16 rounds remains at the statistical noise floor. These empirical observations are bounded screens and do not establish cryptanalytic equivalence to AES or resistance to trail-optimized attacks. Revision-phase testing identified a structured-input correlation in the original 12-round counter-mode output; a round-count sweep placed decorrelation by approximately 14 rounds, and HESPN is therefore specified with 16 rounds. Algebraic degree reaches the maximum observable degree in the restricted-variable experiment by round 8. The branch-number condition B ≥ 4 is formally proved for all admissible matrices. Among the designs compared, HESPN appears to be distinctive in combining rotor-scheduled GF(2) diffusion, a proved local branch-number floor, AES S-boxes, and Argon2id key derivation in one construction.

Article
Computer Science and Mathematics
Computer Science

Çağatay Ersin

Abstract: This study presents an enhanced unified feature representation for classical one-class anomaly detection in industrial image inspection using the MVTec AD dataset. Experiments were conducted on all fifteen categories of the MVTec AD dataset, including object and texture categories. Unlike raw pixel-based representations, the proposed feature representation integrates raw intensity information, histogram of oriented gradients (HOG), local binary pattern (LBP), gray-level co-occurrence matrix (GLCM)-based texture descriptors, and basic statistical intensity features. These complementary descriptors were evaluated with classical one-class anomaly detection methods, including PCA-based reconstruction, kNN-density, Local Outlier Factor, One-Class SVM with RBF kernel, Isolation Forest, and Mahalanobis-PCA. Considering the class imbalance inherent in anomaly detection tasks, Average Precision (AP) was used as the primary evaluation metric, while ROC-AUC and F1@95%-training-quantile were reported as supporting metrics. The results showed that PCA-based reconstruction achieved the strongest overall performance with a mean ROC-AUC of 0.8289, mean AP of 0.9254, and mean F1@95%-training-quantile of 0.8343. Category-based analyses showed particularly strong discrimination in the bottle, toothbrush, leather, transistor, and zipper categories, whereas carpet, screw, and grid remained more challenging. The findings indicate that carefully designed classical feature representations can provide an explainable and CPU-friendly baseline for industrial anomaly detection, particularly at the model-inference stage, while feature extraction remains the dominant computational cost.

Article
Computer Science and Mathematics
Computer Science

Boris Galitsky

Abstract: Consumer rights protection remains fragmented across customer support channels, banking dispute systems, regulatory complaint mechanisms, legal escalation procedures, and public reputation platforms. Consumers frequently encounter asymmetry of information, emotional exhaustion, procedural complexity, and strategic disadvantages when interacting with corporations. We present Complaint Warrior, a synthesized human–AI-agentic framework for consumer rights protection that combines large language models (LLMs), multi-agent orchestration, automated negotiation, strategic reasoning, legal workflow management, and behavioral modeling into a unified dispute-resolution ecosystem.The proposed system integrates autonomous and human-supervised workflows across complaint intake, evidence gathering, company negotiation, credit-card chargeback initiation, social-media escalation, and small-claims litigation preparation. A central contribution is the use of LLM-based reasoning about mental states and organizational intent to predict likely peer behavior during negotiation. Instead of treating customer support interactions as isolated messages, the system models disputes as evolving strategic games involving beliefs, incentives, emotional states, procedural constraints, legal exposure, and reputational risk. Complaint Warrior employs a multi-agent architecture in which specialized agents coordinate through a shared dispute state representation. These agents include complaint-analysis agents, negotiation-strategy agents, legal agents, financial-dispute agents, social-media escalation agents, behavioral-prediction agents, and company-side moderation agents. The system supports both consumer and company workflows, including a subscription mechanism whereby participating companies gain structured negotiation interfaces and AI-assisted compromise optimization in exchange for reduced escalation risk. We describe the full operational pipeline, implementation architecture, reasoning framework, conflict-resolution strategies, and deployment infrastructure. We further discuss safety mechanisms, human oversight, negotiation ethics, explainability, and future directions toward autonomous dispute mediation ecosystems.

Article
Computer Science and Mathematics
Computer Science

Chihhsiong Shih

,

Cheng-Hsu Chen

,

Xiuyuan Yeah

Abstract: Taiwan has one of the highest dialysis prevalences worldwide, making safe and reliable hemodialysis monitoring a critical sensor-based healthcare challenge. Modern hemodialysis machines integrate heterogeneous multimodal sensors (pressure, flow, conductivity, temperature, and cardiovascular signals), but differences in machine brands, data formats, and privacy constraints hinder centralized learning and robust complication prediction. This work proposes a Medical IoT–oriented federated learning framework, PSOFed-HD, that performs dual-layer Particle Swarm Optimization (PSO) to enhance heterogeneous sensor fusion for predicting dialysis-related hypotension and discomfort events. Each hemodialysis machine is paired with an edge gateway acting as an FL client, where local PSO optimizes CNN feature weights over non-IID sensor subsets, while the central server applies PSO-driven aggregation to adaptively weight client models according to validation performance. Experiments on real-world hemodialysis datasets with 17 physiological features demonstrate that standard FedAvg yields an accuracy of 65.24% and F1-score of 0.5318, server-side PSO improves accuracy to 75.11%, and client-side PSO further raises accuracy to 81.97%. The proposed dual-layer PSO framework achieves the best performance, with 90.56% accuracy and an F1-score of 0.8533, along with superior ROC characteristics (AUC = 0.908) and stable cross-validation across 11 folds. These results confirm that jointly optimizing local feature representations and global aggregation weights enables effective fusion of heterogeneous hemodialysis sensor data under privacy-preserving Medical IoT constraints, providing a practical decision-support approach for early complication prediction in dialysis units.

Review
Computer Science and Mathematics
Computer Science

Anastasios Nikiforos

,

Christos Manolas

,

Panos Raptis

Abstract: Background: Higher-immersion virtual reality (VR) systems are increasingly deployed across education, clinical care, gaming, and immersive journalism, but the extent, range, and nature of the peer-reviewed evidence linking system immersion to user outcomes has not been systematically mapped. Objective: This scoping review maps the peer-reviewed journal literature comparing higher- to lower-immersion VR conditions, characterising how the evidence is distributed across application domains, outcome constructs, measurement instruments, and study designs, and identifying knowledge gaps. Methods: Following the PRISMA Extension for Scoping Reviews (PRISMA-ScR), we searched five databases and screened 135 full-text reports. Eligibility was restricted to peer-reviewed journal articles (excluding preprints, conference proceedings, book chapters, and theses) reporting an empirical higher- vs. lower-immersion contrast with extractable quantitative outcome data. Data were charted descriptively; no inferential meta-analytic pooling or certainty rating was performed, consistent with scoping review methodology. Results: Forty-six journal articles met the inclusion criteria, of which 45 provided chartable effect-size data spanning 2019–2026. Evidence concentrated in Gaming and Entertainment (k = 16) and Education and Training (k = 14), with smaller bodies in Journalism and Prosocial Communication (k = 8) and Clinical and Rehabilitation (k = 7). Across studies, the standardised mean difference (Hedges’ g) had a median of 0.79 (range −0.98 to 3.79); 41 of 45 studies favoured the higher-immersion condition and 36 did so with confidence intervals excluding zero. Presence and immersion were the most frequently measured constructs (26/45); measurement instruments were heterogeneous. Conclusions: Peer-reviewed evidence consistently points toward higher system immersion enhancing user outcomes, especially presence-related constructs, but is methodologically fragmented, dominated by small between-subjects studies, and sparse in clinical contexts. The map identifies priorities for future confirmatory synthesis.

Article
Computer Science and Mathematics
Computer Science

Faisal Abdulaziz Almisned

,

Ibrahim Ibrahim Shuaibu

Abstract: Background: External validation performance drops are routinely reported in medical artificial intelligence (AI) literature, yet the study-level metadata features that systematically predict this degradation remain poorly characterised. Existing systematic reviews have catalogued performance metrics without subjecting the explanatory value of reported metadata to rigorous empirical audit. Objective: To quantify the explanatory ceiling of progressively richer metadata tiers on observed AUC degradation at external validation, and to identify which specific features carry replicable predictive signal across 1,000 non-parametric bootstrap resamples. Methods: A systematic search of PubMed/MEDLINE (January 2016 – May 2026), supplemented by Scopus and IEEE Xplore, identified 100 peer-reviewed studies reporting both internal performance and quantitative external validation outcomes for medical imaging AI. Thirteen metadata features were extracted and organised into four progressive complexity tiers. Explanatory power was assessed by the 10-fold cross-validated R² metric from a random-forest regression fitted on each tier. Replicable feature-selection frequency was estimated by 1,000 non-parametric bootstrap resamples with a 5% alpha threshold. Modality-stratified sensitivity analyses were conducted across three clinical imaging domains: ophthalmology (fundus photography), chest radiography, and computed tomography. Results: The full 13-feature model explained only −60.5% of variance in AUC degradation (10-fold cross-validated R² = −0.6053), with the deficit deepening monotonically across all four metadata tiers (Tier 1: R² = −0.3438; Tier 2: −0.4002; Tier 3: −0.5025). Non-parametric bootstrap resampling identified augmentation_applied as the single feature with selection frequency exceeding the 5% alpha threshold (frequency ≈65%), followed by class_balance_ratio (~23%) and architecture_type (~15%). All other features remained below threshold. Modality sensitivity profiles were broadly comparable across the three imaging domains, though CT studies exhibited the widest range of observed performance drops (ΔAUC range: −0.05 to +0.21). Conclusions: Collectively reported metadata—including dataset size, architecture, and AUC—leaves the overwhelming majority of observed generalisation variance unexplained. The consistent explanatory deficit across all metadata tiers signals a structural reporting gap rather than a predictive modelling limitation. Augmentation strategy, class-balance handling, and architecture type constitute the minimum replicable predictive signal currently available. Standardised metadata reporting frameworks, covering training data provenance, preprocessing pipelines, and demographic covariates, are required before comparative generalisation benchmarks can be meaningfully established.

Article
Computer Science and Mathematics
Computer Science

Vladimir Rotkin

Abstract: The rapid development of large language models and autonomous intelligent agents has significantly expanded the capabilities of natural language processing and decision support. However, practical implementation reveals fundamental limitations, particularly in tasks requiring computational robustness, reproducibility of results, and strict information consistency. These issues are particularly critical in fields such as engineering, geometry, and educational systems, where plausible but inaccurate responses ("hallucinations") and unstable behavior undermine system trust. This paper proposes a hybrid intelligent architecture with a deterministic core to address these challenges. Unlike fully autonomous systems, the proposed approach decouples functions: an adaptive agent handles user interaction and its interpretation, while a stationary deterministic core provides robust computation, logical consistency, and graphical display. The architecture introduces a clear distinction between the development phase, which allows for iterative improvement, and the operational phase, characterized by a fixed core that guarantees reproducible and verifiable results. By providing protocol-based interaction between the agent and the deterministic core, the system ensures that all generated output—text, computational, and graphical—remains consistent and adheres to the underlying domain model. This hybrid structure combines the flexibility of modern intelligent agents with the precision and reliability of formal deterministic models, offering a robust foundation for mission-critical intelligent applications.

Article
Computer Science and Mathematics
Computer Science

Fatma Zehra Aytaş

,

Safiyye Kalemci

,

Elif Özel Ay

Abstract: Hyperstructures, which have gained increasing importance in the field of Abstract Algebra in recent years, differ from classical algebraic structures in that the result of an operation can be a set rather than a single element. This property gives them the potential to produce multiple outcomes and offer a higher level of complexity This study focuses specifically on the algorithmic modeling of multiplicative hyperrings. For this purpose, a web-based interactive platform has been developed where users can define and test their own hyper-operation rules on both finite and infinite sets. The developed platform consists of two main sections: an "Education Module" and a "Cybersecurity Module". The Education Module illustrates abstract algebraic concepts using visual tools such as Cayley tables, while the Cybersecurity Module demonstrates the practical applications of hyperring structures in symmetric encryption algorithms similar to Diffie-Hellman key exchange and Caesar cipher. Additionally, a system has been integrated that provides users with AI-assisted constructive suggestions for structures that do not satisfy the axioms. This study aims to build a bridge between the theoretical world of hyperstructures and their applications in computational mathematics and cybersecurity.

Concept Paper
Computer Science and Mathematics
Computer Science

Sheetal Temara

Abstract: While traditional cybersecurity has focused on networks, endpoints, applications and data, the rise of autonomous AI systems is shifting adversarial activity toward human cognition, trust and decision-making. The emergence of systems such as Anthropic’s Mythos signals a broader inflection point in which machines can autonomously identify vulnerabilities, generate exploits and accelerate cyber operations with minimal human input. The human layer becomes increasingly vulnerable to AI-enabled manipulation including synthetic identities, deepfake personas, hyper-personalized phishing and real-time social engineering as technical attack timelines compress from weeks to hours. Defending the human layer requires moving beyond passive awareness training and static security controls toward active and edge-computed autonomous defense. Such systems must be capable of detecting suspicious interactions, verifying identity claims, interrupting coercive or deceptive workflows and preserving user agency without becoming opaque or paternalistic. Cybersecurity must defend human agency against rapid threats instead of merely protecting systems and data.

Article
Computer Science and Mathematics
Computer Science

Cristofer Tamaral

,

Raquel Hijón-Neira

Abstract: Basic Vocational Training (FPB) students face significant learning difficulties in electronics, which leads to a high dependency on the instructor and creates bottlenecks in practical workshops. This study aims to evaluate whether generative artificial intelligence (AI) can act as an effective support tutor during workshop practice to mitigate these issues. To this end, a quasi-experimental design was implemented with Basic Vocational Training students divided into an AI-assisted group and a traditional control group, using pre-test and post-test assessments to measure cognitive progress. The results reveal that the use of AI allows for maintaining performance in conceptual learning despite tasks of increasing difficulty, while significantly increasing students' motivation and perceived autonomy. Although a drastic reduction in the frequency of technical assembly errors was not recorded, AI proved to be an effective support for resolving procedural questions in real-time. It is concluded that the integration of generative AI offers positive implications for Vocational Training, functioning as a supplementary tutor that fosters student independence and optimizes technical classroom dynamics.

Article
Computer Science and Mathematics
Computer Science

Jaime Dionisio Burillo

,

Raquel Hijón Neira

,

Oriol Borrás-Gené

Abstract: This study presents the design, implementation, and exploratory classroom evaluation of a web-based educational assistant built on Retrieval-Augmented Generation (RAG) and Large Language Models (LLMs) for Vocational Education and Training (VET). The platform was designed to generate responses based on teacher-provided course materials, preserve source traceability, and return an abstention message when the retrieved evidence is insufficient. The assistant was deployed in an authentic classroom setting within the Higher Vocational Training programme in Network and Information Systems Administration (ASIR). Nineteen students and one instructor used the system during a practical session and completed a post-session questionnaire combining Likert-scale items with open-ended questions. The findings indicate positive student perceptions of usability, response clarity, perceived reliability, and learning support. Participants particularly valued the ability to obtain focused answers aligned with the instructional materials. The evaluation also revealed a relevant trade-off: restricting the assistant to a controlled corpus reinforced curricular consistency and perceived trustworthiness but limited its capacity to address questions insufficiently covered by the available resources. The absence of conversational memory emerged as the most frequently requested improvement. These preliminary findings suggest that course-constrained RAG assistants may constitute valuable complementary tools for transparent and pedagogically supervised AI-supported learning in technical VET contexts.

Article
Computer Science and Mathematics
Computer Science

Eka Prasetyaningrum

,

Ahmad Zainul Fanani

,

Dwi Wahyu Prabowo

Abstract: Breastfeeding experiences are closely linked to postnatal mental health, yet population-level studies have largely treated mothers as a uniform group. This study applies an unsupervised machine learning approach to identify distinct maternal subgroups within a large UK breastfeeding and mental health dataset. Secondary analysis was conducted on the open-access survey dataset compiled by Braithwaite et al. (2025), comprising 2,010 postpartum mothers who had breastfed their first child. Thirty variables were selected across four domains: prenatal/postnatal social pressure, psychosocial breastfeeding impact (0–10 scale), demographic characteristics, and mental health outcomes (EPDS and GAD-7). After median imputation and StandardScaler normalisation, Principal Component Analysis (PCA) was applied for variance decomposition, followed by K-Means clustering. The optimal cluster count (k=4) was selected based on the elbow curve, Silhouette Score, and Davies-Bouldin Index, validated by Hierarchical Clustering (Ward linkage). Four clinically distinct subgroups emerged is High-Risk (n=303, EPDS M=11.45, 62.7% at-risk), Low-Pressure (n=663, longest breastfeeding duration), High-Healthcare-Professional Pressure (n=517, paradoxically shortest breastfeeding duration), and Resilient (n=527, EPDS M=7.98, mean duration=15.18 months). One-way ANOVA confirmed highly significant between-cluster differences across all psychosocial variables (F=34.59–787.55, p<0.001), with no significant age or education differences. PCA identified guilt impact and maternal identity as the primary axes of variation. Findings indicate that internal psychological burden not demographic profile is the principal differentiator, with direct implications for designing targeted postnatal mental health interventions.

Article
Computer Science and Mathematics
Computer Science

Boris Galitsky

Abstract: Large language models (LLMs) frequently generate fluent chain-of-thought (CoT) reasoning that appears coherent while containing unsupported inferences, omitted alternatives, or logically invalid conclusions, creating significant challenges for trustworthy and explainable AI in high-stakes domains such as healthcare. While prior work has primarily focused on factual hallucinations, the structural characteristics of hallucinated reasoning remain insufficiently understood. This paper investigates diagnostic hallucinations through the lens of discourse structure and introduces a synthetic benchmark of ambiguous patient complaints paired with diagnoses, reasoning traces, discourse-tree representations, and hallucination labels. We hypothesize that hallucinated CoT differs from grounded reasoning not only in factual correctness but also in discourse organization. Our analysis shows that hallucinated reasoning tends to elevate weak or speculative clues into central discourse nuclei, suppress contradictory evidence, and use discourse satellites for post hoc reinterpretation rather than evidence integration. In contrast, grounded reasoning preserves alternative hypotheses, maintains explicit contrast and concession relations, and resolves evidential conflicts at higher discourse levels. Based on these observations, we formulate discourse-level indicators for hallucination detection and reasoning verification. The resulting framework provides interpretable structural explanations of reasoning failures and supports neuro-symbolic validation of LLM-generated reasoning traces, contributing to the development of more trustworthy, explainable, and verifiable language models.

of 70

Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings