Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Stamatis Mastromichalakis

Abstract: Muon orthogonalizes the momentum matrix, granting every singular direction of the update equal trust. EchoMuon prices that trust by each direction's echo: its support in a second, slower momentum buffer. This paper reports what that buys, and what measuring such a gate honestly costs. The gate is a contraction, not a reallocation. Its multiplier never exceeds one in 5,760 logged layer-steps, and the arm carrying it sits higher on the learning-rate axis: over four cells it is penalized 3.7x less than ungated Muon for a step twice too large, and 1.6x more for one half too small, without exception in either direction. A shared learning-rate grid therefore does not compare the two fairly: when the grid sits above the joint optimum, the contracting arm is rescued and the other is not, by an amount comparable to the margins being measured. Our initial hypotheses about the gate were wrong on this point, so we re-ran the full suite under corrected selection. All eight CIFAR arm-cells had picked the bottom edge of a single shared grid; widening it and selecting on a held-out split cuts those margins by 42 to 80% and leaves none individually significant. Tiny ImageNet, whose optimum the original grid did contain, is unchanged at +1.65 pp (t = +7.4) pooled over its two cells, and on the clean cell a norm-matched scalar control reproduces only 30% of the gain, so what remains is directional rather than a step-size effect. Three further boundaries are measured. The FineWeb advantage falls from -0.031 nats at 0.3 tokens per parameter to -0.003 at 0.9, verified not to be a learning-rate artifact. Across ten vision cells the margin scales with the baseline's error rate and crosses zero near 90% accuracy. On byte-level text EchoMuon wins on clean enwik8 (t=-10.1) and loses only under input corruption, which replaces the tokenization boundary we first proposed. We release 1,815 runs, the re-run suite behind these numbers, and the reporting practice we recommend: publish the margin by which each learning rate beat its runner-up. Here a pick that won by 0.0026 in validation loss reversed under an equally valid split, with a 1.45 pp consequence on test accuracy.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Songrun Li

,

Weiran Zhang

,

Yichen Wang

,

Qi Lei

Abstract: Validation-based machine-learning model selection in temporally dependent, weak-signal data can produce a clear numerical winner without sufficient evidence that the choice will remain reliable out of sample. This study develops PC-Audit (Prequential-Calibration Audit), a decision-reliability framework for auditing the evidence behind validation-selected model choices. The framework combines target-date partitioning, prequential non-negative calibration, Model Confidence Sets (MCS), transfer diagnostics, dependence-aware inference, cost and execution checks, temporal stress testing, and fallback simulation. It is evaluated on nine Chinese index ETFs in 45 walk-forward cases from 2020 to 2024 under a controlled CPU-only setting. For signed returns, mean validation-to-test rank correlation is 0.090, PC-Select identifies the lowest-test-MAE candidate in 15 of 45 cases, the label-permutation value is p=0.0521, and the prequential MCS retains 4.80 of five candidates. No Holm-adjusted MAE reduction is found relative to naive references or individual-model baselines. Volatility-like targets show stronger main-period ranking transfer, but rolling/EWMA references remain competitive and 2025 stress weakens the evidence. Audit-triggered fallback does not improve mean MAE but modestly reduces worst-year MAE. PC-Audit therefore provides an auditable and reproducible reliability layer for validation-based time-series model selection rather than a forecasting-accuracy enhancement method.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Rakesh Kumar Agrawal

,

Wasim Mohammed Amin Tambe

,

Nihar Karra

Abstract: Autonomous, goal-directed AI agents are increasingly deployed on consumer edge devices — smart-home hubs, wearables, and ambient environments — where they observe context, reason over objectives, plan multi-step actions, and invoke tools with reduced human involvement. This shift exposes assurance gaps that periodic, organization-level governance frameworks were not designed to close: goal misgeneralization, unsafe or irreversible actuation, prompt injection via untrusted content, unauthorized tool invocation, inferential privacy exposure, and insufficient runtime human oversight. We present TOAIAF, a runtime AI Assurance Control Plane comprising seven components — a Policy Enforcement Engine, an Agent Runtime Governance Layer, a Trust Evaluation Engine, an Explainability Engine, a Privacy and Security Assurance Layer, a Human Oversight Boundary, and a Continuous Learning Assurance Loop — that mediates every proposed agent action through a formally specified, three-stage decision procedure (Algorithm 1): a hard capability/scope check, a non-compensatory safety gate, and a risk-weighted composite trust score. We operationalize six trust metrics (Agent Reliability, Goal Alignment, Privacy Risk, Explainability Confidence, Safety Compliance, and a newly introduced Human Oversight Effectiveness Score). We evaluate the control plane's decision logic through a synthetic, template-generated benchmark of 100 scenarios spanning smart-home, wearable, and ambient-intelligence contexts across six risk categories (benign, unsafe-policy, privacy-risk, goal-misaligned, prompt-injection, tool-misuse), against an ungoverned baseline and a keyword-based guardrail baseline, in seven experiments. TOAIAF's decision procedure achieved 100% flagging on safety-policy-violation, prompt-injection, and tool-authorization scenarios and 92.3% on privacy-risk and goal-misalignment scenarios, versus 100%/0%/75%/27.3% and 0%/0%/0%/0% for the guardrail and ungoverned baselines, respectively, with zero false restrictions across 40 benign scenarios. We report two ablation results transparently rather than selectively: removing the safety gate produced no measurable change in this dataset (23/24 flagged either way), traced to a confound in the experimental design rather than a finding about the gate's value; and including the Human Oversight Effectiveness term produced a small, mathematically-traceable negative effect on sensitivity (47/49 vs. 48/49 flagged), a concrete metric-calibration finding reported transparently. A companion reference-implementation prototype further surfaced a genuine design gap: the WARN verdict currently allows an action to execute, which for physically irreversible actions is a disclosed, unresolved limitation. We map each control-plane component to NIST AI RMF, ISO/IEC 42001, IEEE 7001–7003, and the EU AI Act, and provide a forward mapping to five external agent-safety benchmarks (AgentHarm, AgentDojo, R-Judge, SmartBench, SMH-Bench) for the real-agent evaluation this package does not yet perform. The benchmark, source code, simulation artifacts, and supplementary materials will be made publicly available through a permanent repository upon publication.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Gordana Dodig-Crnkovic

Abstract: This paper provides a presentation and structural overview of the book Evolving Intelligence. The Generative Path from Information and Computation to Cognition and Intelligence. The monograph is forthcoming with World Scientific in 2026. Evolving Intelligence develops a naturalistic, generative framework for understanding how intelligence emerges from the organization of nature. Rather than beginning with human reasoning or with artificial intelligence, the book traces a continuous path from physical and chemical organization in nature through life and cognition to human symbolic culture, artificial intelligence, and hybrid human-AI systems. Its central framework, ICON (Information-Computation-Cognition Naturalism), connects these domains through morphology, interaction, observer-agency, information, computation, cognition, consciousness, and intelligence.The guiding question is not simply “What is intelligence?” but “How can intelligence arise at all within nature?” The book therefore shifts attention from intelligence as an isolated human faculty to the generative processes through which increasingly complex forms of organization, sensing, memory, anticipation, agency, meaning, and problem solving emerge.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Georgi Rusev

,

Hugo Lafaye de Micheaux

,

Fabien Sauter-Starace

,

Petia Koprinkova-Hristova

,

Nikola Kasabov

Abstract: In our previous works we developed a neuromorphic decoder of intended movements of tetraplegic patients using ECoG recordings from brain motor cortex composed by a Motor Control Decoder (MCD) and a Neural Response Decoder (NRD). It was an actor-critic structure able to adapt via reinforcement learning the MCD (actor) based on NRD (critic) predictions. In this paper, we continue the development of novel neuromorphic methods for BMI aiming at further improvement of their functionality. First of all, feature extraction from ECoG data was improved by introduction of additional filtration of raw data. Second, the auto adaptive ability of NRD was upgraded using Intrinsic Plasticity (IP) tuning mechanism. Third, the mechanism of MCD decision improvement using NRD predictions was changed considering labeling approach of NRD training data. The MCD-NRD training and testing was done on a bigger data base and its improved accuracy was demonstrated on new test data sets as well.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Darya Sergienko

,

Roman Parovik

Abstract: The paper presents a parallel implementation of an algorithm for modeling a single dislocation source based on an ensemble of physics‑informed neural networks (Physics‑Informed Neural Networks, PINN). The task of approximating the Berlage pulse, which describes the displacement of a dislocation source under rock deformation, is considered; this is a key element in the study of high‑frequency geoacoustic emission arising during the failure of geomaterials under the influence of mechanical stresses. A method for decomposing the time domain into an arbitrary number of subdomains (\( n=1,\dots,8 \)) is proposed, each of which is trained on a separate computational process. To combine the predictions of the subdomains, the MultiDomainPINN ensemble architecture is used, which ensures smooth coordination of solutions at the boundaries of the overlapping regions. The quality of the approximation is assessed using statistical metrics: mean squared error (MSE), root mean squared error (RMSE), normalized root mean square error (NRMSE), coefficient of determination (\( R^2 \)), and maximum absolute error, which allows for a quantitative comparison of the PINN solution with the analytical Berlage impulse. To quantitatively assess the computational efficiency of the parallel implementation, TAECO metrics are used, which allow for evaluating the algorithm’s efficiency. The assessment was carried out at the inference stage of trained models for parameters that were not involved in the training, by comparing them with successive runs of the 4th‑order Rosenbrock method. MultiDomainPINN provides high performance, with a 70--76‑fold speedup compared to the Rosenbrock method while maintaining comparable accuracy (MSE at the level of \( 10^{-10} \)\( 10^{-8} \)). The developed software is scalable and can be adapted for a wide range of tasks related to modeling seismic signals and wave processes, including machine learning tasks in geophysics, seismology, and solid mechanics.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Kai Wu

,

Hao Lyu

,

Zhen Luo

,

Chaofan Wang

,

Siyu Ye

,

Jinghao Lin

,

Xiaozhong Ji

,

Boyuan Jiang

,

Shengzhi Wang

,

Zihan Wang

+13 authors

Abstract: Can AI improve AI? AI agents can now run experiments, modify code and training pipelines, and iteratively improve AI artifacts—bringing a long-standing idea closer to an empirical research question. Yet the rapidly expanding literature remains fragmented across long-horizon agents, AI for AI (AI4AI), self-improvement, and recursive self-improvement, making it difficult to tell how much progress has actually been made. This survey synthesizes evidence from hundreds of studies to answer this question. We organize the emerging AI4AI landscape around a simple question: how far can an AI system reliably carry an improvement process from idea to verified result? Across model design, agent harnesses, benchmarks, automated research, and self-modifying systems, we examine what current systems can do, how progress should be evaluated, and where claims of self-improvement remain unsupported. A striking pattern emerges in the information flow within the taxonomy– benchmark-model-harness structure of this survey. AI systems increasingly excel at the work of improvement—planning, coding, experimentation, optimization, and repair—but humans still largely determine the goals, evaluation criteria, and what ultimately counts as progress. Moreover, strong performance on individual components rarely translates into reliable end-to-end AI improvement. We call this composition gap. Today’s systems can already produce impressive improvements under bounded conditions, but evidence for reliable research judgment, causal experimentation, persistent gains, and compounding improvement remains limited. By separating demonstrated capability from extrapolated autonomy, this survey maps what AI4AI can do today—and what must change before AI can reliably improve AI itself. The eve ends the moment a system, for the first time, reliably strengthens its successors without humans specifying the goals or evaluation criteria. We hope this taxonomy and survey will help that moment arrive a little sooner.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Ada Fang

,

Kevin Li

,

Ayush Noori

,

Lukas Fesser

,

Marinka Zitnik

Abstract: AI scientists generate hypotheses, propose experiments, and analyze datasets, and several have produced findings that were confirmed in the laboratory. Although they are opening a path to autonomous discovery, the loop from hypothesis to experiment to revised hypothesis is still often anecdotal. Here, we consider how that loop could be closed to enable AI-driven biomedical discovery. Closing the loop requires AI scientists that reason over biomedical knowledge and data across horizons of days to months. It also requires verifying hypotheses with computational tools and experiments in biological laboratories and clinical environments. Agentic and reasoning methods need to learn over time from sparse, delayed, and noisy feedback that verifiers return. In mathematics and program search, where candidate solutions can be verified at negligible cost and often as soon as they are generated, methods that generate and score large numbers of candidates have improved bounds on problems that had been open for decades. In biomedical sciences, comparable progress requires closing the loop where each test is slow, costly, and partially informative. Closed-loop AI-driven biomedical discovery can enable a mode of research in which AI continuously reviews hypotheses against accumulating experimental results. Scientists choose which experiments to pursue, and AI maintains a record that links each result to the hypothesis that produced it.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Joshua A. Senne

Abstract: The data-centric artificial intelligence (AI) paradigm posits that improvements to training data quality yield greater performance increases compared to changes to model architecture; however, the claim is currently supported by case studies and demonstrations rather than controlled experiments. The experiments used five corpora with varying amounts of label noise (0.0% to 54.7%) and three distinct model architectures (LinearSVC, DistilBERT, and RoBERTa), resulting in 135 total experimental conditions based on the combination of two remediation strategies applied across five thresholds. All conditions were assessed via five-fold stratified cross-validation. Label errors within the five corpora were identified using confident learning (a probabilistic approach to flagging potentially mislabeled training examples by comparing predicted class probabilities with observed labels). The identification of label errors were validated independently against the CIFAR-N benchmark prior to conducting downstream analyses. To aid in connecting the results of this research to practice, cost-performance Pareto frontiers were created for each of four distinct levels of annotation costs. Lastly, a theoretical framework was developed to explain the observed architecture-noise interaction patterns.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Que Yang

,

Yukun Du

Abstract: User behavior, content attributes, and real-time contextual features are deployed continuously and in a large scale search and advertising ranking systems. Current fixed-ratio gray scale rollouts are unable to change the degree of validation as a function of feature quality, service status or model impact and some failures are only identified once traffic volume is higher. This paper analyzes the history of deployment of features on an online major U.S. service. It uses the deployment increment, known as a complete production deployment, as the unit of analysis, and manual or automatic (within 24 hours of deployment) rollbacks as the primary outcome, forming a hierarchical Bayesian discrete-time survival model. There are four types of model inputs: data quality metrics (missed value rate, outlier proportion, distribution gap between online and offline data, default value trigger rate); service performance metrics (read latency, cache hit rate, request failure rate, version discrepancies between data centers); ranking quality metrics (RCE changes, calibration error, candidate score offset); and traffic stability metrics (exposure differences between user segments, regional request fluctuations). The research examined 8, 640 production release records from 1, 286 online features during an 18-month period, of which 612 were identified as rollbacks, and matched these with some 7.4 billion online prediction results. The data was split into training, validation and independent test sets according to the release time of the features, and the versions of neighbouring features were only kept within a single data set. Results of the test demonstrated a AUROC of 0.901, AUPRC of 0.714, Brier Score of 0.064, and calibration slope of 0.96, and a risk ratio for the high-risk release group of 3.47 (95%CI 2.89–4.18) after a 24-hour rollback. This was followed by a shadow check on 164 new releases. With the phased rollout of 1%, 5%, 10% or 25% of traffic based on the predicted risk, the median time to detect anomalies was reduced from 29 min to 12 min, the average time spent in the validation cycle was lowered from 5.8 days to 2.7 days, unnecessary rollback was decreased from 21 to 10, and fluctuations in online RCE were controlled within 0.20 percentage points. Captures data anomalies, service degradation and variations in ranking performance in one overall risk assessment, delivering a calibrated, quantitative basis for decisions on scaling up and rollback of features online.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Peng Ye

,

Yichen Jiang

,

Suorong Yang

,

Shengji Tang

,

Shenghe Zheng

,

Jianjian Cao

,

Siyue Ren

,

Zhihao Dou

,

Weida Wang

,

Haonan He

+12 authors

Abstract: Collective intelligence (CI) has emerged as a cornerstone for advancing toward Artificial General Intelligence (AGI), with Large Language Models (LLMs) now serving as the primary driver. In this study, we present the first systematic framework to characterize the evolutionary trajectory of CI, and conduct a comprehensive review of the transformative role of LLMs in driving CI. Historically, CI went through two phases: the initial Collective Intelligence 1.0 Era ("Collective but Unintelligent") introduces various bio-inspired algorithms which demonstrates emergent behaviors but no true cognitive capabilities, while the subsequent Collective Intelligence 2.0 Era ("Specialized yet Disconnected") achieves domain-specific expertise within multi-agent systems but remained isolated across applications. Currently, Collective Intelligence 3.0 Era ("Collectively to Enhance Intelligence") is being shaped by LLMs, with remarkable breakthroughs achieved in both foundation-oriented and application-oriented researches. At the foundational level, by enhancing data, model, and the inference process, LLM-driven CI has made remarkable progress in terms of performance optimization and efficiency improvement. In applications, LLM-driven CI excels in complex, long-horizon tasks, such as scientific discovery and autonomous research, delivering innovative solutions to real-world problems. In the future, in order to solve more intricate and even unexplored real-world tasks and scenarios, there is an urgent necessity to transition to Collective Intelligence 4.0 Era ("Collectively to Create Intelligence"). To achieve this, the following aspects can be considered including fundamental capability construction, capability differentiation, cognitive fusion, dynamic organization, and co-evolution, and several challenges such as unclear emergence mechanisms, bottlenecks in continuous learning, interoperability issues among systems, and high demand for computational resources should be solved. The collaborative efforts of the research communities are expected to accelerate the maturation of LLM-driven CI, create immense value for society and drive leapfrog development across various fields, and steadily march towards the AGI goal.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Marian Pompiliu Cristescu

,

Ioana Petrea

Abstract: Enterprise information systems generate extensive transaction traces, but transforming these records into trustworthy AI-enabled decision intelligence requires more than pattern discovery: data quality, interpretability, statistical qualification, temporal robustness, reproducibility, and human oversight must be integrated. This study develops and evaluates a reproducible, interpretable machine-learning architecture designed as an analytical layer between transaction-generating systems of record and human decision support. A point-of-sale subsystem is used as the empirical validation environment rather than as a full ERP implementation. The architecture combines deterministic data-quality controls, sparse basket reconstruction, normalized Shannon entropy, unsupervised association-rule learning, false-discovery screening, nonparametric inference, generalized linear models, temporal validation, and transparent rule prioritization. The empirical evaluation covers 31,157 transaction baskets, 139,396 valid item lines, 710 products, and 762 active sales days. The resulting incidence representation contains 22.12 million potential positions at 0.630% density, and 1,174 directional rules satisfy the prespecified screening criteria. Out-of-period validation shows that 93.6% of discovery rules retain lift above one, while 49.3% again satisfy all selection thresholds and rank correlations for support, confidence, and lift remain between 0.80 and 0.85. The contribution is an auditable enterprise analytics architecture that separates machine-learned discovery, statistical qualification, and human decision responsibility. The architecture provides a reproducible analytical foundation that can complement enterprise systems of record and business-intelligence or decision-support environments without assuming production ERP integration, cloud deployment, or autonomous decision making.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Yutao Mou

,

Dingyao Yu

,

Xiaotian Luan

,

Zhe Yin

,

Zhangchi Xue

,

Peiyang Liu

,

Pengfei Yang

,

Tong Zhang

,

Shikun Zhang

,

Wei Ye

Abstract: Generalist agents are evolving from task-oriented systems into autonomous, self-evolving systems that acquire reusable skills, maintain long-term memory, interact with external environments, and refine themselves through recursive workflows. This shift creates a lifecycle security challenge beyond conventional model-centric or component-wise perspectives: vulnerabilities may arise from acquired external capabilities, be amplified through the orchestration of skills, memory, and tools, materialize as consequential actions, and persist through feedback across future evolution. Security for generalist agents is therefore inherently a lifecycle problem rather than a collection of isolated component- level risks. Motivated by these observations, we present a lifecycle survey organized around three stages: provenance, orchestration, and execution. We trace how risks originate, propagate, accumulate over long-horizon interactions, and persist through feedback. Across this lifecycle, we map four coupled attack surfaces: skill supply chains, user inputs, long-term memory, and external environments, and organize evaluations and defenses by where safety evidence is collected and interventions occur. Evaluations span interactive execution and offline trajectory auditing, while defenses cover input and context filtering, decision and control integrity, and runtime monitoring and enforcement. Finally, motivated by the growing importance of recursive and self-evolving AI, we identify adversarial risk generation, fine-grained safety attribution, and adversarial-feedback-driven continual safety evolution as three connected directions that form a closed loop of risk discovery, diagnosis, and mitigation toward safe recursive self-improvement agents. The latest papers and repositories are maintained at Pandora-Agent.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Junior Charlie

Abstract: Best-of-N pipeline selection, running N candidate prompts on an m-item dev set and reporting the winner's dev score, is the default way an LLM practitioner picks and reports a prompt. That dev score is optimistic: the winner was chosen in part because it got lucky on the dev items, and its true accuracy on fresh data is lower. We measure this winner's curse directly on four correctness matrices spanning three tasks (SST-2, Subj, AG News) and two models (Qwen2.5-1.5B-Instruct, Llama-3.2-3B-Instruct), sweeping N from 1 to 100 and m from 20 to 200. Bias grows with N and shrinks roughly as 1/sqrt(m), reaching 0.088 accuracy points at N = 100, m = 100 on Subj. We then characterize why real candidate pools inflate this bias beyond the independent-candidates textbook case: two effective-candidate-count estimators, a spectral participation ratio and a moment-matching inversion of the observed selection gap, agree that a 105-to-111-candidate pool behaves like roughly 1.5 to 3.5 independent draws. Breaking that count down by candidate provenance turns up a result opposite our own working hypothesis: LLM-generated (APE-style) candidates carry far less effective diversity (N_eff/K ~ 0.017-0.041) than either hand-written seed prompts or paraphrase/mutation candidates (N_eff/K ~ 0.14-0.24 for both), consistently across all three tasks. We compare six ways to report a best-of-N winner (no correction, a naive dev split, two extreme-value union-bound corrections keyed to K or to N_eff, cross-fitting, and a bootstrap plug-in) on bias, mean squared error, and 90% interval coverage; correcting against N_eff dominates the independent-candidates union bound on every axis and matches or beats every other method on mean squared error, while the naive K-based correction overshoots badly enough to be worse than reporting nothing at all. Resampling the stored matrices at the exact (N, m) of a real, live EvoPromptGA trajectory reproduces its measured bias on all three tasks, inside a 90% predictive band in every case, which supports treating a one-time correctness matrix plus free resampling as a faithful stand-in for repeated live search. Finally, we derive budget-allocation curves for a fixed evaluation budget B = N * m: on tasks already near their accuracy ceiling, the optimal split spends almost the whole budget on m; on a task with real headroom left, it keeps paying to search wider even as N_eff/K stays in the low single digits.

Review
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Yuangang Li

,

Jiechao Gao

,

Nan Hao

,

Xinhe Xu

,

Kecheng Liu

,

Yingzhou Lu

,

Bohao Xu

,

Chenhao Li

,

Jintai Chen

,

Minjie Shen

+12 authors

Abstract: Digital twin technology, a cutting-edge approach that creates dynamic digital replicas of physical systems, is increasingly integral to various industrial applications. However, the effective deployment of digital twins often faces challenges associated with data availability, quality, and interoperability, which can impede the development and operational efficacy of these virtual counterparts. This paper presents a systematic review of the application of artificial intelligence (AI) in enhancing and managing digital twin technologies across multiple domains such as industry, healthcare, urban planning, business, education, technology, and others. It specifically examines how machine learning models, particularly neural networks and deep generative models, are being employed to optimize the creation, maintenance, and functionality of digital twins. The review explores the role of artificial intelligence in overcoming data-related limitations by providing robust, scalable data solutions for digital twins. Additionally, it addresses critical considerations of real-time data processing and system interoperability within these applications. This survey not only identifies the prevailing challenges and opportunities within this emerging field but also highlights potential future research directions that could further the integration of AI with digital twin technology. Through a detailed exploration of the intersection between AI and digital twins, this paper aims to contribute significantly to the knowledge base and encourage further innovations in this interdisciplinary area.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Peng Ye

,

Zhuo Liu

,

Jingqi Ye

,

Fangchen Yu

,

Shengji Tang

,

Yichen Jiang

,

Haonan He

,

Zongsheng Cao

,

Tao Chen

,

Bo Zhang

+3 authors

Abstract: Scientific intelligence, which requires long-horizon reasoning and research across diverse and highly specialized domains, remains a key challenge for artificial intelligence. Although large language models (LLMs) have achieved strong performance on general tasks, their single-model paradigm faces fundamental limitations in scientific domains, including reduced effectiveness under specialization, high scaling costs with diminishing returns, and insufficient capacity for fine-grained long-horizon tasks. In this paper, we introduce an extended-mind-inspired paradigm, extending intelligence to an agentic system composed of LLM, interactive objects, and autonomous interaction processes, enabling deep specialization, new scaling dimensions, and stronger representational capacity. Building on this, we introduce ExoMind, the first extended-mind-inspired agentic system, which integrate systematic data engineering, scientific interaction framework, and systematic training strategy. Across diverse scientific reasoning and research tasks, our method achieves substantial and consistent performance improvements using less data, small model, and low-cost training. Notably, ExoMind achieves leading performance among frontier open-source and closed-source models on challenging scientific benchmarks (as shown in Figure 1), while also improving general-purpose capabilities. These results suggest a customizable, cost-efficient, and high-performance pathway toward democratizing scientific intelligence.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Hisham Zargar

,

Laszlo T. Koczy

Abstract: Offline Handwritten Text Recognition (HTR) for the Urdu Nastaliq script remains a significant hurdle in the field of document analysis. The intricate, cascading nature of the script, coupled with the reliance on localized diacritics, challenges standard sequence modeling. In this work, we propose a modified Convolutional Recurrent Neural Network (CRNN) that adapts a deep ResNet34 backbone to maintain high temporal resolution for Connectionist Temporal Classification (CTC). Our model achieves a baseline Character Error Rate (CER) of 32.38%. To contextualize this performance, we provide an extensive gap analysis comparing our robust CTC baseline with state-of-the-art autoregressive Transformers. We identify critical disparities in decoding mechanisms, data augmentation strategies, and computational resources, providing a roadmap for overcoming current plateaus in unconstrained Nastaliq recognition.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Florentin Smarandache

,

Maikel Y. Leyva-Vázquez

Abstract: A recent proposal supplements the neutrosophic triple with a fourth component N and partitions the paraconsistent region into three regimes—strong, weak and very weak—by thresholds on component sums (Smarandache and Leyva-Vázquez, 2026). We ask whether the ladder is measurable, which is prior to whether it is diagnostic, and we test it twice: over the eight statements of the corpus that motivated it (1,440 elicitations), and then over a 110-item bank crossed with six frontier model families and three competing glosses of the undefined fourth component, treated as a factor rather than as a hidden choice in the prompt (1,980 elicitations). The bank changes every headline number the pilot reported. What survives is a separation at both ends. The strong rung reaches 0.223 on unmarked ethical-conflict items against 0.046 elsewhere, separated under an item-clustered bootstrap and ordered identically by all six models; the very weak rung mirrors it, at 13.9% on epistemic ignorance and exactly zero on ethical conflict. What does not survive is the strength of that reading. The strong rung is not separated from logical paradox specifically (Delta = 0.107, [-0.060, 0.272]), so it marks material sustaining both verdicts rather than value conflict as such; its emptiness on all 176 anchors is a property of the strict inequality, since T+F equals exactly 1.00 in a third of evaluations and a non-strict comparison would place 77.8% of anchors in the strong rung; and no rung rate is a property of a statement, the between-item standard deviation being the size of the mean. The clearest result concerns the fourth component. Elicited N is distinct from I, but in the weak rung it enters as a disjunct that is present in 59.3% of eligible evaluations while deciding 0.4% of them. The refined form of the framework already separates the two as named types, so what we report is the cost of the coarse disjunction: a co-occurrence discarded in three of every five contested evaluations. Because the elicitation tells models that the degrees need not sum to one, we ran the manipulation: deleting that permission leaves the signature essentially where it was, while deleting the neutrosophic framing removes it entirely, so what we report is a property of models answering this question rather than of the instruction inside it. All prompts, items, raw generations and code are released, and three of the corrections above came from an adversarial re-analysis of them.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Tsehang Dorjee

,

Duo La

,

Lengben Zhaxi

Abstract: Cross-lingual voice cloning must preserve speaker identity without allowing reference speech to override the language or dialect requested by target text. We investigate this conflict in dots.tts across Amdo, U-Tsang, and Kham Tibetan, Mandarin, and English. A 4147-row display matrix was deduplicated into 120 unique reference–target conditions and re-scored with a three-seed, speaker-disjoint five-class ensemble. The evaluator achieved 60.83% accuracy and 0.774 macro-F1 on 9008 held-out utterances; language-level accuracy and macro-F1 were 0.978 and 0.976 after collapsing Tibetan dialects, whereas macro-F1 within Tibetan dialects was 0.644. We therefore treat language-level leakage as the stronger proxy. On five fixed reference prompts, disabling prompt latent prefill reduced the Base system's language-level leakage rate from 0.262 to 0.095, but the crossed-cluster 95% confidence interval for the difference [–0.081, 0.450] crossed zero. Reference x-vector re-pairing showed no stable benefit over a budget-matched continuation. Fixed-ratio prompt gating failed prespecified criteria, and classifier-guided decoding did not reduce dialect leakage; strong guidance reduced speaker similarity to 0.602. An x-vector probe and a 20-item evaluation by one native Tibetan listener remain exploratory. The resulting protocol separates target control, reference-side leakage, intelligibility, and speaker similarity, and identifies English-reference-to-Tibetan synthesis as the most concentrated failure route.

Article
Computer Science and Mathematics
Artificial Intelligence and Machine Learning

Brady D. Lund

,

Kinza Alizai

,

Eunice Amoje

,

Anuradha Chandrasekaran

,

Stefan Darvischi

,

Jeanne Denmark

,

Nishanth Joseph Paulraj

,

Lalitha Nallamothula

,

Antonio Paes

,

Bavya Sri Vemulapalli

Abstract: As artificial intelligence systems increasingly serve as a gateway for access to information, economic opportunity, and civic life, higher education programs training AI developers must evolve to prepare practitioners who are not only technically proficient but also socially accountable. This paper proposes a conceptual framework for integrating explainable AI (XAI) and social accountability into data science, computer science, and information science curricula. Drawing on accountability theory, information science, and recent XAI research, the proposed framework is organized around four interrelated pillars: answerability, responsibility, enforcement, and reflexivity. These pillars are further situated within technical, social, organizational, and political dimensions of XAI implementation, with particular focus on how XAI techniques such as LIME, SHAP, model cards, and counterfactual explanations can be operationalized as instruments of meaningful accountability. The paper then proposes a multi-level governance framework that links interpretability methods to institutional oversight, regulatory literacy, and participatory design, illustrated through a concrete scenario grounded in graduate data science education. Together, these elements represent a new pedagogical approach that can equip future AI developers to design and deploy AI systems that are accurate as well as transparent, justifiable, and responsive to the communities they serve.

of 274