Submitted:
19 July 2026
Posted:
21 July 2026
You are already at the latest version
Abstract
Multi-source fault and anomaly isolation has emerged as a critical research paradigm for diagnosing complex engineered infrastructures, where abnormal behaviors are distributed across heterogeneous sensors and coupled subsystems. Despite a sharp surge in literature, tracking this interdisciplinary domain is hindered by the massive volume of global publications and semantic fragmentation across computational boundaries. To address this challenge, this study aims to conduct a quantitative bibliometric analysis of the field while critically evaluating the reliability of OpenAlex as an open-access metadata source. Methodologically, a dataset of 2,385 records was compiled and analyzed. To reduce data dimensionality and uncover hidden thematic structures, the FP-growth association rule mining algorithm and the Mt-KaHyPar hypergraph partitioning framework were sequentially deployed to cluster a normalized controlled vocabulary of 3,741 unique keywords. The structural mapping revealed that the field experienced a sharp increase in research interest starting in 2025, heavily dominated by Chinese institutions and a high prevalence of preprint publications. Furthermore, high-utility search phrases such as Multi-source Root Cause Analysis were identified as optimal for literature aggregation. Crucially, the methodological audit exposed significant data quality anomalies: automated AI-driven indexing in OpenAlex introduced severe semantic detachment, frequently assigning erroneous descriptors like Fault (geology) to strictly technical papers, which directly degraded the internal consistency of the algorithmic clusters. Consequently, it is recommended that future science mapping frameworks extract keywords directly from publication titles and abstracts rather than relying on global index tags, while accounting for missing abstract data during extraction.
Keywords:
fault and anomaly isolation
; multi-source
; OpenAlex records
; bibliometric analysis
; FP-growth association rule
; Mt-KaHyPar hypergraph partitioning
Introduction
Multi-source fault and anomaly isolation has emerged as a critical research paradigm in the diagnostics of complex engineered systems. In modern industrial infrastructures, abnormal behaviors rarely manifest within a single isolated signal; rather, they are distributed across interacting components, heterogeneous sensors, and coupled subsystems [1,2,3]. Consequently, the primary objective of contemporary diagnostic frameworks extends beyond mere fault detection to encompass source isolation, the disambiguation of overlapping failure modes, and the structural tracing of disturbance propagation through multi-channel data streams. This challenge is particularly acute in cyber-physical systems, power grids, process industries, and autonomous platforms, where telemetry is abundant but system dynamics are inherently non-linear, partially observed, and confounded by stochastic uncertainty [4]. Thus, multi-source isolation has evolved into a distinct academic and practical domain situated at the intersection of fault diagnosis, anomaly detection, sensor fusion, and intelligent monitoring [5].
The methodological evolution of this domain has proceeded along several complementary trajectories. Classical model-based techniques, including state observers, parity-space methods, and analytical redundancy formulations, established the mathematical foundations for fault detection and isolation (FDI), primarily within linear and structured dynamical systems [6]. Nonetheless, the deterministic application of these approaches is frequently constrained by model incompleteness, parametric drifts, and poor scalability when applied to high-dimensional, heterogeneous configurations. To circumvent these limitations, data-driven methodologies have gained prominence [7,8]. These encompass multivariate statistical process monitoring, machine learning classifiers, deep neural networks, representation learning, and advanced graph- or attention-based architectures. Concurrently, hybrid frameworks that integrate physical first-principles with data-driven inference have been developed to enhance robustness, interpretability, and generalization in multi-source environments [9]. While seminal reviews track the maturation of intelligent diagnostics, they also highlight a field increasingly fragmented across disparate engineering and computational boundaries.
This disciplinary dispersion underscores the utility of bibliometric analysis. Unlike traditional narrative syntheses, bibliometric mapping provides a scalable quantitative evaluation of the macro-structural properties of a research domain, including publication trajectories, citation networks, institutional collaborations, and temporal thematic evolution. For a multi-disciplinary subject such as multi-source fault isolation—where literature is scattered across automation, mechanical, electrical, and computer science venues—bibliometric methods clarify how the field has evolved and locate its primary intellectual centers. Furthermore, algorithmic mapping can systematically identify emerging conceptual clusters, such as fault isolation in sensor networks, unsupervised anomaly detection in industrial assets, explainable diagnostics, and cross-domain transfer learning, which are frequently underrepresented in traditional surveys.
However, the validity of bibliometric outcomes remains strictly contingent upon the quality, completeness, and systemic biases of the underlying scholarly metadata. OpenAlex [10,11] has recently been adopted as an open-access bibliographic infrastructure, providing extensive coverage of publications, authors, institutions, and citation networks to facilitate transparent and reproducible bibliometric workflows. While its open-source data schema mitigates dependency on proprietary databases and supports strict methodological replication, it simultaneously introduces specific metadata challenges. Inconsistent institutional affiliations, algorithmic text parsing anomalies, semantic ambiguities within automated categorization pipelines, and uneven indexing depth between peer-reviewed journals, conferences, and rapidly expanding preprint repositories can distort publication counts, collaboration network metrics, and thematic clustering outputs [12]. These indexing phenomena are not merely technical; they directly impact the analytical interpretation of a scientific domain that inherently spans diverse disciplinary boundaries and distinct publication cultures.
Research Objectives and Contributions
Against this background, the present study undertakes a quantitative bibliometric analysis of the multi-source fault and anomaly isolation literature using OpenAlex metadata, structured around two complementary objectives.
First, it maps the global intellectual and structural configuration of the field—including its temporal evolution, core publication venues, leading research institutions, and prominent thematic patterns—while providing a critical methodological evaluation of the opportunities and limitations associated with OpenAlex metadata when analyzing an interdisciplinary engineering domain.
Second, it sequentially applies the FP-growth and Mt-KaHyPar algorithms to reduce data dimensionality and identify balanced thematic blocks, thereby establishing a structured framework to enable specialists to navigate and retrieve literature.
Materials and Methods
Bibliometric records were retrieved from the OpenAlex database for the period 2022–2026. Only journal articles and preprints were included in the analysis. The search queries applied to the title_and_abstract fields, targeting the topic "Multi-Source Fault and Anomaly Isolation", are detailed in Table 1.
The exported records were consolidated into a single dataset (Multi-source.csv), comprising 2,809 entries with the following metadata fields: title, author(s), publication year, citation count, open-access status, keywords, topic, source, institution, DOI, and abstract. Following deduplication, 2,385 unique records were retained for subsequent analysis.
Dimensionality reduction of the hypergraph constructed from the Keyword field was performed using the FP-growth algorithm as implemented by Borgelt [13].
Partitioning of the resulting hypergraph into balanced blocks was carried out using the Mt-KaHyPar hypergraph partitioning framework [14].
Results and Discussion
Queries to the OpenAlex database using the strings specified in Table 1 were conducted for the 2022–2026 period, identifying the occurrence of these strings within titles and abstracts. The analysis was restricted exclusively to research articles and preprints. The column N denotes the query results based on exact string matching, while N+ represents the results based on the co-occurrence of all words (i.e., the terms were linked via the boolean AND operator).
The results presented in the table clearly illustrate the reason for discarding the literature search based on the exact matching of the listed strings. Consequently, data collection for this study targeted the broadest possible context. It is also possible to employ queries with intermediate constraints; for example, the query 'Multi-sensor "Anomaly Localization"' yields 11 articles and preprints, whereas no publications are found via exact matching, and the query 'Multi-sensor AND Anomaly AND Localization' returns 201 results.
The term Multi-source Root Cause Analysis was not only identified exactly 5 times via literal string matching, but it also yielded the highest volume of retrieved publications based on the co-occurrence of its constituent words; consequently, it can be recommended for literature aggregation purposes.
Bibliometric Characteristics of the Selected Publication Dataset
Chronological Distribution of Publications: Year (Publication Count), 2026 (772), 2025 (702), 2024 (265), 2023 (176), 2022 (129). The data are current as of July 2, 2026. Academic interest in the subject matter experienced a sharp increase in 2025; furthermore, by mid-2026, the number of indexed publications had already exceeded the total volume recorded for the entirety of 2025.
The majority of the analyzed papers are published under Open Access models—Open Access (1,231), Non-Open Access (813)—thereby facilitating more detailed content evaluation by domain experts.
The substantial volume of preprint publications largely accounts for this high proportion of Open Access materials. The following list itemizes publication venues with a minimum representation of 8 papers during the studied period: Source (Publication Count), Zenodo (CERN European Organization for Nuclear Research) (134), SSRN Electronic Journal (105), arXiv (Cornell University) (83), ArXiv.org (58), Sensors (24), Research Square (22), IEEE Access (21), Open MIND (19), IET conference proceedings (17), Scientific Reports (16), Applied Sciences (16), DOAJ (DOAJ: Directory of Open Access Journals) (14), Preprints.org (13), Measurement Science and Technology (11), Mechanical Systems and Signal Processing (11), Expert Systems with Applications (10), Reliability Engineering & System Safety (10), IEEE Sensors Journal (10), IEEE Transactions on Instrumentation and Measurement (10), Energies (10), Engineering Applications of Artificial Intelligence (9), Measurement (9), Electronics (9), Journal of Systems and Software (8), IEEE Internet of Things Journal (8), Remote Sensing (8). Preprint repositories are italicized in this list. A significant share of the publications is attributed to highly reputable journals, including IEEE Access, Open MIND, Scientific Reports, and Applied Sciences. The relevance of systemic problems and measurement issues is underscored by their coverage across virtually all listed venues, alongside a strong presence within IEEE publications.
An analysis of the leading institutions with more than 14 publications highlights a clear dominance of Chinese research entities: Institution (Publication Count), Centre National de la Recherche Scientifique (24), Tsinghua University (24), Chinese Academy of Sciences (23), Chongqing University (23), Shanghai Electric (China) (21), State Grid Corporation of China (China) (21), Xi'an Jiaotong University (20), Northwestern Polytechnical University (20), China Southern Power Grid (China) (20), Zhejiang University (18), Beihang University (18), Harbin Institute of Technology (18), North China Electric Power University (17), Shanghai Jiao Tong University (15), Shandong University (15), Power Grid Corporation (India) (15), University of Chinese Academy of Sciences (14).
The Centre National de la Recherche Scientifique (CNRS) stands out as the largest French public research organization, consolidating state institutions specialized in both fundamental and applied research.
Methodological Note on Data Anomalies: It is probable that the erroneous institutional affiliation 'Film Independent' instead of 'independent researcher' stems from an algorithmic error in parsing text strings within OpenAlex. For instance, in the publication indexed via DOI 10.56975/tijer.v12i5.158200, the author Kayalvizhi Rajagopal is explicitly designated as an 'independent researcher'. Venues such as the TIJER - International Research Journal raise valid concerns regarding peer-review rigor, whereas preprints inherently lack comprehensive validation, undergoing only minimal technical screening. Consequently, researchers leveraging such sources must exercise manual verification to ensure data credibility.
The Topic field in OpenAlex (Academic subject area or research topic) reflects the core research domains. The following list details the topics associated with at least 10 publications within the analyzed bibliometric records: Topic (Publication Count), Anomaly Detection Techniques and Applications (234), Software System Performance and Reliability (137), Machine Fault Diagnosis Techniques (79), Fault Detection and Control Systems (71), Power Systems Fault Detection (50), Software Engineering Research (32), Smart Grid Security and Resilience (25), Advanced Battery Technologies Research (22), Software Testing and Debugging Techniques (22), Robotics and Sensor-Based Localization (21), HVDC Systems and Fault Protection (20), Network Security and Intrusion Detection (19), earthquake and tectonic studies (19), Structural Health Monitoring Techniques (18), Geochemistry and Geologic Mapping (16), Photovoltaic System Optimization Techniques (15), Cosmology and Gravitation Theories (14), Water Systems and Optimization (13), Advanced Graph Neural Networks (13), Industrial Vision Systems and Defect Detection (13), Power Transformer Diagnostics and Insulation (12), Indoor and Outdoor Localization Technologies (11), Electrical Fault Detection and Protection (10), Time Series Analysis and Forecasting (10), Risk and Safety Analysis (10), Infrastructure Maintenance and Monitoring (10).
These terms strongly align with the core theme of 'Anomaly and Fault Detection'. Using this field as a filter allows researchers to significantly narrow down search results. For instance, applying the filter Topic='Anomaly Detection Techniques and Applications' isolates approximately 30,000 papers and preprints published between 2022 and 2026, effectively narrowing down the global research corpus, which comprises an estimated 35,670,000 total works for the same timeframe.
The Keyword metadata field serves as a secondary useful tool for facilitating the search for relevant literature. Keywords in OpenAlex are selected from a controlled vocabulary, which ensures a normalized spelling of terms. The most frequent keywords assigned to the analyzed records include: Keyword (Publication Count), Computer science (637), Anomaly detection (492), Artificial intelligence (388), Fault (geology) (343), Engineering (329), Pattern recognition (psychology) (236), Anomaly (physics) (225), Fault detection and isolation (222), Geology (172), Physics (166), Data mining (163), Mathematics (143), Machine learning (139), Feature (linguistics) (139), Graph (135), Root cause (128), Robustness (evolution) (117), Root cause analysis (110), Identification (biology) (107), Reliability (semiconductor) (102), Electrical engineering (102), Sensor fusion (101), Artificial neural network (101). While this list contains many high-level or generic terms, it also includes keywords directly related to the subject matter of this study, such as Fault detection and isolation. Consequently, by using this specific keyword as a filter, a domain expert can locate the exact list of publications to which OpenAlex has assigned this value.
The total number of papers and preprints published between 2022 and 2026 that carry this keyword is approximately 21,620 works. This relatively large volume is due to the fact that it encompasses papers on general fault detection, which is a broad topic. A direct check reveals that in a significant number of publications, the term Fault detection appears, but without the term isolation. If the search for the term isolation is added to the keyword filter, the number of works drops to approximately 1,737. Furthermore, in the general keyword index, one can find phrases that include this word—for example, Galvanic isolation—but they are unrelated to the subject matter under consideration.
Note: Applying the same filters for years and publication types to a comprehensive query:
((cross-modal+OR+multi-sensor+OR+multi-source+OR+multivariate+OR+hypergraph)+AND+(localization+OR+isolation+OR+causal)+AND+(fault+OR+anomaly))
yielded 2,227 works, which is comparable to the 2,385 records used in this study. While these two queries are not completely identical, the comparability of their results is what remains significant.
The preceding sections examined the core bibliometric characteristics of the publication dataset compiled using the specific search terms itemized in Table 1.
To contextualize these findings within the broader academic landscape, OpenAlex provides comprehensive data dumps that characterize various metadata fields across its entire global repository. The subsequent sections outline the baseline characteristics of the Keyword and Topics fields derived from these global datasets.
Baseline Characteristics of the 'Keyword' Field
A comprehensive characterization of the Keyword metadata field can be obtained by retrieving the part_0000.parquet data dump from the OpenAlex repository (available at https://openalex.s3.amazonaws.com/browse.html#data/parquet/keywords/updated_date=2026-06-26/), with data current as of July 12, 2026.
This file contains a total of 59,033 distinct keywords. Within this dataset, Fault detection and isolation is assigned to 87,846 records. The next most frequent related keywords are Fault tolerance (68,162 records) and Fault injection (7,645 records). Based on these distributions, when deploying these keywords as search filters, Fault detection and isolation emerges as the most appropriate choice for the specific scope of this study.
Regarding the term Anomaly, the most contextually relevant keywords within the global master list are Anomaly detection (101,332 records), Anomaly-based intrusion detection system (8,254 records), and Anomaly (physics) (127,809 records). The keyword Anomaly detection can therefore serve as a broad filter for aggregating publications relevant to the subject matter analyzed in this work.
Within our selected dataset, the keyword Fault detection and isolation occurs 229 times. Notably, there is a high frequency of the term Fault (geology), which appears 357 times. Furthermore, in these same records, the keyword Anomaly detection occurs 568 times, while Anomaly (physics) is recorded 262 times.
Baseline Characteristics of the 'Topics' Field
A comprehensive characterization of the Topics metadata field can be obtained by retrieving the part_0000.parquet data dump from the OpenAlex repository (available at https://openalex.s3.amazonaws.com/browse.html#data/parquet/topics/updated_date=2026-06-26/). This dataset contains a total of 4,482 distinct records, with data current as of July 12, 2026.
The names of the topics containing the term Fault, along with their respective global record counts, are as follows: Power Systems Fault Detection (39,677 records), HVDC Systems and Fault Protection (45,835 records), Machine Fault Diagnosis Techniques (67,958 records), Fault Detection and Control Systems (143,972 records), Distributed systems and fault tolerance (94,367 records), and Electrical Fault Detection and Protection (28,925 records).
Only a single topic containing the term Anomaly was identified within the global list: Anomaly Detection Techniques and Applications (77,110 records).
Based on this distribution, Fault Detection and Control Systems emerges as the most appropriate topic for the specific scope of this study.
Within our analyzed dataset, the most frequently occurring topic is Anomaly Detection Techniques and Applications, appearing 270 times. This is followed by: Machine Fault Diagnosis Techniques (83 occurrences), Fault Detection and Control Systems (73 occurrences), Power Systems Fault Detection (51 occurrences), HVDC Systems and Fault Protection (22 occurrences), and Electrical Fault Detection and Protection (10 occurrences).
Association Rule Mining via the FP-Growth Algorithm
Typically, a research topic is adequately defined by 3 to 5 general keywords. Author-provided keywords usually consist of 5 to 7 terms, where 1 or 2 are highly specific to the core contribution of the paper, and the remaining 3 to 5 terms establish the broader context of the publication. Given that the controlled vocabulary utilized by OpenAlex is limited to 59,033 terms, its assigned keywords are more appropriately categorized as high-level, general descriptors.
Using the FP-growth algorithm, the co-occurrence of itemsets containing 3 and 5 keywords across the 2,385 records was evaluated with a minimum support threshold of 1%.
At this support level, a total of 202 three-term itemsets were identified; the 5 itemsets with the highest support values are presented in Table 2.
Given that the term artificial_intelligence is a subdiscipline within computer_science, the data in this table indicate that the combination of artificial_intelligence AND anomaly_detection represents the most frequent thematic pairing.
The frequent five-term itemsets uncovered by the algorithm are shown in Table 3. A total of 22 five-term itemsets met the 1% support threshold.
In these results, computer_science serves as the most generic high-level term. The combination machine_learning AND anomaly_detection AND data_mining AND artificial_intelligence represents the most prominent general theme, whereas the combination seismology AND artificial_intelligence AND fault_(geology) emerges as the most specific thematic cluster, which initially appears to indicate a promising direction for future research.
However, a direct search within the dataset's keyword metadata revealed a discrepancy: fault_(geology) appears in 357 records, dropping to 31 records when intersected with seismology, and further reducing to only 14 records when artificial_intelligence is added. A closer examination of the titles and abstracts of these 14 publications revealed that the terms seismology and geology are completely absent from the actual text. For example, in a recent and highly cited paper [15] that is highly relevant to the core technical scope of our study, OpenAlex assigned the following keywords: Inference|Anomaly detection|Dual (grammatical number)|Deep learning|Causal inference|Artificial intelligence|Computer science|Anomaly (physics)|Fault (geology)|Channel (broadcasting)|Machine learning|Geology|Telecommunications|Seismology|Mathematics. Despite this assignment, the publication contains no discussion whatsoever regarding seismology or geology. This serves as another compelling example highlighting the critical necessity of treating OpenAlex-generated metadata with caution.
Analysis of the 'Topics' Field
In contrast to keywords, each publication is assigned to only a single topic within the OpenAlex classification framework.
To evaluate this distribution, we selected the two publications with the highest assignment values for the topics Fault Detection and Control Systems and Anomaly Detection Techniques and Applications from the 2,385 retrieved records. These records are structured in the following format: Title [DOI] → Year → Citation Count → Keywords.
Root cause diagnosis for process faults based on multisensor time-series causality discovery [16] → 2022 → 43 → Process (computing)|Causality (physics)|Root cause|Computer science|Fault (geology)|Data mining|Time series|Path (computing)|Artificial intelligence|Series (stratigraphy)|Root cause analysis|Relevance (law)|Machine learning|Pattern recognition (psychology)|Algorithm|Engineering|Reliability engineering.
Note: As of July 12, 2026, this publication is indexed on ScienceDirect with a 2023 publication year and 46 citations. This discrepancy suggests that OpenAlex indexed the article upon its initial acceptance, as confirmed by the publisher's metadata: Accepted December 13, 2022; Available online December 21, 2022. The highlighted keywords justify the thematic classification of the paper, while the phrase "faults based on multisensor time-series" in the title highlights the multi-source nature of the underlying data. The original keywords provided by the journal are: Root cause diagnosis; Time-series; Causality discovery; Dilated convolutional neural network; Layerwise relevance propagation. A full-text analysis of the paper revealed no occurrences of the term "geology."
Deep Learning-Based Fault Prediction in Wireless Sensor Network Embedded Cyber-Physical Systems for Industrial Processes [17] → 2022 → 41 → Computer science|Cyber-physical system|Fault (geology)|Process (computing)|Machine learning|Wireless sensor network|Artificial intelligence|Deep learning|Data mining|Real-time computing|Computer network.
Note: The original keywords provided by the journal IEEE Access are: {Sensors; Wireless sensor networks; Solid modeling; Wireless communication; Uncertainty; Predictive models; Machine learning; Cyber-physical systems; deep learning; fault classification; LSTM; multivariate multi-step prediction; predictive maintenance; TCN; time-series; imbalanced data; uncertainty propagation}. A clear discrepancy is observed here: OpenAlex assigned the keyword Fault (geology) to the publication, whereas the journal specified fault classification. A full-text analysis of the paper revealed no occurrences of the term "geology."
These examples demonstrate that the keywords assigned by OpenAlex must be treated with caution, which also qualifies the assumptions regarding the highly promising research directions identified in the previous section. OpenAlex utilizes BGE M3-Embedding models to compute semantic similarity scores between each candidate term from its controlled vocabulary and the actual text of the article's title and abstract. This case highlights a broader methodological principle: while AI-driven approaches enable the scaling of tasks that are difficult to implement via classical methods, their automated outputs require systematic verification. For comparison, the current version of the IEEE Thesaurus contains approximately 12,760 domain-specific terms for engineering and technology alone, whereas the OpenAlex controlled vocabulary contains 59,033 terms (as of July 12, 2026) to cover all academic disciplines. However, as the OpenAlex keyword index is expanding rapidly, improvements in automated keyword assignment accuracy can be expected in future iterations.
Hypergraph Partitioning via the Mt-KaHyPar Algorithm for Keyword Thematic Clustering
In this study, publication themes are analyzed using the metadata from the Keyword field. These keywords are generated based on a controlled vocabulary co-developed by OpenAlex and the Centre for Science and Technology Studies (CWTS) at Leiden University. Notably, CWTS is a premier institution in the field of bibliometrics and the developer of the widely used VOSviewer software.
Following data deduplication, a dataset of 2,385 unique records was used. The total number of keyword instances across these records is 21,685, yielding an average of approximately 9 terms per publication. The number of unique keywords within this corpus stands at 3,741.
The keyword co-occurrence patterns within these records can be modeled as hyperedges in a hypergraph framework. A hypergraph composed of 2,385 hyperedges and 3,741 nodes is computationally small and easily processed analytically; however, conducting a thematic analysis to identify dominant research trends presents significant challenges. For instance, visually rendering such a hypergraph remains difficult, even when partitioned into balanced subgraphs. Furthermore, if this hypergraph were partitioned into 4 distinct blocks, each block would still contain approximately 935 terms. This density is overly redundant for clearly characterizing the thematic focus of a block based on its constituent keywords, and it limits the practical utility of a graphical representation for domain experts seeking to aggregate literature on a selected topic.
Consequently, it is necessary to reduce the dimensionality of the hypergraph, filtering it to retain only the core subset of semantically significant terms. Empirical evidence indicates that a combination of 3 to 5 keywords is generally sufficient to effectively filter and retrieve relevant literature.
One method to reduce the dimensionality of the analyzed hypergraph is to deploy the FP-growth algorithm, which identifies the most frequent itemsets within a hyperedge. Furthermore, by restricting the size of co-occurring term sets to 3–5 keywords, we generate weighted hyperedges. Consequently, during subsequent hypergraph partitioning into balanced blocks using the Mt-KaHyPar algorithm, the cutting of more frequent itemsets is penalized more heavily.
To prepare the Keyword fields for further processing, a normalization pipeline was applied: text was converted to lowercase, spaces within keywords were replaced with underscores, and parentheses were replaced with standard four-character mnemonic abbreviations, specifically LPAR (Left Parenthesis) and RPAR (Right Parenthesis). Following this, the pipe separator used in OpenAlex records was replaced with a space character to match the input format required by the fpgrowth utility.
At a 1% support threshold (fpgrowth -s1m3n5), the execution yielded 324 records (hyperedges) containing 42 unique keywords (hypergraph nodes).
Methodological Note: Rationale for Excluding Two-Term Co-occurrences
In addition to the previously discussed principle that the co-occurrence of 3 to 5 keywords establishes a well-balanced thematic context for a publication, the following empirical comparisons support this approach:
Vocabulary Size Constraints: At a 1% support threshold, the 3–5 term constraint (fpgrowth -s1m3n5) identifies 42 unique keywords across all itemsets, whereas allowing two-term pairs (fpgrowth -s1m2n2) increases this number to 70. This represents a substantial expansion of the vocabulary required to describe the thematic blocks, increasing semantic complexity.
Term Frequency Distributions:
Under the first configuration (-s1m3n5), the term occurrence frequencies within the itemsets are distributed as follows: computer_science (205), artificial_intelligence (147), engineering (105), anomaly_detection (86), data_mining (80), machine_learning (59), fault_(geology) (46), anomaly_(physics) (42), pattern_recognition_(psychology) (39), mathematics (31), multivariate_statistics (30), physics (29), geology (26), electrical_engineering (23), reliability_engineering (22), seismology (18), and fault_detection_and_isolation (15).
Under the second configuration (-s1m2n2), the frequencies drop markedly: computer_science (53), anomaly_detection (35), engineering (31), artificial_intelligence (28), fault_(geology) (21), pattern_recognition_(psychology) (16), data_mining (15), anomaly_(physics) (13), geology (13), mathematics (12), machine_learning (12), and fault_detection_and_isolation (12). Due to the structural properties of larger itemsets (3–5 terms), the total frequency counts of individual keywords are higher in the first variant, causing high-level, generic terms to appear more consistently.
Itemset Weight Distributions: Analysing the relative support metrics at the itemset level rather than individual terms reveals the following proportions:
The support weight share generated by fpgrowth -s1m3n5 for itemsets containing fault_detection_and_isolation is 22.89 / 521.43 (where 22.89 is the cumulative support of itemsets containing this term, and 521.43 is the total support sum of all itemsets).
The support weight share generated by fpgrowth -s1m2n2 for itemsets containing fault_detection_and_isolation is 24.99 / 492.83.
For itemsets containing fault_detection, the support share under fpgrowth -s1m3n5 is 142.51 / 521.43.
For itemsets containing fault_detection, the support share under fpgrowth -s1m2n2 is 84.53 / 492.83.
These comparative metrics demonstrate that the primary advantage of the fpgrowth -s1m3n5 configuration lies in its capacity to significantly restrict the unique keyword vocabulary while maintaining a 1% support threshold. Itemsets composed of 3 to 5 terms at a 1% support level generate a higher number of meaningful combinations from a more compact set of unique terms.
A hypergraph constructed based on the results of the fpgrowth -s1m3n5 execution was partitioned into four balanced blocks using the Mt-KaHyPar hypergraph partitioning algorithm with the following configuration parameters: -k 4 -e 0.03 -o km1 --preset-type highest_quality. The partitioning results are presented in Table 4.
Block 0, which contains the most frequent terms that also bear a more generic nature, appears consistent. However, Block 1—which incorporates the previously identified problematic keywords such as Fault (geology), Geology, and Seismology—exhibits excessive heterogeneity. The fundamental limitations within these underlying keywords directly degraded the quality of the thematic hypergraph partitioning. Due to these metadata inconsistencies, the deployment of Mt-KaHyPar to generate balanced thematic blocks from OpenAlex keywords must be approached with a high degree of caution. Consequently, it is advisable for future research to conduct a separate study utilizing keywords extracted directly from the actual text of publication titles and abstracts.
Implementing a controlled vocabulary, as demonstrated by OpenAlex, represents a highly effective approach to indexing global scientific literature. However, like any methodological framework, it possesses inherent trade-offs. Terms derived from a controlled vocabulary often exhibit low frequency within the actual text of publication titles and abstracts—a challenge common to keywords in general, including author-provided ones. Conversely, domain experts primarily rely on the explicit semantic content of these titles and abstracts when searching for relevant literature. Given the massive, multidisciplinary scope of OpenAlex, constructing a universally precise thesaurus is inherently complex. More targeted, domain-specific vocabularies—such as those maintained by IEEE and ACM—traditionally yield higher precision, yet scaling such specialized frameworks to encompass the vast disciplinary breath of OpenAlex remains a significant challenge. Furthermore, the index keyword frameworks of other major platforms, such as Scopus, do not appear to be publicly accessible, and their specific assignment methodologies remain proprietary. Consequently, navigating the trade-offs between broad multidisciplinary coverage and localized semantic precision remains a critical consideration for automated bibliometric systems.
Conclusion
The quantitative bibliometric analysis of OpenAlex records within the domain of "Multi-Source Fault and Anomaly Isolation" revealed a sharp surge in research interest starting in 2025. Institutional affiliations are heavily dominated by Chinese entities, and a substantial portion of the global literature is distributed via preprint repositories.
Based on our empirical retrieval metrics, the most high-utility target phrases for future publication search and aggregation include:
- Multi-source Root Cause Analysis
- Multi-source Fault Localization
- Causal Anomaly Localization
- Multi-sensor Fault Localization
- Multi-sensor Anomaly Localization
Furthermore, association rule mining identified that the most contextually cohesive and frequent OpenAlex keyword sets for this interdisciplinary domain are:
- machine_learning, anomaly_detection, data_mining, artificial_intelligence, computer_science
- physics, anomaly_detection, anomaly_(physics), artificial_intelligence, computer_science
Crucially, this study underscores that metadata automatically assigned by OpenAlex can occasionally exhibit significant semantic detachment from the actual manuscript content, as exemplified by the frequent and erroneous assignment of the fault_(geology) descriptor to strictly engineered systems research.
The proposed framework—which sequentially deploys the FP-growth and Mt-KaHyPar algorithms to reduce keyword hypergraph dimensionality and isolate balanced thematic blocks—demonstrated the technical viability of algorithmic data pruning. However, the observed inaccuracies in automated keyword tagging directly degraded the internal consistency of the resulting clusters.
Consequently, future research should focus on validating this dual-algorithm framework using keywords extracted directly from publication titles and abstracts rather than relying on global index tags. A potential limitation for this approach remains the inconsistent completeness of the Abstract field within raw data exports from OpenAlex, which warrants further data-cleaning protocols.
Funding
The work was funded by the Ministry of Science and Higher Education of the Russian Federation (State Assignment No. 125021302095-2).
References
- Venkatasubramanian, V.; Rengaswamy, R.; Yin, K.; Kavuri, S.N. A review of process fault detection and diagnosis: Part I: Quantitative model-based methods. Comput. Chem. Eng. 2003, 27(3), 293–311. [Google Scholar] [CrossRef]
- A review of process fault detection and diagnosis Part II: Qualitative models and search strategies 10.1016/S0098-1354(02)00161-8
- A Review of Process Fault Detection and Diagnosis Part III: Process History Based Methods March. Comput. Chem. Eng. 2003, 27(3), 327–346. [CrossRef]
- Piardi, L.; Leitão, P.; de Oliveira, A. S. Fault-Tolerance in Cyber-Physical Systems: Literature Review and Challenges. 2020 IEEE 18th International Conference on Industrial Informatics (INDIN), Warwick, United Kingdom, 2020; pp. 29–34. [Google Scholar] [CrossRef]
- Cao, Yu. Intelligent Anomaly Detection with Attention-Based Fusion of Multi-Source Heterogeneous Data. In Proceedings of the 2025 6th International Conference on Computer Science and Management Technology (ICCSMT '25), Association for Computing Machinery, New York, NY, USA, 2026; pp. 590–595. [Google Scholar] [CrossRef]
- Non-Analytical Approaches to Model-Based Fault Detection and Isolation. [CrossRef]
- Data-driven sensor fault diagnosis systems for linear feedback control loops. [CrossRef]
- Yin, S.; Li, X.; Gao, H.; Kaynak, O. Data-Based Techniques Focused on Modern Industry: An Overview. IEEE Trans. Ind. Electron. 2015, vol. 62(no. 1), 657–667. [Google Scholar] [CrossRef]
- Physics-based and data-driven hybrid modeling in manufacturing: a review. [CrossRef]
- The OpenAlex database in review: Evaluating its applications, capabilities, and limitations. [CrossRef]
- OPEN BIBLIOGRAPHIC DATABASES: IN SEARCH OF AN ALTERNATIVE TO SCOPUS AND THE WEB OF SCIENCE. [CrossRef]
- Mitchell, M.; Johnson, S.L.; Porter, M.D. Affiliation errors and distorted researcher mobility: evidence from OpenAlex and Scopus. In Scientometrics; 2026. [Google Scholar] [CrossRef]
- Borgelt, C. An implementation of the FP-growth algorithm. In Proceedings of the 1st International Workshop on Open Source Data Mining: Frequent Pattern Mining Implementations; 2005; Chicago, IL, USA, ACM: New York, 2005; pp. 1–5. [Google Scholar] [CrossRef]
- Gottesbüren, L.; Heuer, T.; Sanders, P.; Schulz, C.; Seemaier, D. Deep multilevel hypergraph partitioning. In Proceedings of the Platform for Advanced Scientific Computing Conference; 2023, Davos, Switzerland; ACM: New York, 2023; pp. 1–11. [Google Scholar] [CrossRef]
- Xing, S.; Wang, Y.; Liu, W. Multi-dimensional anomaly detection and fault localization in microservice architectures: a dual-channel deep learning approach with causal inference for intelligent sensing. Sensors 2025, 25(11), 3396. [Google Scholar] [CrossRef] [PubMed]
- Wang, S.; Zhao, Q.; Han, Y.; Wang, J. Root cause diagnosis for process faults based on multisensor time-series causality discovery. J. Process Control. 2023, 122, 27–40. [Google Scholar] [CrossRef]
- Ruan, H.; Dorneanu, B.; Arellano-Garcia, H.; Xiao, P.; Zhang, L. Deep learning-based fault prediction in wireless sensor network embedded cyber-physical systems for industrial processes. IEEE Access. 2022, 10, 10867–10879. [Google Scholar] [CrossRef]
Table 1.
OpenAlex query results.
| Query string | N+ | N |
| Causal Anomaly Isolation | 86 | 0 |
| Causal Anomaly Localization | 281 | 0 |
| Causal Fault Isolation | 56 | 0 |
| Causal Fault Localization | 163 | 1 |
| Cross-modal Anomaly Isolation | 39 | 0 |
| Cross-modal Anomaly Localization | 148 | 0 |
| Cross-modal Fault Isolation | 18 | 0 |
| Cross-modal Fault Localization | 44 | 0 |
| hypergraph Anomaly Localization | 8 | 0 |
| hypergraph Fault Localization | 10 | 0 |
| Multi-sensor Anomaly Localization | 201 | 0 |
| Multi-sensor Fault Localization | 250 | 1 |
| Multi-source Anomaly Isolation | 144 | 0 |
| Multi-source Fault Isolation | 129 | 0 |
| Multi-source Fault Localization | 334 | 0 |
| Multi-source Root Cause Analysis | 381 | 5 |
| Multi-source Root Cause Identification | 80 | 0 |
| Multivariate Anomaly isolation | 163 | 0 |
| Multivariate Anomaly Localization | 175 | 0 |
| Root Cause Node Identification | 99 | 0 |
Table 2.
Most frequent three-keyword itemsets in the 'Keyword' field.
| Keyword Itemset | Support (%) |
| engineering artificial_intelligence computer_science | 7.25367 |
| data_mining artificial_intelligence computer_science | 5.66038 |
| machine_learning artificial_intelligence computer_science | 5.45073 |
| artificial_intelligence anomaly_detection computer_science | 5.15723 |
| mathematics artificial_intelligence computer_science | 4.02516 |
Table 3.
Most frequent five-keyword itemsets in the 'Keyword' field.
| Keyword Itemset | Support (%) |
| machine_learning anomaly_detection data_mining artificial_intelligence computer_science | 1.84486 |
| physics anomaly_detection anomaly_(physics) artificial_intelligence computer_science | 1.80294 |
| data_mining pattern_recognition_(psychology) anomaly_detection computer_science artificial_intelligence | 1.42558 |
| data_mining anomaly_(physics) artificial_intelligence computer_science anomaly_detection | 1.38365 |
| seismology artificial_intelligence fault_(geology) geology computer_science | 1.38365 |
Table 4.
Distribution of keywords across blocks and their occurrence frequencies within the analyzed bibliometric records.
Table 4.
Distribution of keywords across blocks and their occurrence frequencies within the analyzed bibliometric records.
| Keyword | CountOfKeyword | NoBlock |
| Computer Science | 205 | 0 |
| Artificial Intelligence | 147 | 0 |
| Engineering | 105 | 0 |
| Anomaly Detection | 86 | 0 |
| Data Mining | 80 | 0 |
| Machine Learning | 59 | 0 |
| Anomaly (physics) | 42 | 0 |
| Pattern Recognition (psychology) | 39 | 0 |
| Mathematics | 31 | 0 |
| Multivariate Statistics | 30 | 0 |
| Physics | 29 | 0 |
| Fault (geology) | 46 | 1 |
| Geology | 26 | 1 |
| Reliability Engineering | 22 | 1 |
| Seismology | 18 | 1 |
| Fault Detection and Isolation | 15 | 1 |
| Root Cause | 14 | 1 |
| Root Cause Analysis | 14 | 1 |
| Algorithm | 12 | 1 |
| Root (linguistics) | 9 | 1 |
| Power (physics) | 5 | 1 |
| Real-Time Computing | 5 | 1 |
| Feature (linguistics) | 4 | 2 |
| Graph | 4 | 2 |
| Theoretical Computer Science | 4 | 2 |
| Control Theory (sociology) | 2 | 2 |
| Telecommunications | 2 | 2 |
| Artificial Neural Network | 1 | 2 |
| Sensor Fusion | 1 | 2 |
| Materials Science | 1 | 2 |
| Computer Vision | 1 | 2 |
| Electrical Engineering | 23 | 3 |
| Voltage | 7 | 3 |
| Electronic Engineering | 7 | 3 |
| Time Series | 5 | 3 |
| Statistics | 4 | 3 |
| Transformer | 3 | 3 |
| Series (stratigraphy) | 3 | 3 |
| Acoustics | 2 | 3 |
| Process (computing) | 1 | 3 |
| Benchmark (surveying) | 1 | 3 |
| Deep Learning | 1 | 3 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.