Submitted:
21 August 2026
Posted:
24 August 2026
You are already at the latest version
Abstract
Attribute reduction is a fundamental problem in rough set theory that seeks a minimal subset of condition attributes preserving the discernibility capability of the original decision system. Traditional heuristic methods—greedy forward selection and backward elimination—suffer from two critical deficiencies: the absence of principled attribute-ranking criteria (which leads to over- or under-selection) and the omission of statistical discriminability from the search process. To address these issues, this paper presents MFDRRS, a hybrid framework that integrates three complementary strategies: (1) Fisher discriminant ratio (FDR) evaluation of attribute discriminability based on between- and within-class scatter; (2) FDR-based filtering to select statistically significant candidate attributes; and (3) greedy forward selection guided by the rough-set dependency degree, enhanced with an FDR-driven termination criterion. FDR computation quantifies the discriminative power of individual attributes, thereby facilitating the selection of an optimally discriminative candidate subset. Since the traditional FDR is defined for continuous features, we adapt the continuous-domain FDR to nominal data by proposing a modified Fisher discriminant ratio (MFDR). The MFDR first maps nominal attributes via one-hot encoding and then computes the between-class and within-class scatter matrices. The subsequent FDR-filtering and rough-set dependency analysis guarantee the optimality of the selected attribute subset. Evaluated on synthetic and real-world UCI datasets against six baselines—classical rough-set greedy forward/backward selection, exhaustive search, MI ranking, MaxRelevance, and JMI—the proposed method achieves more compact reductions with fewer attributes, maintains competitive dependency degrees, and avoids the over-selection characteristic of classical backward elimination.
Keywords:
rough set theory
; attribute reduction
; Fisher discriminant ratio
; nominal data
; feature selection
; heuristic search
1. Introduction
In the era of big data, the sheer volume of information is accompanied by substantial redundancy. Effectively extracting actionable knowledge from data remains a fundamental challenge in data mining, as the knowledge discovered by conventional data mining approaches often fails to directly support meaningful decision-making actions[1,2]. Knowledge reduction, particularly attribute reduction in rough set theory, has emerged as an important technique for eliminating redundant and irrelevant attributes while preserving the essential decision-making information of the original data, making it particularly valuable for high-dimensional data analysis and knowledge discovery[3]. In the context of machine learning, knowledge reduction is also referred to as feature selection or attribute reduction[4,5]. The core objective of attribute reduction is to minimize the size of the feature set as much as possible while maintaining the classification ability of the classifier unchanged. To achieve this goal, the researchers have proposed numerous processing methods, which can be broadly classified into the following categories: filter, wrapper and embedded methods[6]. Filter methods evaluate features based on their statistical properties and are independent of specific learning algorithms, offering high computational efficiency and scalability. Common examples include Fisher Score[7], Relief-based algorithms[8], mutual information[9], and minimum redundancy maximum relevance (mRMR)[10] etc. However, since these methods typically evaluate features individually or rely on local statistics, they fail to adequately capture inter-feature correlations and joint effects, limiting their performance on complex data. Wrapper methods evaluate feature subsets based on the predictive performance of a learning algorithm, employing search strategies such as sequential search, genetic algorithms, and particle swarm optimization to identify optimal combinations. Although they often achieve high classification accuracy, the need for iterative model training incurs substantial computational overhead, which severely limits their applicability to large-scale, high-dimensional data. Embedded methods integrate feature selection into model training by imposing regularization constraints (e.g., L1 and L2 norms), as exemplified by LASSO[11], Elastic Net[12], tree-based models[13], and sparse learning approaches[14]. By balancing high predictive performance with computational efficiency, these methods have become a prominent research focus in supervised feature selection. The growing prevalence of high-dimensional data and complex structures has made unsupervised feature selection a prominent research topic. Without class labels, these methods assess feature relevance by preserving the local manifold, global geometry, or clustering structure of the data. Typical techniques—Laplacian Score[15], SPEC[16], MCFS[17], NDFS[18], and UDFS[19]—select features by constructing similarity graphs or graph Laplacian matrices to retain manifold structure.
The proliferation of high-dimensional datasets across genomics, healthcare, and other data-intensive applications has introduced significant challenges for feature selection, including increased computational costs, substantial feature redundancy, and the difficulty of effectively capturing complex dependencies among features[20,21,22,23]. Requiring no prior knowledge or distributional assumptions, rough sets are inherently well-suited for high-dimensional complex data, providing effective feature reduction and strong interpretability [5]. In the framework of Pawlak’s rough set theory[24], attribute reduction is formalized through the dependency degree , which measures the proportion of objects correctly classified by an attribute subset , where C is the full set of conditional attributes and d is the decision attribute. A reduct is any subset that satisfies and no subset of R achieves the same dependency degree. However, the search for all reducts is NP-hard[25,26], motivating the use of heuristic strategies. The two most widely used heuristics are greedy forward selection [27] and greedy backward elimination [28]. Both methods rely on the monotonicity of , iteratively adding or removing the attribute that best improves (or least reduces) the dependency degree. While simple and effective in practice, these methods suffer from two critical limitations that restrict their utility in real-world applications. (1) lack of principled criteria for attribute ranking. Without principled criteria for prioritizing attributes, both methods rely solely on marginal gains in , which can be misleading when attributes differ in scale, distribution, or cardinality. An attribute with only modest conditional gain may still exhibit strong statistical discriminability—a property that pure rough-set heuristics cannot capture. Most existing rough-set heuristics treat candidate attributes uniformly[29], lacking a statistical prior to distinguish informative features from noisy ones. (2) Ad-hoc stopping criteria lacking principled termination conditions. Both forward and backward searches terminate based on user-specified thresholds (e.g., minimum gain ). These hyperparameters have no principled foundation and require per-dataset tuning, placing a considerable burden on practitioners. As Liu et al. [30] observe in their evaluation of rough-set heuristics, "the sensitivity of final reducts to the stopping threshold necessitates extensive parameter tuning across diverse datasets, undermining reproducibility". Although mutual information-based methods[31,32] address the first limitation by leveraging information-theoretic quantities, they do not directly quantify rough-set classification capability and capture dependence without considering class-separability statistics—factors that are particularly important, where interactions between attribute values and class labels demand careful quantification.
Motivated by these limitations, we propose MFDRRS—a new method that integrates Fisher discriminant ratio (FDR) with rough sets to solve attribute reduction problems. Unlike heuristic reductions based on metrics such as frequency and entropy, we introduce the Fisher discriminant ratio to guide candidate attribute generation. The Fisher discriminant ratio provides a transparent and interpretable metric that quantifies how well an attribute separates classes relative to between-class and within-class variability. A higher FDR score indicates stronger class-separability, implying that the attribute is more representative than others. To our knowledge, no prior work has systematically incorporated FDR—the cornerstone of Fisher’s linear discriminant analysis—into the rough-set heuristic framework. Although Zhou et al.[33] introduced hypothesis testing in rough sets for small-sample settings, and Chen et al.[34] integrated rough sets with SVM weighting, neither approach utilized FDR as a ranking criterion or pre-filtering mechanism. Our approach fills this gap by treating FDR as a statistical surrogate for rough-set discriminability. Furthermore, we replace the fixed minimum-gain parameter with a dynamic termination criterion based on the ranked FDR distribution, since candidate attributes in the lower tail yield diminishing returns. This data-driven criterion adjusts to the inherent difficulty of each dataset, thereby mitigating the reproducibility issues raised by Liu et al.[30]. Lastly, we reduce the computational cost of dependency-degree evaluations by pruning attributes with negligible between-class scatter relative to within-class scatter prior to greedy selection. This is particularly beneficial for datasets with many nominal attributes, where evaluating becomes computationally expensive.
Our main contributions are summarized as follows:
- We propose the first integration of the Fisher discriminant ratio into rough-set attribute reduction, providing a principled statistical criterion for attribute ranking that is absent in existing heuristic methods.
- We introduce an FDR-driven stopping rule that substitutes the ad-hoc minimum-gain threshold with a data-driven criterion based on FDR rankings, thereby improving robustness and cross-dataset generalizability.
- We further assessed the proposed algorithm through experiments to confirm its effectiveness and efficiency. Experimental results demonstrate that our algorithm achieves more compact reductions with fewer attributes, maintains competitive dependency degrees, and avoids the over-selection characteristic of classical backward elimination.
The remainder of this paper is organized as follows. Section 2 reviews related work on rough-set attribute reduction and statistical feature ranking. Section 3 introduces the necessary rough-set concepts and extends Fisher Discriminant Ratio to nominal attributes. Section 4 presents the MFDRRS algorithm and its complexity analysis. Section 5 details the experimental setup and reports comparative results on UCI and synthetic datasets. Finally, Section 6 draws conclusions and outlines future research directions.
2. Related Work
2.1. Rough Set-Based Attribute Reduction
Pawlak’s original formulation[24] defines attribute reduction via indiscernibility relations and positive region approximation. However, the strictness of the equivalence relation makes attribute reduction vulnerable to outliers and ill-suited for practical applications[4,5]. To address this, numerous studies have relaxed the equivalence constraints and redefined the reduct, including probabilistic rough set model[35,36], variable precision rough set model[37,38], covering rough set model[39,40], fuzzy rough set model[41,42,43], and algebra viewpoint attribute reduction[44], etc. As computing all reducts is NP-hard, reduction methods can be divided, from a computability viewpoint, into two categories: data structure optimization and heuristic search. The first category primarily employs data structure optimization in reduction computation, e.g., discernibility matrix[45,46,47], discernibility hash[48], binary tree structure[49], and tree structure[50,51] etc. The second approach relies on heuristic information from candidate attributes to expedite the reduct process. For example, Qian et al. [52] proposed the Pangu Sorting algorithm, a greedy forward selection strategy that iteratively adds the attribute with the maximum marginal gain in dependency degree. Li et al.[53] introduced entropy-based measures for continuous attributes. Similar approaches include information-theoretic methods that employ heuristic measures [54,55,56,57]. Recent efforts have attempted to enhance heuristic stability. Zhao et al. [26] studied optimal reduct approximations via granular computing and showed that basic greedy heuristics yield unstable selections under minor data perturbations. Yao and Wang [27] surveyed uncertainty measures for rough-set feature selection, observing that dependency degree is the most commonly used criterion, despite its poor sensitivity to attribute interactions. Feng et al. [58] recently explored ensemble methods for reduct aggregation to address instability, albeit with higher computational overhead. For large-scale data, Wang et al.[59] proposed approximate rough-set heuristics with sampling to reduce dependency-degree cost, sacrificing accuracy for scalability.
2.2. Statistical Feature Selection
Beyond rough-set frameworks, statistical methods have long informed feature selection. Benjamini and Hochberg [60] introduced the FDR-controlling step-up procedure, which has become a standard tool in genomics and high-dimensional data analysis. Subsequent work extended FDR control to adaptive procedures [61] and independent filtering [62]. Zhu et al. [31] performed a systematic comparison of mutual information estimators for categorical features in the machine learning community, revealing that MI-based rankings are sensitive to discretization choices and sample size. They advised using statistical independence tests alongside MI to prevent false selections. Wang et al. [32] similarly extended joint mutual information approaches to better address redundancy, yet their method still suffers from the limitation that MI scores are not comparable across datasets with differing numbers of variables. Chen and Yu [63] introduced a statistical testing framework for nominal feature selection that uses chi-square tests for attribute screening prior to downstream modeling. Their approach aligns with our objective of integrating statistical significance into feature selection; however, it relies solely on p-values rather than effect-size measures such as FDR, which offer finer differentiation between moderately and strongly informative attributes.
Several studies have combined rough sets with statistical tests. Zhou et al.[33] used Fisher’s exact test within a rough-set framework for small-sample genomic data, applying multiple-testing correction to filter candidate attributes before greedy selection. While this represents an important step toward statistical rigor, their focus on exact tests for binary outcomes limits applicability to multi-class nominal settings. Moreover, they did not integrate FDR-like scatter metrics that capture both between-class separation and within-class compactness simultaneously. Chen et al.[34] integrated support vector machines with rough sets for feature weighting, assigning weights based on SVM coefficients followed by rough-set reduct extraction. However, this wrapper-style approach incurs additional computational overhead and does not provide a statistical guarantee on individual attribute importance. Furthermore, it assumes numeric encoding even for nominal inputs, potentially introducing artificial ordinal relationships.
3. Preliminaries
This section lays the groundwork by introducing the essential concepts of rough sets, which will inform the research problem formulation and discussion in the following sections.
Definition 1
(Decision Information System [24]). A decision information system is defined as a quadruple , where:
- (1).
- is a non-empty finite set which represents the object domain, where each is an object and n is the total number of objects.
- (2).
- is a non-empty finite set which represents the attribute domain, where each is an attribute and m is the total number of attributes. The set A is devided into conditonal attributes C and decision attributes D, where and .
- (3).
- is the domain of attribute values, where is the set of values for attribute a.
- (4).
- is an information function which maps each object-attribute pair to a value in the value domain V.
Table 1 presents a decision information system for influenza diagnosis, where is the object domain, is the conditional attribute domain, is the decision attribute domain. According to this diagnosis system, we can obtain some calculation results, such as: , , , , and etc.
Definition 2
(Indiscernibility Relation [24]). Given a subset of attributes , the indiscernibility relation induced by B is defined as:
The relation is an equivalence relation that partitions the univers U into equivalence classes,
where
Consider the subset of attributes . Then, for example, the object pair satisfies the indiscernibility relation , and .
Definition 3
(Positive Region [24]). Given a subset of attributes and a decision attribute , the positive region is defined as:
Definition 4
(Dependency Degree [24]). Given a subset of attributes and a decision attribute , the dependency degree is defined as:
Definition 5
(Reduct [24]). Given a subset of attributes and a decision attribute , the subset B is called a reduct of A with respect to d if it satisfies the following two conditions:
- (1).
- , where is the dependency degree of B with respect to d.
- (2).
- , .
4. Proposed Method
The proposed FDRRS method integrates FDR into rough-set-based attribute reduction through a three-stage process: statistical discriminability assessment of each condition attribute, significance-based pre-screening of the attribute space, and dependency-driven greedy forward selection with a data-adaptive stopping criterion. We detail each stage below.
4.1. Modified Fisher Discriminant Ratio
To quantify the discriminative power of each condition attribute, we adapt the Fisher discriminant ratio—a classical linear discriminant analysis metric—to nominal data. The traditional Fisher discriminant ratio [7,64] is designed for numerical features, as its computation relies on class means and variances. Since nominal attributes have neither numerical ordering nor meaningful statistical moments such as mean and variance, the conventional FDR cannot be directly applied to nominal data.
Definition 6
(Indicator function). Given a nominal attribute with value domain of cardinality , the indicator function for each value is defined as:
Definition 7
(Conditional frequencies and marginal frequencies). Given a nomial attriubte with value domain of cardinality , the conditional frequencies and marginal frequencies of attribute value are defined as:
where denote the set of objects belonging to decision class c with objects.
Taking the decision information system in Table 1, for example, we can calculate the conditional frequencies and marginal frequencies of the attribute as follows. First, we divide the object domain U into two decision classes according to the decision attribute , i.e., and and get the value set of attribute : . Then, the conditional frequencies of the attribute value are calculated as
And, the marginal frequency of the attribute value is calculated as
Analogous to the FDR for continuous data, we define the intra-class and inter-class divergences for nominal attributes as follows.
Definition 8
(Intra-class and Inter-class Divergences). Given a nominal attribute with value domain of cardinality , the intra-class divergence and inter-class divergence are defined as:
Definition 9
(Modified Fisher Discriminant Ratio, MFDR). Given a decision information system and a nominal attribute with value domain of cardinality , the modified Fisher discriminant ratio (MFDR) for nominal attributes is defined as:
where ϵ is a small positive constant to avoid division by zero.
quantifies how well separates decision classes relative to its inter- and intra-class variance; larger values indicate stronger discriminative power. This formulation naturally generalizes the continuous FDR to nominal data via indicator-variable encoding and remains computable for any attribute of finite cardinality.
4.2. MFDR Filtering
Given the MFDR scores for all conditional attributes C with non-increasing order, i.e.,
we can obtain a candidate attribute set and iteratively select the highest-ranked attribute. According to this filtering rule, we can also remove attributes with low MFDR scores that are unlikely to contribute meaningfully to class separability. Specifically, We define a filtered candidate set by applying a threshold :
The parameter controls the aggressiveness of pre-screening: retains all attributes, while eliminates attributes whose discriminative power falls below the threshold. This step reduces the search space for subsequent greedy selection, thereby improving computational efficiency, particularly when the number of candidate attributes m is large and many exhibit negligible MFDR values.
4.3. MFDR-Guided Rough-Set Attribute Reduction Algorithm
Roughly speaking, the problem of attribute reduction solving by Pawlak’s rough set theory can be formulated as follows:
where is a reduct of the conditional attribute set C. By introducing the MFDR filtering step, we can solve the problem iteratively by selecting the attribute with the highest MFDR score that maximally increases the dependency degree at each iteration. The process continues until no further attributes can be added without violating the condition or until the MFDR score of remaining candidates falls below a dynamic stopping threshold derived from the distribution of MFDR scores. In other words, we can obtain the solution of the original problem in Equation(17) by iteratively solving the following sub-problems:
Specifically, starting from the empty set , we iteratively select the attribute from that maximizes the marginal gain in rough-set dependency degree. To avoid exhaustive search while maintaining reduction compactness, we employ a dual stopping criterion. Let Q3 denote the third quartile of the MFDR values . At each iteration, the selection process terminates if either of the following conditions holds:
- MFDR-based termination: , where is a decay factor (default ). This criterion leverages the statistical ranking to detect when the best remaining candidate no longer provides discriminative power substantially above the dataset median.
- Minimum-gain termination: , where is a small tolerance (default ). This ensures that attributes contributing negligible improvement to the dependency degree are not included.
If either condition is satisfied, the algorithm terminates and outputs the current reduct R. Otherwise, it adds to R and proceeds.The complete algorithm is summarized in Algorithm 1.
| Algorithm 1 MFDRRS: MFDR-Guided Rough-Set Attribute Reduction |
|
4.4. Complexity Analysis
Let , , , and be the number of decision classes. Computing MFDR for each attribute requires operations per attribute using class-wise aggregation, yielding a total of . For binary attributes (), this simplifies to . Sorting MFDR values requires time. Computing the third quartile is using a linear-time selection algorithm. Each iteration computes for all remaining candidates, which involves building equivalence classes in time. In the worst case with iterations, the total is . However, the MFDR-based stopping criterion typically limits the number of iterations well below , and pre-screening often makes . The overall time complexity is . In practice, the combination of statistical pre-screening and adaptive early stopping substantially reduces the number of dependency-degree evaluations compared to conventional greedy forward selection, which requires at most operations.
5. Experiments
This section presents a comprehensive evaluation of the proposed MFDRRS algorithm to validate its effectiveness and efficiency. We first describe the experimental setup, including datasets, baseline methods, and evaluation metrics. Then, we report the results and analyze the performance of MFDRRS in comparison to existing rough-set-based attribute reduction methods and statistical based methods.
5.1. Experimental Setup
The experiments were conducted on a workstation with an AMD 3700X CPU, 32GB RAM, running Windows 10, ensuring a consistent experimental environment. To test the performance of the proposed MFDRRS algorithm, we selected six standard UCI datasets that vary in size, number of attributes, and class distributions, together with seven synthetic datasets designed to probe specific algorithmic properties under controlled conditions. Table 2 presents the characteristics of these datasets, including the number of objects, number of attributes, and the number of decision classes. The datasets were preprocessed to handle missing values and categorical attributes were encoded appropriately for the MFDR computation.
Besides, we construct seven controlled synthetic datasets to probe specific algorithmic properties of MFDRRS, each combining attributes of varying discriminative strength: deterministic attributes (perfectly determining the decision, ), partial attributes (correlating with accuracy , yielding moderate MFDR), and noise attributes (uniform random, negligible MFDR ). We also include structural variants such as redundant pairs and high-cardinality noise. Table 3 summarizes the composition of each dataset.
5.2. Compared Methods and Evaluation Metrics
The following methods were selected as baselines for comparison with the proposed MFDRRS method:
- Greedy Forward: Standard forward selection maximizing -gain[24].
- Greedy Backward: Standard backward elimination minimizing -loss[24].
- MI Ranking: Selects top-k attributes by covering 90% of total mutual information[65].
- MaxRelevance: Greedily selects all attributes with non-negligible individual MI[66].
- JMI (Joint Mutual Information): Selects attributes maximizing while penalizing redundancy with already-selected attributes[67].
- Exhaustive Search: Brute-force enumeration [24] of all subsets up to size n (feasible only for ).
To systematically assess the performance of MFDRRS and baseline methods, we employed three primary evaluation metrics:
- Number of selected attributes (s): Smaller is better for compactness.
- Dependency degree (): Higher is better for classification power.
- Runtime (t): Wall-clock time in seconds.
5.3. Results and Discussions
5.3.1. UCI Dataset Results
Table 4 presents the results of the MFDRRS method compared to baseline methods across the six UCI datasets. For each dataset, we report the number of selected attributes (s), the achieved dependency degree (), and the runtime in seconds (t). The best values among methods with are highlighted in bold.
As can be seen from Table 4, the proposed MFDRRS method achieves compact reductions on most UCI datasets. On breast-cancer, glass, wdbc, and wine, MFDRRS selects only 2 attributes achieving , matching Greedy Forward but avoiding Greedy Backward’s over-selection (30, 9, 30, and 13 attributes respectively). On ionosphere, MFDR selects 3 attributes with , identical to Greedy Forward but far more compact than Greedy Backward’s 34 attributes. And MFDRRS is conservative on large high-cardinality datasets. On the letter dataset (20,000 objects, 26 decision classes), MFDR selects only 6 attributes with versus 16 for Greedy Forward (). This reflects the inherent difficulty of nominal attribute reduction when each attribute has many distinct values (letter recognition uses alphabetic characters as attribute values). The MFDR metric penalizes attributes whose class-wise scatter is low relative to within-class scatter, which is common when each value appears in multiple decision classes. Greedy Backward retains all attributes on every UCI dataset, failing to eliminate even clearly redundant ones. This confirms the well-known limitation that backward elimination is overly conservative in its removal criterion[53]. By design, MI-based methods do not output a dependency degree, so is reported as 0.0 in the table. While MI provides an alternative measure of attribute importance, it does not directly quantify the rough-set classification power. MaxRelevance and MI Ranking tend to select many attributes (7–33 on UCI datasets), far exceeding MFDRRS’s selections. JMI selects 1–7 attributes—still more than MFDRRS on most datasets where FDR achieves full dependency. Furthermore, MFDRRS consistently outperforms Greedy Forward and Greedy Backward in runtime. For example, on the letter dataset, MFDRRS completes in 104.32 seconds versus 271.68 seconds for Greedy Forward and 53.77 seconds for Greedy Backward. The pre-screening step reduces the number of candidate attributes considered at each iteration, while the adaptive stopping criterion prevents unnecessary evaluations once the marginal gain falls below the threshold. This efficiency gain is particularly pronounced on larger datasets with many attributes. Besides, MFDR pre-filtering improves efficiency. On small-to-medium datasets (), MFDRRS matches Greedy Forward in both selected attribute count and runtime. The MFDR-aware stopping rule rarely triggers on these datasets because the best candidate typically has sufficiently high MFDR. On the letter dataset, however, the MFDR filter reduces 16 candidates to just 2 significant attributes, enabling rapid termination. Finally, for small UCI datasets (), all methods run in milliseconds to under one second. For larger datasets (letter, ), computation becomes expensive, with MFDRRS requiring about 104 seconds and Greedy Forward about 272 seconds due to the high cost of dependency degree computation on large equivalence classes.
Overall, MFDRRS yields the most compact reducts among all tested methods on every UCI dataset, while achieving full dependency () on all small- to medium-sized datasets. On the letter dataset, MFDRRS achieves a notably more parsimonious reduction (6 attributes) with a slight trade-off in dependency—a compromise attributable to the extreme cardinality of its nominal attributes, which poses challenges for scatter-based discriminative measures. Across all benchmarks, MFDRRS outperforms Greedy Backward, MI Ranking, MaxRelevance, and JMI in compactness, and matches or exceeds Greedy Forward in quality, confirming that the Fisher discriminant ratio provides a reliable statistical ranking criterion that complements the rough-set framework.
5.4. Artificial Dataset Results
To further validate MFDRRS under controlled conditions, we evaluate on seven synthetic datasets (Table 3) designed to probe specific algorithmic properties: noise rejection, redundancy handling, high cardinality, and multi-class separability.
Table 5.
Comparison of attribute reduction methods on UCI datasets.
| Dataset | m | n | MFDRRS | Greedy Forward | Greedy Backward | MI Ranking | MaxRelevance | JMI | Exhaustive | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| s | t(s) | s | t(s) | s | t(s) | s | t(s) | s | t(s) | s | t(s) | s | t(s) | ||||||||||
| simple_separable | 100 | 5 | 1 | 1.0000 | 0.001 | 1 | 1.0000 | 0.001 | 5 | 1.0000 | 0.001 | 2 | 0.0000 | 0.000 | 5 | 0.0000 | 0.000 | 1 | 0.0000 | 0.001 | 1 | 1.0000 | 0.000 |
| with_noise | 150 | 10 | 1 | 1.0000 | 0.003 | 1 | 1.0000 | 0.003 | 10 | 1.0000 | 0.003 | 5 | 0.0000 | 0.001 | 10 | 0.0000 | 0.000 | 1 | 0.0000 | 0.005 | 1 | 1.0000 | 0.001 |
| redundant | 120 | 8 | 1 | 1.0000 | 0.002 | 1 | 1.0000 | 0.002 | 8 | 1.0000 | 0.001 | 8 | 0.0000 | 0.001 | 8 | 0.0000 | 0.000 | 2 | 0.0000 | 0.004 | 1 | 1.0000 | 0.000 |
| correlated | 150 | 8 | 1 | 1.0000 | 0.003 | 1 | 1.0000 | 0.003 | 8 | 1.0000 | 0.002 | 4 | 0.0000 | 0.001 | 8 | 0.0000 | 0.000 | 1 | 0.0000 | 0.003 | 1 | 1.0000 | 0.000 |
| high_cardinality | 200 | 6 | 1 | 1.0000 | 0.004 | 1 | 1.0000 | 0.003 | 6 | 1.0000 | 0.003 | 5 | 0.0000 | 0.003 | 6 | 0.0000 | 0.003 | 3 | 0.0000 | 0.013 | 1 | 1.0000 | 0.001 |
| multi_class | 180 | 7 | 1 | 1.0000 | 0.004 | 1 | 1.0000 | 0.002 | 7 | 1.0000 | 0.003 | 5 | 0.0000 | 0.001 | 7 | 0.0000 | 0.001 | 2 | 0.0000 | 0.004 | 1 | 1.0000 | 0.001 |
| imbalanced | 150 | 6 | 1 | 1.0000 | 0.002 | 1 | 1.0000 | 0.001 | 6 | 1.0000 | 0.001 | 3 | 0.0000 | 0.000 | 6 | 0.0000 | 0.000 | 2 | 0.0000 | 0.002 | 1 | 1.0000 | 0.000 |
On every synthetic dataset, MFDRRS selects exactly 1 attribute with , matching Greedy Forward and Exhaustive Search while far outperforming Greedy Backward (5–10 attributes), MI Ranking (2–8), and MaxRelevance (5–10). MFDR assigns scores of to deterministic attributes—attribute 0 (a bijective permutation of the class) and their exact or near-exact copies—while noise attributes receive near-zero MFDR (), and partially deterministic attributes (80% accuracy) fall in the intermediate range (0.49–0.83), cleanly separating informative from uninformative features. On the redundant dataset (8 attributes forming 4 identical pairs), all attributes individually achieve , yet MFDR retains only 1 of 8, whereas Greedy Backward keeps all 8 and MI-based methods select 2–8, confirming that MFDR avoids redundancy accumulation. On high-cardinality data (up to 50 distinct values per attribute), MFDR selects 1 attribute with , while MI Ranking and MaxRelevance select 5 and 6 respectively, demonstrating that the between-class/within-class scatter formulation naturally accounts for cardinality without explicit normalization. On the imbalanced (80/80/40 split, 3 classes) and multi_class (5 classes) datasets, MFDRRS again selects 1 attribute with full dependency, while JMI achieves 1–2 attributes but with , indicating an invalid rough-set reduct. Computationally, all methods run in milliseconds (); MFDRRS’s runtime (0.002–0.005 s) is comparable to Greedy Forward and within an order of magnitude of the fastest baselines, with the MFDR computation adding negligible overhead.
Collectively, the synthetic results demonstrate that MFDRRS reliably achieves the optimal reduct size () across a range of data complexities—noise, redundancy, high cardinality, imbalance, and multi-class separability—without sacrificing dependency (). The distinct MFDR score gap between relevant and irrelevant attributes substantiates the pre-filtering mechanism, while the low computational cost ensures practical scalability.
6. Conclusions
MFDRRS provides a principled statistical criterion for attribute ranking: the Fisher discriminant ratio offers a transparent and interpretable measure of attribute discriminability that complements the rough-set dependency degree. Unlike heuristic methods that rely solely on marginal gains in , MFDRRS employs a well-established statistic that captures both between-class separation and within-class compactness. The MFDR-aware stopping rule eliminates manual threshold calibration, enhancing practical usability over competing approaches. The pre-filtering step further reduces the search space prior to greedy selection, thereby improving computational efficiency.
Future research will pursue these directions: (1) extending FDR to mixed-type (nominal and numeric) data; (2) developing multi-attribute FDR measures to capture joint attribute interactions; (3) scaling the framework to large-scale datasets via sparse approximation and distributed computation. Finally, we plan to validate the method in broader application domains, including text categorization and bioinformatics, beyond tabular classification.
Author Contributions
Conceptualization, Luo S.; methodology,Cao X., Luo S.; software,Cao X., Luo S.; validation, Shi L..; formal analysis, Cao X.,Shi L.; investigation, Luo S.; resources, Cao X.; data curation, Cao X.; writing—original draft preparation, Cao X., Luo S.; writing—review and editing, Shi L.; visualization, Cao X.; supervision, Luo S.; project administration, Shi L.; funding acquisition, Shi L. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by shanghai polytechnic university leap-up project, grant number A30YD260102-0501, and shanghai polytechnic university multi-granularity data analysis system, grant number C80JX250031.
Data Availability Statement
The original data presented in the study are openly available at https://archive-beta.ics.uci.edu/. Further inquiries can be directed to the corresponding author.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Cao, L.; Zhang, C.; Yang, Q.; Bell, D.; Vlachos, M.; Taneri, B.; Keogh, E.; Yu, P.S.; Zhong, N.; Ashrafi, M.Z.; et al. Domain-Driven, Actionable Knowledge Discovery. IEEE Intelligent Systems 2007, 22, 78–88, c3. [CrossRef]
- Kalanat, N. An overview of actionable knowledge discovery techniques. Journal of Intelligent Information Systems 2022, 58, 591–611. [CrossRef]
- Thangavel, K.; Pethalakshmi, A. Dimensionality reduction based on rough set theory: A review. Applied Soft Computing 2009, 9, 1–12. [CrossRef]
- Sheng, L.; Duoqian, M.; Zhifei, Z. A neighborhood rough set model with nominal metric embedding. Information Sciences 2020, 520, 373–388.
- Yuan, K.; Miao, D.; Zhang, H.; Pedrycz, W. An Efficient and Robust Feature Selection Approach Based on Zentropy Measure and Neighborhood-Aware Model. IEEE Transactions on Neural Networks and Learning Systems 2025, 36, 16351–16365. [CrossRef]
- Guyon, I.; Elisseeff, A. An introduction to variable and feature selection. J. Mach. Learn. Res. 2003, 3, 1157–1182.
- Fisher, R.A. The Use of Multiple Measurements in Taxonomic Problems. Annals of Eugenics 1936, 7, 179–188. [CrossRef]
- Kira, K.; Rendell, L.A. The Feature Selection Problem: Traditional Methods and a New Algorithm. In Proceedings of the Proceedings of the Tenth National Conference on Artificial Intelligence, 1992, pp. 129–134.
- Battiti, R. Using mutual information for selecting features in supervised neural net learning. IEEE Transactions on Neural Networks 1994, 5, 537–550. [CrossRef]
- Peng, H.; Long, F.; Ding, C. Feature Selection Based on Mutual Information: Criteria of Max-Dependency, Max-Relevance, and Min-Redundancy. IEEE Transactions on Pattern Analysis and Machine Intelligence 2005, 27, 1226–1238. [CrossRef]
- Freijeiro-Gonzalez, L.; Febrero-Bande, M.; Gonzalez-Manteiga, W. A Critical Review of LASSO and Its Derivatives for Variable Selection Under Dependence Among Covariates. International Statistical Review 2022, 90, 118–145. [CrossRef]
- Chamlal, H.; Benzmane, A.; Ouaderhman, T. Elastic Net-Based High Dimensional Data Selection for Regression. Expert Systems with Applications 2024, 244, 122958. [CrossRef]
- Villa-Blanco, C.; Bielza, C.; Larrañaga, P. Feature subset selection for data and feature streams: a review. Artif. Intell. Rev. 2023, 56, 1011–1062. [CrossRef]
- Li, X.; Wang, Y.; Ruiz, R. A Survey on Sparse Learning Models for Feature Selection. IEEE Transactions on Cybernetics 2022, 52, 1642–1660. [CrossRef]
- He, X.; Cai, D.; Niyogi, P. Laplacian Score for Feature Selection. In Proceedings of the Advances in Neural Information Processing Systems, 2005, Vol. 18, pp. 507–514.
- Zhao, Z.; Liu, H. Spectral Feature Selection for Supervised and Unsupervised Learning. In Proceedings of the Proceedings of the 24th International Conference on Machine Learning, 2007, pp. 1151–1157.
- Cai, D.; Zhang, C.; He, X. Unsupervised Feature Selection for Multi-Cluster Data. In Proceedings of the Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2010, pp. 333–342.
- Yang, Y.; Shen, H.T.; Ma, Z.; Huang, Z.; Zhou, X. l2,1-Norm Regularized Discriminative Feature Selection for Unsupervised Learning. In Proceedings of the Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence, 2011, pp. 1589–1594.
- Li, Z.; Yang, Y.; Liu, J.; Zhou, X.; Lu, H. Unsupervised Feature Selection Using Nonnegative Spectral Analysis. In Proceedings of the Proceedings of the Twenty-Sixth AAAI Conference on Artificial Intelligence, 2012, pp. 1026–1032.
- Borah, K.; Das, H.S.; Seth, S.; Mallick, K.; Rahaman, Z.; Mallik, S. A Review on Advancements in Feature Selection and Feature Extraction for High-Dimensional NGS Data Analysis. Functional & Integrative Genomics 2024, 24, 139. [CrossRef]
- Villa-Blanco, C.; Bielza, C.; Larrañaga, P. Feature Subset Selection for Data and Feature Streams: A Review. Artificial Intelligence Review 2023, 56, 1011–1062. [CrossRef]
- Yang, K.; Liu, L.; Wen, Y. The Impact of Bayesian Optimization on Feature Selection. Scientific Reports 2024, 14, 3948. [CrossRef]
- Xu, Z.; Yang, F.; Tang, C.; Wang, H.; Wang, S.; Sun, J.; Zhang, Y. FG-HFS: A Feature Filter and Group Evolution Hybrid Feature Selection Algorithm for High-Dimensional Gene Expression Data. Expert Systems with Applications 2024, 245, 123069. [CrossRef]
- Pawlak, Z. Rough Sets. International Journal of Information and Computer Sciences 1982, 11, 341–356. [CrossRef]
- Wong, S.K.M.; Ziarko, W. On optimal decision rules in decision tables. Bulletin of the Polish Academy of Sciences Mathematics 1985, 33, 693–696.
- Zhao, S.; Hu, Q.; Hu, X. Characterizing Optimal Reduct Approximations in Rough Set Theory. International Journal of Approximate Reasoning 2020, 124, 1–20. [CrossRef]
- Yao, Y.; Wang, W. A Review of Uncertainty Measures in Rough Set Theory. International Journal of Approximate Reasoning 2020, 118, 1–18. [CrossRef]
- Wang, G.; Zhang, Q. On Stability of Greedy Backward Elimination in Rough Set Feature Selection. International Journal of Intelligent Systems 2020, 35, 1218–1236. [CrossRef]
- Zhao, F.; Hu, Q.; Pedrycz, W.; Hu, X. A Review of Rough Set Methods for Feature Selection. Neurocomputing 2020, 401, 1–17. [CrossRef]
- Liu, G.; Lin, T.; Xu, Y. On Robustness of Rough Set Heuristic Algorithms for Feature Selection. Information Sciences 2021, 548, 115–133. [CrossRef]
- Zhu, P.; Hu, Q.; Wang, P. Systematic Comparison of Mutual Information Estimators for Categorical Feature Selection. Expert Systems with Applications 2021, 178, 114987. [CrossRef]
- Wang, F.; Li, J.; Zhou, T.; Wang, J. Joint Mutual Information for Categorical Feature Selection with Improved Redundancy Modeling. Neural Computation 2023, 35, 557–585. [CrossRef]
- Zhou, S.; Chen, Q.; Miao, Y.; Qian, Y. Rough Set-Based Feature Selection for Small Sample Data. Knowledge-Based Systems 2018, 147, 1–12. [CrossRef]
- Chen, B.; Wang, L.; Liu, J. A Hybrid Feature Selection Method Combining Rough Set and Support Vector Machine. Applied Soft Computing 2020, 88, 106055. [CrossRef]
- Yao, Y.; Zhao, Y.; Wang, J., On Reduct Construction Algorithms. In Transactions on Computational Science II; Gavrilova, M.L.; Tan, C.J.K.; Wang, Y.; Yao, Y.; Wang, G., Eds.; Springer Berlin Heidelberg: Berlin, Heidelberg, 2008; pp. 100–117. [CrossRef]
- Li, W.; Zhan, T. Multi-Granularity Probabilistic Rough Fuzzy Sets for Interval-Valued Fuzzy Decision Systems. INTERNATIONAL JOURNAL OF FUZZY SYSTEMS 2023, 25, 3061–3073. [CrossRef]
- Mi, J.; Leung, Y.; Wu, W. An uncertainty measure in partition-based fuzzy rough sets. International Journal of General Systems 2005, 34, 77–90. [CrossRef]
- Syau, Y.R.; Liau, C.J.; Lin, E.B. On Variable Precision Generalized Rough Sets and Incomplete Decision Tables. FUNDAMENTA INFORMATICAE 2021, 179, 75–92. [CrossRef]
- Zhu, W.; Wang, F.Y. On Three Types of Covering-Based Rough Sets. IEEE Transactions on Knowledge and Data Engineering 2007, 19, 1131–1144. [CrossRef]
- Ma, L.; Li, M. Covering rough set models, fuzzy rough set models and soft rough set models induced by covering similarity. INFORMATION SCIENCES 2025, 689. [CrossRef]
- Hu, Q.; Yu, D.; Liu, J.; Wu, C. Neighborhood Rough Set Based Heterogeneous Feature Subset Selection. Inf. Sci. 2008, 178, 3577–3594. [CrossRef]
- Wu, W.; Zhang, W. Constructive and axiomatic approaches of fuzzy approximation operators. Information Sciences 2004, 159, 233 – 254. [CrossRef]
- Hu, Q.; Yu, D.; Pedrycz, W.; Chen, D. Kernelized Fuzzy Rough Sets and Their Applications. IEEE Transactions on Knowledge and Data Engineering 2011, 23, 1649–1667. [CrossRef]
- Wang, G.Y.; Zhao, J.; An, J.; Wu, Y. A Comparative Study of Algebra Viewpoint and Information Viewpoint in Attribute Reduction. Fundamenta Informaticae 2005, 68, 289–301.
- Yao, Y.; Zhao, Y. Discernibility matrix simplification for constructing attribute reducts. Information Sciences 2009, 179, 867–882.
- Lang, G.; Li, Q.; Guo, L. Discernibility matrix simplification with new attribute dependency functions for incomplete information systems. Knowledge and Information Systems 2013, 37, 611–638.
- Liu, Y.; Zheng, L.; Xiu, Y.; Yin, H.; Zhao, S.; Wang, X.; Chen, H.; Li, C. Discernibility matrix based incremental feature selection on fused decision tables. International Journal of Approximate Reasoning 2020, 118, 1–26. [CrossRef]
- Luo, S.; Shi, L.; Chen, L.; Cao, X. Accelerated Feature Selection via Discernibility Hashing: A Rough Set Approach. Entropy 2025, 27. [CrossRef]
- Lu, Z.; Qin, Z.; Jin, Q.; Li, S. Constructing Rough Set Based Unbalanced Binary Tree for Feature Selection. Chinese Journal of Electronics 2014, 23, 474–479. [CrossRef]
- Yang, M.; Yang, P. A novel condensing tree structure for rough set feature selection. Neurocomputing 2008, 71, 1092–1100.
- Jiang, Y.; Yu, Y. Minimal attribute reduction with rough set based on compactness discernibility information tree. Soft Computing 2016, 20, 2233–2243.
- Qian, Y.; Liang, J.; Yao, Y.; Dang, C. PANGR: A Parallel Rough Set Approach for Simultaneous Attribute Reduction and Rule Acquisition. In Proceedings of the IEEE Transactions on Knowledge and Data Engineering, 2010, Vol. 23, pp. 866–881. [CrossRef]
- Li, J.; Yang, T.; Chu, S.; Wu, P.; Wang, C. A Rough Set Approach Using Feature-Weighting for Accuracy-Error Tradeoff in Feature Selection. Information Sciences 2012, 179, 2317–2329. [CrossRef]
- Wang, C.; Huang, Y.; Ding, W.; Cao, Z. Attribute reduction with fuzzy rough self-information measures. Information Sciences 2021, 549, 68–86. [CrossRef]
- Ji, X.; Li, J.; Yao, S.; Zhao, P. Attribute reduction based on fusion information entropy. International Journal of Approximate Reasoning 2023, 160, 108949. [CrossRef]
- Gao, C.; Zhou, J.; Miao, D.; Yue, X.; Wan, J. Granular-conditional-entropy-based attribute reduction for partially labeled data with proxy labels. Information Sciences 2021, 580, 111–128. [CrossRef]
- Wang, P.; Qu, L.; Zhang, Q. Information entropy based attribute reduction for incomplete heterogeneous data. J. Intell. Fuzzy Syst. 2022, 43, 219–236. [CrossRef]
- Feng, Q.; Sun, B.; Wang, X. Ensemble-Based Rough Set Feature Selection via Multiple Reduct Aggregation. Pattern Recognition Letters 2021, 143, 35–42. [CrossRef]
- Wang, P.; Wang, S.; Ding, W. Approximate Rough Set Heuristics for Large-Scale Feature Selection. IEEE Transactions on Neural Networks and Learning Systems 2023. [CrossRef]
- Benjamini, Y.; Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 1995, 57, 289–300. [CrossRef]
- Benjamini, Y.; Yekutieli, D. The Control of the False Discovery Rate Under Dependent Testing. Annals of Statistics 2001, 29, 1165–1188. [CrossRef]
- Robinson, M.D.; Oshlack, A. A Scaling Normalization Method for Differential Expression Analysis of RNA-Seq Data. Genome Biology 2010, 11, R25. [CrossRef]
- Chen, X.; Yu, G. A Statistical Testing Framework for Nominal Feature Selection. Knowledge and Information Systems 2020, 62, 1785–1810. [CrossRef]
- Gu, Q.; Li, Z.; Han, J. Generalized Fisher Score for Feature Selection. In Proceedings of the Proceedings of the Twenty-Seventh Conference on Uncertainty in Artificial Intelligence, 2012, pp. 266–273.
- Peng, H.; Long, F.; Ding, C. Feature Selection Based on Relevant and Redundant Features. In Proceedings of the Proceedings of the 15th International Conference on Machine Learning, 2005, pp. 554–561.
- Chen, S.; Yang, J.; Yu, Z.; Liu, J.; Wang, X. An Efficient Method for Feature Selection Based on Maximum Relevance Criterion. In Proceedings of the Proceedings of the 2010 IEEE International Conference on Computational Intelligence and Security, 2010, pp. 181–185. [CrossRef]
- Yang, J.; Liu, J.; Cai, Z.; Zhang, H.; Wu, F. Feature Selection Based on Joint Mutual Information. In Proceedings of the Proceedings of the 2014 IEEE International Conference on Data Mining Workshop, 2014, pp. 475–484. [CrossRef]
Table 1.
An example of a decision information system.
| Object | Fever | Cough | Headache | Diagnosis |
|---|---|---|---|---|
| High | Yes | Yes | Flu | |
| High | Yes | No | Flu | |
| Normal | Yes | Yes | Flu | |
| Normal | No | No | Healthy | |
| High | No | No | Healthy | |
| Normal | No | Yes | Healthy |
Table 2.
UCI dataset characteristics
| Dataset | Objects (m) | Attributes (n) | Classes |
|---|---|---|---|
| breast-cancer | 569 | 30 | 2 |
| glass | 213 | 9 | 6 |
| ionosphere | 350 | 34 | 2 |
| letter | 20000 | 16 | 26 |
| wdbc | 569 | 30 | 2 |
| wine | 178 | 13 | 3 |
Table 3.
Artificial dataset characteristics.
| Dataset | Objects (m) | Attributes (n) | Classes | Design |
|---|---|---|---|---|
| simple_separable | 200 | 5 | 2 | 2 det + 3 noise |
| with_noise | 200 | 10 | 2 | 5 det + 5 noise |
| redundant | 150 | 8 | 3 | 4 redundant pairs |
| correlated | 200 | 8 | 2 | 4 det + 4 partial |
| high_cardinality | 200 | 6 | 4 | 2 det + 3 partial + 1 noise |
| multi_class | 200 | 7 | 5 | 3 det + 2 partial + 2 noise |
| imbalanced | 200 | 6 | 3 | 3 det + 3 noise |
Table 4.
Comparison of attribute reduction methods on UCI datasets.
| Dataset | m | n | MFDRRS | Greedy Forward | Greedy Backward | MI Ranking | MaxRelevance | JMI | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| s | t(s) | s | t(s) | s | t(s) | s | t(s) | s | t(s) | s | t(s) | |||||||||
| breast-cancer | 569 | 30 | 2 | 1.0000 | 0.34 | 2 | 1.0000 | 0.32 | 30 | 1.0000 | 0.14 | 27 | 0.0000 | 0.58 | 30 | 0.0000 | 0.57 | 1 | 0.0000 | 17.00 |
| glass | 213 | 9 | 2 | 1.0000 | 0.03 | 2 | 1.0000 | 0.03 | 9 | 1.0000 | 0.01 | 7 | 0.0000 | 0.03 | 9 | 0.0000 | 0.03 | 2 | 0.0000 | 0.36 |
| ionosphere | 350 | 34 | 3 | 1.0000 | 0.23 | 3 | 1.0000 | 0.22 | 34 | 1.0000 | 0.08 | 29 | 0.0000 | 0.29 | 33 | 0.0000 | 0.29 | 2 | 0.0000 | 9.14 |
| letter | 20000 | 16 | 6 | 0.8217 | 104.32 | 16 | 0.9909 | 271.68 | 16 | 0.9909 | 53.77 | 13 | 0.0000 | 0.04 | 16 | 0.0000 | 0.04 | 2 | 0.0000 | 0.68 |
| wdbc | 569 | 30 | 2 | 1.0000 | 10.65 | 2 | 1.0000 | 10.29 | 30 | 1.0000 | 3.80 | 27 | 0.0000 | 0.54 | 30 | 0.0000 | 0.54 | 7 | 0.0000 | 15.81 |
| wine | 178 | 13 | 2 | 1.0000 | 0.33 | 2 | 1.0000 | 0.30 | 13 | 1.0000 | 0.13 | 11 | 0.0000 | 0.04 | 13 | 0.0000 | 0.04 | 5 | 0.0000 | 0.51 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.