Preprint
Article

This version is not peer-reviewed.

Comparative Concordance and Divergence of Volume-Based Parameter-Integrated Sharma-Mittal Entropy Indices with Conventional Feature Selection Methods

Submitted:

20 July 2026

Posted:

22 July 2026

You are already at the latest version

Abstract
Feature selection is one of the fundamental steps in producing simpler, more interpretable, and computationally efficient models while maintaining predictive power. In this study, rather than fixing the uncertainty relationship between the target variable and each candidate feature to a single entropy parameter, three volume-based indices are proposed that integrate the two-parameter structure of the Sharma-Mittal entropy over a defined parameter region. The indices, named Parameter-Integrated Conditional Sharma-Mittal Entropy, Parameter-Integrated Sharma-Mittal Entropy Information Gain, and Normalized Parameter-Integrated Sharma-Mittal Entropy Information Gain, represent, respectively, the conditional entropy volume, the information gain volume, and the form of this volume normalized relative to the total entropy volume of the target, respectively. Densities were estimated using Gaussian kernel density estimation; the alpha and beta parameters were numerically integrated over the range [0.05; 0.95] × [0.05; 0.95]. The methods were compared across six different regression datasets-Airfoil Self-Noise, AirQualityUCI, BodyFat, meteorology-based reference evapotranspiration, Concrete, and WineQualityWhite-using absolute Pearson correlation, absolute Spearman rank correlation, Shannon information gain, mutual information, and random forest variable importance. The comparison was conducted using raw importance scores, derived feature rankings, Spearman’s rho and Kendall’s tau rank correlations, bivariate rank scatter plots, and Mann–Whitney U test results. The results show that the proposed indices produced high or very high rank agreement with correlation- and information-based reference methods in the Airfoil Self-Noise, AirQualityUCI, BodyFat, and meteorological datasets. Specifically, the absolute Pearson correlation, absolute Spearman rank correlation, Shannon information gain, and mutual information, along with the average Spearman rho values across all datasets, were obtained as 0.884, 0.874, 0.883, and 0.903, respectively; the average agreement with random forest importance scores remained at 0.538. The findings reveal that the parameter-integrated Sharma–Mittal framework offers a density-based and discretization-independent filtering perspective for continuous data; however, they also highlight the need for additional sensitivity and out-of-sample prediction performance evaluations regarding absolute score scales, negative information-gain values, and the rank equivalence of the three proposed indices.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Feature selection is one of the strategic components of the modeling process, particularly in regression problems where a large number of predictors explain the same target variable with varying strengths and different dependency structures. An effective filtering mechanism not only eliminates unnecessary variables but also reduces the computational cost of the model, minimizes the risk of overfitting, and enhances interpretability. Therefore, feature selection serves the combined objectives of improving predictive performance, computational efficiency, and a better understanding of the data-generation mechanism [1,2,3,4].
Classical filtering approaches often rely on Pearson correlation, which measures linear relationships; Spearman’s rank correlation, which summarizes monotonic dependence; or information-theoretic criteria based on the reduction of uncertainty. While these methods are widely accepted, they may be sensitive to a single form of dependence, a single scale, or a specific discretization design. In particular, the fact that entropy and mutual information calculations for continuous variables are influenced by details such as probability density estimation, discretization rules, and numerical integration necessitates careful method design to ensure the robustness of information-based rankings [5,6,7,8,9,10,11,12,13,14,15].
The Sharma-Mittal entropy provides a broader framework of generalized entropy related to the Rényi and Tsallis families, thanks to its two-parameter structure. However, fixing the alpha and beta parameters at a single point in applied parameter selection studies may lead to the metric becoming dependent on parameter selection. The central idea of this study is to evaluate the conditional entropy surface and the information-gain surface volumetrically across the entire selected alpha–beta region, rather than selecting a single parameter pair. Thus, variable importance is defined not through local behavior at a specific parameter setting, but through the holistic structure of entropic behavior in parameter space [16,17,18,19,20,21,22].
The original contribution of this study can be summarized in three dimensions. First, a density-based Sharma-Mittal entropy calculation framework has been established for continuous target and explanatory variables. Second, conditional uncertainty and information gain have been converted into volumetric scores via two-dimensional numerical integration in the alpha-beta parameter plane. Third, the behavior of the proposed scores across six heterogeneous datasets was evaluated in detail using five reference approaches—not only through primary rankings but also via rank correlations, pairwise rank scatter plots, and distributional comparisons. This scope aims to highlight in which types of dependency structures the proposed method converges with classical criteria and in which cases it diverges from model-based variable importance.

2. Materials and Methods

2.1. Parameter-Integrated Sharma–Mittal Entropy Framework

For the continuous random variable Z in the interval 0 z H ^ α , β Y X j , the Sharma-Mittal entropy was calculated using the density function f Z ( z ) . In this study, the alpha and beta parameters were evaluated within a closed range of 0.05 to 0.95, respectively; the upper limit was set at 0.95 to avoid singularities that might occur at the boundaries alpha = 1 and beta = 1. Figure 1 shows the selected parameter space and its relationship with the Rényi, Tsallis, and Shannon special cases of the Sharma-Mittal family.
f ^ Y ( y ) f ^ Y , X j ( y , x ) f ^ { Y | X j } ( y | x ) The density estimation was performed using Gaussian kernel density estimation. The bandwidth was determined using Scott’s automatic rule rather than being set manually for all datasets. For each candidate variable, the marginal density of the target, , and the bivariate joint density, , were estimated; from these, the conditional density, , was derived (see Equation 1). By applying the same computational procedure across all datasets and all variables, the goal was to ensure that any differences in scores stemmed as much as possible from the variable-target relationship rather than from methodological factors [23,24,25,26,27].
j C E = α , β , z : α , β D ,   0 z H ^ α , β Y X j ,     D = α L o w e r , α U p p e r × β L o w e r , β U p p e r
H ^ α , β Y = f ^ Y ( y ) α d y 1 β 1 α 1 1 β
H ^ α , β Y X j = f ^ X j ( x ) f ^ Y , X j ( y , x ) f ^ X j ( x ) α d y 1 β 1 α 1 1 β d x
Figure 1. The location of the Sharma–Mittal entropy in the alpha-beta parameter plane and the integration region Ω = [0.05; 0.95] × [0.05; 0.95] used in this study.
Figure 1. The location of the Sharma–Mittal entropy in the alpha-beta parameter plane and the integration region Ω = [0.05; 0.95] × [0.05; 0.95] used in this study.
Preprints 224148 g001

2.1.1. PICSME (Parameter-Integrated Conditional Sharma-Mittal Entropy)

The Parameter-Integrated Conditional Sharma-Mittal Entropy (PICSME) is defined as the total volume in the alpha-beta plane of the remaining uncertainty regarding the target variable Y when the variable X j is known. A low value in this metric indicates that the conditional uncertainty regarding Y is smaller when X j is observed, and thus the variable has higher explanatory power (see Equation 2). Figure 2 shows the conditional entropy surface from different perspectives.
V j C E = 1 d V = α L o w e r α U p p e r β L o w e r β U p p e r 0 H ^ α , β Y X j 1 d z d β d α
P I C S M E j = V j C E D = 1 D α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y X j d β d α
D = ( α U p p e r α L o w e r ) ( β U p p e r β L o w e r )

2.1.2. PIGSME (Parameter-Integrated Gain of Sharma-Mittal Entropy)

Parameter-Integrated Sharma-Mittal Entropy Information Gain (PIGSME) is the integral over the parameter space of the difference between the marginal entropy surface of the target and the conditional entropy surface of X j . Therefore, PIGSME represents the total volumetric reduction in uncertainty that occurs when the candidate variable is observed. A high PIGSME value is interpreted as indicating that the corresponding predictor provides a higher information gain regarding the target and is expressed as shown in Equation 3.
P I G S M E j = V j I G D = 1 D α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y H ^ α , β Y X j d β d α

2.1.3. NIGSME (Normalized Integrated Gain of Sharma-Mittal Entropy)

Normalized Parameter-Integrated Sharma-Mittal Entropy Information Gain (NIGSME) scales the PIGSME score by dividing it by the total marginal entropy volume of the target variable within the same parameter space (see Equation 4). This normalization aims to reduce the scale effect that may arise when directly comparing the magnitudes of raw information-gain volumes across different datasets (see Figure 4). For NIGSME, a higher score also corresponds to a higher relative information gain.
N I G S M E j = V j I G V Y = α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y H ^ α , β Y X j d β d α α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y d β d α
N I G S M E j = V j I G V Y = 1 α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y X j d β d α α L o w e r α U p p e r β L o w e r β U p p e r H ^ α , β Y d β d α
Figure 3. Parameter-Integrated Sharma–Mittal Entropy Information Gain: A conceptual illustration of the volume of the difference between marginal and conditional entropy surfaces.
Figure 3. Parameter-Integrated Sharma–Mittal Entropy Information Gain: A conceptual illustration of the volume of the difference between marginal and conditional entropy surfaces.
Preprints 224148 g003
Figure 4. Normalized Parameter-Integrated Sharma–Mittal Entropy Information Gain: normalization of the information-gain volume relative to the total Sharma–Mittal entropy volume of the target.
Figure 4. Normalized Parameter-Integrated Sharma–Mittal Entropy Information Gain: normalization of the information-gain volume relative to the total Sharma–Mittal entropy volume of the target.
Preprints 224148 g004

2.2. Datasets and Pre-Processing

The experimental evaluation was conducted on six regression datasets that differed in terms of sample size, the number of explanatory variables, and the application domain. Airfoil Self-Noise represents the acoustic/aerodynamic context, AirQualityUCI represents air quality, BodyFat represents body composition, Meteorology represents hydro-meteorology, Concrete represents materials engineering, and WineQualityWhite represents food/wine quality. This diversity ensures that the proposed indices can be evaluated within an experimental framework suitable for examining similarities and differences between methods in both low-dimensional and high-dimensional physical, sensor-based, and compositional data generation processes. The sample sizes of the datasets, the number of explanatory variables used, the target variables, and the data processing steps are summarized in Table 1 [28,29,30,31,32,33,34,35,36,37,38,39].
The analytical structure between the independent variables and the target variable used in each dataset is shown in Figure 5. In the Airfoil Self-Noise dataset, variable names were defined manually because the source file lacked a header row; in the AirQualityUCI dataset, invalid target observations, along with the date, time, and NMHC(GT) columns, were excluded. In the Meteorology dataset, the date, sunrise, sunset, and weather_code variables were excluded from the analysis; in the AirQualityUCI and Meteorology datasets, a deterministic subsampling of 5,000 observations was applied to limit the cost of two-variable density estimation. In the other datasets, all available observations and continuous explanatory variables were directly included in the analysis [28,29,30,31,32,33,34,35,36,37,38,39].

3. Results

The findings are presented using the same reporting format for each of the six datasets. In each subsection, the raw scores of the reference methods are presented first, followed by the parameter-integrated Sharma-Mittal scores, and finally the feature rankings. The corresponding figures show, in panel (a), the inter-method rank correlations; in panel (b), the geometry of the pairwise ranking relationships; and in panel (c), a comparison of the raw score distributions using the Mann–Whitney U test. Across all datasets, PICSME, PIGSME, and NIGSME produced the same variable rankings. This consistency is consistent with the fact that the target’s marginal entropy volume is constant across variables and that the proposed scores are derived from the same conditional entropy component. Therefore, in this experimental setup, the three indices should be interpreted not as three different rankers, but rather as complementary scales of the same underlying ranking [14,15,40,41,42,43].

3.1. Data 1- Airfoil Self-Noise Dataset

Table 2. Raw importance scores of the reference methods in the Airfoil Self-Noise dataset.
Table 2. Raw importance scores of the reference methods in the Airfoil Self-Noise dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
Frequency_Hz -0,3907 -0,3408 0,2668 0,1729 0,3906
AngleOfAttack_deg -0,1561 -0,1408 0,2019 0,0656 0,0451
ChordLength_m -0,2362 -0,2430 0,0866 0,0664 0,0920
FreeStreamVelocity_m_s 0,1251 0,1162 0,0178 0,0285 0,0434
SuctionSideDisplacementThickness_m -0,3127 -0,2798 0,2443 0,1464 0,4289
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. In the rank analysis, the absolute values of both correlation coefficients were used.
Table 3. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the Airfoil Self-Noise dataset.
Table 3. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the Airfoil Self-Noise dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
Frequency_Hz 1,8447 0,1471 0,0739
AngleOfAttack_deg 1,9634 0,0284 0,0142
ChordLength_m 1,9377 0,0541 0,0272
FreeStreamVelocity_m_s 1,9895 0,0023 0,0012
SuctionSideDisplacementThickness_m 1,9135 0,0783 0,0393
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
In the proposed methods, while the PICSME value for frequency was 1.8447, the PIGSME and NIGSME values were obtained as 0.1471 and 0.0739, respectively. This result indicates that frequency is the variable that reduces conditional uncertainty the most. Suction-side displacement thickness ranked second, with PICSME = 1.9135 and PIGSME = 0.0783. Free-flow velocity, with PICSME = 1.9895 and PIGSME = 0.0023, ranked last across all methods. Thus, the proposed indices produced the same decision structure as the reference filtering methods at both the strong and weak relationship extremes.
When Table 4 and Figure 6 are evaluated together, it can be seen that the three-parameter-integrated Sharma–Mittal index produces exactly the same ranking, and that this ranking is fully consistent with the absolute Pearson correlation, the absolute Spearman rank correlation, and mutual information. For these three reference approaches, rho = 1.000 and tau = 1.000. The relationship between Shannon information gain and random forest importance and the proposed ranking is also strong (rho = 0.900; tau = 0.800); in both cases, only the first two variables swap positions. The fact that the points in Panel (b) cluster largely along a 45-degree line visually confirms this convergence.
Table 4. Rankings of derived features for the eight methods in the Airfoil Self-Noise dataset.
Table 4. Rankings of derived features for the eight methods in the Airfoil Self-Noise dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
Frequency_Hz 1 1 1 1 2 1 1 1
AngleOfAttack_deg 4 4 3 4 4 4 4 4
ChordLength_m 3 3 4 3 3 3 3 3
FreeStreamVelocity_m_s 5 5 5 5 5 5 5 5
SuctionSideDisplacementThickness_m 2 2 2 2 1 2 2 2
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
In Panel (c), the p-values for the raw distributions of the proposed scores compared to those of the reference methods were found to be p < 0.001 for Pearson, p < 0.001 for Spearman, p = 0.037 for Shannon information gain, p < 0.001 for mutual information, and p = 0.037 for random forest importance. While these p-values indicate that the raw score scales are not identical, the rank agreement is extremely high. Therefore, the Airfoil Self-Noise finding demonstrates that the volume-based metric reproduces the same variable importance hierarchy as classical linear and information-based dependency indicators, albeit on a different numerical scale.
Figure 6. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 6. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g006

3.2. Data 2- AirQualityUCI Dataset

In the AirQualityUCI dataset, the target is the CO(GT) concentration, and sensor measurements were evaluated in conjunction with meteorological variables. Table 5 shows that the C6H6(GT) and PT08.S2(NMHC) variables share the top two positions in most of the reference methods. For C6H6(GT), the absolute Pearson correlation coefficient is 0.8955, the absolute Spearman correlation coefficient is 0.9066, the Shannon information gain is 0.9473, and the mutual information is 0.9921; PT08.S2(NMHC), on the other hand, has the highest value in terms of mutual information at 0.9935, by a very narrow margin. In terms of random forest importance, PT08.S2(NMHC) ranks first with 0.4137, followed by C6H6(GT) with 0.3973.
Table 5. Raw importance scores of the reference methods in the AirQualityUCI dataset.
Table 5. Raw importance scores of the reference methods in the AirQualityUCI dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PT08.S1(CO) 0,8472 0,8596 0,7408 0,7494 0,0208
C6H6(GT) 0,8955 0,9066 0,9473 0,9921 0,3973
PT08.S2(NMHC) 0,8834 0,9066 0,9473 0,9935 0,4137
NOx(GT) 0,7849 0,7499 0,6154 0,6392 0,0883
PT08.S3(NOx) -0,6819 -0,7927 0,5552 0,5621 0,0099
NO2(GT) 0,6704 0,7053 0,4471 0,4510 0,0212
PT08.S4(NO2) 0,6099 0,5794 0,3152 0,3327 0,0083
PT08.S5(O3) 0,8236 0,8369 0,6721 0,6763 0,0106
T 0,0199 0,0661 0,0552 0,0374 0,0155
RH 0,0475 -0,0055 0,0526 0,0480 0,0074
AH 0,0458 0,0509 0,0417 0,0366 0,0068
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. In the rank analysis, the absolute values of both correlation coefficients were used.
Table 6. Raw scores of the proposed parameter-integrated Sharma–Mittal entropy-based methods in the AirQualityUCI dataset.
Table 6. Raw scores of the proposed parameter-integrated Sharma–Mittal entropy-based methods in the AirQualityUCI dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
PT08.S1(CO) 1,0981 0,7887 0,4180
C6H6(GT) 0,8994 0,9873 0,5233
PT08.S2(NMHC) 0,8998 0,9869 0,5231
NOx(GT) 1,2223 0,6645 0,3522
PT08.S3(NOx) 1,3483 0,5384 0,2854
NO2(GT) 1,4315 0,4553 0,2413
PT08.S4(NO2) 1,5271 0,3597 0,1906
PT08.S5(O3) 1,1357 0,7510 0,3980
T 1,8947 -0,0079 -0,0042
RH 1,8942 -0,0074 -0,0039
AH 1,9129 -0,0261 -0,0138
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
Parameter-integrated entropy scores ranked C6H6(GT) first and PT08.S2(NMHC) second. For C6H6(GT), PICSME = 0.8994, PIGSME = 0.9873, and NIGSME = 0.5233; for PT08.S2(NMHC), these values are 0.8998, 0.9869, and 0.5231, respectively. The very small difference between the two variables suggests that the common information structure between the sensor responses and the CO(GT) target is nearly equivalent. In contrast, the variables of temperature, relative humidity, and absolute humidity rank lower in the proposed methods; AH drops to eleventh place among all proposed scores.
Table 7. Rankings of derived features from the eight methods in the AirQualityUCI dataset.
Table 7. Rankings of derived features from the eight methods in the AirQualityUCI dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
PT08.S1(CO) 3 3 3 3 5 3 3 3
C6H6(GT) 1 1 1 2 2 1 1 1
PT08.S2(NMHC) 2 1 1 1 1 2 2 2
NOx(GT) 5 6 5 5 3 5 5 5
PT08.S3(NOx) 6 5 6 6 8 6 6 6
NO2(GT) 7 7 7 7 4 7 7 7
PT08.S4(NO2) 8 8 8 8 9 8 8 8
PT08.S5(O3) 4 4 4 4 7 4 4 4
T 11 9 9 10 6 10 10 10
RH 9 11 10 9 10 9 9 9
AH 10 10 11 11 11 11 11 11
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
The results in Figure 7(a) show that the proposed ranking is in very strong agreement with the reference filtering methods. For the proposed methods, the absolute Pearson correlation and mutual information yielded rho = 0.991 and tau = 0.964; Shannon information gain yielded rho = 0.989 and tau = 0.954; and the absolute Spearman rank correlation yielded rho = 0.961 and tau = 0.881. This finding indicates that the proposed entropy volumes effectively reflect not only the linear association but also the information-based relationship between the sensor variables and the target.
The relationship with the random forest is lower but still positive and significant (rho = 0.773; tau = 0.636). The divergence is particularly pronounced for NOx(GT), NO2(GT), PT08.S5(O3), and temperature; while the random forest ranked some of these variables higher, the proposed indices produced a more conservative level of importance based on the direct conditional uncertainty relationship with the target. The Mann–Whitney U test showed that the proposed scores differed from the raw score distributions of the reference methods at the p < 0.001 level for Pearson, Spearman, Shannon information gain, and mutual information; and at the p = 0.005 level for the random forest. This distinction stems from the scoring scale; the high agreement in panels (a) and (b) demonstrates that the underlying importance hierarchy is strongly preserved.
Figure 7. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 7. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g007

3.3. Data 3-BodyFat Dataset

The BodyFat dataset was used to evaluate the relationship between anthropometric measurements and body fat percentage. Table 8 shows that the “abdomen” variable had by far the highest significance across all reference methods. For the abdomen, the absolute Pearson correlation is 0.8077, the absolute Spearman rank correlation is 0.8125, the Shannon information gain is 0.8401, the mutual information is 0.5379, and the random forest importance is 0.7159. Chest, hip, weight, and thigh measurements rank among the variables with the next highest importance.
Table 8. Raw importance scores of the reference methods in the BodyFat dataset.
Table 8. Raw importance scores of the reference methods in the BodyFat dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
Age 0,2881 0,2689 0,5387 0,1030 0,0246
Weight 0,5987 0,6041 0,6207 0,2614 0,0272
Height -0,1089 -0,0272 0,4241 0,0000 0,0439
Neck 0,4752 0,4798 0,5244 0,0855 0,0259
Chest 0,6938 0,6674 0,6956 0,2850 0,0225
Abdomen 0,8077 0,8125 0,8401 0,5379 0,7159
Hip 0,6136 0,6037 0,6796 0,2470 0,0195
Thigh 0,5431 0,5348 0,6257 0,1784 0,0189
Knee 0,4928 0,4791 0,5606 0,1276 0,0231
Ankle 0,2516 0,2886 0,4900 0,0180 0,0198
Biceps 0,4756 0,4825 0,6073 0,1972 0,0152
Forearm 0,3418 0,3797 0,5159 0,0611 0,0160
Wrist 0,3278 0,3000 0,5463 0,0593 0,0275
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. In the rank analysis, the absolute values of both correlation coefficients were used.
Table 9. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the BodyFat dataset.
Table 9. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the BodyFat dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
Age 1,9079 0,0516 0,0263
Weight 1,6693 0,2902 0,1481
Height 1,9588 0,0006 0,0003
Neck 1,7981 0,1614 0,0824
Chest 1,5072 0,4523 0,2308
Abdomen 1,1950 0,7644 0,3901
Hip 1,6296 0,3299 0,1683
Thigh 1,7387 0,2208 0,1127
Knee 1,7782 0,1812 0,0925
Ankle 1,9185 0,0409 0,0209
Biceps 1,7967 0,1628 0,0831
Forearm 1,8644 0,0950 0,0485
Wrist 1,8844 0,0750 0,0383
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
The proposed methods ranked the abdomen variable first, with PICSME = 1.1950, PIGSME = 0.7644, and NIGSME = 0.3901. The chest ranked second, the hips third, weight fourth, and the thighs fifth. This ranking consistently reflects the strong relationship between body fat percentage and circumferential body measurements in terms of both conditional uncertainty and information gain. Notably, the height variable ranked second in random forest importance despite being last in the reference correlation and information criteria; the proposed methods, however, ranked height thirteenth.
Table 10. Derived feature rankings of the eight methods in the BodyFat dataset.
Table 10. Derived feature rankings of the eight methods in the BodyFat dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
Age 11 12 9 8 6 11 11 11
Weight 4 3 5 3 4 4 4 4
Height 13 13 13 13 2 13 13 13
Neck 8 7 10 9 5 8 8 8
Chest 2 2 2 2 8 2 2 2
Abdomen 1 1 1 1 1 1 1 1
Hip 3 4 3 4 10 3 3 3
Thigh 5 5 4 6 11 5 5 5
Knee 6 8 7 7 7 6 6 6
Ankle 12 11 12 12 9 12 12 12
Biceps 7 6 6 5 13 7 7 7
Forearm 9 9 11 10 12 9 9 9
Wrist 10 10 8 11 3 10 10 10
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
Figure 8(a) shows perfect agreement between the proposed ranking and the absolute Pearson correlation (rho = 1.000; tau = 1.000). For the absolute Spearman rank correlation, rho = 0.973 and tau = 0.897; for both Shannon information gain and mutual information, rho = 0.945 and tau = 0.846 were obtained. In Panel (b), the clustering of points associated with these four methods around the diagonal indicates that the proposed indices largely preserve both the linear and information-based importance structure in the context of body fat percentage.
Random forest variable importance, however, exhibited a pattern nearly independent of the proposed ranking (rho = −0.044; tau = −0.026). This sharp divergence can be explained by the random forest’s use of multivariate tree splits and the correlation structure among variables. For example, while height and ankle circumference had low importance in terms of direct univariate uncertainty reduction, they were assigned higher importance for contextual splits in the random forest structure. The fact that the p-value was 0.887 in the comparison of the proposed scores with the raw distribution using the random forest in Panel (c) indicates that there is no discernible difference in the numerical score distributions between these two methods; however, the very low rank correlation suggests that the same distribution position corresponds to the same importance hierarchy.

3.4. Data 4-Meteorology Dataset

The meteorological dataset, with 9,497 daily observations and 45 explanatory variables, is the largest sample in the study. The target variable is FAO-based reference evapotranspiration. Therefore, the dataset presents a complex assessment environment in which multiple physical processes—such as temperature, radiation, humidity, pressure, wind, soil temperature, and soil moisture—simultaneously influence the same target variable. In Table 11, total shortwave radiation ranks first among all reference methods; it is followed by maximum vapor pressure deficit, maximum air temperature, and mean air temperature.
Table 11. Raw importance scores of reference methods in the meteorological dataset.
Table 11. Raw importance scores of reference methods in the meteorological dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
temperature_2m_max 0,9127 0,9189 0,8894 0,9195 0,0204
temperature_2m_min 0,8527 0,8491 0,6466 0,6442 0,0004
precipitation_sum -0,3036 -0,3842 0,1385 0,1389 0,0000
rain_sum -0,2759 -0,3467 0,1137 0,1184 0,0000
snowfall_sum -0,1776 -0,3401 0,0607 0,0710 0,0000
precipitation_hours -0,4036 -0,4000 0,1622 0,1756 0,0000
sunshine_duration 0,8074 0,9224 0,9604 1,0007 0,0003
daylight_duration 0,8802 0,8851 0,7958 0,7935 0,0016
wind_speed_10m_max 0,1050 0,1319 0,0603 0,0450 0,0008
wind_gusts_10m_max 0,0211 0,0712 0,0999 0,0903 0,0004
wind_direction_10m_dominant -0,0734 -0,0753 0,0826 0,0649 0,0001
shortwave_radiation_sum 0,9491 0,9624 1,2317 1,2979 0,8415
temperature_2m_mean 0,9013 0,9006 0,8236 0,8367 0,0043
cloud_cover_mean -0,6565 -0,6649 0,3841 0,3734 0,0002
cloud_cover_max -0,5789 -0,6128 0,2754 0,2694 0,0001
cloud_cover_min -0,4599 -0,5383 0,2343 0,2488 0,0000
dew_point_2m_mean 0,5742 0,6064 0,2890 0,2791 0,0001
dew_point_2m_max 0,6768 0,7110 0,3845 0,3752 0,0001
dew_point_2m_min 0,4320 0,4528 0,1797 0,1659 0,0001
relative_humidity_2m_mean -0,8605 -0,8661 0,7143 0,7278 0,0064
relative_humidity_2m_max -0,7267 -0,7171 0,3982 0,3885 0,0005
relative_humidity_2m_min -0,8055 -0,8449 0,6525 0,6761 0,0001
pressure_msl_mean -0,5239 -0,5202 0,2747 0,2710 0,0001
pressure_msl_max -0,5819 -0,5832 0,2986 0,2951 0,0001
pressure_msl_min -0,4856 -0,4832 0,2652 0,2544 0,0001
surface_pressure_mean -0,0740 -0,0996 0,1464 0,1449 0,0001
surface_pressure_max -0,1573 -0,1664 0,1321 0,1227 0,0001
surface_pressure_min -0,0188 -0,0596 0,1557 0,1461 0,0001
wind_speed_10m_mean 0,0679 0,1040 0,0851 0,0831 0,0070
wind_speed_10m_min -0,0562 -0,0657 0,0545 0,0440 0,0003
wind_gusts_10m_mean -0,0104 0,0591 0,1426 0,1368 0,0018
wind_gusts_10m_min -0,0924 -0,0727 0,0688 0,0566 0,0003
vapour_pressure_deficit_max 0,9146 0,9408 1,0343 1,0910 0,1105
wet_bulb_temperature_2m_mean 0,8072 0,8261 0,5803 0,5684 0,0001
wet_bulb_temperature_2m_max 0,8167 0,8412 0,6053 0,6010 0,0001
wet_bulb_temperature_2m_min 0,7823 0,7959 0,5208 0,5178 0,0001
soil_temperature_0_to_7cm_mean 0,9019 0,8975 0,8092 0,8241 0,0008
soil_temperature_7_to_28cm_mean 0,8886 0,8854 0,7523 0,7635 0,0003
soil_temperature_28_to_100cm_mean 0,8332 0,8359 0,6190 0,6303 0,0002
soil_temperature_0_to_100cm_mean 0,8562 0,8561 0,6684 0,6830 0,0002
soil_moisture_0_to_7cm_mean -0,4303 -0,5312 0,3110 0,3691 0,0002
soil_moisture_7_to_28cm_mean -0,3155 -0,4164 0,2738 0,3217 0,0001
soil_moisture_28_to_100cm_mean -0,0919 -0,1380 0,1761 0,2564 0,0001
soil_moisture_0_to_100cm_mean -0,1771 -0,2576 0,1942 0,2527 0,0001
snowfall_water_equivalent_sum -0,1776 -0,3401 0,0607 0,0730 0,0000
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. In the rank analysis, the absolute values of both correlation coefficients were used.
Table 12. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the meteorological dataset.
Table 12. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the meteorological dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
temperature_2m_max 0,6099 1,0656 0,6360
temperature_2m_min 0,9231 0,7524 0,4490
precipitation_sum 1,6074 0,0681 0,0406
rain_sum 1,6247 0,0508 0,0303
snowfall_sum 1,6771 -0,0016 -0,0010
precipitation_hours 1,5445 0,1310 0,0782
sunshine_duration 0,7803 0,8952 0,5343
daylight_duration 0,7992 0,8763 0,5230
wind_speed_10m_max 1,6714 0,0041 0,0024
wind_gusts_10m_max 1,6583 0,0172 0,0103
wind_direction_10m_dominant 1,6768 -0,0014 -0,0008
shortwave_radiation_sum 0,3953 1,2802 0,7641
temperature_2m_mean 0,6895 0,9859 0,5885
cloud_cover_mean 1,3336 0,3418 0,2040
cloud_cover_max 1,4485 0,2270 0,1355
cloud_cover_min 1,5101 0,1654 0,0987
dew_point_2m_mean 1,4760 0,1995 0,1191
dew_point_2m_max 1,3466 0,3289 0,1963
dew_point_2m_min 1,5998 0,0757 0,0452
relative_humidity_2m_mean 0,8863 0,7892 0,4710
relative_humidity_2m_max 1,2753 0,4002 0,2389
relative_humidity_2m_min 0,9810 0,6945 0,4145
pressure_msl_mean 1,4611 0,2143 0,1279
pressure_msl_max 1,4320 0,2435 0,1453
pressure_msl_min 1,4723 0,2032 0,1213
surface_pressure_mean 1,6270 0,0485 0,0289
surface_pressure_max 1,6368 0,0387 0,0231
surface_pressure_min 1,6206 0,0549 0,0328
wind_speed_10m_mean 1,6492 0,0263 0,0157
wind_speed_10m_min 1,6917 -0,0162 -0,0097
wind_gusts_10m_mean 1,6273 0,0482 0,0287
wind_gusts_10m_min 1,6839 -0,0084 -0,0050
vapour_pressure_deficit_max 0,5139 1,1615 0,6933
wet_bulb_temperature_2m_mean 1,0562 0,6193 0,3696
wet_bulb_temperature_2m_max 1,0203 0,6552 0,3911
wet_bulb_temperature_2m_min 1,1383 0,5372 0,3206
soil_temperature_0_to_7cm_mean 0,7373 0,9382 0,5600
soil_temperature_7_to_28cm_mean 0,8162 0,8593 0,5128
soil_temperature_28_to_100cm_mean 1,0334 0,6421 0,3832
soil_temperature_0_to_100cm_mean 0,9549 0,7206 0,4301
soil_moisture_0_to_7cm_mean 1,4961 0,1794 0,1071
soil_moisture_7_to_28cm_mean 1,5789 0,0966 0,0577
soil_moisture_28_to_100cm_mean 1,6794 -0,0039 -0,0024
soil_moisture_0_to_100cm_mean 1,6667 0,0088 0,0053
snowfall_water_equivalent_sum 1,6771 -0,0016 -0,0010
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
The proposed indices ranked the total shortwave radiation first, with PICSME = 0.3953, PIGSME = 1.2802, and NIGSME = 0.7641. Maximum vapor pressure deficit ranked second; maximum air temperature, third; mean air temperature, fourth; and soil temperature at 0–7 cm, fifth. These results point to a physically consistent pattern of importance, confirming that evapotranspiration is driven by energy input, atmospheric aridity, and temperature regime. Sunshine duration ranked sixth in the proposed ranking, while it rose as high as third in some reference methods. This difference indicates that, while duration variables highly correlated with radiation share common information, the proposed method performs a more balanced decomposition through the conditional uncertainty volume.
Table 13. Derived feature rankings of the eight methods in the meteorological dataset.
Table 13. Derived feature rankings of the eight methods in the meteorological dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
temperature_2m_max 3 4 4 4 3 3 3 3
temperature_2m_min 10 11 12 12 12 10 10 10
precipitation_sum 30 30 35 34 43 30 30 30
rain_sum 31 31 37 37 42 32 32 32
snowfall_sum 32 32 42 41 45 41 41 41
precipitation_hours 28 29 31 30 41 27 27 27
sunshine_duration 13 3 3 3 16 6 6 6
daylight_duration 7 8 7 7 8 7 7 7
wind_speed_10m_max 36 37 44 44 9 39 39 39
wind_gusts_10m_max 43 42 38 38 13 37 37 37
wind_direction_10m_dominant 40 40 40 42 23 40 40 40
shortwave_radiation_sum 1 1 1 1 1 1 1 1
temperature_2m_mean 5 5 5 5 6 4 4 4
cloud_cover_mean 19 19 19 19 19 18 18 18
cloud_cover_max 21 20 23 25 27 21 21 21
cloud_cover_min 25 23 27 29 40 26 26 26
dew_point_2m_mean 22 21 22 23 32 24 24 24
dew_point_2m_max 18 18 18 18 26 19 19 19
dew_point_2m_min 26 27 29 31 28 29 29 29
relative_humidity_2m_mean 8 9 9 9 5 9 9 9
relative_humidity_2m_max 17 17 17 17 11 17 17 17
relative_humidity_2m_min 15 12 11 11 25 12 12 12
pressure_msl_mean 23 25 24 24 37 22 22 22
pressure_msl_max 20 22 21 22 34 20 20 20
pressure_msl_min 24 26 26 27 31 23 23 23
surface_pressure_mean 39 39 33 33 39 33 33 33
surface_pressure_max 35 35 36 36 38 35 35 35
surface_pressure_min 44 44 32 32 33 31 31 31
wind_speed_10m_mean 41 38 39 39 4 36 36 36
wind_speed_10m_min 42 43 45 45 14 45 45 45
wind_gusts_10m_mean 45 45 34 35 7 34 34 34
wind_gusts_10m_min 37 41 41 43 17 44 44 44
vapour_pressure_deficit_max 2 2 2 2 2 2 2 2
wet_bulb_temperature_2m_mean 14 15 15 15 30 15 15 15
wet_bulb_temperature_2m_max 12 13 14 14 22 13 13 13
wet_bulb_temperature_2m_min 16 16 16 16 29 16 16 16
soil_temperature_0_to_7cm_mean 4 6 6 6 10 5 5 5
soil_temperature_7_to_28cm_mean 6 7 8 8 15 8 8 8
soil_temperature_28_to_100cm_mean 11 14 13 13 20 14 14 14
soil_temperature_0_to_100cm_mean 9 10 10 10 18 11 11 11
soil_moisture_0_to_7cm_mean 27 24 20 20 21 25 25 25
soil_moisture_7_to_28cm_mean 29 28 25 21 24 28 28 28
soil_moisture_28_to_100cm_mean 38 36 30 26 35 43 43 43
soil_moisture_0_to_100cm_mean 34 34 28 28 36 38 38 38
snowfall_water_equivalent_sum 32 32 42 40 44 41 41 41
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
The rank correlations in Figure 9(a) show that the proposed methods are highly consistent with correlation- and information-based filters. For the absolute Pearson correlation, rho = 0.950 and tau = 0.840; for the absolute Spearman rank correlation, rho = 0.956 and tau = 0.854; Shannon information gain yielded rho = 0.970 and tau = 0.889; and mutual information yielded rho = 0.958 and tau = 0.872. This pattern indicates that the proposed entropy measures capture both linear and more general information relationships, particularly highlighting the fundamental physical drivers in high-dimensional meteorological structures.
The agreement with random forest importance scores is moderate (rho = 0.497; tau = 0.411). This divergence is evident for the mean wind speed and the mean wind gust: while the random forest ranked these variables fourth and seventh, respectively, the proposed indices ranked them thirty-sixth and thirty-fourth, respectively. In contrast, although the average relative humidity ranked high in both approaches, it was placed fifth by the random forest and ninth by the proposed method. The fact that the proposed scores in Panel (c) differ from all reference methods at the raw scale level with p < 0.001 indicates that the methods’ numerical outputs are not on the same scale; meanwhile, the high rho and tau values in Panel (a) show that the ranking mechanism largely captures a common physical signal.
Figure 9. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 9. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g009

3.5. Data 5-Concrete Dataset

The Concrete dataset evaluates the effect of mix components and sample age on compressive strength. In Table 14, the cement variable ranks near the top in terms of absolute Pearson correlation, Shannon information gain, and random forest importance; while the age variable ranks first in terms of absolute Spearman correlation, mutual information, and random forest importance. These two variables demonstrate that the target strength is influenced by both the material composition and the curing process.
Table 14. Raw importance scores of the reference methods in the Concrete dataset.
Table 14. Raw importance scores of the reference methods in the Concrete dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
Cement 0,4978 0,4776 0,4372 0,2201 0,3251
Blast Furnace Slag 0,1348 0,1625 0,2234 0,1405 0,0808
Fly Ash -0,1058 -0,0780 0,1388 0,0760 0,0170
Water -0,2896 -0,3084 0,3757 0,2783 0,1046
Superplasticizer 0,3661 0,3476 0,2648 0,1671 0,0727
Coarse Aggregate -0,1649 -0,1835 0,3440 0,1448 0,0290
Fine Aggregate -0,1672 -0,1800 0,3115 0,1516 0,0370
Age (day) 0,3289 0,5960 0,3505 0,3194 0,3338
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. In the rank analysis, the absolute values of both correlation coefficients were used.
Table 15. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the Concrete dataset.
Table 15. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the Concrete dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
Cement 1,7511 0,2057 0,1051
Blast Furnace Slag 1,9443 0,0125 0,0064
Fly Ash 1,9411 0,0157 0,0080
Water 1,8298 0,1270 0,0649
Superplasticizer 1,8745 0,0822 0,0420
Coarse Aggregate 1,9449 0,0119 0,0061
Fine Aggregate 1,9428 0,0140 0,0072
Age (day) 1,7889 0,1679 0,0858
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
The parameter-integrated scores ranked cement first, age second, and water third. For cement, PICSME = 1.7511, PIGSME = 0.2057, and NIGSME = 0.1051; for age, PICSME = 1.7889, PIGSME = 0.1679, and NIGSME = 0.0858. Water ranks third with PICSME = 1.8298 and PIGSME = 0.1270. Superplasticizer ranks fourth, fly ash fifth, fine aggregate sixth, blast furnace slag seventh, and coarse aggregate eighth.
Table 16. Rankings of derived features for the eight methods in the Concrete dataset.
Table 16. Rankings of derived features for the eight methods in the Concrete dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
Cement 1 2 1 3 2 1 1 1
Blast Furnace Slag 7 7 7 7 4 7 7 7
Fly Ash 8 8 8 8 8 5 5 5
Water 4 4 2 2 3 3 3 3
Superplasticizer 2 3 6 4 5 4 4 4
Coarse Aggregate 6 5 4 6 7 8 8 8
Fine Aggregate 5 6 5 5 6 6 6 6
Age (day) 3 1 3 1 1 2 2 2
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
Figure 10(a) shows that the agreement between methods in this dataset is at a moderate level compared to previous datasets. For the proposed methods, a correlation of rho = 0.762 for absolute Pearson correlation and mutual information, and tau = 0.571; and a correlation of rho = 0.738 for absolute Spearman rank correlation and random forest importance were obtained. The agreement with Shannon information gain is lower, at rho = 0.619 and tau = 0.429. This result suggests that the linear, non-monotonic, and interactive effects of mixture components on concrete strength are represented differently depending on the method used.
Ranking differences are particularly evident for superplasticizers, aggregate types, and fly ash. For example, the random forest ranked age first, while the proposed methods ranked it second; Shannon information gain, however, ranked water second. The fact that the comparison of the proposed scores with Shannon information gain in Panel (c) yielded a p-value of 0.102 indicates that there is no significant difference in the positions of the raw score distributions. Nevertheless, the fact that the rank correlation remains at a moderate level reaffirms that similarity in the scale distribution does not imply equivalence in the rankings across variables. P-values for other reference methods ranged from 0.028 to 0.037.

3.6. Data 6-WineQualityWhite Dataset

The WineQualityWhite dataset is designed to explain the quality score of white wines based on their physicochemical measurements. According to Table 17, alcohol ranks first in both absolute Pearson and absolute Spearman correlations and in Shannon information gain; it ranks second in mutual information and first in random forest importance. The density variable ranks second in terms of correlation and Shannon information gain and first in terms of mutual information. Chloride ranks third in the correlation and Shannon information gain metrics.
The proposed methods ranked alcohol first, density second, and chloride third. Free sulfur dioxide ranked fourth, total sulfur dioxide fifth, and volatile acid sixth. This ranking preserves the general trend where alcohol and density variables dominate quality classification while significantly elevating the importance of free sulfur dioxide compared to correlation-based filters. Table 18 shows that the PIGSME and NIGSME values are negative for all variables. This indicates that the magnitude of information gain on the absolute scale of the density-based continuous Sharma–Mittal functional and numerical integration can fall below zero; therefore, the sign of PIGSME and NIGSME in this dataset should be interpreted not as absolute information gain but as a relative ranking score within the dataset.
Table 19. Derived feature rankings of the eight methods on the WineQualityWhite dataset.
Table 19. Derived feature rankings of the eight methods on the WineQualityWhite dataset.
Variables Rank Distribution
Comparison Methods’ Ranks Proposed Methods’ Ranks
Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
PICSME PIGSME NIGSME
fixed acidity 6 7 10 11 10 9 9 9
volatile acidity 4 5 6 7 2 6 6 6
citric acid 10 11 4 6 11 7 7 7
residual sugar 8 8 7 4 6 8 8 8
chlorides 3 3 3 5 7 3 3 3
free sulfur dioxide 11 10 8 8 3 4 4 4
total sulfur dioxide 5 4 5 3 5 5 5 5
density 2 2 2 1 8 2 2 2
pH 7 6 11 9 4 11 11 11
sulphates 9 9 9 10 9 10 10 10
alcohol 1 1 1 2 1 1 1 1
Note: Rank 1 corresponds to the highest importance. PICSME is listed in ascending order; PIGSME and NIGSME are listed in descending order. The rank equivalence of the three proposed scores under the same target variable and parameter region was also evaluated.
In Figure 11(a), the highest agreement was achieved using Shannon information gain (ρ = 0.873; τ = 0.745). This result indicates that the proposed volumetric scores in the WineQualityWhite dataset produce a similarity structure particularly aligned with the logic of uncertainty reduction based on disaggregation. The correlation with mutual information is rho = 0.764 and tau = 0.527; with absolute Spearman, rho = 0.618 and tau = 0.491; and with absolute Pearson, rho = 0.600 and tau = 0.455. The agreement with random forest importance scores is weaker (rho = 0.364; tau = 0.200). In particular, free sulfur dioxide ranks fourth in the proposed ranking but tenth or eleventh with absolute Pearson and Spearman, and third with the random forest.
In the comparison of the proposed scores with Pearson’s correlation in Figure 11.c, p = 0.051 was obtained, and p = 0.272 was obtained in the comparison with the random forest; in both cases, no significant difference in the distributions of the raw scores was observed. In contrast, p = 0.043 was found for Spearman, p < 0.001 for Shannon information gain, and p = 0.006 for mutual information. The WineQualityWhite results point to two important observations: first, the proposed framework consistently identifies dominant variables such as alcohol and intensity; second, negative PIGSME/NIGSME values and moderate rank consistency indicate that the continuous intensity-based information gain interpretation requires more comprehensive testing in terms of axiomatic and numerical robustness.

4. Discussion

4.1. Evolution of Entropy Measures in Information Theory

Entropy is one of the fundamental concepts in information theory, statistics, physics, and machine learning. Introduced by Shannon to measure uncertainty in a probability distribution, entropy quantifies the average information content of a random variable and forms the foundation of modern information theory. Although Shannon entropy is still widely used, generalized measures of entropy have been developed because a single formulation is not always sufficient to account for the various types of uncertainty that arise in complex systems. In particular, systems involving long-range interactions, non-equilibrium states, multifractal structures, power-law behavior, or non-additive information composition can be more appropriately represented by alternative measures such as the Rényi, Tsallis, Kaniadakis, Abe, and Sharma-Mittal entropies [7].
One of the first parametric generalizations was proposed by Rényi, who defined α-order entropy for finite probability distributions. This formulation incorporates the α deformation parameter into the model while preserving some fundamental properties of Shannon entropy—such as symmetry and additivity—for independent distributions. Shannon entropy is obtained in the limit as α→1 [17].
Another important development is Tsallis entropy, proposed as a generalization of the Boltzmann-Gibbs statistic. This entropy includes the deformation parameter q; in the limit as q → 1, it reduces to the Boltzmann-Gibbs/Shannon form and has become one of the fundamental components of non-expansive statistical mechanics. Due to its non-additive structure, it is well-suited for modeling systems involving strong correlations, non-ergodic behavior, and long-range interactions. Tsallis-based approaches have been applied to many complex phenomena, such as turbulence, anomalous diffusion, biological systems, financial markets, and physical systems with long-range interactions [18].
Despite their importance, the Rényi and Tsallis entropies represent different paths of generalization, and neither can be considered a direct extension of the other. The Sharma-Mittal entropy overcomes this limitation by providing a two-parameter framework that encompasses both formulations as limit cases. Specifically, the Sharma-Mittal entropy reduces to the Rényi entropy as r→1, to the Tsallis entropy as r→q, and to the Shannon entropy at the corresponding classical parameter limits [16,17,18,19].
The importance of the Sharma-Mittal entropy does not stem solely from its unifying nature; it also offers a broader capacity for generalization. Beck [20] considers the Sharma-Mittal entropy to be among the important generalized measures of information used in characterizing complex systems and notes that this framework encompasses various well-known entropy measures, such as the Tsallis, Kaniadakis, and Abe entropies, as special cases. Similarly, Ilić et al. [22] classify the Sharma-Mittal entropy within a broad family of generalized entropic forms and emphasize its relationship with other deformed logarithmic and trace-form entropy structures.
Recent studies have also examined the mathematical behavior of generalized entropies under different probability distributions. Bodnarchuk et al. analyzed Shannon, Rényi, generalized Rényi, Tsallis, Sharma-Mittal, and modified Shannon entropies for various distributions, such as the gamma, chi-square, exponential, Laplace, and log-normal distributions. Their findings show that entropy values and parameter-dependent behaviors vary depending on both the distribution families and the entropy formulations; therefore, they emphasize the need to consider the probabilistic structure of the data when selecting an entropy measure [44].
In general, entropy theory has evolved from the classical Shannon/Boltzmann-Gibbs formulation into a broad family of generalized information-theoretic tools. Within this development, the Sharma-Mittal entropy stands out as a comprehensive formulation that, thanks to its two-parameter structure, can encompass major measures such as the Rényi, Tsallis, and Shannon entropies in appropriate limits and offers flexibility in modeling complex probabilistic systems [7,16,17,18,19,20,22,44].

4.2. Sharma-Mittal Entropy as a Unifying Two-Parameter Framework

The Sharma-Mittal entropy family provides a comprehensive framework that unifies various generalized entropy measures under a common mathematical structure. Unlike single-parameter entropy formulations, the Sharma-Mittal entropy is determined by two parameters. This structure allows for the control of both the deformation of the probability distribution and the degree of non-expansion within the same framework. Therefore, the Sharma-Mittal entropy can be regarded not only as an alternative entropy measure but also as a general framework that explains the relationships among different entropy families [16].
The fundamental significance of the Sharma-Mittal entropy lies in its ability to encompass well-known entropy measures as special or limiting cases. Under appropriate parameter choices, it can be reduced to the Rényi entropy, the Tsallis entropy, and ultimately the Shannon/Boltzmann-Gibbs entropy. In this respect, it serves as a two-parameter mathematical bridge between entropy measures that were previously treated as separate generalizations. However, its thermostatsistical interpretation has also been critically discussed in the literature [7,16,17,18,19,20,21].
Masi has shown that the Sharma-Mittal entropy can be interpreted as a natural extension of the generalized logarithmic and exponential formalisms and that it provides a consistent framework for explaining the relationships between the Rényi, Tsallis, and Shannon/Boltzmann-Gibbs entropies [19]. In contrast, Aktürk et al. have argued that the physical interpretation of this entropy must be handled with caution. The authors have shown that a free-energy-based interpretation of the Sharma-Mittal relation may not always be valid unless specific thermodynamic assumptions are satisfied [21].
The broader theoretical significance of the Sharma-Mittal entropy has also been highlighted in classifications of generalized entropies. Ilić et al. evaluated the Sharma-Mittal entropy among the two main families of two-parameter entropies in statistical physics and information theory and demonstrated that various entropy formulations can be derived from this family through appropriate parameter choices [22]. Similarly, Bodnarchuk et al. analyzed the Shannon, Rényi, generalized Rényi, Tsallis, Sharma-Mittal, and modified Shannon entropies under different probability models and showed that parameter-dependent entropy behavior changes according to the distribution family [44].
This two-parameter structure provides the Sharma-Mittal entropy with significant flexibility in characterizing uncertainty. While Shannon entropy assigns a single fixed uncertainty value to a probability distribution, the Sharma-Mittal entropy defines a parametric entropy surface that varies depending on the parameters α and β. This allows the entropy measure to be sensitive in different ways to distributional properties such as concentration, tail behavior, low-probability events, and the dominance of high-probability events [16,17,18,19,20,21,22,44].
For this reason, the Sharma-Mittal entropy provides a flexible theoretical foundation for modeling uncertainty in complex data structures. Its two-parameter structure provides theoretical support for the use of Sharma-Mittal-based measures of information in data analysis, feature selection, and machine learning applications.

4.3. Extensions of Sharma-Mittal Entropy in Information Theory

The importance of the Sharma-Mittal entropy does not stem solely from its role as a generalized measure of uncertainty. Thanks to its two-parameter structure and its ability to encompass various classical entropy measures as limiting cases, it provides a flexible foundation for extending information-theoretic concepts beyond the Shannon framework. From this perspective, the Sharma-Mittal entropy provides a theoretical foundation for generalized forms of concepts such as uncertainty reduction, information transfer, and dependency measurement.
One of the key research areas is the generalization of measures of information transfer. In classical information theory, mutual information expresses the extent to which uncertainty about a random variable is reduced by observing another variable. Since this concept is closely related to information gain, it is also of central importance in feature selection, decision trees, and other learning models. Extending this concept beyond Shannon entropy allows for the evaluation of feature importance under more flexible uncertainty structures, particularly in cases where the data distribution cannot be adequately represented by a single, fixed measure of entropy.
One of the significant contributions in this area was made by Ilić and Djordjević. The authors defined α-q mutual information and α-q channel capacity as Sharma-Mittal-based measures of information transfer [45]. The proposed framework satisfies fundamental information-theoretic properties such as non-negativity, boundedness by input and output Sharma-Mittal entropies, consistency with perfect transmission, and zero information transfer in completely degrading channels [45]. Furthermore, with appropriate parameter choices, the measure can be reduced to the classical mutual information, Rényi mutual information, or Tsallis mutual information. These properties indicate that Sharma-Mittal entropy is not only a descriptive uncertainty measure but also a potential basis for generalized dependence and information-transfer metrics.
Another important extension relates to generalized distance/disjointness measures, which are used to quantify differences between probability distributions and are widely used in information theory, statistics, and information geometry. Nielsen and Nock derived closed-form expressions for Rényi and Tsallis entropies and divergence measures for distributions belonging to exponential families and showed that these measures can be expressed in terms of the log-normalizing function of the relevant family [46]. Although this study focuses on Rényi and Tsallis families, it is theoretically important because both are special cases or limiting cases of the Sharma-Mittal framework.
This divergence-based perspective extends entropy-based thinking from a measure of uncertainty to a distributional comparison. In a Sharma-Mittal-based framework, the two-parameter form can enable the evaluation of differences between probability distributions under varying sensitivity conditions. Sharma-Mittal entropy has also been extended to generalized entropy measures, and its two-parameter framework has been used to investigate concavity properties in diffusion processes [47].
Overall, these developments indicate that the Sharma-Mittal entropy has evolved into a more comprehensive information-theoretic framework that extends many concepts related to Shannon entropy. The existence of generalized mutual information, divergence measures, and entropy powers provides a theoretical foundation for research into Sharma-Mittal-based learning algorithms and information-driven feature evaluation approaches.

4.4. Applications of Sharma-Mittal Entropy Across Scientific Domains

The theoretical flexibility of the Sharma-Mittal entropy has enabled its use in various scientific fields. Although it was initially developed in the context of generalized information theory and statistical physics, it has increasingly been applied in areas where uncertainty, distributional complexity, and nonclassical information structures are significant.
In machine learning, Koltcov et al. proposed a framework based on Sharma-Mittal entropy to evaluate topic modeling performance [48]. Their work addressed two fundamental problems in probabilistic topic modeling: determining the optimal number of topics and selecting appropriate hyperparameters. The findings demonstrated that Sharma-Mittal entropy can simultaneously reflect both model quality and semantic stability, and that it offers a more comprehensive evaluation metric compared to traditional perplexity-based assessments.
Another application can be found in the field of engineering design. Ahmed et al. used Sharma-Mittal entropy to measure design diversity and assess the extent to which the design space was explored during the concept generation process [49]. The authors proposed a family of Sharma-Mittal-based design diversity metrics and introduced the Herfindahl-Hirschman Index for Design (HHID) metric. Experimental results demonstrated that the HHID is more consistent with human evaluations and exhibits higher sensitivity to changes in design variety.
Sharma-Mittal entropy has also been used in decision tree and ensemble learning methods. Ignatenko et al. investigated information gain criteria based on Rényi, Tsallis, and Sharma-Mittal entropies in random forest algorithms for classification and regression tasks [50]. By replacing the classical Shannon-based split criterion with generalized entropy-based information gain formulations, they improved prediction performance on six benchmark datasets. Among the entropy measures tested, the Sharma-Mittal information gain frequently yielded the best results.
This result is particularly important in terms of feature selection, as information gain is closely related to the assessment of variable importance in tree-based learning. Although Ignatenko et al. focused on random forest architecture rather than a direct, independent feature selection method, their findings provide direct evidence that Sharma-Mittal-based information gain can serve as a practical alternative to Shannon-based metrics in machine learning systems [50].
Applications can also be found in reliability theory and lifetime analysis. Mohamed and Sakr investigated the cumulative residual Sharma–Taneja–Mittal entropy and developed nonparametric estimation procedures for it [51]. Their work analyzed properties related to stochastic comparisons, hazard rate functions, mean residual life functions, equilibrium random variables, and record values [51]. Similarly, Sfetcu et al. proposed an alternative measure of cumulative residual Sharma–Taneja–Mittal entropy and examined its mathematical properties [52]. These studies indicate that Sharma-Mittal-derived entropy measures are applicable in stochastic modeling and reliability analysis, not only in information theory.
Beyond statistics and machine learning, the Sharma-Mittal entropy has also been studied in the fields of thermodynamics and cosmology. Abreu and Ananias Neto analyzed the generalized second law of thermodynamics for various non-Gaussian entropy formalisms, including the Sharma-Mittal entropy, in the context of visible horizon thermodynamics [53]. Their findings suggest that parameter constraints may be necessary to preserve thermodynamic consistency in extended entropy formulations [53]. Similarly, Naeem and Bibi developed a corrected Friedmann equation based on the Sharma-Mittal entropy, contributing to the interpretation of generalized entropy structures in cosmological models [54].
In general, the literature shows that the Sharma-Mittal entropy has evolved from a theoretical generalized entropy formulation into a multidisciplinary analytical framework. Application areas include topic modeling and random forest learning in machine learning; the measurement of design diversity in engineering; residual life and reliability analysis; and thermodynamic and cosmological modeling. However, despite this growing body of literature, applications to feature selection remain relatively limited; there is a notable research gap, particularly regarding studies involving continuous probability distributions and information measures based on density estimation [48,49,50,51,52,53,54].

4.5. Entropy-Based Feature Selection and Information Gain

Feature selection is one of the fundamental preprocessing steps in machine learning, pattern recognition, data mining, and statistical learning. Its primary purpose is to identify relevant variables while removing irrelevant, redundant, or noisy variables. This reduces dimensionality, decreases computational complexity, improves prediction performance, and enhances the model’s interpretability [1,2,3,4].
The increasing prevalence of high-dimensional datasets has heightened the importance of feature selection methods. In fields such as gene expression analysis, text classification, image processing, network security, and sensor- or web-based data, there may be a large number of variables, and a significant portion of these variables may be irrelevant or redundant. Including such variables in the model can reduce prediction accuracy, increase computational cost, and lead to overfitting [1,2,3,4].
Feature selection methods are generally classified as filter, wrapper, and embedded approaches. Filter methods evaluate features using statistical or data-driven criteria, independent of a specific learning algorithm. Wrapper methods evaluate subsets of features based on model performance, while embedded methods perform feature selection during model training. Among these approaches, filter methods are particularly important because they are computationally efficient, model-independent, and suitable for pre-selecting features in high-dimensional datasets [1,2,3,4].
Criteria based on information theory constitute an important group of suitability measures within filter-based feature selection. Measures based on information gain, mutual information, entropy reduction, and dissimilarity evaluate how much information a candidate feature provides about the target variable. Unlike criteria based solely on correlation, these methods rely on uncertainty and statistical dependence structures; therefore, they can also capture informative relationships that cannot be adequately represented by linear correlation [7,8,9,10,11,12,13].
Traditionally, information gain is derived from Shannon entropy and measures the reduction in uncertainty regarding the target variable after a predictor variable has been observed. This approach has been widely used in decision-tree-based learning, and feature splits have been selected based on entropy-based information gain criteria. However, Shannon entropy represents only a specific formulation of uncertainty. For this reason, generalized entropy families such as those proposed by Rényi, Tsallis, and Sharma-Mittal have recently been investigated as alternative foundations for information gain in machine learning [7,8,9,10,11,12,13,14,15,50].
Ignatenko et al. proposed information gain formulations based on the Rényi, Tsallis, and Sharma-Mittal entropies for random forest structures. Their findings showed that generalized entropy-based information gain metrics can improve prediction performance compared to classical Shannon-based approaches in both classification and regression tasks. In particular, the Sharma-Mittal information gain has demonstrated promising performance across multiple benchmark datasets and has shown that parametric entropy families can provide more flexible partitioning criteria for complex data structures [50].
These findings are significant in that they demonstrate that Sharma-Mittal information gain is practically applicable in machine learning algorithms. In previous applications, Sharma-Mittal entropy has mostly been used as a descriptive, evaluative, or theoretical measure in fields such as subject modeling, design diversity, reliability analysis, thermodynamics, and cosmology. In contrast, its use in the context of information gain represents a more direct integration of this entropy family into prediction-oriented learning systems [48,49,50,51,52,53,54].
However, existing Sharma-Mittal information gain approaches are largely limited to decision tree construction and random forest splitting procedures. Furthermore, entropy values depend on the selected Sharma-Mittal parameter combinations. Since different parameter settings can highlight different aspects of uncertainty, the resulting information gain scores may also vary depending on these choices. This parameter sensitivity remains a significant methodological challenge for generalized entropy-based feature evaluation [16,17,18,19,20,21,22,50].
Therefore, although the Sharma-Mittal information gain has shown promising predictive potential, there is still a need for methods that reduce dependence on fixed parameter choices and extend entropy-based feature evaluation to information measures based on continuous probability distributions and density estimation.

4.6. Continuous Entropy Estimation and Density-Based Approaches

When variables are continuous, estimating quantities based on entropy and information theory becomes more difficult. In discrete structures, entropy can be calculated from observed probability masses or their empirical estimates. In contrast, continuous entropy measures depend on probability density functions; therefore, estimating entropy-based quantities requires either an estimate of the density or a direct estimate of the integral functionals associated with the density.
For this reason, entropy estimation has become an important research topic in statistics, information theory, and machine learning. The primary goal is to develop reliable estimators that can directly approximate entropy-related functionals from observed data, particularly for continuous variables.
Källberg et al. examined statistical inference procedures for Rényi entropy functionals and related entropy-like measures for both discrete and continuous distributions. In their work, estimators based on overlapping or ε-close observations in independent samples were used, and asymptotic properties such as consistency and asymptotic normality were established. This study is significant in that it demonstrates that entropy-type quantities can be treated as integral functionals of probability densities and can be estimated directly from continuous observations [55].
Källberg and Seleznjev further developed this approach by examining entropy-type integral functionals of densities and proposed U-statistic estimators based on ε-near vector observations. They obtained results regarding consistency and asymptotic normality under conditions of weak integrability and smoothness. These findings demonstrate that quantities related to generalized entropy can be estimated for continuous probability distributions without discretizing the data [56].
Alternative approaches have also been proposed to avoid explicit density estimation. Sánchez-Giraldo et al. defined entropy measures from the data using infinitely divisible kernels and positive-definite matrices. This framework defines entropy-like functionals on normalized Gram matrices and extends this approach to kernel-based conditional entropy and mutual information. In this respect, it is important for learning problems involving continuous data representations [57].
In general, these studies indicate a shift from discretization-based approaches in entropy calculations toward continuous and data-driven entropy estimation frameworks. These approaches are particularly important for datasets with naturally continuous variables, as they avoid arbitrary binning and better preserve the data’s inherent distributional structure. Therefore, density-functional and kernel-based entropy estimation methods provide a useful theoretical foundation for developing entropy-based feature selection criteria for continuous variables.

4.7. Limitations of Discretization-Based Entropy Methods

Although entropy-based feature selection methods are widely used, many approaches based on classical information theory are naturally defined for discrete probability distributions. Therefore, when analyzing continuous variables, these variables are often converted into discrete categories before calculating entropy, mutual information, or information gain.
Discretization divides the range of values of a continuous variable into a finite number of intervals and represents the observations using these interval labels. While this process facilitates probability estimation, it introduces significant limitations. First, discretization may alter the original distributional structure of the variable. Fine-grained information in continuous observations can be lost when the number of intervals is small, whereas too many intervals may lead to sparse probability estimates and unstable entropy values.
Second, entropy estimates may become sensitive to the chosen binning strategy. Different choices regarding the number of bins, bin boundaries, or binning algorithms can yield different information gain values on the same dataset. This creates an additional source of variability that is not directly related to the intrinsic information content of the variables.
Third, discretization often requires subjective design decisions. Researchers must choose a partitioning scheme before entropy calculations, and these choices may affect feature rankings and reduce the reproducibility of entropy-based analyses.
For these reasons, continuous entropy estimation methods have sparked increased interest in the development of non-discretization information theory frameworks. Density-functional estimators based on ε-near observations and U-statistics allow for the direct estimation of entropy-related quantities from continuous distributions. Kernel-based approaches, on the other hand, offer another alternative that yields entropy-like measures directly from data representations without requiring explicit probability density estimates. Taken together, these developments support the development of entropy-based learning criteria that can operate on continuous variables without relying on discretization procedures [55,56,57,58,59,60].

4.8. Research Gap and Positioning of the Proposed Framework

The reviewed literature shows that the Sharma-Mittal entropy transforms into a flexible, generalized entropy framework that includes the Rényi, Tsallis, and Shannon entropies as limits or special cases. Furthermore, this framework has been extended to various information theory concepts, such as divergence-based formulations, mutual information, channel capacities, and entropy powers. In addition, the Sharma-Mittal entropy has been applied in diverse fields such as machine learning, engineering design, reliability analysis, thermodynamics, and cosmology [16,17,18,19,20,21,22,44,45,46,47,48,49,50,51,52,53,54].
Continuous entropy estimation has provided theoretical support for density-functional and data-driven entropy estimation methods. More recently, Sharma-Mittal information gain has been applied in random forest algorithms and has demonstrated its potential in prediction-oriented learning tasks [50,55,56,57,58,59,60].
Despite these developments, the integration of Sharma-Mittal entropy, generalized information gain, and continuous entropy estimation within a unified feature selection framework remains limited. In particular, there is no systematic approach in the literature that jointly addresses density-based entropy estimation, non-discretized computation, conditional Sharma-Mittal entropy, Sharma-Mittal-based information gain, and parameter space integration for the purpose of feature ranking on continuous datasets.
The positioning of the proposed framework relative to representative studies is summarized in Table 20.
Accordingly, the proposed Density-Estimated Parameter-Integrated Sharma–Mittal Entropy Feature Selection (DE-PISME-FS) framework aims to bridge the gap between generalized entropy theory, continuous entropy estimation, and feature selection. By integrating Sharma–Mittal-based information measures across the parameter space and operating directly on continuous density estimates, the framework provides parameter-robust and non-discretized feature suitability scores for continuous datasets.

4.9. A Discussion of Inferences Drawn from Experimental Data

A joint evaluation of the six datasets shows that the relationship between the parameter-integrated Sharma–Mittal framework and the reference methods varies depending on the data structure. In the Airfoil Self-Noise and BodyFat datasets, the proposed methods distinguished the fundamental physical-acoustic drivers and anthropometric body composition drivers, respectively, with near-perfect accuracy. In the AirQualityUCI and meteorology datasets, despite the high dimensionality and interdependence of sensor/atmospheric variables, the proposed ranking showed high agreement with correlation- and information-based filters. The Concrete and WineQualityWhite datasets, on the other hand, represent boundary cases where inter-method agreement decreases due to more heterogeneous, interactive, and scale-dependent relationships.
Table 21 summarizes the Spearman’s rho and Kendall’s tau values for the proposed ranking and each reference method, broken down by dataset. When the simple arithmetic mean of the six datasets is calculated, mutual information (rho = 0.903) emerges as the reference metric closest to the proposed methods. This is followed by absolute Pearson correlation (rho = 0.884), Shannon information gain (rho = 0.883), and absolute Spearman rank correlation (rho = 0.874). For random forest variable importance, an average rho of 0.538 and an average tau of 0.432 were found. This pattern is closer to the model-independent filtering methods of the proposed framework; it reflects the variable importance of a multivariate, partition-based, and interaction-sensitive learning algorithm.
This study also clearly demonstrates the relationship among the three proposed indices. While PICSME operates based on the volume of conditional uncertainty, PIGSME calculates the difference between the same information structure and the target entropy surface, and NIGSME normalizes this difference by the total volume of the target. The fact that the three scores produce the same ranking across all datasets in the current experiments indicates that, in terms of the ranking function, the proposed framework represents a single fundamental piece of information at different scales. This is methodologically consistent; however, cross-dataset calibration, variable grouping, or multi-objective selection scenarios that could differentiate these three forms should be investigated in the future. Otherwise, the claim that the three indices are distinct feature selectors would be overly strong.
Caution is also warranted when interpreting the results of the Mann–Whitney U test. The p-values summarized in Table 22 are influenced by the fact that raw scores are generated on different numerical scales. For example, although p-values in the Airfoil Self-Noise and AirQualityUCI datasets mostly appear significant, the rank correlations are above 0.90. Similarly, in the Concrete dataset, although the Shannon information gain yields a p-value of 0.102, the rank agreement is only moderate. Therefore, p-values should be viewed as a secondary finding describing the differences in the score distributions of the methods; primary conclusions regarding the similarity or difference in rankings across methods should be based on rho, tau, and pairwise ranking plots.
The strongest aspect of the proposed framework is its ability to integrate the behavior within a selected region of the generalized entropy family without relying on a single parameter pair. In contrast, the negative PIGSME and NIGSME values observed in WineQualityWhite indicate that the absolute scale of continuous density-based entropy calculations is not sufficiently calibrated for a direct interpretation of information gain. Therefore, follow-up studies should thoroughly test the robustness of the score and ranking under different density estimators, different alpha–beta ranges, numerical integration resolutions, and bootstrap-based uncertainty intervals. Additionally, the practical value of the ranking during the filtering stage should be assessed by evaluating the error and explanatory performance of independent prediction models built using the selected top k variables under nested cross-validation.

5. Conclusions

This study proposed a volume-based, parameter-integrated Sharma–Mittal entropy framework for feature selection in continuous regression datasets. The main motivation was to avoid representing the relationship between a candidate feature and the target variable through a single fixed entropy parameter or through a discretized form of continuous data. Instead, the proposed approach evaluates the Sharma–Mittal entropy surface over a two-dimensional alpha–beta parameter domain and transforms this surface into feature relevance scores through numerical integration. In this respect, the study contributes to entropy-based feature selection by shifting the focus from pointwise entropy evaluation to a broader volumetric interpretation of uncertainty and information gain.
Within this framework, three related indices were defined: Parameter-Integrated Conditional Sharma–Mittal Entropy (PICSME), Parameter-Integrated Sharma–Mittal Entropy Information Gain (PIGSME), and Normalized Parameter-Integrated Sharma–Mittal Entropy Information Gain (NIGSME). PICSME measures the remaining conditional uncertainty volume of the target variable after observing a candidate feature, whereas PIGSME measures the corresponding reduction in uncertainty across the same parameter domain. NIGSME expresses this reduction relative to the total marginal entropy volume of the target variable. Although these indices are expressed on different numerical scales, the empirical results showed that they produced the same feature rankings within each dataset, indicating that they represent complementary interpretations of the same underlying density-based uncertainty structure.
The proposed indices were evaluated on six regression datasets and compared with absolute Pearson correlation, absolute Spearman rank correlation, Shannon information gain, mutual information, and random forest variable importance. The results showed that the proposed Sharma–Mittal-based rankings generally exhibited high or very high agreement with classical filter-based and information-theoretic methods, particularly in the Airfoil Self-Noise, AirQualityUCI, BodyFat, and meteorology datasets. Across all datasets, the average Spearman rank correlations between the proposed ranking and absolute Pearson correlation, absolute Spearman correlation, Shannon information gain, and mutual information were 0.884, 0.874, 0.883, and 0.903, respectively. The agreement with random forest importance was more moderate, with an average Spearman correlation of 0.538, suggesting that the proposed filter-based indices and model-dependent importance measures capture different aspects of feature relevance.
Overall, the proposed framework offers a new contribution to entropy-based feature selection by combining continuous density estimation, Sharma–Mittal entropy, parameter-domain integration, and volumetric feature relevance scoring. The findings indicate that the method can recover feature importance structures that are largely consistent with established filter-based criteria while preserving a continuous, density-based, and discretization-independent formulation. At the same time, the differences in raw score scales and the occurrence of negative PIGSME or NIGSME values in some datasets indicate that further methodological research is needed. Future studies should examine bandwidth sensitivity, parameter-region selection, grid resolution, alternative density estimators, classification settings, high-dimensional datasets, and out-of-sample predictive validation to assess the practical contribution of the proposed framework to model performance, robustness, and interpretability.

6. Limitations

The methodological limitations of the study and the quality assurance measures implemented to mitigate these limitations are summarized in Figure 12 at both general and dataset-specific levels. The kernel density estimate was used only in the calculation of the three proposed indices, while Shannon information gain was obtained through adaptive quantile-based discretization of continuous variables. In the Gaussian kernel density estimation, the bandwidth was automatically determined using Scott’s rule; the search range for the alpha and beta parameters was kept constant at [0.05, 0.95] across all datasets; and numerical integration was performed on a parameter grid with a resolution of 60×60. To keep the computational cost of bivariate kernel density estimation under control for datasets with large sample sizes, deterministic subsampling with 5,000 observations was applied. The lower panel of Figure 12 also presents data processing details for each dataset, including sample size, number of variables, removal of outliers or invalid target records, exclusion of columns with no predictive power, and computation time. While these design choices enhance experimental comparability, it should be noted that absolute index values—and particularly the rankings of variables at the boundary—may be sensitive to a certain extent to bandwidth, parameter range, grid resolution, and subsample size [14,15,23,24,25,26,27,40,41,42,43].
The fact that the importance of random forest variables can vary depending on the tree splitting mechanism used with correlated predictors is one of the key factors explaining why this method diverges from the proposed filter-based indices in some datasets. In addition, the signed raw values of the Pearson and Spearman correlation coefficients were distinguished from the absolute coefficients used to generate the importance rankings. In the comparative analyses, the primary evidence regarding the consistency of rankings among the methods was obtained from Spearman’s rho, Kendall’s tau, and pairwise rank scatter plots. The Mann–Whitney U test was used to complementarily assess the differences in the raw score distributions among the methods; however, because the scales and descriptive ranges of the compared scores were not the same, the results obtained from this test should be interpreted in the context of the divergence in score distributions rather than as an absolute method superiority [14,15,23,24,25,26,27,40,41,42,43].

Author Contributions

Conceptualization, N.O.Ü., M.G. and D.Y.; methodology, N.O.Ü., M.G.; software, N.O.Ü., M.G.; validation, N.O.Ü. and M.G., D.Y.; formal analysis, N.O.Ü. and M.G.; investigation, N.O.Ü. and M.G.; resources, N.O.Ü. and M.G.; data curation, N.O.Ü. and M.G.; writing—original draft preparation, N.O.Ü. and M.G.; writing—review and editing, N.O.Ü. and M.G.; visualization, N.O.Ü. and M.G.; supervision, D.Y.; project administration, N.O.Ü. and M.G. The interpretation of the analytical findings N.O.Ü. and M.G. and D.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new primary data were created in this study. The datasets analyzed in this study are publicly available from their original repositories. The Airfoil Self-Noise, Air Quality, Concrete Compressive Strength, and Wine Quality datasets are available from the UCI Machine Learning Repository. The BodyFat dataset is publicly available from the StatLib/Journal of Statistics Education body fat dataset repository. The meteorological dataset was compiled from publicly available historical weather records obtained through the Open-Meteo Historical Weather API. The preprocessing steps applied to the datasets are described in the Materials and Methods section. The processed data files, source code, and generated feature-ranking results are available from the corresponding author upon reasonable request [28,29,30,31,32,33,34,35,36,37,38,39].

Acknowledgments

In this section, you can acknowledge any support given that is not covered by the author contribution or funding sections. This may include administrative and technical support, or donations in kind (e.g., materials used for experiments). Where GenAI has been used for purposes such as generating text, data, or graphics, or for study design, data collection, analysis, or interpretation of data, please add “During the preparation of this manuscript/study, the author(s) used [tool name, version information] for the purposes of [description of use]. The authors have reviewed and edited the output and take full responsibility for the content of this publication.”.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
DE-PISME-FS Density-Estimated Parameter-Integrated Sharma–Mittal Entropy Feature Selection
IG Information Gain
KDE Kernel Density Estimation
MI Mutual Information
NIGSME Normalized Parameter-Integrated Sharma–Mittal Entropy Information Gain
PICSME Parameter-Integrated Conditional Sharma–Mittal Entropy
PIGSME Parameter-Integrated Sharma–Mittal Entropy Information Gain
RF Random Forest
SME Sharma–Mittal Entropy

Appendix A

Appendix A.1. Rank-Equivalence Relationship Among PICSME, PIGSME, and NIGSME

Preprints 224148 i001
Note: In each dataset, the marginal entropy surface of the target variable is the same for all candidate variables. Therefore, PIGSME is a version of PICSME from which the constant target volume has been removed; NIGSME is a version of this information-gain volume scaled by a constant target volume. Under the applied parameter range and the density estimation method used, these transformations yielded the same variable ranking across all datasets. While this observation demonstrates that the three scores do not produce conflicting selections, it also suggests that, for the three metrics to provide distinct decision information in future studies, they need to be extended through cross-dataset standardization, target entropy uncertainty analysis, or the use of different objective functions.

References

  1. Guyon, I.; Elisseeff, A. An introduction to variable and feature selection. J. Mach. Learn. Res. 2003, 3, 1157–1182.
  2. Kumar, V.; Minz, S. Feature selection: A literature review. Smart Comput. Rev. 2014, 4, 211–229. [CrossRef]
  3. Venkatesh, B.; Anuradha, J. A review of feature selection and its methods. Cybern. Inf. Technol. 2019, 19, 3–26. [CrossRef]
  4. Bolón-Canedo, V.; Sánchez-Maroño, N.; Alonso-Betanzos, A. Recent advances and emerging challenges of feature selection in the context of big data. Knowl.-Based Syst. 2015, 86, 33–45. [CrossRef]
  5. Pearson, K. Notes on regression and inheritance in the case of two parents. Proc. R. Soc. Lond. 1895, 58, 240–242. [CrossRef]
  6. Spearman, C. The proof and measurement of association between two things. Am. J. Psychol. 1904, 15, 72–101. [CrossRef]
  7. Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423, 623–656.
  8. Cover, T.M.; Thomas, J.A. Elements of Information Theory, 2nd ed.; Wiley-Interscience: Hoboken, NJ, USA, 2006.
  9. Vergara, J.R.; Estévez, P.A. A review of feature selection methods based on mutual information. Neural Comput. Appl. 2014, 24, 175–186. [CrossRef]
  10. Peng, H.; Long, F.; Ding, C. Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy. IEEE Trans. Pattern Anal. Mach. Intell. 2005, 27, 1226–1238. [CrossRef]
  11. Battiti, R. Using mutual information for selecting features in supervised neural net learning. IEEE Trans. Neural Netw. 1994, 5, 537–550. [CrossRef]
  12. Brown, G.; Pocock, A.; Zhao, M.-J.; Luján, M. Conditional likelihood maximisation: A unifying framework for information theoretic feature selection. J. Mach. Learn. Res. 2012, 13, 27–66.
  13. Kraskov, A.; Stögbauer, H.; Grassberger, P. Estimating mutual information. Phys. Rev. E 2004, 69, 066138.
  14. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32.
  15. Louppe, G.; Wehenkel, L.; Sutera, A.; Geurts, P. Understanding variable importances in forests of randomized trees. In Advances in Neural Information Processing Systems 26; Curran Associates: Red Hook, NY, USA, 2013; pp. 431–439.
  16. Sharma, B.D.; Mittal, D.P. New non-additive measures of entropy for a discrete probability distribution. J. Math. Sci. 1975, 10, 28–40.
  17. Rényi, A. On measures of entropy and information. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability; University of California Press: Berkeley, CA, USA, 1961; Volume 1, pp. 547–561.
  18. Tsallis, C. Possible generalization of Boltzmann–Gibbs statistics. J. Stat. Phys. 1988, 52, 479–487. [CrossRef]
  19. Masi, M. A step beyond Tsallis and Rényi entropies. Phys. Lett. A 2005, 338, 217–224. [CrossRef]
  20. Beck, C. Generalised information and entropy measures in physics. Contemp. Phys. 2009, 50, 495–510. [CrossRef]
  21. Aktürk, E.; Bağcı, G.B.; Sever, R. Is Sharma–Mittal entropy really a step beyond Tsallis and Rényi entropies? arXiv 2007, arXiv:cond-mat/0703277.
  22. Ilić, V.M.; Korbel, J.; Gupta, S.; Scarfone, A.M. An overview of generalized entropic forms. Europhys. Lett. 2021, 133, 50005. [CrossRef]
  23. Silverman, B.W. Density Estimation for Statistics and Data Analysis; Chapman and Hall: London, UK, 1986.
  24. Scott, D.W. Multivariate Density Estimation: Theory, Practice, and Visualization, 2nd ed.; Wiley: Hoboken, NJ, USA, 2015.
  25. Wand, M.P.; Jones, M.C. Kernel Smoothing; Chapman and Hall: London, UK, 1995.
  26. Sheather, S.J.; Jones, M.C. A reliable data-based bandwidth selection method for kernel density estimation. J. R. Stat. Soc. Ser. B Methodol. 1991, 53, 683–690. [CrossRef]
  27. Turlach, B.A. Bandwidth Selection in Kernel Density Estimation: A Review; CORE and Institut de Statistique, Université Catholique de Louvain: Louvain-la-Neuve, Belgium, 1993.
  28. Brooks, T.F.; Pope, D.S.; Marcolini, M.A. Airfoil Self-Noise and Prediction; NASA Reference Publication 1218; National Aeronautics and Space Administration: Hampton, VA, USA, 1989.
  29. Dua, D.; Graff, C. UCI Machine Learning Repository: Airfoil Self-Noise Data Set; University of California, Irvine, School of Information and Computer Sciences: Irvine, CA, USA, 2019.
  30. De Vito, S.; Massera, E.; Piga, M.; Martinotto, L.; Di Francia, G. On field calibration of an electronic nose for benzene estimation in an urban pollution monitoring scenario. Sens. Actuators B Chem. 2008, 129, 750–757. [CrossRef]
  31. Dua, D.; Graff, C. UCI Machine Learning Repository: Air Quality Data Set; University of California, Irvine, School of Information and Computer Sciences: Irvine, CA, USA, 2019.
  32. Johnson, R.W. Fitting percentage of body fat to simple body measurements. J. Stat. Educ. 1996, 4. [CrossRef]
  33. Allen, R.G.; Pereira, L.S.; Raes, D.; Smith, M. Crop Evapotranspiration: Guidelines for Computing Crop Water Requirements; FAO Irrigation and Drainage Paper 56; Food and Agriculture Organization of the United Nations: Rome, Italy, 1998.
  34. Open-Meteo. Historical Weather API. Available online: https://open-meteo.com/en/docs/historical-weather-api (accessed on 10 July 2026).
  35. Hersbach, H.; Bell, B.; Berrisford, P.; Hirahara, S.; Horányi, A.; Muñoz-Sabater, J.; Nicolas, J.; Peubey, C.; Radu, R.; Schepers, D.; et al. The ERA5 global reanalysis. Q. J. R. Meteorol. Soc. 2020, 146, 1999–2049. [CrossRef]
  36. Yeh, I.-C. Modeling of strength of high-performance concrete using artificial neural networks. Cem. Concr. Res. 1998, 28, 1797–1808. [CrossRef]
  37. Dua, D.; Graff, C. UCI Machine Learning Repository: Concrete Compressive Strength Data Set; University of California, Irvine, School of Information and Computer Sciences: Irvine, CA, USA, 2019.
  38. Cortez, P.; Cerdeira, A.; Almeida, F.; Matos, T.; Reis, J. Modeling wine preferences by data mining from physicochemical properties. Decis. Support Syst. 2009, 47, 547–553. [CrossRef]
  39. Dua, D.; Graff, C. UCI Machine Learning Repository: Wine Quality Data Set; University of California, Irvine, School of Information and Computer Sciences: Irvine, CA, USA, 2019.
  40. Mann, H.B.; Whitney, D.R. On a test of whether one of two random variables is stochastically larger than the other. Ann. Math. Stat. 1947, 18, 50–60.
  41. Kendall, M.G. A new measure of rank correlation. Biometrika 1938, 30, 81–93. [CrossRef]
  42. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830.
  43. Virtanen, P.; Gommers, R.; Oliphant, T.E.; Haberland, M.; Reddy, T.; Cournapeau, D.; Burovski, E.; Peterson, P.; Weckesser, W.; Bright, J.; et al. SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nat. Methods 2020, 17, 261–272. [CrossRef]
  44. Bodnarchuk, I.; Mishura, Y.; Ralchenko, K. Properties of the Shannon, Rényi and other entropies: Dependence in parameters, robustness in distributions and extremes. arXiv 2024, arXiv:2411.15817.
  45. Ilić, V.M.; Djordjević, I.B. On the α-q-mutual information and the α-q-capacities. Entropy 2021, 23, 702. [CrossRef]
  46. Nielsen, F.; Nock, R. On Rényi and Tsallis entropies and divergences for exponential families. arXiv 2011, arXiv:1105.3259.
  47. Bukal, M. The concavity of generalized entropy powers. IEEE Trans. Inf. Theory 2022, 68, 7054–7059. [CrossRef]
  48. Koltcov, S.; Ignatenko, V.; Koltsova, O. Estimating topic modeling performance with Sharma–Mittal entropy. Entropy 2019, 21, 660. [CrossRef]
  49. Ahmed, F.; Ramachandran, S.K.; Fuge, M.; Hunter, S.; Miller, S. Design variety measurement using Sharma–Mittal entropy. J. Mech. Des. 2021, 143, 061702. [CrossRef]
  50. Ignatenko, V.; Surkov, A.; Koltcov, S. Random forests with parametric entropy-based information gains for classification and regression problems. PeerJ Comput. Sci. 2024, 10, e1775. [CrossRef]
  51. Mohamed, M.S.; Sakr, H.H. Properties of residual cumulative Sharma–Taneja–Mittal model and its extensions in reliability theory with applications to human health analysis and mixed coherent mechanisms. Entropy 2026, 28, 32. [CrossRef]
  52. Sfetcu, R.C.; Robe-Voinea, E.G.; Şerban, F. An alternate measure of the cumulative residual Sharma-Taneja-Mittal entropy. An. Şt. Univ. Ovidius Constanţa Ser. Mat. 2025, 33, 125–142. [CrossRef]
  53. Abreu, E.M.C.; Ananias Neto, J. Statistical approaches on the apparent horizon entropy and the generalized second law of thermodynamics. Phys. Lett. B 2022, 824, 136803. [CrossRef]
  54. Naeem, M.; Bibi, A. Correction to the Friedmann equation with Sharma–Mittal entropy: A new perspective on cosmology. Ann. Phys. 2024, 462, 169618. [CrossRef]
  55. Källberg, D.; Leonenko, N.; Seleznjev, O. Statistical inference for Rényi entropy functionals. In Conceptual Modelling and Its Theoretical Foundations: Essays Dedicated to Bernhard Thalheim on the Occasion of His 60th Birthday; Hameurlain, A., Küng, J., Wagner, R., Eds.; Springer: Berlin/Heidelberg, Germany, 2012; pp. 36–51.
  56. Källberg, D.; Seleznjev, O. Estimation of entropy-type integral functionals. arXiv 2012, arXiv:1209.2544.
  57. Sánchez Giraldo, L.G.; Rao, M.; Principe, J.C. Measures of entropy from data using infinitely divisible kernels. IEEE Trans. Inf. Theory 2015, 61, 535–548. [CrossRef]
  58. Paninski, L. Estimation of entropy and mutual information. Neural Comput. 2003, 15, 1191–1253. [CrossRef]
  59. Kozachenko, L.F.; Leonenko, N.N. Sample estimate of the entropy of a random vector. Probl. Inf. Transm. 1987, 23, 95–101.
  60. Leonenko, N.; Pronzato, L.; Savani, V. A class of Rényi information estimators for multidimensional densities. Ann. Stat. 2008, 36, 2153–2182. [CrossRef]
Figure 2. Parameter-Integrated Conditional Sharma–Mittal Entropy: volumetric interpretation of the conditional entropy surface in the alpha–beta plane, obtained using the selected density estimation approach.
Figure 2. Parameter-Integrated Conditional Sharma–Mittal Entropy: volumetric interpretation of the conditional entropy surface in the alpha–beta plane, obtained using the selected density estimation approach.
Preprints 224148 g002
Figure 5. Experimental flow of the six datasets: the relationship between the dataset, the independent variables, and the dependent target variable.
Figure 5. Experimental flow of the six datasets: the relationship between the dataset, the independent variables, and the dependent target variable.
Preprints 224148 g005
Figure 8. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 8. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g008
Figure 10. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 10. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g010
Figure 11. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Figure 11. Inter-method comparison in the Airfoil Self-Noise dataset: (a) Spearman’s rank correlation (ρ) and Kendall’s tau (τ) matrix, (b) pairwise rank scatter plots, and (c) p-values from the Mann–Whitney U test for raw score distributions.
Preprints 224148 g011
Figure 12. Methodological limitations, assurance measures, and data processing design of the study. Note: (a) Methodological limitations common to all six datasets and the standardized assurance measures implemented to mitigate these limitations. (b) Information specific to each dataset regarding sample size, number of independent variables, data cleaning, removal of invalid target observations, exclusion of non-predictive variables, deterministic subsampling, and computation time.
Figure 12. Methodological limitations, assurance measures, and data processing design of the study. Note: (a) Methodological limitations common to all six datasets and the standardized assurance measures implemented to mitigate these limitations. (b) Information specific to each dataset regarding sample size, number of independent variables, data cleaning, removal of invalid target observations, exclusion of non-predictive variables, deterministic subsampling, and computation time.
Preprints 224148 g012
Table 1. Analytical profile of the datasets and the data-processing steps applied.
Table 1. Analytical profile of the datasets and the data-processing steps applied.
Dataset Data Size (N) Feature Size Target Variable Application Area Data-Processing Summary
Airfoil Self-Noise 1.503 5 SoundPressureLevel_dB Acoustic/aerodynamic Five continuous aerodynamic variables and the target variable SoundPressureLevel_dB were used directly in the analysis.
AirQualityUCI 7.674 11 CO(GT) Weather Quality Observations with an invalid target variable were removed; the date, time, and NMHC(GT) columns were excluded from the predictor set. During the density estimation phase, 5,000 observations were used with deterministic subsampling.
BodyFat 249 13 BodyFat Body Composition The analysis file, which consists of body composition measurements, included 13 continuous explanatory variables.
Meteorology 9.497 45 et0_fao_evapotranspiration Hydrometeorology The variables “date,” “sunrise,” “sunset,” and “weather_code” have been excluded from the model. For the density estimate, a deterministic subsampling of 5,000 observations was applied.
Concrete 1.030 8 Concrete compressive strength (MPa) Materials Engineering Eight mixture/age variables were used as explanatory variables for the compressive strength target.
WineQualityWhite 4.898 11 quality Food/wine quality Physicochemical parameters specific to white wines were used; all observations were directly incorporated into the density-based calculation.
Note: From the initial 9,357 observations in the AirQualityUCI dataset, records with invalid target values were removed, leaving 7,674 observations. In both the AirQualityUCI and meteorological datasets, a sub-sampling of 5,000 observations was applied only during the KDE phase to limit the cost of density estimation.
Table 17. Raw importance scores of the reference methods in the WineQualityWhite dataset.
Table 17. Raw importance scores of the reference methods in the WineQualityWhite dataset.
Variables Comparison Methods’ Values
Pearson
Correlation
Spearman
Correlation
Shannon
Information
Gain
Mutual
Information
Random
Forest
Importance
fixed acidity -0,1137 -0,0845 0,0197 0,0237 0,0615
volatile acidity -0,1947 -0,1966 0,0401 0,0505 0,1253
citric acid -0,0092 0,0183 0,0460 0,0541 0,0580
residual sugar -0,0976 -0,0821 0,0346 0,0805 0,0692
chlorides -0,2099 -0,3145 0,0651 0,0571 0,0631
free sulfur dioxide 0,0082 0,0237 0,0338 0,0427 0,1161
total sulfur dioxide -0,1747 -0,1967 0,0432 0,0879 0,0695
density -0,3071 -0,3484 0,0907 0,1535 0,0622
pH 0,0994 0,1094 0,0184 0,0288 0,0698
sulphates 0,0537 0,0333 0,0223 0,0238 0,0615
alcohol 0,4356 0,4404 0,1393 0,1421 0,2436
Note: The raw coefficients marked in the Pearson and Spearman rows are reported. The absolute values of both correlation coefficients were used in the rank analysis.
Table 18. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the WineQualityWhite dataset.
Table 18. Raw scores of the proposed methods based on parameter-integrated Sharma–Mittal entropy in the WineQualityWhite dataset.
Variable Proposed Methods’ Values
PICSME PIGSME NIGSME
fixed acidity 1,7828 -0,2455 -0,1597
volatile acidity 1,7570 -0,2197 -0,1429
citric acid 1,7629 -0,2256 -0,1468
residual sugar 1,7785 -0,2413 -0,1569
chlorides 1,7407 -0,2034 -0,1323
free sulfur dioxide 1,7563 -0,2190 -0,1425
total sulfur dioxide 1,7566 -0,2193 -0,1427
density 1,7163 -0,1790 -0,1165
pH 1,7875 -0,2502 -0,1628
sulphates 1,7852 -0,2479 -0,1613
alcohol 1,6540 -0,1167 -0,0759
Note: For PICSME, a lower score indicates a greater reduction in conditional uncertainty; for PIGSME and NIGSME, a higher score indicates a greater relative gain in information. The scales are interpreted within the dataset.
Table 20. Comparative positioning of the proposed framework relative to representative studies.
Table 20. Comparative positioning of the proposed framework relative to representative studies.
Study Sharma-Mittal
Entropy
Information Gain Continuous Density Estimation Parameter
Integration
Feature
Selection
Koltcov et al. (2019)
Ahmed et al. (2021)
Ilić & Djordjević (2021) Mutual
Information
Ignatenko et al. (2024)
Källberg et al. (2011)
Källberg & Seleznjev (2013)
Sánchez-Giraldo et al. (2014)
Proposed DE-PISME-FS Framework
Table 21. Summary of the data-set-based agreement between the rankings of the proposed methods and those of the reference methods (Spearman’s rho / Kendall’s tau).
Table 21. Summary of the data-set-based agreement between the rankings of the proposed methods and those of the reference methods (Spearman’s rho / Kendall’s tau).
Dataset Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information Gain
Mutual
Information
Random
Forest
Airfoil Self-Noise 1,000 / 1,000 1,000 / 1,000 0,900 / 0,800 1,000 / 1,000 0,900 / 0,800
AirQualityUCI 0,991 / 0,964 0,961 / 0,881 0,989 / 0,954 0,991 / 0,964 0,773 / 0,636
BodyFat 1,000 / 1,000 0,973 / 0,897 0,945 / 0,846 0,945 / 0,846 -0,044 / -0,026
Meteorology 0,950 / 0,840 0,956 / 0,854 0,970 / 0,889 0,958 / 0,872 0,497 / 0,411
Concrete 0,762 / 0,571 0,738 / 0,500 0,619 / 0,429 0,762 / 0,571 0,738 / 0,571
WineQualityWhite 0,600 / 0,455 0,618 / 0,491 0,873 / 0,745 0,764 / 0,527 0,364 / 0,200
Simple arithmetic mean 0,884 / 0,805 0,874 / 0,771 0,883 / 0,777 0,903 / 0,797 0,538 / 0,432
Note: Each cell is presented in the form of rho/tau. The average values were calculated using equal weighting based on the number of data sets; no weighting was applied based on sample sizes.
Table 22. Comparison of the raw score distributions of the proposed methods and the reference methods using the Mann–Whitney U test.
Table 22. Comparison of the raw score distributions of the proposed methods and the reference methods using the Mann–Whitney U test.
Dataset Absolute
Pearson
Correlation
Absolute
Spearman
Correlation
Shannon
Information Gain
Mutual
Information
Random
Forest
Airfoil Self-Noise < 0,001 < 0,001 0,037 < 0,001 0,037
AirQualityUCI < 0,001 < 0,001 < 0,001 < 0,001 0,005
BodyFat < 0,001 < 0,001 < 0,001 < 0,001 0,887
Meteorology < 0,001 < 0,001 < 0,001 < 0,001 < 0,001
Concrete 0,028 0,037 0,102 0,028 0,037
WineQualityWhite 0,051 0,043 < 0,001 0,006 0,272
Note: The p-values were obtained from a Mann–Whitney U test comparing the raw scores of the proposed triplet score group with those of the reference method for the relevant dataset. Because there is a scale difference between the methods, the results should not be interpreted as a test of ordinal consistency.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings