6. Evaluating Clustering Algorithms for AI Adoption Analysis in the EU: A Multimetric Approach
To assess relative performance of various clustering techniques in capturing large European Union firm artificial intelligence (AI) adoption patterns, standardized evaluation measures were employed to assess six different algorithms, including Density-Based, Fuzzy C-Means, Hierarchical, Model-Based, Neighborhood-Based, and Random Forest clustering. These measures—ranging from explanatory power (R²) to statistical efficiency (AIC, BIC), from measures of geometric cohesion (Silhouette Score, Dunn Index) to cluster structure (Entropy, Maximum Diameter, Calinski-Harabasz Index)—allow the relative merits and demerits of each algorithm to be assessed in detail. The aim of this is to identify the algorithm that best achieves balance between model fit, interpretability, and the geometric integrity of the resulting clusters and thereby offers the best of all possible instruments to analyze AI diffusion along macroeconomic patterns (
Table 6).
Comparison of six clustering techniques—Density-Based Clustering, Fuzzy C-Means Clustering, Hierarchical Clustering, Model-Based Clustering, Neighborhood-Based Clustering, and Random Forest Clustering—has different performance profiles on various standardized evaluation measures. These measures are R², AIC, BIC, Silhouette Score, Maximum Diameter, Minimum Separation, Pearson’s Gamma, Dunn Index, Entropy, and the Calinski-Harabasz Index, all standardized to between 0 and 1 to enable direct comparison. The objective of the analysis here is to identify the algorithm with the best balance between statistical quality and geometrical clustering quality. Beginning with R², which is the ratio of the amount of the variance in the data that is explained by the clustering model, to the total amount of variance in the data, we have the best possible score by Neighborhood-Based Clustering, reflecting excellent explanation of data. Hierarchical Clustering is next with the best possible score, followed by moderate scores from Random Forest Clustering. Lower in the ranks are Model-Based and Fuzzy C-Means, and lowest in the ranks is Density-Based Clustering, implying failure to explain the data’s variance structure. These are in line with the observations by Sarmas, Fragkiadaki, and Marinakis (2024), who highlighted the superiority of ensemble and neighborhood-aware clustering to capturing subtle consumer behavior to be used in demand response in transport systems. With regards to criteria in selecting models such as the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), measuring both the goodness of fit and the complexity of models, Hierarchical Clustering and Neighborhood-Based Clustering get the best possible scores, implying optimal performance. In turn, Density-Based Clustering and Fuzzy C-Means get the worst possible scores, implying low efficiency of the models and possible overfit or lack of parsimony. When comparing the Silhouette Score, which is how similar an object is to its own cluster in contrast to other clusters, we have the best possible score by Density-Based Clustering, implying forming well-separated and well-defined clusters. Hierarchical Clustering is next best, followed by moderate cohesion by Neighborhood-Based Clustering. Fuzzy C-Means, Model-Based, and Random Forest get poor scores in this dimension, meaning that their cluster boundaries are not well defined. These findings are in line with general trends found in comparative clustering research such as that of Thamrin and Wijayanto (2021), who illustrated different kinds of performance trade-off between soft and hard clustering models based on the data structure and population homogeneity.
Looking in particular at Maximum Diameter, which measures the greatest intra-cluster distance and ideally would be minimized, Model-Based Clusters, Density-Based Clusters, and Neighborhood-Based Clusters exhibit the tightest clusters with the lowest diameters. Conversely, Fuzzy C-Means measures the largest value, reflecting large and perhaps poor clusters. This trend is in line with the application of clustering observed in Elkahlout and Elkahlout (2024), wherein spatial clustering of groundwater wells necessitated diligent consideration of intra-cluster variability to obtain meaningful geographic boundaries. Hierarchical Clusters and Random Forest Clusters are in the middle of the spectrum. Minimum Separation, which is the measure of the minimum distance between cluster centers and optimally would be large, positions Density-Based Clusters on top, with excellent cluster separation. Hierarchical Clusters perform in the middle, and while Neighborhood-Based Clusters scores low, this is perhaps suggestive of overlapping or close clusters. Fuzzy C-Means ranks lowest, further evidence of the former's poor intra- and extra-class definability. Pearson’s Gamma, reflecting data distance correlations with cluster assignments, places Density-Based Clusters in top position, with Hierarchical Clusters and Neighborhood-Based Clusters performing reasonably well. Random Forest and Fuzzy C-Means are lowest on this list, and imply poor spatial correspondence. Dunn Index, which integrates both the cluster compactness and separation and serves as a strong measure of overall cluster quality, yet again positions Density-Based Clusters on top, with Hierarchical Clusters and Neighborhood-Based Clusters immediately in second and third positions. This measure is in keeping with observations from Silhouette, Separation, and Pearson’s Gamma. Fuzzy C-Means and Model-Based Clusters lag behind, reflecting poor intra-class compactness and inter-class distinctness. This is consistent with observations by Da Silva, Melton, and Wunsch (2020), who highlighted the importance of dynamic and incremental measures of validity to rank hard partitioning techniques, particularly where clusters undergo changes or update in the online setting. Entropy, reflecting here the degree of disorder or randomness in cluster assignments and optimally would be low, further penalizes Fuzzy C-Means, which measures the greatest value, and suggests overlapping and noisy clusters. In contrast, Density-Based Clusters, Neighborhood-Based Clusters, and Model-Based Clusters obtain the lowest entropies and more ordered cluster assignments. These findings confirm the warning uttered by Gagolewski, Bartoszuk, and Cena (2021) that cluster validity indexes can differ in significant ways between and among different algorithms and are best interpreted in their specific contexts and not comparatively in isolation. Lastly, the Calinski-Harabasz Index, the variance ratio measure that penalizes low between-cluster and within-cluster dispersion, ranks Hierarchical Clustering in first position, and Fuzzy C-Means next. This is partially at odds with the rest of the measures but suggests that Hierarchical Clustering works exceptionally well if viewed from a variance-based dimension. On this measure, the lowest rank is occupied by Density-Based Clustering and it is possible to speculate that although spatially well-defined, such clusters will not meet traditional expectations of statistical variance—a difference expressing the model-agnostic findings highlighted by Sarmas, Fragkiadaki, and Marinakis (2024) in their research on explainable ensemble clustering on the modeling of complex systems.
Together, the results demonstrate that no algorithm excels the rest on all measures but that different patterns are clear. Density-Based Clustering behaves well in clustering quality in terms of geometry with leading scores in measures of structure, separation, and coherence such as Silhouette Score, Dunn Index, Pearson’s Gamma, and Minimum Separation. These findings are in line with those of Auliani, Novita, and Afdal (2024), who demonstrated the superiority of the former in creating well-separate clusters in car sales data, especially in the data with noise. But the weak behavior of Density-Based Clustering in statistical measures such as R², AIC, BIC, and Calinski-Harabasz Index identifies it as lacking in explanation and statistical efficiency in pursuit of robust model-based inference. Hierarchical Clustering, on the other hand, has top-performing overall behavior with top ratings in R² and Calinski-Harabasz coupled with decent performance in structural measures such as Dunn Index and Pearson’s Gamma. This is in line with findings by Azkeskin and Aladağ (2025), who viewed hierarchical clustering to be effective in identifying regional energy patterns with statistical cohesiveness. Hierarchical Clustering is thus found to be a balanced algorithm with the potential to produce statistically sound and geometrical meaningful clusters. Neighborhood-Based Clustering has the best statistical profile with leading results in R², AIC, and BIC and decent results in diameter, entropy, and compactness. It does not have the lead in measures of geometrical separation, but is strong enough on all sides to be a serious runner. The balanced statistical foundation and decent structure of the models provide it with the potential to bridge the gap between interpretability and performance. Random Forest Clustering is found in the middle ground with decent behavior in all sides except in excelling in any specific area. Similarly, Model-Based Clustering has mixed results with some decent statistical behavior but poor geometrical cluster properties—a trend observed by Ambarsari et al. (2023) in comparing fuzzy versus probabilistic clustering methods in population welfare segmentation. Fuzzy C-Means Clustering, on the other hand, performs mixed results on nearly all measures, especially in terms of cohesion, separation, entropy, and statistical fit. This is consistent with findings by Sarmas, Fragkiadaki, and Marinakis (2024), who demonstrated fuzzy clustering methods to be lacking in situations where clear delineation and strong interpretability is needed. Considering all of these findings collectively as a whole, Neighborhood-Based Clustering is the best performer overall. Its balance of strong statistical fit, computational efficiency, simple cluster shape, and moderate but sufficient structural preservation makes it the best overall and most consistent algorithm to use to cluster in this context. While Density-Based Clustering generates well-separate and spatially coherent clusters, the lack of statistical stability decreases the utility of this algorithm in contexts that require both interpretability and inferability. Hierarchical Clustering is still another strong option, particularly under the application of the use of variance-based measures or hybrid approaches. Ultimately, whichever algorithm to employ would best be dictated by the specific aims of the analysis—whether statistical explanation, geometric simplicity, and/or implementation ease is of utmost importance. But with the application of the normalization measures here, Neighborhood-Based Clustering provides the best overall and strongest balance of performance in all of the measures of evaluation (
Table 7).
Clustering outcomes here, based on macroeconomic indicators, attempt to provide explanations of patterns of adoption of artificial intelligence (AI) technologies—reflected in ALOAI—among large EU companies in different industrial and country contexts. Such explanation is based on standardized macroeconomic indicators such as current health expenditures (HEAL), domestic credit to the non-financial sector (DCPS), exports (EXGS), GDP per capita (GDPC), gross fixed capital formation (GFCF), inflation (INFD), and trade openness (TRAD). Derived clusters of seven are quite dissimilar in size, within-cluster homogeneity/similarity, and silhouette score, reflecting great heterogeneity in how macroeconomic environments are related to adoption of AI among European countries. Cluster 5 is the largest (n = 58), and with moderate within-cluster heterogeneity proportion (0.384), reasonably large within-cluster sum of squares (124.42), and moderate silhouette (0.346). Rather strikingly, it has negative ALOAI center of –0.837, reflecting below-average use of AI despite containing the largest number of countries. Its economic profile of uniformly negative or near-zero on salient variables such as GDP per capita (–0.821), domestic credit (–0.826), and trade openness (–0.074) reflects countries that are perhaps economically constrained, locked into traditional systems, or less integrated with the world, and lag behind on spread of AI. This is consistent with Popović, Todorović, and Milijić (2024), who illustrate how adoption of AI is positively linked with circular use of material and innovation-driven economies—factors which Cluster 5 countries could be lacking. Furthermore, Brey and van der Marel (2024) suggest the strategic role of human capital in enabling the integration of AI, and that Cluster 5 underperformance can also be traced to educational infrastructure and digital preparedness deficits. On the opposite side, Cluster 2, among better-defined clusters (n = 25, var. exp. 0.187), is characterized by very-high ALOAI center of 1.407, reflecting above-average enterprise level use of AI. Its macroeconomic profile of strong GDP per capita, moderate management of inflation, and healthy and favorable levels of both domestic and external credit and trade reflect dynamic economies. Such countries are also bound to be privileged with more developed financial and strategic digital systems and more exposure to international markets and innovation systems. Czeczeli et al. (2024) note that such countries are more likely to be resistant to inflation and policy flexible—two properties that foster economic stability and support investment in AI (
Figure 2).
Their economic profile is comprised of favorable values on nearly all of the indicators, with special characteristics including strong home credit (1.512), moderate exports (–0.453), and respectable GDP per capita (0.907). Such a cluster is expected to be comprised of developed, mid-sized EU economics with stable access to capital and balanced external trade profiles that support moderate to high AI adoption. Such findings are supported by Bosna et al. (2024), who used clustering and ANFIS analysis to reveal macroeconomic balance to be the primary determinant of supporting growth and innovation following eurozone membership. Cluster 6 is small (n=24) but shares comparable structural characteristics with Cluster 2 with the exception of low ALOAI centre (0.379), meaning that, although macroeconomic fundamentals are reasonably favorable such as health spending (0.39), trade openness (0.837), and exports (0.857), other variables such as labor market rigidity, policy gaps, or low industrial digital maturity are likely to curb AI diffusion. These structural barriers are likely to be symptomatic of institutional preparedness challenges as discovered in the regression EU inflation study of Czeczeli et al. (2024), where clustering untangled different preparedness profiles to economic shocks. Cluster 3 is the largest low-ALOA cluster with large silhouette score (n = 35, silhouette = 0.334, ALOAI = 0.018). It has marginally positive health and credit indicators but negative exports (–0.741), trade openness (–0.776), and GDP per capita (–0.095), signifying internal economic development with minimal external market integration. Such findings are in consonance with observations by Arora et al. (2024), who showed by correlation and clustering that macroeconomic groupings of variables tend to divide along lines of internal vs. external orientation with implications on preparedness to innovate. Cluster 4 is small (n = 6) but is different in having high silhouette score (0.894) and above-mean ALOAI (0.693). It is marked by exceptionally strong exports (3.619), trade openness (3.579), and very high GDP per capita (2.933) but poor health spending (–1.533) and GFCF (–1.215). This is indicative of a group of high-income, export-dependent economies where dynamism of the private sector is capable of compensating poor public investment and infrastructure in health. Such configurations are representative of those influenced by industrial competitiveness rather than by institutional support, and also by Merkulova and Nikolaeva (2022) within their cluster membership of EU taxes indicators and fiscal capacity. Cluster 1, small in number (n = 2), has highly elevated measures of GDP per capita (1.809), trade (1.693), exports (1.653), and GFCF (5.168), but with very low health spending (–0.897) and domestic credit (–1.156). ALOAI is flat (0.018), inferring under-adoption of AI due to underdeveloped policy ecosystems or mismatch between financial and innovation systems. Nenov et al. (2023) see similar mismatch in their neural model predictions, noting how successful economies have low innovation outcomes if institutional or behavioral factors are not appropriately in balance with structural capabilities. Cluster 7 includes the extreme dataset in isolation, with highly elevated inflation (10.237) and negative scores in credit, GDP per capita, and trade. Its negative ALOAI (–0.527) is evidence of systemic economic volatility and infers best to be interpreted as representing simply an extreme (outlier) or abnormal macroeconomic regime not reflecting wider tendencies. Such extremes are in support of the application of unsupervised clustering analysis to reveal macroeconomic outliers, as previously demonstrated in multidimensional cluster research such as Bosna et al. (2024) and Czeczeli et al. (2024).
Figure 3.
Cluster Membership Visualization in Two-Dimensional Projection with Case Labels.
Figure 3.
Cluster Membership Visualization in Two-Dimensional Projection with Case Labels.
Conversely, the least adopter clusters (Clusters 5 and 3) are characterized by poor access to finance, low productivity, and low international integration. This is in line with the cross-EU country analysis by Popović et al., which revealed how extremely sensitive AI adoption is to material use strategies and economic environment, especially in environments with restricted access to material inputs. The evidence supports the suggestion that economic sophistication, access to finance, and external orientation (through exports and international trade) are positively associated with AI adoption in large enterprises. There are exceptions, though—like Cluster 1's very macro indicators with low adoption and Cluster 4's external orientation and high GDP with low public spending—highlighting that economic factors are not sufficient to secure innovation adoption. Rather, as emphasized by Uren & Edwards (2023), organizational maturity and technology readiness mediate the role. Preparedness of the institution, sector patterns, and the prevailing digital cultural environment mediate crucially whether economic slack is turned into technological adoption. Such influences are evidenced in the work by Kochkina et al. (2024), who found that industry-specific strategic fit, enhanced by sector-matched application of AI and preparedness assessments, plays an influent role in shaping successful integration of AI—even within technologically developed environments. The silhouette scores also verify the heterogeneity of these clusters. Cluster 4, with a score of 0.894, is the internally best-coherent cluster and is marked by stable and replicable profile features—i.e., distinct macro indicators and adoption of AI. Cluster 2 and Cluster 6, on the other hand, although prospective in economic orientation, have poor silhouette scores, exemplifying more internal heterogeneity and perhaps more intricate dynamics. Cluster 5 and Cluster 3, although with the number of entities, are low-adopting domains and require targeted intervention in policy. Such clusters are likely to enjoy the greatest benefits from strategic intervention in the form of targeted investment in infrastructure; digital skills and education programs; and international competitiveness-enhancing programs. In essence, such cluster analysis reveals that large EU firm adoption of AI is positively associated with access to financing, external orientation, and GDP per capita, though they are not determinant factors. Institutional power, technological readiness, and strategic fit—through and especially public-private investment systems—are essential to the macroeconomic levers' translation into successful digital transformation.
Table 8.
Cluster Centroids for Standardized Macroeconomic Variables.
Table 8.
Cluster Centroids for Standardized Macroeconomic Variables.
| |
ALOAI |
HEAL |
DCPS |
EXGS |
GDPC |
GCFG |
INFD |
TRAD |
| Cluster 1 |
0.018 |
-1.156 |
1.653 |
5.168 |
1.809 |
-0.897 |
-0.504 |
1.693 |
| Cluster 2 |
1.407 |
1.512 |
-0.453 |
0.365 |
0.907 |
0.762 |
-0.080 |
-0.512 |
| Cluster 3 |
0.018 |
0.415 |
-0.741 |
-0.588 |
-0.095 |
0.797 |
-0.320 |
-0.776 |
| Cluster 4 |
0.693 |
0.849 |
3.619 |
-1.215 |
2.933 |
-1.533 |
-0.191 |
3.579 |
| Cluster 5 |
-0.837 |
-0.826 |
-0.131 |
0.157 |
-0.821 |
-0.739 |
0.167 |
-0.074 |
| Cluster 6 |
0.379 |
-0.277 |
0.857 |
-0.130 |
0.326 |
0.390 |
-0.191 |
0.837 |
| Cluster 7 |
-0.527 |
-0.576 |
-0.747 |
2.406 |
-0.814 |
-2.450 |
10.237 |
-0.714 |
The data analysis of the result of the implementation of the K-Nearest Neighbors (KNN) clustering algorithm on the set of macroeconomic and financial variables provides insightful observations regarding the patterns of artificial intelligence (AI) adoption—through the ALOAI indicator—among large enterprises (250+ staff) with European Union economies. Not accounting for agriculture, mining, and finance, the ALOAI indicator records the proportion of enterprises utilizing any of the AI technologies such as machine learning or recognition of images. The standard variables on which the clustering is performed are current health spending (HEAL), domestic credit to the non-financial sector (DCPS), exports of goods and services (EXGS), Gross Domestic Product (GDP) per capita (GDPC), gross fixed capital formation (GFCF), inflation (INFD), and trade openness (TRAD). The seven cluster centroids represent the average standardized figures of each of the variables from the member countries. Cluster 2 is characterized by the greatest ALOAI indicator (1.407), which validates strong adoption of AI by its constituents. This cluster records high-availability credit (DCPS = 1.512), significant health spending (HEAL = 1.512), robust GDP per capita (GDPC = 0.907), and robust investment in capital formation (GFCF = 0.762). Despite slightly low values in exports and trade, the economies' internal resilience regarding infrastructure, investment, and access to finance seems to be adequate to facilitate digital transformation. The observed patterns are consistent with Iuga & Socol (2024), who emphasize how readiness in the use of artificial intelligence and preventing brain drain are inextricably connected with institutional investment and availability of finance. That such uptake is observed in the cluster suggests collaboration of macroeconomic stability, investment in the provision of social services, and financial capability to produce technological innovation, regardless of whether they have macroeconomic orientation towards international trade. This supports arguments in Czeczeli et al. (2024), who observe that economic resilience and preparedness—especially in situations of macroeconomic volatility—are intricately ingrained in the fiscal and lending architecture of a nation. Cluster 4 also features the ALOAI indicator with a high score (0.693), although with differences in the economic profile. It features the highest levels of exports (EXGS = 3.619) and trade openness (TRAD = 3.579), as well as the highest level of GDP per capita (GDPC = 2.933). On the contrary, it features low health spending (HEAL = –1.533) and negative capital formation (GFCF = –1.215), reflecting low investment in public infrastructure or long-term assets. This reflects that economic models are based on private-sector dynamism, high competitiveness, and international integration. As the analysis by Papagiannis et al. (2021) of intelligent infrastructure and public-private preparedness in Eastern Europe reveals, robust adoption of AI is even possible in market-exposure and innovation-pressure-driven systems lacking public investment. Cluster 6 features an ALOAI of moderate magnitude (0.379) and is a mixed-transitional group. It features mixed signs, with positive values of credit availability (DCPS = 0.857), modest health spending (HEAL = –0.277), and robust trade openness (TRAD = 0.837), but other factors are near- or slightly below-average. The profile identifies emerging and converging economies that have the macroeconomic fundamentals of digital transformation but have not yet translated them into elevated levels of AI adoption. As Iuga & Socol (2024) highlight, such economies tend to require stronger institutional infrastructure, targeted policy instruments, and brain drainage countermasures to leverage their AI preparedness more effectively. Additionally, workforce competences and support structures of innovation may not yet be fully compatible with the demands of digital transformation. Cluster 3, with very low ALOAI (0.018), is characterized by the economic profile of structural weakness. While it features modest health and credit indicators, it features clearly negative values of exports (–0.741), trade (–0.776), and GDP per capita (–0.095). This reflects underdeveloped and weakly integrated economies into international markets, with low external exposure and low national income levels that heavily hamper technological diffusion. These findings are corroborated by Guarascio et al. (2025), who illustrate that regional heterogeneity in exposure to AI and employment is disproportionately driven by macroeconomic underdevelopment and sectoral inflexibility. Even with some government investment in health and/or credit, structural weaknesses prevent firms from rolling out cutting-edge technologies on large scale. Cluster 5 has the lowest ALOAI score (–0.837), and it is characterized by very weak digital transformation. The economic indicators are unambiguously negative or low on average, such as GDP per capita (–0.821), availability of credits (–0.131), low health spending (–0.826), and low capital formation (–0.739). These economies are presumably faced with several systemic barriers—economic, institutional, and infrastructure—that severely impinge on the capabilities of businesses to access digital instruments and invest in AI technologies. As demonstrated by Rađenović et al. (2024) from their cluster analysis of eco-innovation, such underdevelopment is typically an indicator of overall policy inertness and poor coordination of innovation ecosystems. In the absence of targeted fiscal measures, support to private sector digitalization, and inclusion into EU innovation policies, these economies are unlikely to escape low adoption equilibria.
Figure 4.
Cluster-Wise Standardized Means of Macroeconomic Variables with Error Bars.
Figure 4.
Cluster-Wise Standardized Means of Macroeconomic Variables with Error Bars.
Cluster 1 is an intriguing and educational example in which a low ALOAI (0.018) is found with exceptionally favorable macroeconomic indicators. It is the best performer in GDP per capita (1.809), exports (1.653), and trade integration (1.693), and in gross fixed capital formation (GFCF = 5.168), reflecting a structural wealth and integration profile. But it also manifests stark weaknesses in health spending (HEAL = –0.897) and access to credit (DCPS = –1.156). Such dualities imply that macroeconomic prosperity is not in itself enough to provide successful AI adoption. As Uren & Edwards (2023) contend, organisational preparedness in the form of digital competency, strategic alignment, and institutional flexibility is paramount in converting advantageous macro settings into innovation results. Likewise, Baumgartner et al. (2024) note the requirement of essential digital capabilities and transformation competencies on the firm level, which in turn might be scarce even in ostensibly prosperous economies. Hence, the example of Cluster 1 serves to illustrate that the diffusion of AI is demonstrably dependent on the convergence of financial accessability, institutional backing, and technological preparedness. Cluster 7 consists of a single extreme outlier. It is characterized by anomalously high inflation (INFD = 10.237) and negatively skewed values on all of the remaining indicators, including GDP per capita, access to credit, and international integration. The attendant negative ALOAI (–0.527) reinforces the hypothesis that macro dysfunction generates a setting hostile to digital innovation. Such settings are typically associated with brain drain (Iuga & Socol (2024)), policy ambiguity, and low institutional capability, which together constitute a feedback cycle of suboptimality in AI preparedness. Here, any push to support the adoption of AI would not be merely about altering digital policy, but macroeconomic stabilization. Cluster 7 is thus best interpreted as structural abnormality, and presents in itself a cautionary reminder of technological transformation's foundational prerequisites.