4. Discussion
This study analyzed 38 fMRI datasets (19 healthy, 19 tinnitus) and 80 EEG recordings (40 per group) to investigate neuroimaging biomarkers for tinnitus detection. The results demonstrated that CNN-based analysis of resting-state fMRI achieved optimal classification performance with slice 17 showing 99.0% ± 0.4% accuracy, representing the best-performing axial slice among 32 evaluated slices, while hybrid models combining pre-trained CNNs with traditional classifiers (VGG16-DT) reached 98.95% accuracy for automated tinnitus detection. EEG microstate analysis revealed systematic disruptions in neural network dynamics, with the most pronounced alterations observed in gamma-band microstate B occurrence rates (Cohen's d = 2.11, p = 1.58×10⁻¹⁴), where healthy participants exhibited significantly higher rates (56.56 vs 43.81 events/epoch). Additional significant microstate changes included reduced alpha-band microstate A coverage in both 4-state (34.62% vs 29.95%) and 6-state configurations, shortened microstate B durations in 5-state (1.23 vs 1.11 ms) and 7-state (1.12 vs 0.99 ms) configurations, and decreased beta-band microstate D durations in both 5-state (3.31 vs 2.93 ms) and 7-state (2.70 vs 2.34 ms) clustering solutions, collectively indicating that tinnitus pathophysiology involves widespread disruptions particularly affecting high-frequency gamma oscillations and alpha-band resting-state networks with altered temporal dynamics. Machine learning classification of EEG microstate features achieved up to 98.8% accuracy using Random Forest and Decision Tree algorithms, while CWT-transformed EEG analysis with deep learning models reached 95.4% accuracy, with delta and alpha frequency bands providing the most discriminative information for tinnitus detection.
Our findings demonstrate the effectiveness of neuroimaging approaches for automated tinnitus detection, examining fMRI and EEG modalities separately to establish their individual diagnostic capabilities. The superior performance of hybrid CNN models on rs-fMRI data (VGG16-Decision Tree: 98.95% accuracy) aligns with previous work by Xu et al. [
21], who achieved 94.4% AUC using CNN on functional connectivity matrices from 200 participants. However, our study extends beyond connectivity analysis to direct slice-based classification, revealing spatial heterogeneity in discriminative capacity across brain regions, with mid-to-superior axial slices containing the most diagnostically relevant biomarkers. The EEG microstate analysis revealed systematic disruptions in tinnitus patients, particularly in gamma-band microstate B occurrence (Cohen's d = 2.11) and alpha-band coverage reductions, consistent with recent findings by Najafzadeh et al. [
45], who reported significant alterations in beta and gamma band microstates with exceptional classification performance (100% accuracy in gamma band). Our results corroborate their findings regarding gamma-band importance while additionally demonstrating robust performance across multiple frequency bands using traditional machine learning approaches. The observed reductions in microstate durations and occurrence rates support the hypothesis of altered neural network dynamics in tinnitus, extending previous work by Jianbiao et al. [
46], who identified increased sample entropy in δ, α2, and β1 bands using a smaller cohort (n=20). Our CWT-transformed EEG analysis using deep learning architectures yielded competitive results, with VGG16 achieving 95.4% accuracy in the Delta band. This approach differs from the innovative graph neural network methodology employed by Awais et al. [
46], who achieved 99.41% accuracy by representing EEG channels as graph networks with GCN-LSTM architecture. While their graph-based approach showed superior single-metric performance, our multimodal framework provides broader clinical applicability through the integration of both structural (fMRI) and temporal (EEG) neural signatures. The novelty of our approach lies in the separate but comprehensive examination of both fMRI and EEG modalities, providing distinct insights into structural and temporal neural signatures of tinnitus. The spatially-specific fMRI slice analysis revealed heterogeneous discriminative patterns across brain regions, while frequency-specific EEG microstate characterization demonstrated systematic temporal disruptions. Unlike previous studies focusing on single modalities, our framework establishes that beta and alpha frequency bands contain the most discriminative EEG features, while specific fMRI slices corresponding to auditory processing regions show optimal diagnostic performance, each contributing unique diagnostic information. The consistent performance of tree-based classifiers (Random Forest, Decision Tree) across EEG features supports findings by Doborjeh et al. [
47], who emphasized the importance of feature selection in achieving high classification accuracy (98-100%) for therapy outcome prediction. The clinical implications of our findings are substantial, with the high classification accuracies across multiple modalities suggesting potential for objective diagnostic tools in tinnitus assessment. The identification of specific neural signatures, particularly in gamma and alpha frequency bands, provides neurobiological insights that may inform therapeutic target identification and support the development of personalized treatment strategies.
The CNN performance analysis across 32 axial slices of rs-fMRI data (
Figure 10) revealed spatial heterogeneity in discriminative capacity between tinnitus patients and healthy controls, with superior-to-middle axial slices demonstrating higher classification accuracy compared to inferior slices. Slices 17, 26, 14, 12, and 10 achieved accuracies exceeding 98%, while inferior slices (29, 32, 30, 31) showed lower discrimination with accuracies below 61%. This spatial gradient in classification performance aligns with neuroimaging findings demonstrating that superior temporal regions, particularly the superior temporal gyrus and primary auditory cortex, are involved in tinnitus pathophysiology [
48]. Previous rs-fMRI studies have shown that tinnitus patients exhibit altered activity in superior temporal cortex regions, including the middle and superior temporal gyri, which corresponds to the higher performance of CNN models in middle-to-upper axial slices observed in our analysis. Furthermore, superior/middle temporal regions are involved in processing conflicts between auditory memory and signals from the peripheral auditory system, potentially explaining the discriminative capacity of these brain regions for tinnitus classification. The performance difference between superior (accuracy >98%) and inferior slices (accuracy <61%) suggests that tinnitus-associated neural alterations are spatially localized to specific anatomical regions, supporting the approach that analysis of superior temporal and auditory processing areas may provide relevant biomarkers for tinnitus assessment [
49].
Figure 11a provides important insights into the altered neural dynamics observed in individuals with tinnitus by illustrating the spatial distribution of maximum voxel intensity differences across fMRI slices. The increased intensity variations found in the tinnitus group, especially in regions related to auditory processing and sensory integration, suggest disrupted neural activity that may be associated with the persistent perception of phantom sounds. These differences likely reflect mechanisms such as neural hyperactivity or maladaptive plasticity within the auditory cortex and its interconnected networks, both of which have been widely reported in the literature on tinnitus pathophysiology. In comparison, the more uniform patterns in healthy individuals indicate stable neural activity in the absence of such disturbances. The observed distribution of changes also suggests the involvement of non-auditory regions, including components of the limbic system, which may contribute to the emotional and cognitive experiences commonly reported in tinnitus [
50]. These findings support the hypothesis that tinnitus involves not only abnormal auditory processing but also dysfunctional cross-modal and intra-auditory connectivity, reinforcing the presence of widespread alterations in brain networks.
Figure 11b further highlights the differences in neural activity between healthy individuals and tinnitus patients through an analysis of mean voxel intensity values across axial fMRI slices. The consistently higher intensity values observed in the tinnitus group, particularly within slices 5 to 18 and 30 to 32, indicate abnormal neural dynamics that may be linked to increased activity and disrupted connectivity within critical auditory and non-auditory brain regions. These slices likely include structures such as the thalamus, auditory midbrain, and components of the limbic system, all of which play essential roles in auditory perception and the affective response to sound [
51]. Additionally, the higher standard deviation of voxel intensity in the tinnitus group, reaching a value of 0.1572 in slice 31 compared to 0.0522 in the control group, suggests greater neural variability. This variability may reflect disrupted thalamocortical rhythms and maladaptive reorganization of brain activity [
52]. Overall, these results reinforce the understanding that tinnitus is a condition involving widespread neural dysfunction. It affects not only the auditory pathways but also brain areas responsible for attention, emotional regulation, and multisensory integration [
53]. The clear distinctions illustrated in
Figure 11b align with previous neuroimaging studies reporting abnormal functional connectivity and increased activity within both auditory and limbic networks, providing additional support for the role of impaired neural synchronization in the manifestation of tinnitus.
The classification performance demonstrated in
Figure 12 by hybrid models, particularly VGG16-DT (98.95% ± 2.94%), shows the effectiveness of combining pre-trained CNN feature extraction with traditional machine learning classifiers for neuroimaging-based tinnitus diagnosis. These results are consistent with recent developments in medical imaging where hybrid architectures have shown improved performance over standalone deep learning approaches by integrating automated feature extraction with interpretable classification mechanisms [
45]. The performance across all models exceeding the 95% clinical relevance threshold, as illustrated in
Figure 12, indicates the potential clinical utility of fMRI-based automated tinnitus detection, supporting previous observations that resting-state fMRI can distinguish tinnitus patients from healthy controls through altered neural connectivity patterns [
21]. The superior performance of VGG16-based models over ResNet50 variants (98.76% vs. 97.94% average accuracy) may reflect VGG16's architectural suitability for capturing spatial patterns in brain imaging data, while the improved stability observed in hybrid approaches (particularly VGG16-RF with ±1.28% standard deviation) enhances the reliability necessary for clinical implementation. The average improvement of 1.2-2.8% in accuracy achieved by hybrid models over standalone CNNs indicates the value of integrating multiple algorithmic approaches, suggesting that the interpretability and robustness of traditional machine learning methods complement the feature extraction capabilities of deep neural networks in medical diagnostic applications.
The systematic disruptions in EEG microstate dynamics observed in our tinnitus cohort (
Figure 13) provide evidence for neural network alterations underlying phantom auditory perception. The most pronounced finding, reduced gamma-band microstate occurrence rates (Cohen's d = 2.11,
Table 9), aligns with established research demonstrating that gamma band activity in the auditory cortex correlates with tinnitus intensity and that decreased tinnitus loudness is associated with reduced gamma activity in auditory regions [
54,
55]. While previous studies have reported increased local gamma power in tinnitus patients, our findings suggest that global gamma-band network coordination is disrupted, reflecting maladaptive reorganization of auditory cortical networks that may contribute to the persistent phantom perception. The systematic reductions in alpha-band microstate parameters observed in our study (
Figure 13,
Table 9), including decreased coverage (29.95% vs 34.62%) and occurrence rates, are consistent with documented disruptions of default mode network connectivity in tinnitus patients and the established role of alpha oscillations in spatiotemporal organization of brain networks [
56]. These alpha-band alterations likely reflect impaired resting-state network integrity, potentially underlying the intrusion of phantom auditory percepts into consciousness and the cognitive dysfunctions commonly associated with chronic tinnitus. The consistent beta-band microstate disruptions, including reduced coverage and shortened durations, align with neuroimaging evidence showing alterations in multiple resting-state networks in tinnitus patients, including attention networks and sensorimotor integration systems [
49]. Our findings are further supported by recent research from Najafzadeh et al., who reported alterations in beta band microstates with increased microstate A duration and decreased microstate B duration in tinnitus patients, along with elevated occurrence rates in the tinnitus group [
45]. The high statistical power (>0.99) and large effect sizes (Cohen's d = 1.08-2.34) across all frequency bands indicate clinically meaningful differences that collectively support the emerging conceptualization of tinnitus as a network disorder affecting multiple brain systems beyond the traditional auditory processing pathways [
57].
The superior performance of tree-based algorithms, particularly Random Forest (RF) and Decision Tree (DT), achieving 98.8% accuracy with perfect precision (100.0%) for EEG microstate-based tinnitus classification (
Figure 14,
Table 10), aligns with recent findings demonstrating the effectiveness of ensemble methods in neurological disorder detection. These results are consistent with studies showing that Random Forest models excel in EEG-based tinnitus classification due to their robustness and ability to reduce overfitting while identifying key frequency band features [
45]. The perfect precision achieved by RF and DT models indicates their reliability in minimizing false positive diagnoses, which is clinically relevant for avoiding unnecessary interventions. In contrast, the suboptimal performance of Deep Neural Networks (DNN) with 72.5% accuracy and high variability (±16.9%) suggests that traditional machine learning approaches may be more suitable for microstate feature classification in limited sample scenarios, a finding supported by research demonstrating processing speed advantages of tree-based methods over deep learning approaches [
58]. The analysis of CWT-transformed EEG signals revealed distinct frequency-specific patterns, with VGG16 demonstrating the most consistent performance across Delta (95.4%), Theta (93.4%), and Alpha (94.1%) bands (Figure 16). The superior performance of low-frequency components (Delta and Alpha bands) with ROC AUC values exceeding 0.98 for CNN and VGG16 models supports previous research indicating that these frequency bands contain the most discriminative information for automated tinnitus detection [
59]. The statistical validation using DeLong's test confirmed that while CNN and VGG16 performed comparably (p > 0.05), both significantly outperformed ResNet50 in Delta and Theta bands (p < 0.05), suggesting that simpler deep learning architectures may be more effective for EEG-based tinnitus classification than more complex models like ResNet50. These findings collectively demonstrate that both traditional machine learning approaches using microstate features and deep learning methods applied to frequency-transformed EEG data can achieve high classification accuracy, with the choice of method depending on the specific feature representation and computational requirements.
The present study employed parallel unimodal neuroimaging analyses to characterize tinnitus-related neural alterations using independent EEG and fMRI datasets obtained from separate cohorts. Although direct multimodal integration was not feasible, the convergent results across modalities offer complementary perspectives on tinnitus pathophysiology. In fMRI, the highest classification performance was observed in mid-to-superior axial slices, particularly slice 17, achieving 99.0% accuracy. This region corresponds to cortical areas encompassing the auditory cortex and functionally related association networks implicated in tinnitus generation [
52]. This spatial localization aligns with previous neuroimaging evidence demonstrating aberrant functional connectivity within auditory and attention-related networks among tinnitus patients [
60,
61]. In parallel, EEG microstate analysis revealed significant alterations in high-frequency gamma oscillations (Cohen’s d = 2.11 for 7-state microstate B occurrence) and alpha-band network dynamics, consistent with electrophysiological findings of abnormal neural synchronization and resting-state instability in tinnitus [
45,
55]. The gamma-band abnormalities may reflect disrupted local cortical computations within regions corresponding to those identified by fMRI, as gamma oscillations are tightly linked to localized sensory and perceptual processing [
62]. Moreover, the reduced alpha-band microstate coverage and duration observed in the EEG cohort mirror fMRI findings of altered connectivity within the default-mode and attentional networks [
49,
63]. Collectively, these parallel findings suggest that tinnitus involves both spatially localized disruptions in functional connectivity, detectable through hemodynamic imaging, and temporally dynamic instabilities in large-scale electrophysiological networks. Future investigations employing simultaneous EEG–fMRI acquisition in matched cohorts could directly examine the spatiotemporal coupling between these hemodynamic and electrophysiological alterations, thereby elucidating the mechanistic interplay between localized cortical dysfunction and distributed network reorganization in tinnitus pathophysiology [
64,
65].
In our fMRI analysis, individual time points were treated as independent observations to enable voxel-wise spatial pattern classification at fine temporal resolution. This methodological approach was motivated by emerging evidence supporting single-volume decoding frameworks that prioritize instantaneous spatial activation patterns over temporally aggregated connectivity metrics [
66,
67]. While conventional resting-state fMRI analyses typically compute functional connectivity through temporal correlations between brain regions [
68], such approaches primarily capture static or time-averaged network interactions and may overlook transient spatial configurations that encode clinically relevant neurophysiological states [
69]. The time-point-based classification framework employed here aligns with recent developments in dynamic pattern analysis and machine learning-based neuroimaging, where instantaneous spatial representations have demonstrated discriminative capacity for clinical phenotyping [
70]. By analyzing individual volumes, our approach preserves spatial heterogeneity across axial slices and enables deep learning architectures to detect subtle regional activation patterns that may be obscured in connectivity matrices derived from temporal averaging [
71]. The substantial sample size generated through this approach (400 time points per subject per slice) provided sufficient statistical power for the classifier to learn generalizable spatial biomarkers while accounting for temporal variability inherent in resting-state acquisitions [
72]. We acknowledge that this methodology does not explicitly model temporal dependencies or dynamic functional connectivity fluctuations, which represent complementary dimensions of brain network organization [
73]. Traditional connectivity-based approaches excel at characterizing inter-regional synchronization and network-level dysfunction [
12], whereas our spatial pattern classification framework provides orthogonal information regarding localized activation signatures. These methodological paradigms should be viewed as complementary rather than mutually exclusive, each offering distinct insights into tinnitus-related neural alterations [
74]. The high classification accuracy achieved through spatial pattern analysis (up to 99% for optimal slices) suggests that tinnitus manifests detectable spatial activation signatures independent of explicit temporal modeling, potentially reflecting sustained aberrant neural activity within specific cortical territories [
75]. Future investigations will integrate temporal dynamics through hybrid architectures incorporating recurrent neural networks, long short-term memory units, or graph convolutional networks to jointly model spatiotemporal features [
76], thereby providing a more comprehensive characterization of tinnitus-related network dysfunction. Additionally, combining time-point classification with sliding-window connectivity analysis and time-varying graph metrics would enable assessment of whether transient spatial patterns identified here correspond to specific dynamic functional connectivity states [
77]. Such multimodal temporal-spatial integration would enhance both the biological interpretability and clinical utility of neuroimaging-based tinnitus biomarkers [
78]. This slice-wise spatial classification approach is consistent with recent applications in other neuropsychiatric disorders, where similar time-point-based CNN methodologies achieved diagnostic accuracies exceeding 98% for schizophrenia detection using resting-state fMRI data, demonstrating the broader applicability of instantaneous spatial pattern analysis across clinical populations [
79]. Such parallel findings across distinct clinical conditions support the validity of spatial decoding frameworks as complementary tools to traditional connectivity-based analyses in psychiatric neuroimaging.
The use of Global Field Power (GFP) as a spatial summary metric for subsequent time-frequency decomposition represents a deliberate methodological choice that balances computational efficiency with preservation of neurophysiologically relevant signal characteristics. While GFP collapses the 64-channel topographic distribution into a single time series representing the spatial standard deviation of scalp potentials at each time point, this metric specifically captures the overall strength of neural synchronization and global brain state dynamics that are independent of reference electrode selection [
80]. Importantly, our analytical framework did not rely solely on GFP-derived features; the comprehensive microstate analysis explicitly preserved and analyzed topographic information through spatial clustering of multi-channel voltage distributions, extracting features including microstate duration, coverage, occurrence, and spatial configuration patterns across all 64 electrodes [
35]. The GFP-to-CWT transformation was implemented as a complementary analysis stream designed to capture temporal dynamics of global neural synchronization within each frequency band, which has demonstrated particular relevance for characterizing aberrant neural oscillations in tinnitus pathophysiology [
45]. This approach aligns with established frameworks in clinical neurophysiology where GFP serves as a robust marker of overall cortical excitability and network-level synchronization changes, particularly for detecting global alterations in brain state dynamics that characterize neuropsychiatric conditions [
81]. The integration of both topographically-resolved microstate features (preserving spatial information) and GFP-derived time-frequency representations (emphasizing global synchronization dynamics) provided complementary perspectives on tinnitus-related neural alterations, with the former achieving 98.8% classification accuracy through spatial pattern analysis and the latter reaching 95.4% accuracy through temporal dynamics characterization. Future investigations will explore spatially-resolved time-frequency decomposition approaches, such as channel-wise CWT or source-space analysis, to integrate both spatial topography and temporal dynamics within unified multivariate representations [
82].
All experiments were conducted on a system running Windows 11, equipped with an NVIDIA RTX 3050 Ti GPU, an Intel Core i7 processor, and 32 GB of RAM. The programming language used for model implementation was Python, with PyCharm as the integrated development environment (IDE). The deep learning models were developed and evaluated using the TensorFlow and Keras libraries, which provided a robust framework for neural network construction and training. Training times for each model varied significantly. The SVM required 2.025 hours, the RF took 4.63 hours, and the DT completed training in 0.335 hours. In comparison, the more complex DNN took 16.14 hours, and the CNN required 24.87 hours. These results reflect the trade-off between model complexity and computational efficiency.
Several methodological limitations must be acknowledged in interpreting these results. The EEG and fMRI datasets were derived from different participant cohorts, preventing direct multimodal feature integration and limiting our ability to leverage complementary information from both neuroimaging modalities. The fMRI data utilized publicly available datasets with male-only participants experiencing acoustic trauma-induced tinnitus, while EEG data were collected independently from mixed-etiology cohorts with balanced gender representation, resulting in potential differences in acquisition protocols, participant characteristics, and clinical assessment procedures that may affect direct comparisons between modalities. The fMRI dataset's restriction to male participants with acoustic trauma-induced tinnitus limits generalizability to the broader tinnitus population, particularly regarding gender-specific neural responses and non-acoustic etiologies. Furthermore, the different etiologies of tinnitus in our datasets may limit the generalizability of our findings, as acoustic trauma-induced tinnitus often involves specific cochlear damage patterns and may exhibit distinct neural signatures compared to tinnitus from other causes. The cross-sectional study design limits understanding of temporal stability of the identified neural biomarkers and their relationship to tinnitus symptom progression over time. Additionally, the relatively modest sample sizes (40 per group for EEG, 19 per group for fMRI) may limit generalizability across different tinnitus subtypes, severity levels, and demographic populations.
Future investigations should prioritize the collection of matched EEG and fMRI data from identical participant cohorts to enable true multimodal analysis and feature fusion approaches that could potentially improve classification accuracy and provide more comprehensive characterization of tinnitus-related neural alterations. Validation across diverse tinnitus etiologies will be essential to ensure broad clinical applicability of these classification approaches. Longitudinal studies tracking patients over extended periods would help establish the temporal stability of identified biomarkers and their potential utility for monitoring treatment response. The development of real-time processing algorithms and optimization of computational requirements will be essential for clinical translation, along with validation studies across multiple clinical centers using standardized protocols. Integration of additional clinical measures, including detailed tinnitus severity assessments, audiological profiles, and psychological evaluations, would enhance the clinical relevance of these neuroimaging-based classification approaches and support the development of personalized treatment strategies based on individual neural signatures.
The high classification accuracies achieved in this study (98.8% for EEG microstate analysis and 98.95% for fMRI hybrid models) suggest potential clinical utility for objective tinnitus diagnosis. However, successful clinical implementation requires addressing several practical considerations including integration with existing audiological workflows, establishment of standardized acquisition protocols across different clinical centers, and development of user-friendly interfaces for non-technical clinical staff. The computational requirements and processing times observed in our study (ranging from 0.335 to 24.87 hours depending on the model) indicate the need for optimized implementations and appropriate hardware infrastructure in clinical settings. Furthermore, regulatory approval pathways for AI-based diagnostic tools will require validation on larger, more diverse patient populations and demonstration of consistent performance across different scanner types and acquisition parameters.
The identification of objective neural biomarkers for tinnitus represents a significant advancement toward evidence-based diagnosis and treatment monitoring in a condition that has traditionally relied on subjective patient reports. The high classification accuracies achieved across both neuroimaging modalities suggest potential for reducing diagnostic uncertainty and supporting clinical decision-making, particularly in cases where symptom presentation is ambiguous or when objective assessment is required for research or medico-legal purposes. The specific neural signatures identified, particularly in gamma and alpha frequency bands, may inform the development of targeted therapeutic interventions, including neurofeedback protocols and brain stimulation approaches. However, successful clinical implementation will require careful consideration of cost-effectiveness, training requirements for clinical staff, and integration with existing healthcare infrastructure to ensure broad accessibility and practical utility in routine clinical practice.
Taken together, the findings from this study underscore the complementary strengths of EEG and fMRI in elucidating tinnitus-related neural mechanisms. The EEG analyses captured the rapid temporal fluctuations and network instabilities underlying abnormal oscillatory dynamics, whereas the fMRI analyses identified spatially localized alterations in auditory and attentional cortical regions. The convergence of these independent observations supports the hypothesis that tinnitus arises from both localized cortical hyperactivity and disrupted large-scale network coordination. By establishing a comparative framework that integrates temporal and spatial perspectives, this work advances the understanding of tinnitus pathophysiology and highlights the value of multi-perspective neuroimaging for developing objective diagnostic biomarkers. Future research should extend this framework using simultaneous EEG–fMRI acquisition and larger, matched cohorts to directly examine the spatiotemporal interactions between electrophysiological and hemodynamic activity in tinnitus, ultimately paving the way toward personalized and mechanism-based therapeutic interventions.