Preprint
Article

This version is not peer-reviewed.

Comparative Analysis of CNN and Transformer Models for Multi-Class Diabetic Retinopathy Grading Using Fundus Images

Submitted:

12 August 2026

Posted:

13 August 2026

You are already at the latest version

Abstract
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although both convolutional neural network (CNN)-based and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons under a unified experimental framework remain limited. This study aims to systematically compare representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fully fine-tuned and independently evaluated over five runs using different random seeds. Performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. The repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Retinal diseases are responsible for a substantial proportion of vision impairment and blindness worldwide, posing an increasing challenge to healthcare systems [1]. These disorders, including diabetic retinopathy (DR), pathological myopia (PM), age-related macular degeneration (AMD), and glaucoma (GLA), impose a substantial burden on healthcare systems and significantly affect quality of life, particularly among aging and working-age populations [2]. Early detection and regular monitoring of retinal abnormalities play a crucial role in preventing permanent vision loss.
Among these conditions, diabetic retinopathy (DR), a microvascular complication of diabetes, is one of the most prevalent and clinically significant causes of blindness [2,3]. The global burden of DR continues to increase, with the number of affected individuals projected to rise from approximately 126.6 million in 2010 to 191.0 million by 2030. In particular, cases of vision-threatening diabetic retinopathy (VTDR) are expected to grow from 37.3 million to 56.3 million if timely intervention is not achieved [4]. Despite the availability of effective screening programs and treatment strategies, DR remains a leading cause of vision loss among working-age populations [5]. Early detection and accurate grading of disease severity are essential for timely intervention and effective treatment planning. However, manual examination of retinal fundus images is time-consuming, requires expert ophthalmologists, and is subject to inter-observer variability.
The rapid evolution of artificial intelligence (AI) has accelerated the adoption of deep learning (DL) for automated medical image analysis. This technology has significantly improved the accuracy of disease detection across various image modalities [5,6,7,8], including various retinal diseases [9,10]. Transfer learning (TL) and pretrained models have further improved performance in data-limited clinical settings [11]. In particular, convolutional neural networks (CNNs) have demonstrated strong performance in retinal disease classification tasks due to their ability to learn hierarchical spatial representations from fundus images [12]. Despite their success, CNN-based models primarily focus on local feature extraction and may struggle to capture global contextual relationships, which can be critical for distinguishing subtle differences in diabetic retinopathy severity levels. More recently, transformer-based models have emerged as an alternative paradigm in computer vision, offering the capability to model global contextual relationships through self-attention mechanisms [13]. These models have shown competitive performance in various image classification tasks.
While both CNN-based and transformer-based models have demonstrated promising performance, their relative effectiveness for fine-grained multi-class classification tasks, such as diabetic retinopathy grading, remains an active research topic. As deep learning architectures continue to evolve, consistent comparative evaluations conducted under identical experimental settings for understanding their relative strengths, limitations, and practical applicability in automated retinal screening systems. Moreover, relatively fewer studies investigate the trade-off between classification performance and computational efficiency, which is an important consideration for deploying deep learning models in practical retinal screening systems.
To address these gaps, this study presents a comprehensive comparative analysis of six state-of-the-art pretrained deep learning architectures for five-class diabetic retinopathy grading using retinal fundus images. Three representative CNN-based models, namely ResNet50, EfficientNet-B0, and MobileNetV2, are compared with three transformer-based models, including Vision Transformer (ViT), Swin-Tiny Transformer, and Swin Transformer. These models are trained and evaluated under identical preprocessing, augmentation, transfer learning, training protocols, and evaluation criteria to ensure a fair comparison and to provide a consistent experimental protocol for comparative evaluation. In addition to prediction performance, this study analyzes the computational efficiency of each architecture in terms of model complexity and inference speed, providing practical insights into selecting suitable DL models for automated diabetic retinopathy screening.
The key contributions of this study can be summarized as follows:
  • A unified experimental framework for the comparative evaluation of representative CNN-based and transformer-based architectures for five-class diabetic retinopathy grading.
  • A systematic evaluation of six pretrained deep learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer, Swin-Tiny Transformer, and Swin Transformer, using transfer learning.
  • Investigation of multi-class diabetic retinopathy grading under an imbalanced dataset using a unified training strategy based on weighted loss and balanced sampling.
  • A comprehensive comparative analysis consisting of overall and per-class classification performance, QWK, learning and convergence behavior, computational efficiency, and performance stability across five independent runs.

3. Materials and Methods

3.1. The Proposed Framework Overview

In this study, diabetic retinopathy (DR) grading is formulated as a supervised multi-class classification task. The objective is to automatically classify retinal fundus images into five DR severity levels, ranging from no apparent DR to advanced disease stages. The proposed methodological framework, illustrated in Figure 1, consists of seven main steps:
  • Data Acquisition: Retinal fundus images representing the five DR severity classes are collected from publicly available datasets.
  • Data Splitting into training, validation, and testing subsets for model development and evaluation.
  • Data Preprocessing: Standardization of input images to ensure compatibility across different DL architectures.
  • Data Augmentation and Class Imbalance Handling: apply data augmentation and training strategies to improve model generalization and address class distribution imbalance.
  • Transfer Learning using Pretrained DL Models: Adaptation of CNN- and transformer-based pretrained models for multi-class DR grading.
  • Classification: Prediction of DR severity grades using the learned representations and classification layers.
  • Model Evaluation and Analysis: Evaluation and comparison of model performance by the use of standard evaluation metrics and computational analysis.
The following subsections describe each stage of the proposed workflow in detail.

3.2. Dataset Acquisition and Statistics

In this study, the APTOS 2019 blindness detection dataset was utilized for multi-class diabetic retinopathy (DR) grading [34]. The dataset is publicly available through the Kaggle repository and was originally released as part of the Asia Pacific Tele-Ophthalmology Society (APTOS) challenge. It consists of 3,662 RGB retinal fundus images that were clinically examined and annotated by trained ophthalmologists according to the International Clinical Diabetic Retinopathy Disease Severity Scale (ICDRSS) [35]. Each fundus image is assigned to one of five DR severity grades: no diabetic retinopathy (No DR), mild DR, moderate DR, severe DR, and proliferative DR.
The publicly available dataset includes predefined training, validation, and testing subsets, which were directly adopted in this study without performing any additional data splitting to ensure reproducibility. The predefined subsets preserve a similar class distribution across the training (approximately 80%), validation (10%), and test (10%) sets, resulting in an approximately stratified distribution across the five DR severity grades, as summarized in Table 1. The dataset download source is provided in the Data Availability Statement. As shown in Table 1, the dataset exhibits a clear class imbalance, with a substantially larger number of normal retinal images than advanced DR stages. This characteristic was considered during model training by employing class-balanced optimization strategies, as described in Section 4.
Representative retinal fundus images corresponding to the five diabetic retinopathy severity grades are presented in Figure 2. These examples illustrate the progressive visual characteristics of diabetic retinopathy, ranging from normal retinal appearance to proliferative disease.

3.3. Data Preprocessing

All retinal fundus images were processed using a unified preprocessing pipeline to ensure consistency across DL architectures and enable fair comparison between CNN- and transformer-based models. First, all images were converted into RGB format to maintain a consistent three-channel representation. The images were then resized to a fixed resolution of 224x224 pixels to meet the input requirements of the selected pretrained models.
Pixel intensity values were normalized to provide a standardized input distribution and improve training stability. Since the models were initialized with weights pretrained on ImageNet, the images were normalized using the corresponding ImageNet mean and standard deviation. Additionally, the categorical diabetic retinopathy severity labels were converted to numerical class labels corresponding to the five grading levels: No DR (0), Mild DR (1), Moderate DR (2), Severe DR (3), and Proliferative DR (4). The same preprocessing procedure was applied across all evaluated architectures to ensure that performance differences were primarily attributable to model characteristics rather than to variations in data preparation.

3.4. Data Augmentation and Class Imbalance Handling

To improve model generalization and reduce overfitting, data augmentation was applied only to the training set. The augmentation pipeline consisted of random horizontal flipping, small-angle rotation, slight geometric transformations, and mild brightness and contrast adjustments. No augmentation was applied to validation or testing images to ensure unbiased model evaluation. To reduce the effect of class imbalance among DR severity grades, a class-weighted loss function was employed during training, assigning higher penalties to underrepresented classes.

3.5. Selection of Pretrained Deep Learning Architectures

To investigate the performance differences between convolution-based and attention-based architectures, this study selected representative pretrained models from CNN and transformer families. The selected CNN models (ResNet50, EfficientNet-B0, and MobileNetV2) represent complementary convolutional design strategies, namely residual learning, compound model scaling, and lightweight network design. Similarly, the selected transformer models (ViT, Swin-Tiny Transformer, and Swin Transformer) represent complementary attention-based paradigms, including global self-attention and hierarchical window-based attention, while covering both lightweight and higher-capacity architectures. This selection enables a comprehensive comparison of classification performance and computational requirements for multi-class diabetic retinopathy grading.

3.5.1. CNN-Based Models

To evaluate the effectiveness of convolution-based architectures for diabetic retinopathy grading, three representative CNN models were selected: ResNet50, EfficientNet-B0, and MobileNetV2. These models represent different CNN design strategies, including deep residual learning, efficient model scaling, and lightweight architectures. CNN-based models are particularly effective at extracting local spatial patterns from retinal images, which are important for identifying DR-related abnormalities, including lesions, texture variations, and structural changes.
ResNet50:
ResNet50 [36] is a deep CNN architecture based on residual learning, in which shortcut connections are introduced to improve gradient flow and mitigate the degradation problem associated with training deeper networks. Through multiple convolutional layers, ResNet50 learns hierarchical visual representations ranging from low-level features, such as edges and textures, to more complex disease-related patterns. In this study, ResNet50 was selected as a standard CNN baseline to evaluate the capability of deep residual architectures for multi-class DR grading.
EfficientNet-B0:
EfficientNet-B0 [37] is based on compound scaling, which systematically balances network depth, width, and input resolution to improve performance while maintaining computational efficiency. This balanced scaling allows the model to learn representative visual features with fewer parameters compared with conventional CNN architectures. EfficientNet-B0 was included to investigate the relationship between classification performance and computational efficiency in retinal image analysis.
MobileNetV2:
MobileNetV2 [38] is a lightweight CNN architecture designed to reduce computational complexity through depth wise separable convolutions and inverted residual blocks. These mechanisms reduce the number of trainable parameters while preserving effective feature extraction. In this study, MobileNetV2 was evaluated to assess the potential of lightweight CNN architectures for efficient DR grading, particularly in scenarios with limited computational resources.

3.5.2. Transformer-Based Models

To compare convolution-based models with attention-based approaches, three transformer architectures were evaluated: Vision Transformer (ViT), Swin-Tiny Transformer, and Swin Transformer. Unlike CNNs, which primarily focus on local feature extraction, transformer-based models use self-attention mechanisms to capture relationships between different image regions. This capability may be beneficial for retinal image analysis, where disease severity can depend on abnormalities distributed across different areas of the fundus image.
Vision Transformer (ViT):
ViT [39] adapts transformer architectures for image classification by dividing an image into fixed-size patches and modeling relationships among these patches using self-attention mechanisms. This design enables the model to capture global contextual information across the entire retinal image. ViT was selected as the standard transformer baseline to evaluate the effectiveness of global attention-based representation learning for DR grading.
Swin-Tiny Transformer:
Swin-Tiny Transformer [40] is a lightweight hierarchical transformer architecture that introduces shifted window-based self-attention. By limiting attention computation within local windows while enabling information exchange between regions, it reduces computational complexity while maintaining contextual feature learning. Swin-Tiny was included to investigate the performance of efficient transformer architectures compared with lightweight CNN models.
Swin Transformer (Base variant):
Swin Transformer [40] extends the vision transformer concept by incorporating hierarchical feature extraction and shifted-window attention mechanisms, enabling the model to capture multi-scale representations. This architecture combines local feature modeling with broader contextual understanding, which can be valuable for analyzing retinal abnormalities appearing at different scales. In this study, Swin Transformer was selected as a stronger hierarchical transformer architecture for comparison with CNN-based models.
The selected architectures allow analysis of model behavior across different design families, including CNN versus transformer models, standard versus lightweight architectures, and accuracy–efficiency trade-offs.

3.6. Transfer Learning, Fine-Tuning, and Classification Strategy

To leverage knowledge learned from large-scale pretrained architectures, this study adopted a transfer learning approach for multi-class diabetic retinopathy grading from retinal fundus images. As described in Section 3.4, the selected CNN-based and transformer-based models were pretrained on large-scale image datasets and have demonstrated strong capabilities in extracting meaningful visual representations. Through transfer learning, these pretrained representations were adapted to the DR classification task, reducing the need for training models from scratch and improving the learning process when using a limited amount of medical imaging data.
To adapt each pretrained architecture for the proposed classification task, the original classification layer was replaced with a task-specific classification head consisting of five output neurons corresponding to the five DR severity grades: Class 0: No DR, Class 1: Mild DR, Class 2: Moderate DR, Class 3: Severe DR, and Class 4: Proliferative DR. The output of each model was passed through a SoftMax activation function to generate a probability distribution across the five classes, where the class with the highest probability score was selected as the final predicted DR grade.
During fine-tuning, all pretrained model parameters, including the backbone and the task-specific classification head, were optimized using the training data to adapt the learned feature representations to retinal fundus images. The validation set was used to monitor model performance and select the best-performing model checkpoint during training. The same fine-tuning strategy was applied across all evaluated architectures to provide a consistent experimental setting for comparing CNN-based and transformer-based models.

4. Experimental Setup and Evaluation Protocols

4.1. Implementation Environment

The proposed framework was implemented in Python using the PyTorch deep learning framework. Torchvision was employed for image preprocessing and data augmentation, while pretrained CNN and transformer models were loaded using the TIMM library. Model evaluation and visualization were performed using Scikit-learn and Matplotlib, respectively. Additional libraries, including Pillow and TQDM, were used for image handling and monitoring the training process. All experiments were conducted using Google Colab Pro+ with a Python 3 runtime and NVIDIA A100 GPU acceleration. To ensure consistency throughout the experiments, all models were trained and evaluated under the same hardware environment and software configuration. Furthermore, five independent experiments were performed using different random seeds, while the random seeds for Python, NumPy, and PyTorch were fixed within each run to ensure reproducibility.

4.2. Hyperparameter Configuration and Training Setup

All models were trained using the same experimental configuration to ensure a consistent comparison between the evaluated architectures. The models were initialized with ImageNet pretrained weights and fully fine-tuned using the training set, while the validation set was used to monitor model performance, guide early stopping, and select the best-performing checkpoint during training. Several hyperparameters were considered during model optimization, including the learning rate, batch size, optimizer, and number of training epochs. The AdamW optimizer was employed with an initial learning rate of 1×10−4, a weight decay of 1×10−4, and a batch size of 32 for all evaluated models. All models were trained for a maximum of 20 epochs. Early stopping with a patience of six epochs was employed based on the validation loss to reduce overfitting and prevent unnecessary training. The selected training configuration was then applied consistently across all evaluated architectures during model training and evaluation. To address the variability introduced by random initialization and stochastic training, each model was independently trained and evaluated five times using different random seeds. The same training configuration and data splitting were maintained across all runs to ensure a fair comparison. Classification results were aggregated across the five runs and reported as mean ± standard deviation (std), providing a more reliable assessment of model performance and stability than a single experimental run. The resulting run-level performance was also used for the statistical analysis of differences among the evaluated models.
To mitigate class imbalance, Weighted Cross-Entropy Loss together with a WeightedRandomSampler was adopted during training. The final hyperparameter configuration used in this study is summarized in Table 2.

4.3. Evaluation Metrics

To comprehensively evaluate the performance of the selected CNN-based and transformer-based models for multi-class diabetic retinopathy grading, two categories of evaluation criteria were considered: classification performance and computational efficiency.
For classification performance, all models were evaluated using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC-ROC). Since the APTOS dataset exhibits an imbalanced distribution across the five DR severity classes, both macro-average and weighted-average precision, recall, F1-score, and AUC were reported. For multiclass AUC, a one-vs-rest (OvR) strategy was employed, where each DR severity label was evaluated against the remaining labels. The resulting class-specific AUC values were then summarized using macro-average and weighted-average. Macro-average metrics assign equal importance to each class regardless of its size, providing a better assessment of the model’s performance on minority classes. In contrast, weighted-average metrics account for the number of samples in each class, providing an overall performance measure that reflects the class distribution. These complementary metrics offer a more comprehensive evaluation of model performance under class imbalance.
Since diabetic retinopathy grading is an ordinal classification problem, Quadratic Weighted Kappa (QWK) [41] was additionally employed to quantify the agreement between the predicted and true DR severity grades while accounting for the severity of misclassification. In addition, QWK is less affected by class imbalance than conventional classification metrics, making it particularly suitable for evaluating the imbalanced APTOS2019 dataset. Unlike standard classification metrics, QWK assigns larger penalties to predictions that differ by multiple severity grades than to those differing by only one grade. Consequently, it provides a more clinically meaningful evaluation of model performance for ordinal diabetic retinopathy grading.
All classification metrics were computed from the confusion matrix, which consists of true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). The definitions and mathematical formulations of these evaluation metrics are summarized in Table 3. For all classification metrics, values closer to one indicate better predictive performance, whereas lower loss values indicate better model convergence. The reported overall classification results represent the mean and standard deviation obtained across the five independent runs. To further analyze the learning behavior of each architecture, the weighted cross-entropy loss was employed as the optimization objective during training, while training and validation loss curves were monitored to assess model convergence. The model corresponding to the lowest validation loss was selected as the final model for testing.
In addition to prediction performance, the computational efficiency of each architecture was also investigated. The number of trainable parameters was reported as a measure of model complexity, while inference time was measured to assess the time required to generate predictions for unseen retinal fundus images. These complementary measures provide practical insight into the trade-off between predictive performance and computational cost, which is particularly important when selecting models for real-world diabetic retinopathy screening systems.

5. Experimental Results and Comparative Analysis

This section presents the experimental results obtained from the six pretrained deep learning models for five-class diabetic retinopathy grading. The performance of the CNN-based and transformer-based architectures is evaluated using the metrics described in Section 4.3 under the same experimental configuration to ensure a fair comparison. The analysis focuses on three aspects: overall classification performance, learning behavior during training, and computational efficiency. Finally, the strengths and limitations of each architecture are discussed to provide practical insights into selecting appropriate deep learning models for automated DR screening.

5.1. Overall Classification Performance

The overall classification performance of the evaluated CNN-based and transformer-based models on the independent test set using accuracy, precision, recall, F1-score, AUC, and QWK is reported in Table 4 and illustrated in Figure 3. To provide a comprehensive evaluation under the imbalanced class distribution of the APTOS 2019 dataset, both macro-average and weighted-average performance metrics are reported. Macro-average metrics assign equal importance to all diabetic retinopathy severity classes, regardless of sample size, making them more suitable for evaluating recognition of minority classes. In contrast, weighted-average metrics account for the class distribution by assigning greater importance to majority classes, thereby reflecting the overall predictive performance on the complete dataset.
Overall, the transformer-based models outperformed the CNN-based models across most evaluation metrics. Among the CNN architectures, EfficientNet-B0 achieved the strongest performance, with an accuracy with an accuracy of 0.791, QWK of 0.850, macro F1-score of 0.604, and weighted F1-score of 0.793. This indicates that EfficientNet-B0 was the most effective CNN baseline, outperforming both ResNet50 and MobileNetV2. ResNet50 showed the weakest performance, particularly in macro F1-score (0.479), suggesting limited ability to correctly classify minority DR severity grades.
Among the transformer-based models, Swin-Tiny achieved the highest accuracy (0.823), QWK (0.898), macro precision (0.667), macro F1-score (0.644), weighted recall (0.823), and weighted F1-score (0.821), demonstrating the strongest overall classification performance. The highest QWK value further indicates the strongest agreement between the predicted and true DR severity grades while accounting for the ordinal nature of the classification task. Swin Transformer achieved the highest macro recall (0.640), weighted precision (0.832), and weighted AUC (0.966), while obtaining a QWK (0.897) comparable to that of Swin-Tiny. ViT also demonstrated competitive performance, particularly in terms of weighted AUC (0.957) and weighted F1-score (0.809). The relatively small differences between ViT, Swin-Tiny, and Swin Transformer indicate that transformer-based architectures performed comparably, with Swin-based architectures showing the strongest overall balance across metrics. The small standard deviation values across the five separate runs indicate generally stable performance across different random seeds. Particularly, QWK showed low variability among the transformer-based models, with standard deviations ranging from 0.010 to 0.015. Although some variation was observed across individual classification metrics, the overall performance patterns were still consistent across the repeated runs.
A noticeable difference is observed between the macro-average and weighted-average metrics across all evaluated models. In general, the weighted-average values were consistently higher than their corresponding macro-average values. This behavior is expected due to the imbalanced distribution of the APTOS 2019 dataset, where the majority classes contribute more heavily to the weighted evaluation metrics. For example, ResNet50 achieved a weighted F1-score of 0.672, whereas its macro F1-score decreased to 0.479, indicating substantially weaker performance on the minority DR severity grades. Similar trends were observed for EfficientNet-B0 and MobileNetV2. In comparison, the transformer-based models exhibited smaller gaps between the macro and weighted metrics, suggesting more balanced classification performance across both majority and minority classes.
The AUC and QWK results further support these findings. All transformer-based models achieved weighted AUC values above 0.95, substantially higher than those of the CNN-based architectures. Similarly, the transformer models obtained the highest QWK values, indicating stronger agreement between the predicted and true DR severity grades while accounting for the ordinal relationship between disease stages. The consistently higher AUC and QWK values suggest that transformer models possess stronger discriminative power in separating DR severity grades across a wide range of decision thresholds. Furthermore, the relatively small gap between macro and weighted AUC values indicate that transformer-based architectures maintain reliable class discrimination even for minority classes, whereas CNN models exhibit greater sensitivity to class imbalance.
Finally, to determine whether the observed performance differences were statistically meaningful, statistical analysis was conducted across the five independent runs using QWK, macro F1-score, and weighted F1-score. The Friedman test [42] was first applied to assess overall differences among the six models, followed by pairwise Wilcoxon signed-rank tests with Holm correction [42] to account for multiple comparisons. The Friedman analysis identified significant overall differences among the evaluated architectures for all three metrics (Holm-adjusted p < 0.05), confirming that model architecture had a measurable effect on classification performance. Post-hoc analysis showed that the differences among individual leading models were not statistically significant after correction, indicating comparable performance among the strongest architectures despite differences in their mean values. Swin-Tiny nevertheless achieved the highest mean QWK, macro F1-score, and weighted F1-score across the five runs.

5.2. Per-Class Performance Analysis

To further examine the classification behavior of the best-performing CNN-based model (EfficientNet-B0) and transformer-based model (Swin-Tiny), a per-class analysis was conducted using precision, recall, and F1-score for each diabetic retinopathy (DR) severity grade. The per-class performance metrics reported in Table 5 correspond to the mean and standard deviation obtained across five independent runs. To provide a qualitative visualization of the classification behavior, the corresponding confusion matrices from a representative test run are presented in Figure 4.
Overall, Swin-Tiny demonstrated stronger per-class performance than EfficientNet-B0 across nearly all disease severity grades and evaluation metrics. The main exception was Mild DR recall, where EfficientNet-B0 achieved a higher value (0.553 vs. 0.506), while performance for the No DR class was largely comparable between the two models. Both models achieved excellent performance in identifying the majority class “No DR”, with F1-scores exceeding 0.97, indicating reliable discrimination of healthy retinal images. Performance was lower for the Mild and Moderate classes, reflecting the greater difficulty of distinguishing intermediate DR severity levels. For Mild DR, Swin-Tiny achieved higher precision (0.544) and F1-score (0.517), whereas EfficientNet-B0 achieved higher recall (0.553 vs. 0.506), indicating that it correctly identified a larger proportion of actual Mild DR cases and missed fewer cases at this severity level. For Moderate DR, Swin-Tiny showed a clearer advantage, achieving higher precision (0.710), recall (0.793), and F1-score (0.748) than EfficientNet-B0.
The Severe DR class remained the most challenging for both architectures. Swin-Tiny achieved slightly higher precision (0.368), recall (0.400), and F1-score (0.380) than EfficientNet-B0. The variability observed across the independent runs further indicates that performance estimates for this less represented class were less stable than those for the No DR class. Swin-Tiny also achieved higher mean performance for Proliferative DR, with precision of 0.740 compared with 0.647 for EfficientNet-B0, recall of 0.503 compared with 0.430, and F1-score of 0.598 compared with 0.515. However, the relatively large standard deviations for several Proliferative DR metrics indicate notable variation across the five runs. Overall, the standard deviations reported in Table 5 show greater variability for several of the less represented DR severity grades compared with the No DR class. This finding highlights the influence of class representation on the stability of per-class performance estimates and provides additional context for the observed differences between the two architectures.
The representative confusion matrices (see Figure 4) further illustrate the class-specific classification patterns. A substantial proportion of the classification errors involved neighboring DR severity grades, consistent with the gradual progression of diabetic retinopathy and the visual similarity between adjacent stages. Both models correctly identified the majority of No DR cases. Mild and Severe DR were frequently misclassified as Moderate, with Moderate also accounting for several misclassifications of Proliferative DR. Swin-Tiny correctly classified more Moderate (70 vs. 58) and Proliferative DR cases (19 vs. 14) than EfficientNet-B0 in the representative run, whereas EfficientNet-B0 correctly identified slightly more Mild cases (17 vs. 15). These patterns are generally consistent with the mean per-class results reported across the five independent runs in Table 5.

5.3. Learning Behavior and Convergence

To further analyze the training behavior of the evaluated models, the training and validation loss curves were examined throughout the training process. Figure 5 presents the learning curves from a representative experimental run (seed 123). As shown in Figure 5, all models exhibited a rapid decrease in training loss during the first few epochs, indicating successful optimization. The CNN-based models showed relatively smooth validation curves, whereas the transformer-based models exhibited larger fluctuations in validation loss after the initial training stage, indicating greater sensitivity to continued optimization. As training progressed, the gap between the training and validation losses gradually increased, particularly for the transformer-based models, suggesting increased overfitting during later epochs.
Among the CNN-based models, EfficientNet-B0 and MobileNetV2 showed a relatively rapid reduction in training loss, while ResNet50 exhibited a more gradual convergence pattern. The transformer-based models also converged rapidly, reaching very low training losses within the first few epochs. However, unlike the CNN-based models, their validation losses generally reached their minimum early and subsequently fluctuated or increased despite continued reductions in training loss. This pattern indicates that additional training increasingly optimized the training data without corresponding improvements in validation loss, indicating varying degrees of overfitting across transformer-based architectures.
To further characterize convergence behavior, Table 6 reports the minimum validation loss and its corresponding epoch for the representative run. The transformer-based models achieved their minimum validation losses much earlier than the CNN-based models. Specifically, ViT and Swin Transformer reached their minimum validation losses at epoch 3, while Swin-Tiny reached its minimum at epoch 5. In comparison, ResNet50, EfficientNet-B0, and MobileNetV2 reached their minimum validation losses at epochs 16, 9, and 16, respectively. These results indicate that transformer-based models learn discriminative features rapidly but begin to overfit earlier, while CNN-based models require more training epochs to reach their best generalization performance. The observed differences in convergence behavior also reflect the distinct characteristics of the evaluated CNN- and transformer-based architectures, which may respond in a different way to the common training configuration and reach their optimal validation performance at different stages of training. Early stopping was triggered for ViT, Swin-Tiny, and Swin Transformer at epochs 19, 16, and 19, respectively, whereas all CNN-based models completed the maximum 20 epochs.

5.4. Computational Efficiency Analysis

To compare the computational efficiency of the evaluated models, the number of trainable parameters and the inference latency per image were analyzed. Inference latency was measured using a batch size of one under the same hardware environment for all models, following warm-up iterations and repeated measurements. The latency results are reported as mean ± standard deviation across five independent runs. The results are summarized in Table 7.
Among the CNN-based models, MobileNetV2 exhibited the smallest model size, requiring only 2.23 million trainable parameters while achieving the lowest inference latency within the CNN family (7.023 ms/image). EfficientNet-B0 also demonstrated a compact model size, requiring 4.01 M parameters but exhibited a higher inference latency (9.835 ms/image) than both MobileNetV2 and ResNet50. ResNet50 despite requiring a substantially larger model size (23.52 M parameters), achieved an intermediate inference latency of (7.730 ms/image).
The transformer-based models required considerably more parameters than the CNN models. ViT and Swin Transformer contained approximately 86 M parameters, while Swin-Tiny reduced the model size to 27.52 M parameters. Interestingly, ViT achieved the lowest inference latency among all evaluated models despite having one of the largest model sizes. This observation suggests that inference latency is not solely determined by the number of trainable parameters but is also influenced by the underlying network architecture and its execution efficiency on the target hardware.
Overall, the results demonstrate that model size and inference latency are not directly proportional. While lightweight architectures generally reduce the number of trainable parameters, they do not necessarily achieve the lowest inference latency. These findings highlight the importance of considering predictive performance, model complexity, and inference efficiency when selecting models for automated diabetic retinopathy screening.

5.5. Comparison with Related Studies

To provide additional context for the obtained results, the best-performing model in this study was compared with selected studies that also addressed five-class DR grading using the APTOS 2019 dataset. The selected studies employed pretrained or hybrid DL approaches with additional strategies, including Siamese learning [43], patch-based feature extraction [44], and hybrid CNN architectures [45]. The reported accuracy, F1-score, and QWK values are summarized in Table 8, where available. However, several of these studies report a limited set of evaluation metrics, with accuracy frequently serving as the primary performance measure, while F1-score and QWK are not consistently reported. For an imbalanced multi-class dataset such as APTOS 2019, accuracy alone may not fully reflect performance across DR severity grades, particularly the less represented classes. Therefore, Table 8 presents accuracy, F1-score, and QWK for comparison where these measures are available.
As shown in Table 8, the performance of Swin-Tiny is competitive with the selected related studies, although some approaches achieved higher accuracy. These differences may be attributed to the use of additional task-specific strategies, including Siamese learning, patch-based processing, and hybrid architectures, as well as differences in the objectives and experimental designs of the studies. The comparison is also constrained by the limited availability of complementary evaluation metrics in some studies, making it difficult to assess performance beyond overall accuracy. In comparison, the present study considers multiple overall and class-sensitive metrics, including F1-score and QWK, together with per-class performance, computational efficiency, and performance stability across five independent runs. This broader evaluation provides additional insight into model behavior beyond a single aggregate performance measure.

5.6. Discussion and Limitations

The experimental results consistently demonstrate that transformer-based architectures outperform CNN-based models for five-class diabetic retinopathy grading. Among the evaluated architectures, Swin-Tiny achieved the highest accuracy, QWK, macro F1-score, and weighted F1-score, while Swin Transformer achieved the highest macro recall, weighted precision, and AUC. ViT also remained competitive across the evaluated metrics. Among the CNN models, EfficientNet-B0 provided the strongest overall performance. The per-class analysis further supported these findings, with Swin-Tiny achieving higher mean performance than EfficientNet-B0 across most DR severity grades and evaluation metrics. However, greater variability was observed for the less represented classes, particularly Severe DR, highlighting the greater difficulty of distinguishing these disease stages. These improvements can be attributed to the self-attention mechanism, which captures both local lesion characteristics and long-range contextual relationships within retinal fundus images.
The computational analysis further demonstrates that model size and inference latency are not directly proportional. MobileNetV2 had the smallest parameter count and the lowest inference latency among the CNN models, while ViT achieved the lowest overall inference latency despite its substantially larger model size. In contrast, Swin Transformer had the largest computational requirements, whereas Swin-Tiny provided a more favorable balance between predictive performance and model complexity, achieving the strongest overall classification performance with substantially fewer parameters than the larger Swin Transformer. These findings suggest that the choice of architecture should depend on the intended deployment scenario. When the primary objective is classification performance, transformer-based models provide a strong choice, with Swin-Tiny achieving the strongest overall performance in this study. However, when model size and computational resources are critical, lightweight CNN models such as MobileNetV2 provide an effective alternative with fewer parameters and competitive classification performance. The inference results further indicate that latency is architecture-dependent and does not necessarily increase with model size, as ViT achieved the lowest inference latency among the evaluated models.
Despite the comprehensive analysis we presented, this study has some limitations. First, the experiments were conducted using a single benchmark dataset, which limits conclusions regarding generalizability across different populations, imaging devices, and acquisition conditions. Second, the imbalanced distribution of APTOS 2019 resulted in relatively few test samples for some DR severity grades, increasing the variability of class-specific performance estimates. Several solutions can be explored for addressing class imbalance and improving the recognition of underrepresented DR severity grades as future direction. Finally, the study employed a common training configuration across architectures rather than architecture-specific hyperparameter optimization, as the primary objective was to provide a controlled architectural comparison. Finally, the study employed a common training configuration across architectures rather than architecture-specific hyperparameter optimization, as the primary objective was to provide a controlled architectural comparison. Future evaluation across additional retinal datasets and optimized architecture-specific configurations would further establish the generalizability of the observed findings.

6. Conclusion

This study presented a comparative analysis of six pretrained CNN-based and transformer-based models for five-class diabetic retinopathy grading using retinal fundus images. All models were evaluated under the same preprocessing, training, and evaluation settings to ensure a fair comparison. The results demonstrated that transformer-based architectures consistently outperformed CNN-based models in overall classification performance. Among the evaluated models, Swin-Tiny achieved the strongest overall classification performance, while EfficientNet-B0 provided the strongest CNN baseline. In addition, the study highlighted the relationship between predictive performance and computational efficiency, demonstrating that model size and inference latency are not necessarily proportional and that architecture selection should consider both predictive and computational requirements.
The comprehensive evaluation across overall and per-class metrics further demonstrated the challenges associated with distinguishing less represented DR severity grades, particularly Severe DR. Future work will extend the evaluation to additional retinal datasets to evaluate the generalizability of these findings and investigate more recent architectures and architecture-specific optimization. Data augmentation and advanced imbalance-aware learning strategies can also be explored as future direction to improve the recognition of underrepresented DR severity grades. Addressing these aspects can further improve and support the development of reliable automated diabetic retinopathy screening systems.

Author Contributions

Conceptualization, MAT; methodology, MAT; software, MAT; validation, MAT; formal analysis, MAT; investigation, MAT; resources, MAT; data curation, MAT; writing—original draft preparation, MAT; writing—review and editing, MAT; visualization, MAT; supervision, MAT; project administration, MAT. The author has read and agreed to the published version of the manuscript.

Funding

The authors would like to acknowledge the Deanship of Graduate Studies and Scientific Research, Taif University, Saudi Arabia, for funding this work.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The APTOS 2019 dataset used in this study was downloaded from the MariaHerreroT Kaggle repository and is publicly available at: (https://www.kaggle.com/datasets/mariaherrerot/aptos2019). The dataset can also be accessed programmatically using the kagglehub Python package.

Acknowledgments

The authors would like to acknowledge the Deanship of Graduate Studies and Scientific Research, Taif University for funding this work.

Conflicts of Interest

The authors declare no conflicts of interest.:

Abbreviations

The following abbreviations are used in this manuscript:
AI Artificial Intelligence
APTOS Asia Pacific Tele-Ophthalmology Society
AUC Area Under the Receiver Operating Characteristic Curve
CLAHE Contrast Limited Adaptive Histogram Equalization
CNN Convolutional Neural Network
DR Diabetic Retinopathy
DL Deep Learning
FP False Positive
FN False Negative
ICDRSS International Clinical Diabetic Retinopathy Disease Severity Scale
QWK Quadratic Weighted Kappa
ROC Receiver Operating Characteristic
ReLU Rectified Linear Unit
RGB Red, Green, and Blue
TP True Positive
TN True Negative
TL Transfer learning
ViT Vision Transformer

References

  1. Kropp, M.; Golubnitschaja, O.; Mazurakova, A.; Koklesova, L.; Sargheini, N.; Vo, T.-T.K.S.; de Clerck, E.; Polivka, J., Jr.; Potuznik, P.; Polivka, J.; et al. Diabetic Retinopathy as the Leading Cause of Blindness and Early Predictor of Cascading Complications-Risks and Mitigation. EPMA J. 2023, 14, 21–42. [Google Scholar] [CrossRef] [PubMed]
  2. Maneu, V.; Lax, P.; Cuenca, N. Current and Future Therapeutic Strategies for the Treatment of Retinal Neurodegenerative Diseases. Neural Regen. Res. 2022, 17, 103–104. [Google Scholar] [CrossRef] [PubMed]
  3. Gettinger, K.; Lee, D.; Tomita, Y.; Negishi, K.; Kurihara, T. Diabetic Retinopathy, a Comprehensive Overview on Pathophysiology and Relevant Experimental Models. Int. J. Mol. Sci. 2025, 26. [Google Scholar] [CrossRef] [PubMed]
  4. Zheng, Y.; He, M.; Congdon, N. The Worldwide Epidemic of Diabetic Retinopathy. Indian J. Ophthalmol. 2012, 60, 428–431. [Google Scholar] [CrossRef] [PubMed]
  5. Abou Taha, A.; Dinesen, S.; Vergmann, A.S.; Grauslund, J. Present and Future Screening Programs for Diabetic Retinopathy: A Narrative Review. Int. J. Retin. Vitr. 2024, 10, 14. [Google Scholar] [CrossRef] [PubMed]
  6. Tsuneki, M. Deep Learning Models in Medical Image Analysis. J. Oral Biosci. 2022, 64, 312–320. [Google Scholar] [CrossRef] [PubMed]
  7. Hu, C.; Yueyue, W.; Joy, C.; Dongsheng, J.; Xiaopeng, Z.; Qi, T.; Manning, W. Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv 2021. [Google Scholar]
  8. Qari, S.; Thafar, M.A. Brain Stroke Classification Using CT Scans with Transformer-Based Models and Explainable AI. Diagnostics 2025, 15. [Google Scholar] [CrossRef] [PubMed]
  9. Saleh, G.A.; Batouty, N.M.; Haggag, S.; Elnakib, A.; Khalifa, F.; Taher, F.; Mohamed, M.A.; Farag, R.; Sandhu, H.; Sewelam, A.; et al. The Role of Medical Image Modalities and AI in the Early Detection, Diagnosis and Grading of Retinal Diseases: A Survey. Bioengineering 2022, 9. [Google Scholar] [CrossRef] [PubMed]
  10. Zhang, H.-Q.; Arif, M.; Thafar, M.A.; Albaradei, S.; Cai, P.; Zhang, Y.; Tang, H.; Lin, H. PMPred-AE: A Computational Model for the Detection and Interpretation of Pathological Myopia Based on Artificial Intelligence. Front. Med. 2025, 12, 1529335. [Google Scholar] [CrossRef] [PubMed]
  11. Kim, H.E.; Cosa-Linan, A.; Santhanam, N.; Jannesari, M.; Maros, M.E.; Ganslandt, T. Transfer Learning for Medical Image Classification: A Literature Review. BMC Med. Imaging 2022, 22, 69. [Google Scholar] [CrossRef] [PubMed]
  12. Das, D.; Biswas, S.K.; Bandyopadhyay, S. A Critical Review on Diagnosis of Diabetic Retinopathy Using Machine Learning and Deep Learning. Multimed. Tools Appl. 2022, 81, 25613–25655. [Google Scholar] [CrossRef] [PubMed]
  13. Suvalakshmi, S.; Vinoth Kumar, B. Diabetic Retinopathy Classification Using Transformer Models: An Comprehensive Survey. In Lecture Notes in Networks and Systems; Lecture Notes in Networks and Systems; Springer Nature Switzerland: Cham, 2026; pp. 58–72. ISBN 9783032062529. [Google Scholar]
  14. Bhulakshmi, D.; Rajput, D.S. A Systematic Review on Diabetic Retinopathy Detection and Classification Based on Deep Learning Techniques Using Fundus Images. PeerJ Comput. Sci. 2024, 10, e1947. [Google Scholar] [CrossRef] [PubMed]
  15. Lakshminarayanan, V.; Kheradfallah, H.; Sarkar, A.; Jothi Balaji, J. Automated Detection and Diagnosis of Diabetic Retinopathy: A Comprehensive Survey. J. Imaging 2021, 7, 165. [Google Scholar] [CrossRef] [PubMed]
  16. Mutua, E.N.; Kasamani, B.S.; Reich, C. Deep Learning Applications for Diabetic Retinopathy and Retinopathy of Prematurity Diseases Diagnosis: A Systematic Review. Int. J. Ophthalmol. 2025, 18, 1594–1602. [Google Scholar] [CrossRef] [PubMed]
  17. Khan, Z.; Khan, F.G.; Khan, A.; Rehman, Z.U.; Shah, S.; Qummar, S.; Ali, F.; Pack, S. Diabetic Retinopathy Detection Using VGG-NIN a Deep Learning Architecture. IEEE Access 2021, 9, 61408–61416. [Google Scholar] [CrossRef]
  18. Saeed, F.; Hussain, M.; Aboalsamh, H.A. Automatic Diabetic Retinopathy Diagnosis Using Adaptive Fine-Tuned Convolutional Neural Network. IEEE Access 2021, 9, 41344–41359. [Google Scholar] [CrossRef]
  19. Sajid, M.Z.; Hamid, M.F.; Youssef, A.; Yasmin, J.; Perumal, G.; Qureshi, I.; Naqi, S.M.; Abbas, Q. DR-NASNet: Automated System to Detect and Classify Diabetic Retinopathy Severity Using Improved Pretrained NASNet Model. Diagnostics 2023, 13, 2645. [Google Scholar] [CrossRef] [PubMed]
  20. Kommaraju, R.; Anbarasi, M.S. Diabetic Retinopathy Detection Using Convolutional Neural Network with Residual Blocks. Biomed. Signal Process. Control 2024, 87, 105494. [Google Scholar] [CrossRef]
  21. Chawla, M. Enhancing Diabetic Retinopathy Detection Using an Optimized MobileNet Architecture. Biomed. Signal Process. Control 2026. [Google Scholar] [CrossRef]
  22. Chilukoti, S.V.; Maida, A.S.; Hei, X. Diabetic Retinopathy Detection Using Transfer Learning from Pre-Trained Convolutional Neural Network Models. IEEE J Biomed Heal Informatics 2022. [Google Scholar] [CrossRef]
  23. Asia, A.-O.; Zhu, C.-Z.; Althubiti, S.A.; Al-Alimi, D.; Xiao, Y.-L.; Ouyang, P.-B.; Al-Qaness, M.A.A. Detection of Diabetic Retinopathy in Retinal Fundus Images Using CNN Classification Models. Electronics 2022, 11, 2740. [Google Scholar] [CrossRef]
  24. Hemanth, S.V.; Alagarsamy, S.; Rajkumar, T.D. A Novel Deep Learning Model for Diabetic Retinopathy Detection in Retinal Fundus Images Using Pre-Trained CNN and HWBLSTM. J. Biomol. Struct. Dyn. 2024, 43, 1–19. [Google Scholar] [CrossRef] [PubMed]
  25. Sagenela, V.K.; Gidijala, K.; Karanam, S.R.; Sagar, N.S.S. Artificial Intelligence in Diabetic Retinopathy Detection: A Decade of Progress from Machine Learning to Transformer-Based Frameworks. Arch. Comput. Methods Eng. 2026. [Google Scholar] [CrossRef]
  26. Nazih, W.; Aseeri, A.O.; Atallah, O.Y.; El-Sappagh, S. Vision Transformer Model for Predicting the Severity of Diabetic Retinopathy in Fundus Photography-Based Retina Images. IEEE Access 2023, 1–1. [Google Scholar] [CrossRef]
  27. Ohri, K.; Kumar, M.; Sukheja, D. Correction: Self-Supervised Approach for Diabetic Retinopathy Severity Detection Using Vision Transformer. Prog. Artif. Intell. 2024, 13, 185–186. [Google Scholar] [CrossRef]
  28. Ahmed, F.; Uddin, M.D.J. OcuViT: A Vision Transformer-Based Approach for Automated Diabetic Retinopathy and AMD Classification. J. Imaging Inf. Med. 2025. [Google Scholar] [CrossRef] [PubMed]
  29. Yang, Y.; Cai, Z.; Qiu, S.; Xu, P. A Novel Transformer Model with Multiple Instance Learning for Diabetic Retinopathy Classification. IEEe Access 2024, 12, 6768–6776. [Google Scholar] [CrossRef]
  30. Palaniappan, D.; Mohan, N.R.R.; Premavathi, T.; Gulzar, Y.; Ali, M.; Mir, M.S. Swin-DRNet: A Robust Transformer Framework for Diabetic Retinopathy Screening under Heterogeneous Imaging Conditions. Sci. Rep. 2026, 1–26. [Google Scholar] [CrossRef] [PubMed]
  31. Sushith, M.; Lakkshmanan, A.; Saravanan, M.; Castro, S. Attention Dual Transformer with Adaptive Temporal Convolutional for Diabetic Retinopathy Detection. Sci. Rep. 2025, 15, 7694. [Google Scholar] [CrossRef] [PubMed]
  32. Bala, R.; Sharma, A.; Goel, N. CTNet: Convolutional Transformer Network for Diabetic Retinopathy Classification. Neural Comput. Appl. 2024, 36, 4787–4809. [Google Scholar] [CrossRef]
  33. Sekar, P.; Bhoopalan, R.; Nagaprasad, N.; Mamo, T.R.; Dhanabal, S.P.; Krishnaraj, R. D-TNet: A Hybrid Dense Net-Transformer Model for Robust Diabetic Retinopathy Detection. Sci. Rep. 2025, 15, 39594. [Google Scholar] [CrossRef] [PubMed]
  34. APTOS 2019 Blindness Detection. Available online: https://kaggle.com/aptos2019-blindness-detection (accessed on 11 November 2025).
  35. Liu, R.; Wang, X.; Wu, Q.; Dai, L.; Fang, X.; Yan, T.; Son, J.; Tang, S.; Li, J.; Gao, Z.; et al. DeepDRiD: Diabetic Retinopathy-Grading and Image Quality Estimation Challenge. Patterns (N Y) 2022, 3, 100512. [Google Scholar] [CrossRef] [PubMed]
  36. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition.
  37. Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv 2019, 6105–6114. [Google Scholar]
  38. Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.-C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE, June 2018; pp. 4510–4520. [Google Scholar]
  39. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020. [Google Scholar]
  40. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, October 2021; pp. 10012–10022. [Google Scholar]
  41. de la Torre, J.; Puig, D.; Valls, A. Weighted Kappa Loss Function for Multi-Class Classification of Ordinal Data in Deep Learning. Pattern Recognit. Lett. 2018, 105, 144–154. [Google Scholar] [CrossRef]
  42. Rainio, O.; Teuho, J.; Klén, R. Evaluation Metrics and Statistical Tests for Machine Learning. Sci. Rep. 2024, 14, 6086. [Google Scholar] [CrossRef] [PubMed]
  43. Tariq, M.; Palade, V.; Ma, Y. Effective Diabetic Retinopathy Classification with Siamese Neural Network: A Strategy for Small Dataset Challenges. IEEE Access 2024, 12, 182814–182827. [Google Scholar] [CrossRef]
  44. Kobat, S.G.; Baygin, N.; Yusufoglu, E.; Baygin, M.; Barua, P.D.; Dogan, S.; Yaman, O.; Celiker, U.; Yildirim, H.; Tan, R.-S.; et al. Automated Diabetic Retinopathy Detection Using Horizontal and Vertical Patch Division-Based Pre-Trained DenseNET with Digital Fundus Images. Diagnostics 2022, 12, 1975. [Google Scholar] [CrossRef] [PubMed]
  45. Yasashvini; Raja Sarobin M., V.; Panjanathan, R.; Jasmine, G.; Anbarasi, J. Diabetic Retinopathy Classification Using CNN and Hybrid Deep Convolutional Neural Networks. Symmetry 2022, 14, 1932. [Google Scholar] [CrossRef]
Figure 1. Overall workflow of the proposed framework for multi-class diabetic retinopathy grading.
Figure 1. Overall workflow of the proposed framework for multi-class diabetic retinopathy grading.
Preprints 228061 g001
Figure 2. Representative retinal fundus images illustrating the five diabetic retinopathy (DR) severity grades from the APTOS 2019 dataset.
Figure 2. Representative retinal fundus images illustrating the five diabetic retinopathy (DR) severity grades from the APTOS 2019 dataset.
Preprints 228061 g002
Figure 3. Comparison of the overall classification performance of the evaluated CNN-based and transformer-based models on the independent test set in terms of Accuracy, Macro F1-score, and Weighted F1-score. Error bars indicate variability across five independent runs.
Figure 3. Comparison of the overall classification performance of the evaluated CNN-based and transformer-based models on the independent test set in terms of Accuracy, Macro F1-score, and Weighted F1-score. Error bars indicate variability across five independent runs.
Preprints 228061 g003
Figure 4. Confusion matrices from a representative run (seed 42) for EfficientNet-B0 and Swin-Tiny on the independent test set for five-class DR grading.
Figure 4. Confusion matrices from a representative run (seed 42) for EfficientNet-B0 and Swin-Tiny on the independent test set for five-class DR grading.
Preprints 228061 g004
Figure 5. Representative training and validation loss curves from a single experimental run (seed 123) of the evaluated CNN-based and transformer-based models over a maximum of 20 training epochs.
Figure 5. Representative training and validation loss curves from a single experimental run (seed 123) of the evaluated CNN-based and transformer-based models over a maximum of 20 training epochs.
Preprints 228061 g005
Table 1. Distribution of retinal fundus images across the five diabetic retinopathy severity grades and the predefined training, validation, and testing subsets of the APTOS 2019 dataset.
Table 1. Distribution of retinal fundus images across the five diabetic retinopathy severity grades and the predefined training, validation, and testing subsets of the APTOS 2019 dataset.
Class Train Validation Test Total Samples of each class
0 - No DR 1434 172 199 1,805
1- Mild DR 300 40 30 370
2 - Moderate DR 808 104 87 999
3 - Severe DR 154 22 17 193
4- Proliferative DR 234 28 33 295
Total samples in each split 2,930 366 366 3,662
Table 2. Hyperparameter configuration adopted during model training and fine-tuning for all evaluated architectures.
Table 2. Hyperparameter configuration adopted during model training and fine-tuning for all evaluated architectures.
Hyperparameter Value
Input image size 224 × 224
Optimizer AdamW
Learning rate 1×10−4
Batch size 32
Number of Epochs 20 with patience = 6
Loss function Weighted Cross Entropy
Sampling strategy WeightedRandomSampler
Table 3. Definitions and mathematical formulations of the prediction evaluation metrics used in this study.
Table 3. Definitions and mathematical formulations of the prediction evaluation metrics used in this study.
Metrics Definition Mathematical Formula
Accuracy Represents the proportion of correctly classified samples among all evaluated images. Accuracy =
(TP + TN)/(TP + TN + FP + FN)
Precision Measures the reliability of positive predictions by indicating how many samples predicted as positive are actually positive. Precision = TP/(TP + FP)
Recall (Sensitivity) Measures how effectively the model identifies positive samples from all samples belonging to the positive class. It is also called true positive rate (TPR) Recall = TP/(TP + FN)
F1-score Harmonic mean of precision and recall, providing a balanced evaluation of both metrics. F1 = 2 × (Precision × Recall) / (Precision + Recall)
AUC-ROC Measures the model’s ability to discriminate between classes across different classification thresholds using the true positive rate (TPR) and false positive rate (FPR). AUC = 0 1 TPR(FPR) d(FPR)
QWK Measures the agreement between the predicted and actual DR severity grades while accounting for the ordinal relationship between classes using the observed agreement matrix (O), the expected agreement matrix (E), and the quadratic weight matrix (W). QWK = 1 − Σ(WO) / Σ(WE)}
Table 4. Overall classification performance (mean ± standard deviation across five independent runs) of the evaluated CNN-based and transformer-based models on the independent test set. The best mean value for each metric is highlighted in bold.
Table 4. Overall classification performance (mean ± standard deviation across five independent runs) of the evaluated CNN-based and transformer-based models on the independent test set. The best mean value for each metric is highlighted in bold.
Model Accuracy QWK Macro Average Weighted Average
Precision Recall F1 score AUC Precision Recall F1 score AUC
ResNet50 0.684 ± 0.012 0.801 ± 0.013 0.546 ± 0.013 0.578 ± 0.016 0.479 ± 0.008 0.878 ± 0.007 0.804 ± 0.017 0.684 ± 0.011 0.672 ± 0.006 0.927 ± 0.006
EfficientNet-B0 0.791 ± 0.022 0.850 ± 0.032 0.616 ± 0.038 0.605 ± 0.039 0.604 ± 0.040 0.915 ± 0.014 0.800 ± 0.021 0.791± 0.022 0.793 ± 0.022 0.953 ± 0.007
MobileNetV2 0.737 ± 0.009 0.817 ± 0.015 0.565 ± 0.015 0.601 ± 0.019 0.568 ± 0.019 0.904 ± 0.009 0.779 ± 0.006 0.737 ± 0.009 0.750 ± 0.007 0.941± 0.004
ViT 0.807 ± 0.024 0.876 ± 0.010 0.648 ± 0.041 0.635 ± 0.015 0.629 ± 0.027 0.918 ± 0.010 0.825 ± 0.017 0.807 ± 0.024 0.809 ± 0.021 0.957 ± 0.003
Swin-Tiny 0.823 ± 0.015 0.898 ± 0.010 0.667 ± 0.028 0.636 ± 0.013 0.644 ± 0.020 0.931 ± 0.008 0.826 ± 0.010 0.823 ± 0.015 0.821 ± 0.012 0.964 ± 0.005
Swin Transformer 0.810 ± 0.034 0.897 ± 0.015 0.658 ± 0.035 0.640 ± 0.020 0.632 ± 0.033 0.935 ± 0.013 0.832 ± 0.008 0.810 ± 0.034 0.811 ± 0.032 0.966 ± 0.004
Table 5. Per-class classification performance (mean ± standard deviation across five independent runs) of the best-performing CNN-based model (EfficientNet-B0) and transformer-based model (Swin-Tiny) on the independent test set. Bold indicates the best mean value for each metric within each DR severity grade.
Table 5. Per-class classification performance (mean ± standard deviation across five independent runs) of the best-performing CNN-based model (EfficientNet-B0) and transformer-based model (Swin-Tiny) on the independent test set. Bold indicates the best mean value for each metric within each DR severity grade.
DR Severity Grade EfficientNet-B0 Swin-Tiny
Precision Recall F1-score Precision Recall F1-score
No DR 0.978 ± 0.009 0.969 ± 0.004 0.973 ± 0.002 0.974 ± 0.010 0.973 ± 0.002 0.973 ± 0.005
Mild 0.449 ± 0.073 0.553 ± 0.004 0.494 ± 0.058 0.544 ± 0.103 0.506 ± 0.125 0.517 ± 0.031
Moderate 0.660 ± 0.043 0.682 ± 0.048 0.671 ± 0.041 0.710 ± 0.032 0.793 ± 0.202 0.748 ± 0.122
Severe 0.344 ± 0.067 0.388 ± 0.089 0.364 ± 0.072 0.368 ± 0.076 0.400 ± 0.101 0.380 ± 0.037
Proliferative 0.647 ± 0.133 0.430 ± 0.091 0.515 ± 0.10 0.740 ± 0.076 0.503 ± 0.107 0.598 ± 0.074
Table 6. Best validation loss, corresponding training epoch, and early stopping information for the evaluated DL models in the representative run (seed 123).
Table 6. Best validation loss, corresponding training epoch, and early stopping information for the evaluated DL models in the representative run (seed 123).
Model Best Validation Loss Epoch of Best Model Early Stopping
ResNet50 1.1598 16 No
EfficientNet-B0 1.1008 9 No
MobileNetV2 1.0728 16 No
ViT 0.9462 3 Yes (epoch 19)
Swin-Tiny 0.8369 5 Yes (epoch 16)
Swin Transformer 0.9000 3 Yes (epoch 19)
Table 7. Model complexity and inference efficiency of the evaluated CNN-based and transformer-based architectures. Inference latency is reported as mean ± standard deviation across five independent runs.
Table 7. Model complexity and inference efficiency of the evaluated CNN-based and transformer-based architectures. Inference latency is reported as mean ± standard deviation across five independent runs.
Model Number of Parameters Mean Inference Latency (ms/image)
ResNet50 23,518,277 7.730 ± 0.083
EfficientNet-B0 4,013,953 9.835 ± 0.171
MobileNet 2,230,277 7.023 ± 0.181
ViT 85,802,501 6.225 ± 0.012
Swin Tiny 27,523,199 12.341 ± 0.175
Swin Transformer 86,748,349 21.647 ± 0.065
Table 8. Comparison of the proposed approach with selected related studies for five-class diabetic retinopathy grading using the APTOS 2019 dataset. NR: not reported.
Table 8. Comparison of the proposed approach with selected related studies for five-class diabetic retinopathy grading using the APTOS 2019 dataset. NR: not reported.
Study Model/Method Accuracy F1 score QWK
Tariq et al. (2024) [43] Siamese + VGG16 0.810 NR 0.890
Siamese + ResNet50 0.790 NR 0.860
Kobat et al. (2022) [44] Patch Division-Based + DenseNET 0.859 0.721 NR
Yasashvini et al. (2022) [45] CNN 0.756 NR NR
CNN + DenseNet 0.962 NR NR
Best Model in this study Swin-Tiny 0.823 0.821 0.898
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.