Electroencephalography (EEG) is a non-invasive technique for investigating brain activity; however, deep learning analysis remains challenging due to inter-subject variability and heterogeneous recording conditions. Most previous studies have focused on single disorders or datasets, while systematic evaluation across heterogeneous public datasets has received less attention. In this multi-dataset case study, a unified preprocessing, data harmonization, and leakage-free evaluation pipeline was applied to systematically evaluate four established deep learning architectures (CNN-1D, EEGNet, ShallowConvNet, and TCN) for subject-independent five-class EEG classification of healthy controls, Alzheimer’s disease, depression, Parkinson’s disease, and epilepsy. EEG recordings from OpenNeuro and PhysioNet included 144 subjects and 24,354 epochs. Standardized preprocessing, training-data-based channel-wise normalization, and subject-wise stratified partitioning were applied to prevent data leakage and ensure reliable evaluation. Five independent random seeds were used to assess result stability and reproducibility. Mean accuracy ranged from 73.14% to 83.66%, with macro-AUC values between 0.9389 and 0.9785. ShallowConvNet achieved the best performance, with an accuracy of 83.66 ± 4.01%, macro F1-score of 82.46 ± 3.68%, and macro-AUC of 0.9785±0.0102. Gradient-based saliency analysis identified EEG channels contributing to model predictions, providing insight into learned representations and supporting future investigations of potential EEG biomarkers. The findings indicate that the proposed evaluation pipeline enables a fair and systematic comparison of existing deep learning architectures across heterogeneous public EEG datasets and may support future validation on additional independent datasets.