Submitted:
14 August 2025
Posted:
10 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Method: Adaptive Rate Bayesian Dropout (ARB-Dropout)
| Algorithm 1 ARB-Dropout Training |
|
| Algorithm 2 ARB-Dropout Single-Pass Inference with MC Dropout Baseline |
|
4. Experimental Setup
4.1. Datasets
- CIFAR-10 [6]: 60,000 RGB images of size from 10 classes (50,000 train, 10,000 test).
- CIFAR-100 [6]: Same format as CIFAR-10 but with 100 classes (600 images per class).
- SVHN [7]: Over 600,000 RGB digit images (classes 0–9) from real-world house numbers.
- STL-10 [8]: RGB images from 10 classes, resized to for our experiments (5,000 train, 8,000 test).
4.2. Architecture
- Two convolutional layers with ReLU activation and max pooling:
- Fully connected layer: .
- Classification head: (where K is the number of classes).
- Noise head: with activation to produce .
4.3. Training Details
4.4. Evaluation Protocol
- Accuracy: Standard classification accuracy.
-
Normalized Uncertainty:
- -
- For MC Dropout: mean predictive entropywhere is the mean softmax probability vector over T stochastic forward passes.
- -
- For ARB-Dropout: sum of normalized predictive entropy and normalized variance:where is the epistemic variance from analytic propagation, is the aleatoric variance from the noise head, and is a dataset-specific scaling constant.
- Negative Log-Likelihood (NLL): Measures probabilistic calibration via the log-loss on predicted probabilities.
- Brier Score: Mean squared difference between predicted probabilities and one-hot labels.
- Expected Calibration Error (ECE): Mean absolute difference between accuracy and confidence across equal-width probability bins.
- Inference Time: Average wall-clock time per test batch.
4.5. Implementation
5. Results

| Dataset | Method | Accuracy | ECE | Norm. Unc. | Time (s) |
|---|---|---|---|---|---|
| CIFAR-10 | MC Dropout | 0.6917 | 0.1904 | 0.1325 | 20.62 |
| ARB-Dropout | 0.6924 | 0.0760 | 2.0563 | 3.08 | |
| CIFAR-100 | MC Dropout | 0.3427 | 0.3382 | 0.2149 | 20.79 |
| ARB-Dropout | 0.3441 | 0.0878 | 1.6380 | 3.34 | |
| SVHN | MC Dropout | 0.8858 | 0.0793 | 0.0385 | 45.67 |
| ARB-Dropout | 0.8859 | 0.0399 | 2.1495 | 7.93 | |
| STL-10 | MC Dropout | 0.4901 | 0.2626 | 0.2834 | 15.75 |
| ARB-Dropout | 0.4876 | 0.0932 | 1.0735 | 4.79 |
6. Conclusion
References
- David J.C. MacKay. A practical Bayesian framework for backpropagation networks. Neural Computation, 4(3):448–472, 1992. [CrossRef]
- Yarin Gal and Zoubin Ghahramani. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML), pages 1050–1059. PMLR, 2016.
- Yarin Gal, Jiri Hron, and Alex Kendall. Concrete Dropout. In Advances in Neural Information Processing Systems, volume 30, 2017.
- Alex Kendall and Yarin Gal. What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision? In Advances in Neural Information Processing Systems, volume 30, 2017.
- Jeremiah Liu, Yin Lin, Suchismita Padhy, Dustin Tran, Tania Bedrax-Weiss, and Balaji Lakshminarayanan. Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness. In Advances in Neural Information Processing Systems, volume 33, pages 7498–7512, 2020.
- A. Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
- Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A.Y. Ng. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011.
- A. Coates, A. Ng, and H. Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, pages 215–223, 2011.
- D.P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
- C. Guo, G. Pleiss, Y. Sun, and K.Q. Weinberger. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), pages 1321–1330, 2017.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).