Submitted:
21 October 2025
Posted:
22 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- (a)
- Over a fixed wavelet dictionary, linear-head. With explicit rates controlled by the eigenvalues of , we demonstrate global linear convergence to the unique ridge minimizer. This results in useful guidelines for considering () as a conditioning lever instead of just an anti-overfitting knob.
- (b)
- WNNs that are fully trainable (nonconvex). GD converges to stationary points and enjoys linear rates within regions meeting PL under the conditions of natural smoothness and boundedness for wavelets (within a restricted dilation/shift domain). We give implementable step-size limitations and demonstrate how L₂ dampens flat directions to widen PL basins.
- (c)
- regime that is over-parameterized (NTK). We extract rate constants related to the kernel spectrum induced by the wavelet dictionary and demonstrate that L₂ directs GD toward the minimum-RKHS-norm interpolant associated with the WNN-specific NTK.
1. Related Work
3. Preliminaries & Problem Setup
3.1. Wavelet Neural Network (WNN) Model
3.2. Assumptions (Wavelets, Data, and Loss)
3.3. Training Algorithms (GD/SGD with Weight Decay)
3.4. Problem Decompositions: Three Regimes
3.5. PL Inequality and Its Role
3.7. Step-Size and Regularization Prescriptions (Preview)
4. Mmethodology
4.1. Objective and Gradient Updates
4.2. Fixed-Feature (Ridge) Training of the Linear Head
4.3. Fully-Trainable WNN: Block GD, Schedules, and Stability
4.4. Choosing η and λ: Prescriptions and Diagnostics
5. Theoretical Results
5.1. Linear Convergence for the Fixed-Feature (Ridge) Regime
5.2. Fully-Trainable WNN Under a PL Inequality
6. Experiments & Evaluation
6.1. Datasets & Tasks
6.2. Metrics
6.3. Synthetic Regression (Approximation)
6.4. Denoising Robustness
6.5. Sensitivity to Learning Rate and Weight Decay
6.6. Learning Dynamics
6.7. Prediction Fidelity
6.8. Reproducibility Checklist
7. Discussion & Limitations
7.1. Practical Implications
7.2. Sensitivity and Stability
7.3. Robustness Under Distribution Shift
7.4. Limitations
7.5. Future Work
| η | λ | Val MSE | PSNR (dB) |
|---|---|---|---|
| 3e-3 | 3e-4 | 0.032 | 31.2 |
| 1e-3 | 1e-4 | 0.036 | 30.8 |
| 5e-3 | 1e-3 | 0.038 | 30.1 |
8. Conclusion and Future Directions
Future Directions
- (a)
- For canonical wavelet families (Mexican-hat, Morlet, and Daubechies), derive closed-form WNN-specific NTKs and examine their spectra with realistic initializations.
- (b)
- In order to quantify expansion as a function of λ, determine the conditions under which L₂ causes global or broader PL regions for trainable dilations/translations.
- (c)
- Create adaptive controllers with theoretical stability guarantees that simultaneously adjust η and λ utilizing real-time spectral/gradient diagnostics.
- (d)
- Use wavelet priors to expand the analysis to structured outputs (such as graphs and sequences) and classification losses (logistic and cross-entropy).
- (e)
- Examine robustness in the presence of adversarial perturbations and covariate shift, when wavelet localization might provide demonstrable stability benefits.
Author’s contribution
Funding
Acknowledgments
Conflicts of Interest
Appendix A
A.1 Smoothness and Descent
A.2 Ridge Objective: Conditioning and Rates
References
- Wu, J., Li, J., Yang, J. and Mei, S., 2025. Wavelet-integrated deep neural networks: A systematic review of applications and synergistic architectures. Neurocomputing, p.131648. [CrossRef]
- Kio, A.E., Xu, J., Gautam, N. and Ding, Y., 2024. Wavelet decomposition and neural networks: a potent combination for short term wind speed and power forecasting. Frontiers in Energy Research, 12, p.1277464. [CrossRef]
- Wang, P. and Wen, Z., 2024. A spatio-temporal graph wavelet neural network (ST-GWNN) for association mining in timely social media data. Scientific Reports, 14(1), p.31155. [CrossRef]
- Baharlouei, Z., Rabbani, H. and Plonka, G., 2023. Wavelet scattering transform application in classification of retinal abnormalities using OCT images. Scientific reports, 13(1), p.19013. [CrossRef]
- Garrigos, G. and Gower, R.M., 2023. Handbook of convergence theorems for (stochastic) gradient methods. arXiv preprint arXiv:2301.11235. [CrossRef]
- Xia, L., Massei, S. and Hochstenbach, M.E., 2025. On the convergence of the gradient descent method with stochastic fixed-point rounding errors under the Polyak–Łojasiewicz inequality. Computational Optimization and Applications, 90(3), pp.753-799. [CrossRef]
- Galanti, T., Siegel, Z.S., Gupte, A. and Poggio, T., 2022. SGD and weight decay provably induce a low-rank bias in neural networks.
- Tan, Y. and Liu, H., 2024. How does a kernel based on gradients of infinite-width neural networks come to be widely used: a review of the neural tangent kernel. International Journal of Multimedia Information Retrieval, 13(1), p.8. [CrossRef]
- Jacot, A., Gabriel, F. and Hongler, C., 2018. Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems, 31.
- Medvedev, M., Vardi, G. and Srebro, N., 2024. Overfitting behaviour of gaussian kernel ridgeless regression: Varying bandwidth or dimensionality. Advances in Neural Information Processing Systems, 37, pp.52624-52669.
- Somvanshi, S., Javed, S.A., Islam, M.M., Pandit, D. and Das, S., 2025. A survey on kolmogorov-arnold network. ACM Computing Surveys, 58(2), pp.1-35. [CrossRef]
- Sadoon, G.A.A.S., Almohammed, E. and Al-Behadili, H.A., 2025, January. Wavelet neural networks in signal parameter estimation: A comprehensive review for next-generation wireless systems. In AIP Conference Proceedings (Vol. 3255, No. 1, p. 020014). AIP Publishing LLC.
- Wang, P. and Wen, Z., 2024. A spatio-temporal graph wavelet neural network (ST-GWNN) for association mining in timely social media data. Scientific Reports, 14(1), p.31155. [CrossRef]
- Uddin, Z., Ganga, S., Asthana, R. and Ibrahim, W., 2023. Wavelets based physics informed neural networks to solve non-linear differential equations. Scientific Reports, 13(1), p.2882. [CrossRef]
- Imtiaz, T., 2022. Automatic cell nuclei segmentation in histopathology images using boundary preserving guided attention based deep neural network.
- Somvanshi, S., Javed, S.A., Islam, M.M., Pandit, D. and Das, S., 2025. A survey on kolmogorov-arnold network. ACM Computing Surveys, 58(2), pp.1-35. [CrossRef]
- Kilani, B.H., 2025. Convolutional Kolmogorov–Arnold Networks: a survey.
- Xiao, Q., Lu, S. and Chen, T., 2023. An alternating optimization method for bilevel problems under the Polyak-Łojasiewicz condition. Advances in Neural Information Processing Systems, 36, pp.63847-63873.
- Yazdani, K. and Hale, M., 2021. Asynchronous parallel nonconvex optimization under the polyak-łojasiewicz condition. IEEE Control Systems Letters, 6, pp.524-529. [CrossRef]
- Chen, K., Yi, C. and Yang, H., 2024. Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks. arXiv preprint arXiv:2410.02176. [CrossRef]
- Kobayashi, S., Akram, Y. and Von Oswald, J., 2024. Weight decay induces low-rank attention layers. Advances in Neural Information Processing Systems, 37, pp.4481-4510.
- Seleznova, M. and Kutyniok, G., 2022, April. Analyzing finite neural networks: Can we trust neural tangent kernel theory?. In Mathematical and Scientific Machine Learning (pp. 868-895). PMLR.
- Tan, Y. and Liu, H., 2024. How does a kernel based on gradients of infinite-width neural networks come to be widely used: a review of the neural tangent kernel. International Journal of Multimedia Information Retrieval, 13(1), p.8. [CrossRef]
- Tang, A., Wang, J.B., Pan, Y., Wu, T., Chen, Y., Yu, H. and Elkashlan, M., 2025. Revisiting XL-MIMO channel estimation: When dual-wideband effects meet near field. IEEE Transactions on Wireless Communications. [CrossRef]
- Cui, Z.X., Zhu, Q., Cheng, J., Zhang, B. and Liang, D., 2024. Deep unfolding as iterative regularization for imaging inverse problems. Inverse Problems, 40(2), p.025011. [CrossRef]












Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).