Submitted:
22 July 2026
Posted:
23 July 2026
You are already at the latest version
Abstract
Keywords:
MSC: 68T07; 57R35
1. Introduction
2. Related Work
3. Preliminaries
3.1. Real-Analytic Functions and the Identity Theorem
3.2. The Interpolating Manifold
4. Constraint Unlearnability
4.1. All-or-Nothing for Pointwise Constraints
4.2. Measure Zero of the Satisfying Parameters
4.3. Implications for Gradient Descent
5. Discussion
5.1. Inductive Bias in Practice
5.2. Implicit Regularization and Constraints
5.3. Limitations
Activation Functions
Pointwise vs. Functional Constraints
Exact vs. Approximate Satisfaction
Parameterization Condition
Connectivity of M
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Cooper, Y. Global minima of overparameterized neural networks. SIAM J. Math. Data Sci. 2021, 3(2), 676–691. [Google Scholar] [CrossRef]
- Min, Y.; Azizan, N. HardNet: Hard-constrained neural networks with universal approximation guarantees. In Advances in Neural Information Processing Systems (NeurIPS); 2025. [Google Scholar]
- Beucler, T.; Pritchard, M.; Rasp, S.; Ott, J.; Baldi, P.; Gentine, P. Enforcing analytic constraints in neural-networks emulating physical systems. Phys. Rev. Lett. 2021, 126(9), 098302. [Google Scholar] [CrossRef] [PubMed]
- Nguyen, Q.; Hein, M. The loss surface of deep and wide neural networks. Proc. Int. Conf. Mach. Learn. (ICML) 2017, PMLR 70, 2603–2612. [Google Scholar]
- Simsek, B.; Ged, F.; Jacot, A.; Spadaro, F.; Hongler, C.; Gerstner, W.; Brea, J. Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances. Proc. Int. Conf. Mach. Learn. (ICML) 2021, PMLR 139, 9722–9732. [Google Scholar]
- Madden, L.; Thrampoulidis, C. Memory capacity of two layer neural networks with smooth activations. SIAM J. Math. Data Sci. 2024, 6(3), 679–702. [Google Scholar] [CrossRef]
- Madden, L. Interpolation with deep neural networks with non-polynomial activations: necessary and sufficient numbers of neurons. arXiv 2024, arXiv:2405.13738. [Google Scholar]
- Constantinescu, V.-R.; Popescu, I. Approximation and interpolation of deep neural networks. arXiv 2024, arXiv:2304.10552. [Google Scholar]
- Vershynin, R. Memory capacity of neural networks with threshold and rectified linear unit activations. SIAM J. Math. Data Sci. 2020, 2(4), 1004–1033. [Google Scholar] [CrossRef]
- Yun, C.; Sra, S.; Jadbabaie, A. Small ReLU networks are powerful memorizers: A tight analysis of memorization capacity. In Advances in Neural Information Processing Systems (NeurIPS); 2019. [Google Scholar]
- Lu, L.; Pestourie, R.; Yao, W.; Wang, Z.; Verdugo, F.; Johnson, S.G. Physics-informed neural networks with hard constraints for inverse design. SIAM J. Sci. Comput. 2021, 43(6), B1105–B1132. [Google Scholar] [CrossRef]
- Zhong, F.; Fogarty, K.; Hanji, P.; Wu, T.; Sztrajman, A.; Spielberg, A.; Tagliasacchi, A.; Bosilj, P.; Oztireli, C. Neural fields with hard constraints of arbitrary differential order. In Advances in Neural Information Processing Systems (NeurIPS); 2023. [Google Scholar]
- Balestriero, R.; LeCun, Y. POLICE: Provably optimal linear constraint enforcement for deep neural networks. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023. [Google Scholar]
- Donti, P.L.; Rolnick, D.; Kolter, J.Z. DC3: A learning method for optimization with hard constraints. International Conference on Learning Representations (ICLR), 2021. [Google Scholar]
- Sontag, E.D. Critical points for least-squares problems involving certain analytic functions, with applications to sigmoidal nets. Adv. Comput. Math. 1996, 5(2–3), 245–268. [Google Scholar] [CrossRef]
- Crăciun, A.; Ghoshdastidar, D. Non-singularity of the gradient descent map for neural networks with piecewise analytic activations. In Advances in Neural Information Processing Systems (NeurIPS); 2025. [Google Scholar]
- Jacot, A.; Gabriel, F.; Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems (NeurIPS); 2018. [Google Scholar]
- Rahaman, N.; Baratin, A.; Arpit, D.; Draxler, F.; Lin, M.; Hamprecht, F.A.; Bengio, Y.; Courville, A. On the spectral bias of neural networks. Proc. Int. Conf. Mach. Learn. (ICML) 2019, PMLR 97, 5301–5310. [Google Scholar]
- Razin, N.; Cohen, N. Implicit regularization in deep learning may not be explainable by norms. In Advances in Neural Information Processing Systems (NeurIPS); 2020. [Google Scholar]
- Krantz, S.G.; Parks, H.R. A Primer of Real Analytic Functions, 2nd ed.; Birkhäuser: Boston, 2002. [Google Scholar]
- Kuditipudi, R.; Wang, X.; Lee, H.; Zhang, Y.; Li, Z.; Hu, W.; Arora, S.; Ge, R. Explaining landscape connectivity of low-cost solutions for multilayer nets. In Advances in Neural Information Processing Systems (NeurIPS); 2019. [Google Scholar]
- Lee, J.D.; Simchowitz, M.; Jordan, M.I.; Recht, B. Gradient descent only converges to minimizers. Conference on Learning Theory (COLT), 2016; PMLR 49, pp. 1246–1257. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).