Submitted:
06 May 2023
Posted:
09 May 2023
Read the latest preprint version here
Abstract
Keywords:
1. Introduction
- Measure approach for NN training: Probability measures can be used to analyze the behavior of neural networks, both during training and inference. For example, the loss function used to train a neural network can be seen as a divergence between the model distribution and the target distribution, which can be expressed in terms of probability measures, [17]. Similarly, the output of a neural network can be interpreted as a probability distribution over the output space, which can be analyzed using probability theory. This can be useful for understanding the uncertainty of the predictions made by the neural network, such as the Stochastic Deep Networks (SDN) introduced in [25]. The design of a SDN involves a deep architecture that can handle probability distribution inputs and outputs and uses the classical Wasserstein distance as the loss function. Moreover, layers correspond to a sequence of elementary blocks that maps random vectors to random vectors to capture multiple interactions between samples from the distributions.
- Probability measures as a tool to regularize neural networks: Probability measures can be used as a tool to regularize neural networks, by adding a regularization term to the loss function that encourages the output of the neural network to match a prior distribution by the so called entropic measure, such as in Eq. (26). Within this case, we mention the entropic version of Sinkhorn algorithm where an entropic regularization is added, developed, e.g., in [40]. This algorithm can be useful for scaling OT problem in high dimensional setting or for generative modeling tasks, where the goal is to learn a distribution that can generate new data samples.
- Neural networks to approximate empirical measures: Neural networks can be used to approximate probability measures, such as empirical distributions over a state space for mean field function. We specifically consider this feature in Section 3. Analogously, one can consider NN architecture for learning population dynamics. For example, in [33], the authors focus on a recurrent architecture to model the diffusion of a population by means of a NN by injecting random noise.
- Neural Network to learn unknown probability measure for unsupervised learning tasks such as density estimation, where the goal is to learn the underlying distribution of the data. For example, in [21], the authors try to build an algorithm to directly learn a probability measure via an optimal transport metric, also providing some convergence results for the learning algorithm.
2. A continuous idealization of NN for a Mean Field Optimal Control Problem
- denotes the input of the NN;
- denotes the output of the NN:
- the corresponding target.
2.1. Mean Field Optimal Control Problem
- going from layer index T to continuous parameter t;
- passing from discrete set of inputs/output to distribution that represents the joint distribution in modelling the input-label distribution;
- identifying targets y as random variables sampled from the projection over of the distribution ;
- passing from empirical risk minimization to population risk (i.e. minimization over expectation ).
- , , are bounded;
- f, L, are Lipschitz continuous with respect to x with the Lipschitz constants of f and L being independent from parameters ;
- has finite support in
2.2. Empirical measures over controls
3. Learning Mean Field Function
3.1. The McKean-Vlasov SDE
3.2. Mean Field Optimal Transport
- Optimal control via direct approximation of controls v;
- Deep Galerkin Method for solving a forward-backward systems of PDEs;
- Augmented Lagrangian Method with Deep Learning exploiting the variational formulation of MFOT and the primal/dual approach.
3.3. Alternative methods
4. Learning Gradient Flow
4.1. Gradient Flow in Wasserstein space

4.2. Stochastic Optimization with ICNN and JKOnet
5. From OT to AOT
5.1. OT variants
- Martingale Optimal Transport. A first model to encode a temporal adapted structure into OT problems moves within the theory Martingale Optimal Transport (MOT) dealing with transport plans with martingale couplings. The stability of a sequence of transport plans with martingales marginals has been proved to hold in [10], also for the Weak Optimal Transport.
-
Multi marginal optimal transport problem. Multi-marginal optimal transport problems introduced in [4] are formulated as follows. Consider probability measures on and consider the optimization problemwith c being a lower semi-continuous cost function defined on with marginal laws andTrying to approximate this problem by discretizing the state space through N points, chosen and a priori fixed, lead to a linear programming problem of size .
-
From Martingale-constrained to Moment Constrained Optimal Transport . In [5,6], the authors develop algorithms to sample from probability distribution while preserving the convex order enabling to use linear programming solver.The same authors study the same problem also in [4] following another perspective by relaxing the martingale constraint into a finite number of moment constraints by means of N real-valued bounded (test) function defined on .This method is called Moment Constrained Optimal Transport (MCOT). Differently from multi marginal setting, it is not the state space to be discretized but the marginal laws constrained of Eq. (18) are relaxed into a finite number of moment constraints via some given It solves the following:where the infimum is computed over the set of probability measures with values in satisfying for alland
5.2. Adapted Optimal Transport
5.3. Adapted Wasserstein Distance
5.3.1. The class of filtered processes
- is a geodesic space (assuming ) and it is the completion of . The first property suggests the possibility an interpolation, in the sense of McCann, also of stochastic processes that is not possible if one considers the usual Wasserstein space ;
- two filtered processes and are identified if . Moreover, this equivalence relation encodes also information about their filtration.
- the space of filtered stochastic processes is defined as a set of equivalence classes as in the case of spaces. An equivalence class of collects all representatives processes that are identical from a probabilistic perspective.
5.3.2. Adapted Empirical Measure
- in [38], adapted empirical distribution is constructed via convolution of a smooth kernel with the empirical measure on paths;
- in [1] the authors show the convergence of adapted empirical measure in , removing the assumption of compact support for the underlying distribution. Moreover, the authors introduce the non-uniform adapted empirical measures by defining a non-uniform grid on , dealing with cubes of different sizes, i.e. of lower dimension near the origin, bigger far in the margins.
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Sample Availability
Abbreviations
| OT | Optimal Transport |
| COT | Causal Optimal Transport |
| AOT | Adapted Optimal Transport |
| NN | Neural Network |
| MFOCP | Mean Field Optimal Control Problem |
| MFG | Mean Field Games |
| MFC | Mean Field Control |
| ML | Machine Learning |
| DL | Deep Learning |
| MFOT | Mean Field Optimal Transport |
| SDE | Stochastic Differential Equation |
| ICNN | Input-Convex Neural Networks |
| JKO | Jordan-Kinderlehrer-Otto |
| SNN | Stochastic Neural Network |
| SGD | Stochastic Gradient Descent |
| HJB | Hamilton-Jacobi-Bellman |
| ODE | Ordinary Differential Equation |
| MCOT | Moment Constrained Optimal Transport |
| MOT | Martingal Optimal Transport |
References
- Acciaio, B.; Hou, S. Convergence of Adapted Empirical Measures on . 2022 arXiv e-prints. [CrossRef]
- Acciaio B.; Kratsios A.; Pammer G. Metric hypertransformers are universal adapted maps. 2022 Preprint, arXiv:2201.13094.
- Ambrosio, L.; Gigli, N. Savaré, G- Gradient flows: in metric spaces and in the space of probability measures, Springer Science & Business Media, 2008.
- Alfonsi, A.; Coyaud, R.; Ehrlacher, V.; Lombardi, D. Approximation of Optimal Transport problems with marginal moments constraints. Mathematics of Computation, 2020, ⟨10.1090/mcom/3568⟩. ⟨hal-02128374⟩. [CrossRef]
- Alfonsi, A.; Corbetta, J.; Jourdain, B. Sampling of one-dimensional probability measures in the convex order and computation of robust option price bounds. International Journal of Theoretical and Applied Finance, 20190(0):1950002, 0. [CrossRef]
- Alfonsi, A.; Corbetta, J.; Jourdain, B. Sampling of probability measures in the convex order by Wasserstein projection. arXiv e-prints, page arXiv:1709.05287, Sep 2017.
- Alvarez-Melis, D.; Schiff, Y.; Mroueh, Y. Optimizing Functionals on the Space of Probabilities with Input Convex Neural Networks, arXiv e-prints, 2021.
- Archibald, R.; Bao, F.; Yong, J. “An Online Method for the Data Driven Stochastic Optimal Control Problem with Unknown Model Parameters”, arXiv e-prints, 2022.
- Backhoff-Veraguas, J.; Bartl, D.; Beiglblock, M.; and Eder, M. All adapted topologies are equal. Probab. Theory Related Fields, 178(3-4):1125–1172, 2020. [CrossRef]
- Backhoff-Veraguas, J.; Pammer, G. Stability of martingale optimal transport and weak optimal transport. Ann. Appl. Probab. 32 (1) 721 - 752, February 2022. [CrossRef]
- Bartl, D.; Beiglböck, M.; Pammer, G. The Wasserstein space of stochastic processes. arXiv e-prints, 2021. [CrossRef]
- Baudelet, S.; Frénais, B.; Laurière, M.; Machtalay, A.; Zhu, Y. Deep Learning for Mean Field Optimal Transport. arXiv e-prints 2023. [CrossRef]
- Brandon, A.; Xu, L.; Kolter, J.Z. Input Convex Neural Networks Proceedings of the 34th International Conference on Machine Learning, 2017, PMLR 70:146-155.
- Backhoff-Veraguas, J.; Beiglboeck, M.; Lin, Y.; Zalashko, A. Causal transport in discrete time and applications. SIAM Journal on Optimization, 2017, 27(4):2528–2562. [CrossRef]
- Backhoff-Veraguas, J.; Bartl, D.; Beiglblock, M.; Wiesel, J. Estimating processes in adapted Wasserstein distance. Ann. Appl. Probab. 32 (1) 529 - 550, February 2022. [CrossRef]
- Bao, F.; Cao, Y.; Archibald, R.; Zhang, H. Uncertainty quantification for deep learning through stochastic maximum principle. arXiv: 3489122, 2021.
- Bonnet, B.; Cipriani, C.; Fornasier, M.; Huang, H. A measure theoretical approach to the mean-field maximum principle for training NeurODEs, Nonlinear Analysis, Volume 227, 2023, 113161, ISSN 0362-546X. [CrossRef]
- Bunne, C.; Meng-Papaxanthos, L.; Krause, A.; Cuturi, M. Proximal Optimal Transport Modeling of Population Dynamics, arXiv e-prints, 2021.
- Benoît, B. A Pontryagin Maximum Principle in Wasserstein spaces for constrained optimal control problems. ESAIM: COCV, 25 2019 52. [CrossRef]
- Carlier, G. On the linear convergence of the multi-marginal Sinkhorn algorithm. SIAM Journal on Optimization, 2022, 32 (2), pp.786-794. [CrossRef]
- Canas, G.; Rosasco, L. Learning Probability Measures with respect to Optimal Transport Metrics, Advances in Neural Information Processing Systems 25, 2012.
- Chizat, L.; Bach, F. , On the global convergence of gradient descent for overparameterized models using optimal transport. In Advances in neural information processing systems, 2018, pages 3040–305.
- Chizat, L.; Colombo, M. , Fernández-Real, X.; and Figalli, A. Infinite-width limit of deep linear neural networks, arXiv e-prints, 2022. [CrossRef]
- Cuturi, M. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems, 2013 pages 2292–2300.
- de Bie, G.; Peyré, G.; Cuturi, M. Stochastic Deep Networks, Proceedings of the 36th International Conference on Machine Learning, 2019. Long Beach, California, PMLR 97.
- Di Persio, L. , Garbelli M. Deep Learning and Mean-Field Games: A Stochastic Optimal Control Perspective. Symmetry 2021 13(1):14. [CrossRef]
- E, W.; Han, J.; Li, Q. A mean-field optimal control formulation of deep learning. Res Math Sci 6, 10. 2019. [CrossRef]
- Feydy, J.; Séjourné, T.; Vialard, F.X.; Amari, S.; Trouve, A.; Peyré, G. Interpolating between Optimal Transport and MMD using Sinkhorn Divergences. Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, 2019, PMLR 89:2681-2690.
- Fernández-Real, X. : Figalli, A. The Continuous Formulation of Shallow Neural Networks as Wasserstein-Type Gradient Flows. In: Avila, A., Rassias, M.T., Sinai, Y. (eds) Analysis at Large. Springer, Cham. 2022. [CrossRef]
- Gangbo, W.; Mayorga, S.; Swiech, A. Finite Dimensional Approximations of Hamilton-Jacobi Bellman Equations in Spaces of Probability Measures. SIAM Journal on Mathematical Analysis 2021 53:2, 1320-1356. [CrossRef]
- Genevay, A.; Peyre, G.; Cuturi, M. Learning Generative Models with Sinkhorn Divergences. Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics, 2018, PMLR 84:1608-1617.
- Jimenez, C. ; A. Marigonda A.; Quincampoix, M. Dynamical systems and Hamilton-Jacobi-Bellman equations on the Wasserstein space and their L2 representations, 2022. Preprint at https://cvgmt.sns.it/media/doc/ paper/5584/AMCJMQ_HJB_2022-03-30.pdf.
- Hashimoto, T.; Gifford, D.; Jaakkola, T. Learning population-level diffusions with generative rnns. International Conference on Machine Learning, 2016, pages 2417–2426.
- Lassalle, R. Causal transport plans and their Monge–Kantorovich problems, Stochastic Analysis and Applications 2018, 36:3, 452-484. [CrossRef]
- Li, Q.; Lin, T.; Shen, Z. Deep Learning via Dynamical Systems: An Approximation Perspective. 2019 Published in arXiv:1912.10382v1.
- Li, Q., Long, C., Cheng, T.; E, W. Maximum principle based algorithms for deep learning. J. Mach. Learn. Res. 2017, 18, 1, 5998–6026.
- Pham, H.; Warin, X. Mean-field neural networks: learning mappings on Wasserstein space, 2022 arXiv e-prints.
- Pflug, G.; Pichler, A. . A distance for multistage stochastic optimization models. SIAM Journal on Optimization, 2012, 22(1):1-23.
- Pichler, A.; Weinhardt, M. The nested Sinkhorn divergence to learn the nested distance. Comput Manag Sci 2022, 19, 269–293. [Google Scholar] [CrossRef]
- Peyré, G.; Cuturi, M. Computational Optimal Transport: With Applications to Data Science, Foundations and Trends in Machine Learning, 2019. Vol. 11: No. 5-6, pp 355-607. [CrossRef]
- Sinkhorn, R. Diagonal equivalence to matrices with prescribed row and column sums. The American Mathematical Monthly, 1967, 74(4):402–405. [CrossRef]
- Seguy, V.; Cuturi, M. , Principal Geodesic Analysis for Probability Measures under the Optimal Transport Metric, 2015 arXiv e-prints.
- Sirignano, J.; Spiliopoulos, K. Mean Field Analysis of Deep Neural Networks. Mathematics of Operations Research 2021, 47(1):120-152. [CrossRef]
- Xu, T.; Li K., W.; Munn, M.; Acciaio, B. ; COT-GAN: Generating Sequential Data via Causal Optimal Transport. Neural Information Processing Systems (NeurIPS), 2020.
- Zalashko, A. ; Causal optimal transport: theory and applications. PhD Thesis, 2017 University of Vienna.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).