Submitted:
30 December 2022
Posted:
05 January 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Notions of Computational Topology
2.1. Simplicial complexes
- Each face of any simplex of K is also a simplex of K.
- The intersection of any two simplices of K is either empty or a face of both simplices.
2.2. Betti numbers
3. Related work
3.1. Nonlinear dimensionality reduction
3.2. Variational Auto-Encoder
3.3. Topology and Auto-Encoders
4. Implementation details
5. Problem formulation
5.1. Data sets
5.2. Illustration of the problem
6. InvMap VAE
6.1. Method
6.2. Results
7. Witness Simplicial VAE
7.1. Method
7.1.1. Witness complex construction
7.1.2. Witness complex simplicial regularization
- It does not depend on any embedding whereas in [28] the author was relying on a UMAP embedding for his simplicial regularization of the decoder.
- We use only one witness simplical complex built from the input data whereas the author of [28] was using one fuzzy simplicial complex built from the input data and a second one built from the UMAP embedding and both were built via the fuzzy simplicial set function provided with UMAP (keeping only simplices with highest probabilities).
- the simplicial regularization term for the encoder.
- the simplicial regularization term for the decoder.
- e and d respectively the (probabilistic) encoder and decoder.
- K a (witness) simplicial complex built from the input space.
- a simplex belonging to the simplicial complex K.
- the vertex number j of the -simplex which has exactly vertices. is thus a data point in the input space X.
- the Mean Square Error between a and b.
- the expectation for the following a symmetric Dirichlet distribution with parameters and . When , which is what we used in practice, the symmetric Dirichlet distribution is equivalent to a uniform distribution over the -simplex , and as tends towards 0, the distribution becomes more concentrated on the vertices.
7.1.3. Witness Simplicial VAE
- Perform a witness complex filtration of the input data to get a persistence diagram (or a barcode).
- Build a witness complex given the persistence diagram (or the barcode) of this filtration, and potentially any additional information on the Betti numbers which should be preserved according to the problem (number of connected components, 1-dimensional holes...).
- Train the model using this witness complex to compute the loss of equation 5.
7.1.4. Isolandmarks Witness Simplicial VAE
- the loss of the Witness Simplicial VAE.
- l the number of landmarks.
- the Frobenius norm.
- the approximate geodesic distance matrix of the landmarks points in the input space computed once before learning.
- D the Euclidean distance matrix of the encodings of the landmarks points computed at each batch.
- K the Isomap kernel defined as with I the identity matrix and A the matrix composed only by ones.
7.2. Results




8. Discussion
9. Conclusions
Author Contributions
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AE | Auto-Encoder |
| i.i.d. | independant and identically distributed |
| ELBO | Evidence lower bound |
| Isomap | Isometric Mapping |
| k-nn | k-nearest neighbors |
| MDPI | Multidisciplinary Digital Publishing Institute |
| MSE | Mean Square Error |
| s.t. | such that |
| TDA | Topological Data Analysis |
| UMAP | Uniform Manifold Approximation and Projection |
| VAE | Variational Auto-Encoder |
| WC | Witness complex |
Appendix A. Variational Auto-Encoder derivations
Appendix A.1. Derivation of the marginal log-likelihood
Appendix A.2. Derivation of the ELBO
Appendix B. UMAP-based InvMap-VAE results

Appendix C. Illustration of the importance of the choice of the filtration radius hyperparameter for the witness complex construction

Appendix D. Bad neural network weights initialization with Witness Simplicial VAE

References
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Advances in Neural Information Processing Systems; Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K.Q., Eds.; Curran Associates, Inc., 2014; Volume 27. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes; ICLR; Bengio, Y., LeCun, Y., Eds.; 2014. [Google Scholar]
- Rezende, D.J.; Mohamed, S.; Wierstra, D. Stochastic Backpropagation and Approximate Inference in Deep Generative Models. In Proceedings of the 31st International Conference on Machine Learning; Proceedings of Machine Learning Research. Xing, E.P., Jebara, T., Eds.; PMLR: Bejing, China, 2014; Volume 32, pp. 1278–1286. [Google Scholar]
- Medbouhi, A.A. Towards topology-aware Variational Auto-Encoders: from InvMap-VAE to Witness Simplicial VAE. Master thesis, KTH Royal Institute of Technology, Sweden, 2022. [Google Scholar]
- Hensel, F.; Moor, M.; Rieck, B. A Survey of Topological Machine Learning Methods. Frontiers in Artificial Intelligence 2021, 4, 52. [Google Scholar] [CrossRef]
- Ferri, M. Why Topology for Machine Learning and Knowledge Extraction? Machine Learning and Knowledge Extraction 2019, 1, 115–120. [Google Scholar] [CrossRef]
- Edelsbrunner, H.; Harer, J. Computational Topology - an Introduction; American Mathematical Society, 2010; pp. I–XII, 1–241. [Google Scholar]
- Wikipedia, the free encyclopedia. Simplicial complex example, 2009.
- Wikipedia, the free encyclopedia. Simplicial complex nonexample, 2007.
- Wilkins, D.R. Algebraic Topology, Course 421; 1988-2008; Trinity College: Dublin.
- de Silva, V.; Carlsson, G. Topological estimation using witness complexes. IEEE Symposium on Point-based Graphic 2004, 157–166. [Google Scholar]
- Rieck, B. Topological Data Analysis for Machine Learning, Lectures; European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, 2020. Image made available under the Creative Commons. Available online: https://creativecommons.org/licenses/by/4.0/ (accessed on 12 November 2020).
- Pearson, K. LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 1901, 2, 559–572. [Google Scholar] [CrossRef]
- Hotelling, H. Relations Between Two Sets of Variates. Biometrika 1936, 28, 321–377. [Google Scholar] [CrossRef]
- Lee, J.A.; Verleysen, M. Nonlinear Dimensionality Reduction, 1st ed.; Springer Publishing Company, Incorporated, 2007. [Google Scholar]
- Tenenbaum, J.B.; de Silva, V.; Langford, J.C. A Global Geometric Framework for Nonlinear Dimensionality Reduction. Science 2000, 290, 2319. [Google Scholar] [CrossRef] [PubMed]
- Kruskal, J. Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis. Psychometrika 1964. [Google Scholar] [CrossRef]
- Kruskal, J. Nonmetric multidimensional scaling: a numerical method. Psychometrika 1964. [Google Scholar] [CrossRef]
- Borg, I.; Groenen, P. Modern Multidimensional Scaling: Theory and Applications (Springer Series in Statistics); 2005. [Google Scholar] [CrossRef]
- Van der Maaten, L.; Hinton, G. Visualizing data using t-SNE. Journal of Machine Learning Research 2008, 9, 2579–2605. [Google Scholar]
- Hinton, G.E.; Roweis, S. Stochastic Neighbor Embedding. In Advances in Neural Information Processing Systems; Becker, S., Thrun, S., Obermayer, K., Eds.; MIT Press, 2002; Volume 15. [Google Scholar]
- McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, 2018. Available online: http://github.com/lmcinnes/umap.
- Kingma, D.P.; Welling, M. An Introduction to Variational Autoencoders. Foundations and Trends in Machine Learning 2019, 12, 307–392. [Google Scholar] [CrossRef]
- Gabrielsson, R.B.; Nelson, B.J.; Dwaraknath, A.; Skraba, P.; Guibas, L.J.; Carlsson, G.E. A Topology Layer for Machine Learning. CoRR 2019. [Google Scholar]
- Polianskii, V. An Investigation of Neural Network Structure with Topological Data Analysis. Master’s Thesis, KTH Royal Institute of Technology, Sweden, 2018. [Google Scholar]
- Moor, M.; Horn, M.; Rieck, B.; Borgwardt, K.M. Topological Autoencoders. CoRR 2019. [Google Scholar]
- Hofer, C.D.; Kwitt, R.; Dixit, M.; Niethammer, M. Connectivity-Optimized Representation Learning via Persistent Homology. CoRR 2019. [Google Scholar]
- Gallego-Posada, J. Simplicial AutoEncoders: A connection between Algebraic Topology and Probabilistic Modelling. Master’s Thesis, University of Amsterdam, Netherlands, 2018. [Google Scholar]
- Gallego-Posada, J.; Forré, P. Simplicial Regularization. In ICLR 2021 Workshop on Geometrical and Topological Representation Learning; 2021. [Google Scholar]
- Zhang, H.; Cisse, M.; Dauphin, Y.N.; Lopez-Paz, D. mixup: Beyond Empirical Risk Minimization. In Proceedings of the International Conference on Learning Representations; 2018. [Google Scholar]
- Verma, V.; Lamb, A.; Beckham, C.; Courville, A.C.; Mitliagkas, I.; Bengio, Y. Manifold Mixup: Encouraging Meaningful On-Manifold Interpolation as a Regularizer. CoRR 2018. [Google Scholar]
- Khrulkov, V.; Oseledets, I.V. Geometry Score: A Method For Comparing Generative Adversarial Networks. CoRR 2018. [Google Scholar]
- Pérez Rey, L.A.; Menkovski, V.; Portegies, J. Diffusion Variational Autoencoders; 2020; pp. 2676–2682. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; Chintala, S. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32; Wallach, H., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E., Garnett, R., Eds.; Curran Associates, Inc., 2019; pp. 8024–8035. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. In Proceedings of the 3rd International Conference for Learning Representations, San Diego, CA, USA; 2015. [Google Scholar]
- Simon, S. Witness Complex, 2020. Available online: https://github.com/MrBellamonte/WitnessComplex (accessed on 20 December 2022).
- Maria, C.; Boissonnat, J.D.; Glisse, M.; Yvinec, M. The Gudhi Library: Simplicial Complexes and Persistent Homology. In Technical Report; 2014. [Google Scholar]
- Maria, C. Filtered Complexes. In GUDHI User and Reference Manual, 3.4.1 ed.; GUDHI Editorial Board; 2021. [Google Scholar]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; Vanderplas, J.; Passos, A.; Cournapeau, D.; Brucher, M.; Perrot, M.; Duchesnay, E. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 2011, 12, 2825–2830. [Google Scholar]
- Marsland, S. Machine Learning - An Algorithmic Perspective; Chapman and Hall / CRC machine learning and pattern recognition series; CRC Press, 2009; pp. I–XVI, 1–390. [Google Scholar]
| 1 | This paper presents in a more concise way our main work developed during Medbouhi’s master thesis [4] and provides an extension of the Witness Simplicial VAE method. |
| 2 | The reader is invited to look at the mentioned paper [11] for a complete view on witness complexes, because here we make some simplifications and define the witness complex as a particular case of the original definition. |

















Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).