Submitted:
11 October 2024
Posted:
11 October 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Key Euclidean Distance Properties for Computational Optimization
- Symmetry: . In pairwise distance calculations, this property lets you compute each distance once and store it in a symmetric matrix, halving the number of computations [26].
- Non-Negativity: . It provides a basic validation check during distance computations, allowing the algorithm to terminate early or avoid further incorrect computations [1].
- Identity of Indiscernibles: . This allows skipping distance calculations for identical points and filling the main diagonal of pairwise distance matrices with zeros, reducing unnecessary operations [1].
- Triangle Inequality (Subadditivity): This property states that for any points A, B, and C, the direct distance between A and C is less than or equal to the sum of the distances through an intermediate point B, i.e., . Triangle Inequality is crucial for reducing unnecessary distance calculations in clustering and nearest-neighbor algorithms. For example, in nearest-neighbor search, if and are known, you can skip calculating if the inequality shows it is not needed [27]. This property also enables optimizations in hierarchical clustering [28] and network-based applications by using intermediate distances to avoid redundant computations. In KD-trees, the triangle inequality helps prune unnecessary comparisons, speeding up the search process [29].
- Spatial Locality: Points closer in space have smaller Euclidean distances, while distant points have larger distances. This property enables partitioning techniques, like KD-trees or Ball Trees [29,30], to group nearby points, allowing computations to focus on local regions instead of the entire dataset.
- Monotonicity of the Square Root Function: The square root function is monotonically increasing, meaning that if , then . Many distance-based algorithms (e.g., KNN and K-means) only need relative distances. Skipping the square root and using squared Euclidean distance reduces computational complexity while preserving the distance order [1,31].
- Dimensional Independence: Euclidean distance calculations can be performed independently for each dimension. This allows for parallelization, where each dimension’s contribution is computed separately, speeding up computations on multi-core processors [3].
-
Additivity of Independent Dimensions or Subspaces: In high-dimensional spaces, Euclidean distance is the sum of squared differences across dimensions. Low-variance or threshold-exceeding dimensions can be skipped, enabling early termination and reducing computation with minimal accuracy loss [24].Scaling: While positive homogeneity is a property of norms, a similar scaling behavior holds for Euclidean distance under scalar multiplication. Specifically, for a scalar , the distance between two points scales as . This reflects that the distance is scaled by the absolute value of the scalar. After re-scaling, you can utilize previously computed distances and apply the scaling factor, avoiding full recomputation [1].
- Convexity: The squared Euclidean distance is a convex function of its inputs. This property is crucial for optimization, as many algorithms (e.g., gradient descent) leverage convexity to ensure efficient convergence to a global minimum with fewer iterations [4]. While the squared Euclidean distance is convex with respect to one variable when the other is fixed, the overall optimization problem may still be non-convex.
- Continuity: Euclidean distance is continuous, so small changes in point positions cause small changes in distance, . This allows for distance approximations in real-time applications, enabling faster computations in dynamic environments by skipping exact recalculations [29].
- Invariance under Geometric Transformations: Euclidean distance remains unchanged under translations, rotations, or reflections. For any transformation T, , enabling optimization by skipping distance recomputation, improving performance in graphics and simulations [1].
- Minkowski Special Case: Since Euclidean distance is a special case of the Minkowski distance with , other Minkowski norms, such as the -norm () and -norm (), provide upper or lower bounds. These bounds can be used for fast filtering: if an estimate using a simpler norm is already too large, further exact calculations can be skipped, improving performance [1].
- Sparsity Robustness: Euclidean distance computations in sparse vector spaces can be significantly optimized by only computing distances on non-zero dimensions. This property is particularly useful in large-scale machine learning tasks involving high-dimensional but sparse datasets, as it reduces the number of operations [32].
3. Key Approaches for Optimizing Euclidean Distance Computation
3.1. Algorithm-Specific Approaches
Squared Euclidean Distance
Lower Bound Techniques
| Algorithm 1: Snippet of K-means pseudocode with Block Vectors optimization |
![]() |
Triangle Inequality
| Algorithm 2: Snippet of K-means pseudocode with Triangle Inequality optimization |
![]() |
Recursive Distance Updating
| Algorithm 3: Hierarchical clustering algorithm using Lance-Williams formula for recursive distance updating |
![]() |
Advanced Initialization
Precomputing and Caching Distances
Spatial Data Structures and Approximate Methods
Early Exit Strategies
Dimensionality Reduction Techniques
Approximation via Clustering and Dimensional Culling
3.2. Low-Level Optimizations
Approximation with Lower Precision
Loop Unrolling


Machine Code Optimization
Vectorization
| Listing 1: Python code for computing pairwise squared Euclidean distances between two sets of points using vectorization capabilities of NumPy |
![]() |
Parallelization
| Listing 2: Python code for computing pairwise squared Euclidean distances between two sets of points using Numba’s JIT compilation, which combines automatic machine code optimization, vectorization, and parallelism across multiple CPU cores for enhanced performance. |
![]() |
Hardware Acceleration

3.3. Hybrid Approaches
4. Comparative Analysis of Optimization Approaches
4.1. Computation Speedup
4.2. Complexity
4.3. Scalability
4.4. Memory Usage
4.5. Hardware Requirements
4.6. Accuracy Impact
4.7. Best Use Cases
5. Discussion
6. Conclusion
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| ANN | Approximate Nearest Neighbors |
| BLAS | Basic Linear Algebra Subprograms |
| CPU | Central Processing Unit |
| CUDA | Compute Unified Device Architecture |
| FPGA | Field-Programmable Gate Array |
| GPU | Graphics Processing Unit |
| JIT | Just-In-Time |
| KNN | K-Nearest Neighbors |
| KD-tree | K-Dimensional Tree |
| LSH | Locality-Sensitive Hashing |
| MSSC | Minimum Sum-of-Squares Clustering |
| OpenMP | Open Multi-Processing |
| PCA | Principal Component Analysis |
| SIMD | Single Instruction, Multiple Data |
| t-SNE | t-Distributed Stochastic Neighbor Embedding |
References
- Deza, M.M.; Deza, E. Encyclopedia of Distances, 4 ed.; Springer: Berlin, Heidelberg, 2016; p. 756. [Google Scholar] [CrossRef]
- Bottesch, T.; Bühler, T.; Kächele, M. Speeding up k-means by approximating Euclidean distances via block vectors. International conference on machine learning; PMLR, 2016; pp. 2578–2586. [Google Scholar]
- Mussabayev, R.; Mussabayev, R. Superior Parallel Big Data Clustering Through Competitive Stochastic Sample Size Optimization in Big-Means. Intelligent Information and Database Systems; Nguyen, N.T., Chbeir, R., Manolopoulos, Y., Fujita, H., Hong, T.P., Nguyen, L.M., Wojtkiewicz, K., Eds.; Springer Nature Singapore: Singapore, 2024; pp. 224–236. [Google Scholar] [CrossRef]
- Liberti, L.; Lavor, C. Euclidean distance geometry; Springer, 2017; Volume 3. [Google Scholar]
- Croom, F.H. Principles of topology; Courier Dover Publications, 2016. [Google Scholar]
- Braga-Neto, U. Fundamentals of pattern recognition and machine learning; Springer, 2020. [Google Scholar]
- Oyewole, G.J.; Thopil, G.A. Data clustering: application and trends. Artificial Intelligence Review 2023, 56, 6439–6475. [Google Scholar] [CrossRef] [PubMed]
- Varoquaux, G.; Buitinck, L.; Louppe, G.; Grisel, O.; Pedregosa, F.; Mueller, A. Scikit-learn: Machine learning without learning the machinery. GetMobile: Mobile Computing and Communications 2015, 19, 29–33. [Google Scholar] [CrossRef]
- Burger, W.; Burge, M.J. Digital image processing: An algorithmic introduction; Springer Nature, 2022. [Google Scholar]
- Tolebi, G.; Dairbekov, N.S.; Kurmankhojayev, D.; Mussabayev, R. Reinforcement learning intersection controller. 2018 14th International Conference on Electronics Computer and Computation (ICECCO); IEEE, 2018; pp. 206–212. [Google Scholar]
- Fischer, M.M.; Scholten, H.J.; Unwin, D. Geographic information systems, spatial data analysis and spatial modelling: an introduction. In Spatial analytical perspectives on GIS; Routledge, 2019; pp. 3–20. [Google Scholar]
- Tang, Y.; Zhao, L.; Zhang, S.; Gong, C.; Li, G.; Yang, J. Integrating prediction and reconstruction for anomaly detection. Pattern Recognition Letters 2020, 129, 123–130. [Google Scholar] [CrossRef]
- Eiselt, H.A.; Sandblom, C.L. Decision analysis, location models, and scheduling problems; Springer Science & Business Media, 2013. [Google Scholar]
- Carter, C.R.; Rogers, D.S.; Choi, T.Y. Toward the theory of the supply chain. Journal of supply chain management 2015, 51, 89–97. [Google Scholar] [CrossRef]
- Perfilyeva, A.; Bespalova, K.; Kuzovleva, Y.; Mussabayev, R.; et al. . Genetic diversity and origin of Kazakh Tobet Dogs. Scientific Reports 2024, 14, 23137. [Google Scholar] [CrossRef]
- van den Belt, M.; Gilchrist, C.; Booth, T.J.; Chooi, Y.H.; Medema, M.H.; Alanjary, M. CAGECAT: The CompArative GEne Cluster Analysis Toolbox for rapid search and visualisation of homologous gene clusters. BMC bioinformatics 2023, 24, 181. [Google Scholar] [CrossRef]
- Mussabayev, R. Colour-based object detection, inverse kinematics algorithms and pinhole camera model for controlling robotic arm movement system. 2015 Twelve International Conference on Electronics Computer and Computation (ICECCO); IEEE, 2015; pp. 1–9. [Google Scholar]
- Mukhamediev, R.I.; Yakunin, K.; Aubakirov, M.; Assanov, I.; Kuchin, Y.; Symagulov, A.; Levashenko, V.; Zaitseva, E.; Sokolov, D.; Amirgaliyev, Y. Coverage path planning optimization of heterogeneous UAVs group for precision agriculture. IEEE Access 2023, 11, 5789–5803. [Google Scholar] [CrossRef]
- Bojanowski, P.; Grave, E.; Joulin, A.; Mikolov, T. Enriching word vectors with subword information. Transactions of the association for computational linguistics 2017, 5, 135–146. [Google Scholar] [CrossRef]
- Li, J.; Wu, L.; Hong, R.; Hou, J. Random walk based distributed representation learning and prediction on social networking services. Information Sciences 2021, 549, 328–346. [Google Scholar] [CrossRef]
- Altman, N.; Krzywinski, M. The curse (s) of dimensionality. Nat Methods 2018, 15, 399–400. [Google Scholar] [CrossRef]
- Mussabayev, R.; Mladenovic, N.; Jarboui, B.; Mussabayev, R. How to use K-means for big data clustering? Pattern Recognition 2023, 137, 109269. [Google Scholar] [CrossRef]
- Uddin, S.; Haque, I.; Lu, H.; Moni, M.A.; Gide, E. Comparative performance analysis of K-nearest neighbour (KNN) algorithm and its different variants for disease prediction. Scientific Reports 2022, 12, 6256. [Google Scholar] [CrossRef]
- Aggarwal, C.C.; Hinneburg, A.; Keim, D.A. On the surprising behavior of distance metrics in high dimensional space. Database theory—ICDT 2001: 8th international conference, 2001 proceedings 8. London, UK, 4-6 January 2001; Springer, 2001; pp. 420–434. [Google Scholar]
- Maitrey, S.; Jha, C. MapReduce: simplified data analysis of big data. Procedia Computer Science 2015, 57, 563–571. [Google Scholar] [CrossRef]
- Qi, Z.; Xiao, Y.; Shao, B.; Wang, H. Toward a distance oracle for billion-node graphs. Proceedings of the VLDB Endowment 2013, 7, 61–72. [Google Scholar] [CrossRef]
- Elkan, C. Using the triangle inequality to accelerate k-means. In Proceedings of the 20th international conference on Machine Learning (ICML-03); 2003; pp. 147–153. [Google Scholar]
- Contreras, P.; Murtagh, F. Hierarchical clustering; 2015; pp. 103–124. [Google Scholar] [CrossRef]
- Bentley, J.L. Multidimensional binary search trees used for associative searching. Commun. ACM 1975, 18, 509–517. [Google Scholar] [CrossRef]
- Omohundro, S.M. Five Balltree Construction Algorithms. Technical Report TR-89-063; International Computer Science Institute, 1989. [Google Scholar]
- Bock, H.H. Origins and extensions of the k-means algorithm in cluster analysis. Electronic journal for history of probability and statistics 2008, 4, 1–18. [Google Scholar]
- Ying, Y.; Li, P. Distance metric learning with eigenvalue optimization. The Journal of Machine Learning Research 2012, 13, 1–26. [Google Scholar]
- Cormen, T.H.; Leiserson, C.E.; Rivest, R.L.; Stein, C. Introduction to algorithms; MIT press, 2022. [Google Scholar]
- Fränti, P.; Sieranoja, S. How much can k-means be improved by using better initialization and repeats? Pattern Recognition 2019, 93, 95–112. [Google Scholar] [CrossRef]
- Hamerly, G. Making k-means even faster. In Proceedings of the 2010 SIAM International Conference on Data Mining (SDM); SIAM, 2010; pp. 130–140. [Google Scholar] [CrossRef]
- Jeon, Y.; Yoon, S. Multi-Threaded Hierarchical Clustering by Parallel Nearest-Neighbor Chaining. IEEE Transactions on Parallel and Distributed Systems 2015, 26, 2534–2548. [Google Scholar] [CrossRef]
- Arthur, D.; Vassilvitskii, S. k-means++: The advantages of careful seeding. In Proceedings of the eighteenth annual ACM-SIAM symposium on Discrete algorithms; SIAM, 2007; pp. 1027–1035. [Google Scholar]
- Park, H.S.; Jun, C.H. A simple and fast algorithm for K-medoids clustering. Expert systems with applications 2009, 36, 3336–3341. [Google Scholar] [CrossRef]
- Indyk, P.; Motwani, R. Approximate nearest neighbors: towards removing the curse of dimensionality. Proceedings of the Thirtieth Annual ACM Symposium on Theory of Computing; Association for Computing Machinery: New York, NY, USA, 1998; STOC’98; pp. 604–613. [Google Scholar] [CrossRef]
- McNames, J. Rotated partial distance search for faster vector quantization encoding. IEEE Signal Processing Letters 2000, 7, 244–246. [Google Scholar] [CrossRef]
- Beyer, K.; Goldstein, J.; Ramakrishnan, R.; Shaft, U. When is “nearest neighbor” meaningful? Database Theory—ICDT’99: 7th International Conference Jerusalem, Israel, 10-12 January 1999; Proceedings 7. Springer, 1999; pp. 217–235. [Google Scholar]
- Jolliffe, I.T.; Cadima, J. Principal component analysis: a review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 2016, 374, 20150202. [Google Scholar] [CrossRef]
- Gewers, F.L.; Ferreira, G.R.; Arruda, H.F.D.; Silva, F.N.; Comin, C.H.; Amancio, D.R.; Costa, L.d.F. Principal component analysis: A natural approach to data exploration. ACM Computing Surveys (CSUR) 2021, 54, 1–34. [Google Scholar] [CrossRef]
- van der Maaten, L.; Hinton, G. Visualizing Data using t-SNE. Journal of Machine Learning Research 2008, 9, 2579–2605. [Google Scholar]
- Bingham, E.; Mannila, H. Random projection in dimensionality reduction: applications to image and text data. Knowledge Discovery and Data Mining 2001. [Google Scholar] [CrossRef]
- Borodin, A.; Ostrovsky, R.; Rabani, Y. Subquadratic approximation algorithms for clustering problems in high dimensional spaces. Machine Learning 2004, 56, 153–167. [Google Scholar] [CrossRef]
- Rodriguez, A.; Segal, E.; Meiri, E.; Fomenko, E.; Kim, Y.J.; Shen, H.; Ziv, B. Lower numerical precision deep learning inference and training. Intel White Paper 2018, 3, 1–19. [Google Scholar]
- Lee, S.; Gerstlauer, A. Data-dependent loop approximations for performance-quality driven high-level synthesis. IEEE Embedded Systems Letters 2017, 10, 18–21. [Google Scholar] [CrossRef]
- Pikus, F.G. The Art of Writing Efficient Programs: An advanced programmer’s guide to efficient hardware utilization and compiler optimizations using C++ examples; Packt Publishing Ltd, 2021. [Google Scholar]
- Lam, S.K.; Pitrou, A.; Seibert, S. Numba: A llvm-based python jit compiler. In Proceedings of the Second Workshop on the LLVM Compiler Infrastructure in HPC; 2015; pp. 1–6. [Google Scholar]
- Patterson, D.A.; Hennessy, J.L. Computer Architecture: A Quantitative Approach, 5th ed.; The Morgan Kaufmann Series in Computer Architecture and Design, Elsevier Science & Technology; 2011. [Google Scholar]
- Harris, C.R.; Millman, K.J.; Van Der Walt, S.J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N.J.; others. Array programming with NumPy. Nature 2020, 585, 357–362. [Google Scholar] [CrossRef]
- Masek, J.; Burget, R.; Karasek, J.; Uher, V.; Dutta, M.K. Multi-GPU implementation of k-nearest neighbor algorithm. 2015 38th International Conference on Telecommunications and Signal Processing (TSP); IEEE, 2015; pp. 764–767. [Google Scholar]
- Dean, J.; Ghemawat, S. MapReduce: simplified data processing on large clusters. Commun. ACM 2008, 51, 107–113. [Google Scholar] [CrossRef]
- Boikos, K.; Bouganis, C.S. A scalable fpga-based architecture for depth estimation in slam. International Symposium on Applied Reconfigurable Computing; Springer, 2019; pp. 181–196. [Google Scholar]
- Gribel, D.; Vidal, T. HG-means: A scalable hybrid genetic algorithm for minimum sum-of-squares clustering. Pattern Recognition 2019, 88, 569–583. [Google Scholar] [CrossRef]



| Optimization Approach | Computation Speedup | Complexity | Scalability | Memory Usage | Hardware Requirements | Accuracy Impact | Best Use Cases |
|---|---|---|---|---|---|---|---|
| Squared Euclidean Distance | Moderate | Low | High | Minimal | None | None | K-means, KNN, large datasets |
| Lower Bound Techniques | High | Moderate | High | Low | None | None | K-means clustering, nearest neighbor search |
| Triangle Inequality | High | Moderate | High | Minimal | None | None | K-means, KNN, hierarchical clustering |
| Recursive Distance Updating | Moderate | High | Low to Moderate | Moderate | None | None | Hierarchical clustering |
| Precomputing / Caching | High (retrieval) | Moderate to High | Moderate | High | None | None | Static datasets, K-medoids, DBSCAN |
| Spatial Data Structures | High | Moderate | High (low dimensions), Low in high dimensions | Moderate | None | None | Nearest neighbor search, low to moderate dimensionality |
| Approximate Methods | Very High | Moderate | Very High | Moderate | None | Moderate to High | Large-scale nearest neighbor search |
| Dimensionality Reduction | High | High | Moderate to High | Moderate to High | None | Moderate to High | High-dimensional data, exploratory data analysis |
| Advanced Initialization | Moderate to High | Low | High | Minimal | None | Potential Positive Impact | K-means, multi-start clustering, better centroid initialization |
| Early Exit Strategies | High | Moderate | High | Minimal | None | Minimal | Nearest neighbor search, high-dimensional data |
| Clustering and Dimensional Culling | High | Moderate | High | Moderate | None | Moderate to High | High-dimensional data, feature selection |
| Approximation with Lower Precision | High | Low | High | Minimal | None | Low to Moderate | Large-scale distance computations, approximate solutions |
| Loop Unrolling | Moderate | Low | Moderate | Low | None | None | Low-level optimizations, numerical computations |
| Vectorization | High | Low | High | Low | SIMD-enabled CPU | None | Any distance-intensive operation |
| Parallelization | Very High | High | Very High | Low to Moderate | Multi-core CPU, GPU | None | Large datasets, real-time processing |
| Hardware Acceleration (GPU, FPGA) | Very High | High | Very High | Low to Moderate | GPU, FPGA | Potential Minimal | High-throughput, real-time processing |
| Machine Code Optimization | High | Low to Moderate | High | Low | Modern CPU | None | C++, Assembly, Python-based applications with Numba library, dynamic environments |
Short Biography of Authors
![]() |
Rustam Mussabayev is an Associate Professor and the Head of the AI Research Lab at Satbayev University, Kazakhstan. He holds a Candidateof Engineering Sciences degree(equivalent to a PhD in Computer Science) with expertise in datascience, high-performance computing, and operations research. His research interests span a wide range of topics, including clustering, natural language processing, machine learning, and optimization. He has received numerous awards, including the StatePrize “Best Researcher 2023” of the Republic of Kazakhstan and the Best Paper Award at ACIIDS 2024. His work is widely published in top-tier journals and conferences, contributing to advancements in data analysis, algorithm development, and high-performance computing solutions. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).





