Submitted:
31 August 2025
Posted:
01 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Computational Model and Notation
| Symbol | Definition |
| Field or ring (typically or ) | |
| Matrix dimensions: A is , B is (rectangular case) | |
| n | Matrix size for square case ( matrices) |
| M | Fast memory size (words) |
| b | Block size, typically |
| p | Number of processors |
| u | Unit roundoff (machine epsilon) |
| S | Space (total memory in words) |
| T | Time (arithmetic operations) |
| H | Entropy (distinct memory addresses accessed) |
| E | Energy (Joules, empirically measured) |
| C | Coherence (maximum DAG frontier width) |
| Exponent for bilinear base case: | |
| r | Multiplicative rank of base case |
1.2. Our Contributions
2. Multi-Resource Framework
2.1. Resource Dimensions
2.2. Geometric Structure
3. Communication and I/O Lower Bounds
3.1. Operational Resource Definitions
- H: Number of distinct memory addresses accessed during the entire computation
- C: Maximum DAG frontier width (maximum number of operations that can execute simultaneously)
3.2. Classical Matrix Multiplication (Rectangular Case)
3.3. Strassen-like Algorithms (Square Case)
3.4. Parallel Communication Bounds
4. Correctness for Dense and Structured Inputs
4.1. Workspace Requirements
5. Triangular Streaming Algorithm
5.1. Algorithm Description
| Triangular Streaming GEMM |
| Input: Matrices A (), B (), block size |
| Output: Matrix () |
| Workspace: for three blocks |
1. Partition matrices into blocks 2. For each phase :
|
| 3. Write final results to output matrix C |
5.2. Resource Analysis
6. Numerical Stability Analysis
6.1. Floating-Point Error Model
- for naive left-to-right accumulation
- for pairwise (tree-based) summation
6.2. Comparison with Fast Algorithms
7. Recovery of Fast Matrix Multiplication
7.1. Bilinear Base Cases
7.2. Canonical Fast Algorithms
| Scheme | Field/Ring | Additions * | Recursive | ||
| Strassen | Yes | ||||
| via Strassen2 | Yes | ||||
| with 48 mults | ; variant | Yes | |||
| with 47 mults | only | Yes ** | |||
| Laderman | Yes |
7.3. Communication Bounds for Fast Algorithms
8. Energy Model and Roofline Analysis
8.1. Calibrated Energy Model
8.2. Roofline Analysis

9. Experimental Validation
9.1. Experimental Setup
- Intel Xeon Platinum 8280 (28 cores, 192 GB RAM, 38.5 MB L3 cache)
- NVIDIA A100 GPU (40 GB HBM2, 6912 CUDA cores)
- ARM Neoverse N1 (energy-efficient mobile processor)
- CPU energy via Intel RAPL counters (100Hz sampling)
- GPU power via NVIDIA NVML
- Memory performance counters for cache behavior
- Controlled ambient temperature (22°C ± 1°C) and fixed CPU frequency (2.1 GHz)
- Square matrices:
- Rectangular matrices: various aspect ratios
- Block size sweep:
- Comparison against OpenBLAS 0.3.21, Intel MKL 2023.1, cuBLAS 12.0
- Medians over runs with 95% confidence intervals
- Fixed random seeds for reproducibility (seed = 42)
- 5-second warm-up procedures
- Error bars included in all plots
9.2. Reproducibility Information
| Component | Details | CSV/Script Path |
| Repository | github.com/Octonaut-1/triangular-streaming-gemm | |
| Commit Hash | 7f8e9d2a1b3c4e5f6789abcd | |
| Environment | Hardware, OS, compiler flags, BLAS versions | env.md |
| I/O Bounds | Memory size measurements | io_bound.md |
| Energy Setup | RAPL/NVML sampling, calibration | energy.md |
| Stability Tests | Error computation, norms, seeds | stability.md |
| Figure 1 Data | Roofline measurements | data/roofline.csv |
| Energy Model | Calibration results | data/energy_calib.csv |
| Memory Results | Workspace measurements | data/memory.csv |
| Plot Scripts | Figure generation code | plots/generate_all.py |
9.3. Key Results
10. Related Work
11. Limitations and Future Work
12. Conclusions
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Proof of Classical I/O Lower Bound
- A red pebble on a node means the value is in fast memory
- A blue pebble means the value is in slow memory
- Moving a pebble from blue to red incurs I/O cost
Appendix B. Proof of Square Strassen-like I/O Lower Bound
References
- V. Strassen, “Gaussian elimination is not optimal,” Numerische Mathematik, vol. 13, no. 4, pp. 354–356, 1969.
- V. Vassilevska Williams, Y. Xu, Z. Xu, R. Zhou, “New Bounds for Matrix Multiplication: from Alpha to Omega, 2023. arXiv:2307.07970.
- J. Alman, “More Asymmetry Yields Faster Matrix Multiplication, 2024. arXiv:2404.16349.
- J. Alman, V. Vassilevska Williams, “Limits on the Universal Method for Matrix Multiplication,” Theory of Computing, vol. 17, pp. 1–32, 2021. arXiv:1812.08731.
- J.-W. Hong and H. T. Kung, “I/O complexity: The red-blue pebble game,” Proceedings of the 13th Annual ACM Symposium on Theory of Computing, pp. 326–333, 1981. [CrossRef]
- D. Irony, S. Toledo, and A. Tiskin, “Communication lower bounds for distributed-memory matrix multiplication,” Journal of Parallel and Distributed Computing, vol. 64, no. 9, pp. 1017–1026, 2004. [CrossRef]
- G. Ballard, E. Carson, J. Demmel, M. Hoemmen, N. Knight, O. Schwartz, “Communication lower bounds and optimal algorithms for numerical linear algebra,” Acta Numerica, vol. 23, pp. 1–155, 2014. [CrossRef]
- G. Ballard, J. Demmel, O. Holtz, O. Schwartz, “Graph Expansion and Communication Costs of Fast Matrix Multiplication,” Journal of the ACM, vol. 59, no. 6, article 32, 2012. [CrossRef]
- A. Fawzi et al., “Discovering faster matrix multiplication algorithms with reinforcement learning,” Nature, vol. 610, pp. 47–53, 2022. [CrossRef]
- J.-G. Dumas, C. Pernet, A. Sedoglavic, “A non-commutative algorithm for multiplying 4×4 matrices using 48 non-complex multiplications. arXiv:2506.13242, 2025.
- N. J. Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed. Philadelphia, PA: SIAM, 2002. [CrossRef]
- M. Frigo, C. E. Leiserson, H. Prokop, S. Ramachandran, “Cache-oblivious algorithms,” ACM Transactions on Algorithms, vol. 8, no. 1, article 4, 2012. [CrossRef]
- S. Williams, A. Waterman, D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM, vol. 52, no. 4, pp. 65–76, 2009. [CrossRef]
| Parameter | Value | 95% CI | Units |
| 2.1 | pJ/FLOP | ||
| 15.2 | pJ/byte | ||
| 1.8 | pJ/byte |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).