Submitted:
13 July 2026
Posted:
13 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We developed the first method that combines PCA and SVD, along with their quaternion forms, to work together for grey-level and colored image distillation.
- The presentation of the entire training set as a single quaternion matrix, which is first modified by the M-PCA and, for the second time, by block-rotating the left singular matrix of the SVD for grey-level images and QSVD for color images.
2. Related Work
2.1. Gradient-Based Dataset Distillation
2.2. Knowledge Distillation and Teacher-Student Models
2.3. Non-Gradient-Based Dataset Distillation
3. Proposed Method
3.1. Extension of M-PCA to Grey-Level Images
3.2. SVD-Based Distillation
3.3. Block Rotation
3.4. Extension to Color Images via Quaternion PCA
3.5. Distillation from the Modified Training Color Images
4. Experimental Setup
4.1. Datasets
- Digit-MNIST [26]: A grayscale dataset of handwritten digits (0–9), consisting of 60,000 training images and 10,000 test images, each of size pixels, distributed across 10 classes.
- Fashion-MNIST [27]: A grayscale dataset of clothing items (e.g., shirts, shoes, bags), with the same structure as Digit-MNIST: 60,000 training and 10,000 test images at pixels across 10 classes. It is considered a bit more challenging benchmark due to greater intra-class visual variation.
- CIFAR-10 [28]: A color image dataset of pixels comprising 50,000 training and 10,000 test images distributed across 10 classes (e.g., airplanes, birds, and cats). It’s higher visual complexity and the use of color makes the image database a significantly harder classification problem than the MNIST variants.
- CIFAR-100 [29]: A color image dataset sharing the same image size and overall structure as CIFAR-10, but spanning 100 fine-grained classes with 500 training images and 100 test images per class. Its large class count and limited per-class samples make it the most demanding benchmark considered in this work.
- BloodMNIST [30]: A color image dataset of pixels comprising 11,959 training, 1,712 validation, and 3,421 test images distributed across 8 classes (e.g., Basophils, Eosinophils, and Erythroblasts). It consists of microscopic images of human blood cells and is widely used for training and testing machine learning models.
4.2. CNN Architectures
- Baseline CNN [31]: A lightweight network used for classifying Digit-MNIST, Fashion-MNIST, and BloodMNIST. It consists of two convolutional layers followed by a ReLU activation function and a max-pooling operation. The output layer is fully connected and maps to the number of target classes. This compact architecture is a standard low-complexity baseline in the dataset distillation literature and is well suited to the greyscale inputs of the MNIST-family datasets.
- ResNet50V2 [32]: 50-layer residual network for CIFAR-10 and CIFAR-100 classification. ResNet50V2 adopts a pre-activation residual design, placing batch normalization and activation before each convolution, allowing for more stable gradient flow when training on complex datasets. For all experiments, the backbone was kept fixed, and only the fully connected classification head attached to the network’s output was trained on the distilled images. This protocol isolates the quality of the distilled data from any confounding effect of fine-tuning the deep feature extractor.
4.3. Training Setup
- Framework: All models were implemented and trained in Python using TensorFlow/Keras.
- Images per class (IPC): We distilled 10, 20, and 50 images per class for Digit-MNIST and Fashion-MNIST, and 10 and 50 images per class for CIFAR-10, CIFAR-100, and BloodMNIST.
- Epochs: Each model was trained for multiple epochs that varied depending on the dataset. For Digit-MNIST, Fashion-MNIST, and CIFAR-10, the number of epochs used was 100, 200, 300, 600, and 1000; for CIFAR-100, the number of epochs was 100, 200, and 400; and for BloodMNIST, 100, 200, 400, and 600 epochs were conducted as summarized in Table 1.
- Batch size: Set to 32 for all experiments.
- Loss function: Categorical cross-entropy was used since all tasks are multiclass classification problems:where n is the number of training samples, c is the number of classes, and is the ground truth label of the i-th input sample for class . Specifically,and is the predicted probability for the i-th sample to belong to class .
- Optimizer: The Adam optimizer was implemented by the two classifiers selected by us with a fixed learning rate of . No learning rate decay was applied, ensuring consistent and comparable results across all models and datasets.
5. Experimental Results
5.1. Ablation Study
6. Discussion
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| M-PCA | Modified Principal Component Analysis |
| SVD | Singular Value Decomposition |
| M-PCA-SVD | Modified PCA and SVD method for grey-scale images |
| QM-PCA | Quaternion Modified PCA |
| QSVD | Quaternion SVD |
| QM-PCA-QSVD | Quaternion Modified PCA and Quaternion SVD method for colored images |
| NN | Neural Network |
| ML | Machine Learning |
| IPC | Number of Images per class |
| CNN | Convolutional Neural Network |
| AC | Accuracy |
| TP | True Positive |
| TN | True Negative |
| FP | False Positive |
| FN | False Negative |
References
- Lei, S.; Tao, D. A comprehensive survey of dataset distillation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 17–32. [Google Scholar] [CrossRef] [PubMed]
- Sachdeva, N.; McAuley, J. Data distillation: A survey. arXiv 2023, arXiv:2301.04272. [Google Scholar]
- Wang, T.; Zhu, J.Y.; Torralba, A.; Efros, A.A. Dataset distillation. arXiv 2018, arXiv:1811.10959. [Google Scholar]
- Zeng, X.; Ahmed, A.; Tunio, M.H. HFed-MIL: Patch Gradient-Based Attention Distillation Federated Learning for Heterogeneous Multi-Site Ovarian Cancer Whole-Slide Image Analysis. Electronics 2025, 14, 3600. [Google Scholar] [CrossRef]
- Sirakov, N.M.; Shahnewaz, T.; Nakhmani, A. Training Data Augmentation with Data Distilled by Principal Component Analysis. Electronics 2024, 13. [Google Scholar] [CrossRef]
- Li, G.; Togo, R.; Ogawa, T.; Haseyama, M. Importance-aware adaptive dataset distillation. Neural Netw. 2024, 172, 106154. [Google Scholar] [CrossRef] [PubMed]
- Su, D.; Hou, J.; Gao, W.; Tian, Y.; Tang, B. D 4: Dataset distillation via disentangled diffusion model. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 5809–5818. [Google Scholar]
- Ezekwem, N.N.; Sirakov, N.M. Image Distillation with the Machine-Learned Gradient of the Loss Function and the K-Means Method. Mathematics 2025, 13. [Google Scholar] [CrossRef]
- Nguyen, T.; Chen, Z.; Lee, J. Dataset meta-learning from kernel ridge-regression. arXiv 2020, arXiv:2011.00050. [Google Scholar]
- Liu, Y.; Gu, J.; Wang, K.; Zhu, Z.; Jiang, W.; You, Y. Dream: Efficient dataset distillation by representative matching. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 17314–17324. [Google Scholar]
- Sucholutsky, I.; Schonlau, M. Soft-label dataset distillation and text dataset distillation. In Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN); IEEE, 2021; pp. 1–8. [Google Scholar]
- Bohdal, O.; Yang, Y.; Hospedales, T. Flexible dataset distillation: Learn labels instead of images. arXiv 2020, arXiv:2006.08572. [Google Scholar]
- Li, G.; Togo, R.; Ogawa, T.; Haseyama, M. Compressed gastric image generation based on soft-label dataset distillation for medical data sharing. Comput. Methods Programs Biomed. 2022, 227, 107189. [Google Scholar] [CrossRef] [PubMed]
- Vinaroz, M.; Park, M.J. Differentially private kernel inducing points using features from scatternets (DP-KIP-scatternet) for privacy preserving data distillation. arXiv 2023, arXiv:2301.13389. [Google Scholar]
- Ganesh, P.; Chen, Y.; Lou, X.; Khan, M.A.; Yang, Y.; Sajjad, H.; Nakov, P.; Chen, D.; Winslett, M. Compressing large-scale transformer-based models: A case study on bert. Trans. Assoc. Comput. Linguist. 2021, 9, 1061–1080. [Google Scholar] [CrossRef]
- Tan, H.; Wang, W.; Wu, S.; Wu, X.; Sun, Y.T.; Chang, C.; Zhang, S.; Qi, X. Dataset Distillation by Influence Matching. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 19654–19664. [Google Scholar]
- Le, Y.; Yang, X. Tiny imagenet visual recognition challenge. CS 231N 2015, 7, 3. [Google Scholar]
- Xu, Y.; Hu, C.; An, P.; Li, Y.L. Mitigating The Distribution Shift of Diffusion-based Dataset Distillation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 33943–33952. [Google Scholar]
- Sirakov, N.M.; Ngo, L.H. Automatic image distillation with wavelet transform and modified principal component analysis. Electronics 2025, 14, 1357. [Google Scholar] [CrossRef]
- Lowe, S.C.; Fuller, A.; Oore, S.; Shelhamer, E.; Taylor, G.W. BRIDGING GENERATIVE AND PREDICTIVE PARADIGMS VIA HIDDEN-SELF-DISTILLATION. ICLR Access, 2026. [Google Scholar]
- Golub, G.H.; Reinsch, C. Singular value decomposition and least squares solutions. In Linear algebra; Springer, 1971; pp. 134–151. [Google Scholar]
- Chen, M.; Wang, C.; Meng, X.; Wang, Z. Quaternion Principal Component Analysis for Multi-modal Fusion. In Proceedings of the International Conference on Genetic and Evolutionary Computing, 2015; Springer; pp. 11–19. [Google Scholar]
- Ngo, L.H.; Luong, M.; Sirakov, N.M.; Viennet, E.; Le-Tien, T. Skin lesion image classification using sparse representation in quaternion wavelet domain. Signal Image Video Process. 2022, 16, 1721–1729. [Google Scholar] [CrossRef]
- Abdi, H.; Williams, L.J. Principal component analysis. Wiley Interdiscip. Rev. Comput. Stat. 2010, 2, 433–459. [Google Scholar] [CrossRef]
- Chang, J.H.; Ding, J.J.; et al. Quaternion matrix singular value decomposition and its applications for color image processing. In Proceedings of the Proceedings 2003 international conference on image processing (Cat. No. 03CH37429).; IEEE, 2003; Vol. 1, pp. I–805. [Google Scholar]
- LeCun, Y.; Cortes, C.; Burges, C.C. MNIST handwritten digit database. ATT Labs [Online] 2010, 2, 18. [Google Scholar]
- Xiao, H.; Rasul, K.; Vollgraf, R. Fashion-mnist: A novel image dataset for benchmarking machine learning algorithms. arXiv 2017, arXiv:1708.07747. [Google Scholar]
- Krizhevsky, A.; Hinton, G.; et al. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009. [Google Scholar]
- Krizhevsky, A. Learning multiple layers of features from tiny images; Technical report; University of Toronto, 2009. [Google Scholar]
- Yang, J.; Shi, R.; Wei, D.; Liu, Z.; Zhao, L.; Ke, B.; Pfister, H.; Ni, B. MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data 2023, 10, 41. [Google Scholar] [CrossRef] [PubMed]
- Shin, M.; Kim, M.; Kwon, D.S. Baseline CNN structure analysis for facial expression recognition. In Proceedings of the 2016 25th IEEE international symposium on robot and human interactive communication (RO-MAN); IEEE, 2016; pp. 724–729. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Identity mappings in deep residual networks. In Proceedings of the Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016; Springer, 2016; Proceedings, Part IV 14, pp. 630–645. [Google Scholar]
- Shobiry, A.K.F.; Puspitasari, R.; et al. Hybrid Feature Benchmark for Blood Cell Classification Using ResNet50 and EfficientNetV2 Features with SVM and ANN Classifiers via Unsupervised Segmentation. Int. J. Artif. Intell. Med. Issues 2025, 3, 110–126. [Google Scholar]
- Banerjee, A. Computational Complexity of PCA. 2020. Available online: https://alekhyo.medium.com/computational-complexity-of-pca-4cb61143b7e5 (accessed on 2026-06-30).
- Developers, Scikit-learn. Decomposing signals in components (matrix factorization problems). Available online: https://scikit-learn.org/stable/modules/decomposition.html (accessed on 2026-06-30).







| Dataset | Network | Distilled IPC | Epochs Tested | Classes |
|---|---|---|---|---|
| Digit-MNIST | Baseline CNN | 10, 20, 50 | 100, 200, 300, 600, 1000 | 10 |
| Fashion-MNIST | Baseline CNN | 10, 20, 50 | 100, 200, 300, 600, 1000 | 10 |
| CIFAR-10 | ResNet50V2 | 10, 50 | 100, 200, 400, 600, 1000 | 10 |
| CIFAR-100 | ResNet50V2 | 10, 50 | 100, 200, 400 | 100 |
| BloodMNIST | Baseline CNN | 10, 50 | 100, 200, 400, 600 | 8 |
| Dataset | IPC | Epochs | ||||
|---|---|---|---|---|---|---|
| 100 | 200 | 300 | 600 | 1000 | ||
| 10 | 88.31 | 89.46 | 89.92 | 89.26 | 89.59 | |
| Digit-MNIST | 20 | 92.74 | 93.13 | 93.17 | — | — |
| 50 | 96.89 | 97.43 | 97.68 | 97.12 | 98.01 | |
| 10 | 76.87 | 80.43 | 76.29 | 79.33 | 80.63 | |
| Fashion-MNIST | 20 | 81.55 | 82.22 | 82.69 | — | — |
| 50 | 85.18 | 87.38 | 86.53 | 85.44 | 85.21 | |
| Dataset | IPC | Epochs | Time (sec) | Accuracy (%) |
|---|---|---|---|---|
| Digit-MNIST | 50 (distilled) | 300 | 900 | 97.68 |
| 6000 (original) | 50 | 3656 | 99.54 | |
| Fashion-MNIST | 50 (distilled) | 200 | 1060 | 87.38 |
| 6000 (original) | 50 | 5493 | 92.91 |
| Dataset | IPC | Type | Network | Epochs | ||||
|---|---|---|---|---|---|---|---|---|
| 100 | 200 | 400 | 600 | 1000 | ||||
| CIFAR-10 | 10 | Original | ResNet50V2 | 41.42 | 41.63 | 42.13 | 43.68 | 42.17 |
| Distilled | 72.91 | 73.85 | 73.96 | 73.68 | 73.39 | |||
| CIFAR-10 | 50 | Original | ResNet50V2 | 55.53 | 56.28 | 57.55 | 58.35 | 57.37 |
| Distilled | 73.99 | 74.16 | 75.59 | 74.55 | 74.76 | |||
| CIFAR-100 | 10 | Original | ResNet50V2 | 30.05 | 31.10 | 30.95 | — | — |
| Distilled | 39.98 | 41.27 | 41.95 | — | — | |||
| CIFAR-100 | 50 | Original | ResNet50V2 | 44.34 | 45.47 | 45.15 | — | — |
| Distilled | 51.56 | 55.90 | 56.27 | — | — | |||
| BloodMNIST | 10 | Original | Baseline NN | 65.92 | 67.99 | 73.21 | 70.95 | — |
| Distilled | 70.91 | 72.36 | 73.81 | 72.15 | — | |||
| BloodMNIST | 50 | Original | Baseline NN | 80.33 | 81.44 | 83.98 | 80.88 | — |
| Distilled | 80.83 | 82.39 | 85.47 | 85.30 | — | |||
| Method | Scheme | Digit-MNIST | Fashion-MNIST | CIFAR-10 | CIFAR-100 | ||||
|---|---|---|---|---|---|---|---|---|---|
| IPC=10 | IPC=50 | IPC=10 | IPC=50 | IPC=10 | IPC=50 | IPC=10 | IPC=50 | ||
| Random | — | 95.1 | 97.9 | 73.8 | 82.5 | 26.0 | 43.4 | 14.6 | 30.0 |
| Herding | — | 93.7 | 94.8 | 71.1 | 71.9 | 31.6 | 40.4 | 17.3 | 33.7 |
| DC | GM | 94.7 | 98.8 | 82.3 | 83.6 | 44.9 | 53.9 | 26.6 | 32.1 |
| DSA | GM | 97.8 | 99.2 | 86.6 | 88.7 | 52.1 | 60.6 | 32.4 | 38.6 |
| DCC | GM | — | — | — | — | 54.5 | 64.2 | 33.5 | 39.3 |
| DM | DM | 97.3 | 94.8 | — | — | 48.9 | 63.0 | 29.7 | 43.6 |
| CAFE | DM | 97.5 | 98.9 | 83.0 | 88.2 | 50.9 | 62.3 | 31.5 | 42.9 |
| MTT | TM | 97.3 | 98.5 | 87.2 | 88.3 | 65.3 | 71.6 | 40.1 | 47.7 |
| FTD | TM | — | — | — | — | 66.6 | 73.8 | 43.4 | 50.7 |
| TESLA | TM | — | — | — | — | 66.4 | 72.6 | 41.7 | 47.9 |
| KIP | KRR | 97.5 | 98.3 | 86.8 | 88.0 | 62.7 | 68.6 | 28.3 | — |
| FRePo | KRR | 98.6 | 99.2 | 86.2 | 89.6 | 65.5 | 71.7 | 42.5 | 44.3 |
| RFAD | KRR | 98.5 | 98.8 | 87.0 | 88.8 | 66.3 | 71.1 | 33.0 | — |
| RCIG | KRR | 98.9 | 99.2 | 88.5 | 90.2 | 69.1 | 73.5 | 44.1 | 46.7 |
| IDC* | GM | 98.4 | 99.1 | 86.0 | 86.2 | 67.5 | 74.5 | 44.8 | — |
| DREAM* | GM | 98.6 | 99.2 | 86.4 | 86.8 | 69.4 | 74.8 | 46.8 | 52.6 |
| RTP* | BPTT | 99.3 | 99.4 | 90.0 | 91.2 | 71.2 | 73.6 | 42.9 | — |
| HaBa* | TM | — | — | — | — | 69.9 | 74.0 | 40.2 | 47.0 |
| IDM* | DM | — | — | — | — | 58.6 | 67.5 | 45.1 | 50.0 |
| KFS* | DM | — | — | — | — | 72.0 | 75.0 | 40.0 | 50.6 |
| IADD | PM | — | — | — | — | ||||
| Ours (M-PCA-SVD) | — | 89.92 | 98.01 | 80.63 | 87.38 | — | — | — | — |
| Ours (QM-PCA-QSVD) | — | — | — | — | — | 73.96 | 75.59 | 41.95 | 56.27 |
| Configuration | Distillation | Network | IPC | Accuracy (%) |
|---|---|---|---|---|
| Original (full dataset) | None | Baseline CNN | 50 | 81.26 |
| M-PCA only | Partial | Baseline CNN | 50 | 84.16 |
| SVM + Grid Search CV [33] | None | EfficientNetV2 | WTS | 76.80 |
| QM-PCA-QSVD + Rotation (Ours) | Full | Baseline CNN | 50 | 85.47 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).