Submitted:
27 November 2024
Posted:
28 November 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We constructed a Hybrid KAN model that integrates CNNs, Transformers, and KANs into a unified framework, leveraging the strengths of each component. CNNs are used to extract local features, Transformers capture global dependencies, while KANs accurately capture complex patterns and subtle differences in fabric defect images through their flexible nonlinear feature extraction capabilities.
- We proposed a simple KAN Conv Block (KCB) using KANConv to extract local features from images. Compared to traditional convolutional modules, KCB reduces model complexity and parameter count while maintaining similar accuracy.
- We replaced the traditional MLP in the Transformer architecture with KAN, constructing a new KAN Transformer Block (KTB) for extracting global features from images. KTB leverages the characteristics of KAN to enhance the Transformer’s ability to capture global contextual information in images, making it more adaptable when dealing with complex fabric texture images.
2. Related Work
2.1. Semantic Segmentation
2.2. Kolmogorov–Arnold Networks
3. Methodlogy
3.1. Preliminary
3.2. Overview of the Proposed Method
3.3. Hybrid KAN Block
3.3.1. KAN Conv Block
3.3.2. KAN Transformer Block
4. Experiments
4.1. Datasets
- Four Fabric Defects Dataset : Currently, there is no multi-class semantic segmentation dataset specifically for fabric defects. Therefore, we collect and annotate four types of fabric defects: Hole, Oil, Stain, and ThreadError. The four types of fabric defects consist of a total of 325 images, and image enhancement techniques such as rotation, scaling, and random cropping are used to expand the dataset to 1625 images.
- Fabric Dataset [37]: This dataset contains a total of 1600 fabric defect images, which only include defects and the background. The defects in the dataset have weak features and contain false defects, making the detection task quite challenging.
- ZJU-Leaper Dataset [3]: This is currently the largest fabric dataset. It consists of 15 types of fabric texture images, which are divided into 4 groups based on the complexity of the background texture. We randomly selected 500 defect images for each of the 15 types of fabric textures, resulting in a total of 7,500 images to be used as the dataset.
4.2. Implementation Details and Evaluation Metrics
- Implementation Details: The experiments in this paper are deployed on mmsegmentation. The learning rate is set at 0.01, momentum is at 0.9, using the stochastic gradient descent algorithm, with a weight decay of 0.0005. The maximum number of iterations is set at 40,000, with a batch size of 8, on an RTX 4060Ti GPU.
- Evaluation Metrics: We use Pixel Accuracy (PA), Mean Intersection over Union (mIoU), and Mean Dice Coefficient (mDice) as evaluation metrics.
4.3. Comparative Experiments
4.3.1. Results on Four Fabirc Defects Dataset
4.3.2. Results on Fabric Dataset
4.3.3. Results on ZJU-Leaper Dataset
4.3.4. Visualization Results
4.4. Ablation Experiments
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Ngan, H.Y.; Pang, G.K.; Yung, N.H. Automated fabric defect detection—A review. Image and vision computing 2011, 29, 442–458. [Google Scholar] [CrossRef]
- Rasheed, A.; Zafar, B.; Rasheed, A.; Ali, N.; Sajid, M.; Dar, S.H.; Habib, U.; Shehryar, T.; Mahmood, M.T. Fabric defect detection using computer vision techniques: a comprehensive review. Mathematical Problems in Engineering 2020, 2020, 8189403. [Google Scholar] [CrossRef]
- Zhang, C.; Feng, S.; Wang, X.; Wang, Y. Zju-leaper: A benchmark dataset for fabric defect detection and a comparative study. IEEE Transactions on Artificial Intelligence 2020, 1, 219–232. [Google Scholar] [CrossRef]
- Zhao, S.; Yin, L.; Zhang, J.; Wang, J.; Zhong, R. Real-time fabric defect detection based on multi-scale convolutional neural network. IET Collaborative Intelligent Manufacturing 2020, 2, 189–196. [Google Scholar] [CrossRef]
- Huang, Y.; Jing, J.; Wang, Z. Fabric defect segmentation method based on deep learning. IEEE Transactions on Instrumentation and Measurement 2021, 70, 1–15. [Google Scholar] [CrossRef]
- Koulali, I.; Eskil, M.T. Unsupervised textile defect detection using convolutional neural networks. Applied Soft Computing 2021, 113, 107913. [Google Scholar] [CrossRef]
- Chen, M.; Yu, L.; Zhi, C.; Sun, R.; Zhu, S.; Gao, Z.; Ke, Z.; Zhu, M.; Zhang, Y. Improved faster R-CNN for fabric defect detection based on Gabor filter with Genetic Algorithm optimization. Computers in Industry 2022, 134, 103551. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 2016, 39, 1137–1149. [Google Scholar] [CrossRef]
- Lu, B.; Huang, B. A texture-aware one-stage fabric defect detection network with adaptive feature fusion and multi-task training. Journal of Intelligent Manufacturing 2024, 35, 1267–1280. [Google Scholar] [CrossRef]
- Vaswani, A. Attention is all you need. Advances in Neural Information Processing Systems 2017. [Google Scholar]
- Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022.
- Qu, H.; Di, L.; Liang, J.; Liu, H. U-SMR: U-SwinT & multi-residual network for fabric defect detection. Engineering Applications of Artificial Intelligence 2023, 126, 107094. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
- Xu, H.; Liu, C.; Duan, S.; Ren, L.; Cheng, G.; Hao, B. A Fabric Defect Segmentation Model Based on Improved Swin-Unet with Gabor Filter. Applied Sciences 2023, 13, 11386. [Google Scholar] [CrossRef]
- Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-unet: Unet-like pure transformer for medical image segmentation. European conference on computer vision. Springer, 2022, pp. 205–218.
- Hornik, K.; Stinchcombe, M.; White, H. Multilayer feedforward networks are universal approximators. Neural networks 1989, 2, 359–366. [Google Scholar] [CrossRef]
- Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljačić, M.; Hou, T.Y.; Tegmark, M. Kan: Kolmogorov-arnold networks. arXiv 2024, arXiv:2404.19756. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440.
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer, 2015, pp. 234–241.
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2881–2890.
- Xiao, T.; Liu, Y.; Zhou, B.; Jiang, Y.; Sun, J. Unified perceptual parsing for scene understanding. Proceedings of the European conference on computer vision (ECCV), 2018, pp. 418–434.
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European conference on computer vision (ECCV), 2018, pp. 801–818.
- Zhang, H.; Dana, K.; Shi, J.; Zhang, Z.; Wang, X.; Tyagi, A.; Agrawal, A. Context encoding for semantic segmentation. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 7151–7160.
- Kirillov, A.; Girshick, R.; He, K.; Dollár, P. Panoptic feature pyramid networks. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 6399–6408.
- Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.; others. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 6881–6890.
- Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for semantic segmentation. Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 7262–7272.
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems 2021, 34, 12077–12090. [Google Scholar]
- Zhang, W.; Pang, J.; Chen, K.; Loy, C.C. K-net: Towards unified image segmentation. Advances in Neural Information Processing Systems 2021, 34, 10326–10338. [Google Scholar]
- Cheng, B.; Schwing, A.; Kirillov, A. Per-pixel classification is not all you need for semantic segmentation. Advances in neural information processing systems 2021, 34, 17864–17875. [Google Scholar]
- Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1290–1299.
- Li, C.; Liu, X.; Li, W.; Wang, C.; Liu, H.; Liu, Y.; Chen, Z.; Yuan, Y. U-kan makes strong backbone for medical image segmentation and generation. arXiv 2024, arXiv:2406.02918. [Google Scholar]
- Rege Cambrin, D.; Poeta, E.; Pastor, E.; Cerquitelli, T.; Baralis, E.; Garza, P. KAN You See It? KANs and Sentinel for Effective and Explainable Crop Field Segmentation. arXiv e-prints 2024, pp. arXiv–2408.
- Bodner, A.D.; Tepsich, A.S.; Spolski, J.N.; Pourteau, S. Convolutional Kolmogorov-Arnold Networks. arXiv 2024, arXiv:2406.13155. [Google Scholar]
- Yang, X.; Wang, X. Kolmogorov-arnold transformer. arXiv 2024, arXiv:2409.10594. [Google Scholar]
- He, Y.; Xie, Y.; Yuan, Z.; Sun, L. MLP-KAN: Unifying Deep Representation and Function Learning. arXiv 2024, arXiv:2410.03027. [Google Scholar]
- Wang, J.; Xu, G.; Li, C.; Gao, G.; Wu, Q. Sddet: An enhanced encoder–decoder network with hierarchical supervision for surface defect detection. IEEE Sensors Journal 2022, 23, 2651–2662. [Google Scholar] [CrossRef]
- Shi, D. TransNeXt: Robust Foveal Visual Perception for Vision Transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 17773–17783.






| Method | PA (%) | mIoU (%) | mDice (%) |
|---|---|---|---|
| UNet [20] | 96.57 | 70.59 | 83.08 |
| PSPNet [21] | 98.94 | 87.60 | 93.20 |
| DeepLabV3+ [23] | 99.08 | 89.59 | 94.32 |
| Segformer [28] | 99.21 | 90.79 | 94.99 |
| KNet [29] | 99.26 | 90.86 | 95.01 |
| TransNeXt [38] | 99.16 | 89.87 | 94.46 |
| HKAN | 99.33 | 91.60 | 95.46 |
| Method | PA (%) | mIoU (%) | mDice (%) |
|---|---|---|---|
| UNet [20] | 99.67 | 64.14 | 72.16 |
| PSPNet [21] | 99.70 | 67.33 | 75.83 |
| DeepLabV3+ [23] | 99.08 | 68.24 | 76.82 |
| Segformer [28] | 99.70 | 69.74 | 78.38 |
| KNet [29] | 99.69 | 69.35 | 77.99 |
| TransNeXt [38] | 99.69 | 69.57 | 78.21 |
| HKAN | 99.68 | 70.41 | 79.08 |
| Model | Group1 | Group2 | ||||
|---|---|---|---|---|---|---|
| PA (%) | mIoU (%) | mDice (%) | PA (%) | mIoU (%) | mDice (%) | |
| U-Net [20] | 97.26 | 83.17 | 90.18 | 96.69 | 81.46 | 89.02 |
| PSPNet [21] | 97.32 | 84.81 | 91.30 | 97.11 | 83.32 | 90.30 |
| DeepLabV3+ [23] | 97.65 | 85.73 | 91.88 | 97.53 | 85.74 | 91.90 |
| Segformer [28] | 97.75 | 86.40 | 92.32 | 97.45 | 85.45 | 91.71 |
| KNet [29] | 97.73 | 86.49 | 92.38 | 97.47 | 85.42 | 91.69 |
| TransNeXt [38] | 97.57 | 85.45 | 91.71 | 97.18 | 83.92 | 90.71 |
| HKAN | 97.77 | 86.63 | 92.46 | 97.53 | 85.88 | 91.99 |
| Model | Group3 | Group4 | ||||
|---|---|---|---|---|---|---|
| PA (%) | mIoU (%) | mDice (%) | PA (%) | mIoU (%) | mDice (%) | |
| U-Net [20] | 97.60 | 82.57 | 89.72 | 96.21 | 80.47 | 88.36 |
| PSPNet [21] | 98.44 | 87.99 | 93.28 | 97.96 | 88.81 | 93.84 |
| DeepLabV3+ [23] | 98.33 | 87.43 | 92.94 | 97.91 | 88.67 | 93.75 |
| Segformer [28] | 98.39 | 87.67 | 93.09 | 97.89 | 88.44 | 93.61 |
| KNet [29] | 98.19 | 86.45 | 92.31 | 97.68 | 87.81 | 93.23 |
| TransNeXt [38] | 98.21 | 85.80 | 92.54 | 97.61 | 87.27 | 92.89 |
| HKAN | 98.43 | 88.10 | 93.36 | 97.98 | 89.10 | 94.01 |
| Method | PA (%) | mIoU (%) | mDice (%) |
|---|---|---|---|
| CNN Encoder | 96.57 | 70.59 | 83.08 |
| Transformer Encoder | 98.69 | 85.28 | 91.66 |
| KANConv Encoder | 98.09 | 81.60 | 89.15 |
| KANTransformer Encoder | 98.93 | 87.52 | 93.10 |
| Hybrid KAN Encoder | 99.26 | 90.84 | 95.02 |
| HKAN | 99.33 | 91.60 | 95.46 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).