Submitted:
02 January 2025
Posted:
03 January 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Materials and Methods
| Algorithm 1 Advancing colonoscopy analysis through text-to-image synthesis using generative AI | |
| |
| Helper Functions: | |
| FineTune(G, D): | Fine-tune model using DreamBooth and LoRA |
| SynthesizeImages(G, prompts, n): | Generate n images using G |
| GenerateMasks(D, SAM): | Generate masks for images in D using SAM |
| TrainClassifier(M, D): | Train classifier M on dataset D |
| EvaluateClassifier(C, D): | Evaluate classifier C |
| TrainSegmentation(S, D): | Train segmentation model S on dataset D |
| EvaluateSegmentation(S, D): | Evaluate segmentation model S |
4. Results
4.1. Image Generation Results
4.2. Model Comparison
4.3. Image Segmentation Results
4.4. Image Classification Results
-
High PerformanceThe augmented dataset demonstrated exceptional performance, consistently achieving accuracy metrics above 92% in different models and evaluation metrics. This indicates robust and reliable model behavior across various architectural implementations.
-
Performance ConsistencyA notable pattern emerged in the comparative analysis. The original dataset showed inconsistent performance, with only one model achieving high accuracy (EfficientNet at 97% validation accuracy). The synthetic dataset consistently showed lower performance across all models. In contrast, the augmented dataset maintained consistently high performance across multiple architectures, suggesting improved data quality and representation.
-
Strong GeneralizationThe minimal gap between training and validation accuracy in the augmented dataset (typically within 2-3 percentage points) indicates effective knowledge transfer, reduced overfitting, and robust model generalization capabilities.
-
Architecture-Agnostic PerformanceUnlike the original and synthetic datasets, which showed significant performance variations between different architectures, the augmented dataset demonstrated more balanced performance. This suggests reduced architecture-specific bias in the learning process.
-
Enhanced Model ApplicabilityThe consistent high performance across diverse architectural approaches suggests that the knowledge extracted from the augmented dataset is more universally applicable. The learned features are more generalizable across different model architectures, and the dataset provides robust training signals for various deep learning approaches.
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AIGC | AI-Generated Content |
| ArSDM | Arbitrary-Style Diffusion Model |
| AUC | Area Under the Curve |
| BiT | Big Transfer |
| BLIP | Bootstrapping Language-Image Pre-training |
| CADe | Computer-aided Detection |
| CFM | Cross Fusion Module |
| cGAN | Conditional Generative Adversarial Network |
| CIM | Cross Interaction Module |
| CLIP | Contrastive Language-Image Pre-Training |
| CNN | Convolutional Neural Network |
| CRC | Colorectal Cancer |
| DB | DreamBooth |
| DeiT | Data-efficient Image Transformers |
| DL | Deep Learning |
| DR | Adenoma Detecting Rate |
| FCNN | Fully Convolutional Neural Network |
| FID | Fréchet Inception Distance |
| FPN | Feature Pyramid Network |
| GANs | Generative Adversarial Networks |
| HD | High Definition |
| IS | Inception Score |
| LDM | Latent Diffusion Model |
| LinkNet | Link Network |
| LoRA | Low-Rank Adaptation |
| MANet | Multi-scale Attention Network |
| Mask R-CNN | Mask Region-based Convolutional Neural Network |
| mDice | Mean Dice Coefficient |
| mIoU | Mean Intersection over Union |
| ML | Machine Learning |
| Polyp-PVT | Polyp Pyramid Vision Transformer |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| PSPNet | Pyramid Scene Parsing Network |
| SAM | Spatial Attention Module |
| SD | Stable Diffusion |
| SSPP | Single Sample Per Person |
| U-Net | U-shaped Network |
| VQGAN | Vector Quantized Generative Adversarial Network |
References
- Wang, P.; et al. Real-time automatic detection system increases colonoscopic polyp and adenoma detection rates: a prospective randomised controlled study. Gut 2019, 68, 1813–1819. [CrossRef]
- Bernal, J.; et al. Comparative Validation of Polyp Detection Methods in Video Colonoscopy: Results from the MICCAI 2015 Endoscopic Vision Challenge. IEEE Trans Med Imaging 2017, 36, 1231–1249. [CrossRef]
- Kim, J.J.H.; Um, R.S.; Lee, J.W.Y.; Ajilore, O. Generative AI can fabricate advanced scientific visualizations: ethical implications and strategic mitigation framework. AI and Ethics 2024. [CrossRef]
- Videau, M.; Knizev, N.; Leite, A.; Schoenauer, M.; Teytaud, O. Interactive Latent Diffusion Model. In Proceedings of the Genetic and Evolutionary Computation Conference; ACM: New York, NY, USA, 2023; pp. 586–596. [CrossRef]
- Bray, F.; Ferlay, J.; Soerjomataram, I.; Siegel, R.L.; Torre, L.A.; Jemal, A. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2018, 68, 394–424. [CrossRef]
- S. K. Alhabeeb and A. A. Al-Shargabi, "Text-to-Image Synthesis with Generative Models: Methods, Datasets, Performance Metrics, Challenges, and Future Direction," IEEE Access, vol. 1, pp. 1-1, Jan. 2024, . [CrossRef]
- Y. X. Tan, C. P. Lee, N. Mai, K. M. Lim, J. Y. Lim, and A. Alqahtani, "Recent Advances in Text-to-Image Synthesis: Approaches, Datasets and Future Research Prospects," IEEE Access, vol. 11, pp. 88099-88115, Jan. 2023, . [CrossRef]
- G. Iglesias, E. Talavera, and A. Díaz-Álvarez, "A survey on GANs for computer vision: Recent research, analysis and taxonomy," Computer Science Review, vol. 48, pp. 100553-100553, May 2023, . [CrossRef]
- Peter, O.O.E.; Rahman, M.M.; Khalifa, F. Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA. In Conference and Labs of the Evaluation Forum 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:271772323.
- R. Najjar, "Redefining Radiology: A Review of Artificial Intelligence Integration in Medical Imaging," Diagnostics, vol. 13, pp. 2760, Aug. 2023, . [CrossRef]
- O. Abdullah, B. N. Jagadale, A. Naji, M. Ghaleb, Ahmed, A. Aqlan, and D. Esmail, "Efficient artificial intelligence approaches for medical image processing in healthcare: comprehensive review, taxonomy, and analysis," Artificial Intelligence Review, vol. 57, Jul. 2024, . [CrossRef]
- A. Arora et al., "The value of standards for health datasets in artificial intelligence-based applications," Nature Medicine, vol. 29, pp. 1-10, Oct. 2023, . [CrossRef]
- Pengxiao, H.; Changkun, Y.; Jieming, Z.; Jing, Z.; Hong, J.; Xuesong, L. Latent-based Diffusion Model for Long-tailed Recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops 2024.
- Du, Y.; et al. ArSDM: Colonoscopy Images Synthesis with Adaptive Refinement Semantic Diffusion Models. MICCAI 2023. [CrossRef]
- Ku, H.; Lee, M. TextControlGAN: Text-to-Image Synthesis with Controllable Generative Adversarial Networks. Applied Sciences 2023, 13, 5098. [CrossRef]
- Iqbal, M.A.; Jadoon, W.; Kim, S.K. Synthetic Image Generation Using Conditional GAN-Provided Single-Sample Face Image. Applied Sciences 2024, 14, 5049. [CrossRef]
- Shin, Y.; Qadir, H.A.; Aabakken, L.; Bergsland, J.; Balasingham, I. Automatic Colon Polyp Detection using Region based Deep CNN and Post Learning Approaches. 2019. [CrossRef]
- Qadir, H.A.; Shin, Y.; Solhusvik, J.; Bergsland, J.; Aabakken, L.; Balasingham, I. Polyp Detection and Segmentation using Mask R-CNN: Does a Deeper Feature Extractor CNN Always Perform Better? In International Symposium on Medical Information and Communication Technology (ISMICT); IEEE, 2019; pp. 1–6. [CrossRef]
- Dong, B.; Wang, W.; Fan, D.-P.; Li, J.; Fu, H.; Shao, L. Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers. 2021. [CrossRef]
- Repici, A.; et al. Efficacy of Real-Time Computer-Aided Detection of Colorectal Neoplasia in a Randomized Trial. Gastroenterology 2020, 159, 512-520.e7. [CrossRef]
- Kudo, S.; et al. Artificial Intelligence-assisted System Improves Endoscopic Identification of Colorectal Neoplasms. Clinical Gastroenterology and Hepatology 2020, 18, 1874-1881.e2. [CrossRef]
- Zhou, J.; et al. A novel artificial intelligence system for the assessment of bowel preparation (with video). Gastrointest Endosc 2020, 91, 428-435.e2. [CrossRef]
- Mahmood, F.; Chen, R.; Durr, N.J. Unsupervised Reverse Domain Adaptation for Synthetic Medical Images via Adversarial Training. IEEE Trans Med Imaging 2018, 37, 2572–2581. [CrossRef]
- Goceri, E. Medical image data augmentation: techniques, comparisons and interpretations. Artif Intell Rev 2023, 56, 12561–12605. [CrossRef]
- Yang, Z.; Zhan, F.; Liu, K.; Xu, M.; Lu, S. AI-Generated Images as Data Source: The Dawn of Synthetic Era. arXiv 2023. [CrossRef]
- Cao, Y.; et al. A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT. arXiv 2023. [CrossRef]
- Bandi, A.; Adapa, P.V.S.R.; Kuchi, Y.E.V.P.K. The Power of Generative AI: A Review of Requirements, Models, Input–Output Formats, Evaluation Metrics, and Challenges. Future Internet 2023, 15, 260. [CrossRef]
- Bendel, O. Image synthesis from an ethical perspective. AI Soc 2023. [CrossRef]
- Derevyanko, N.; Zalevska, O. Comparative analysis of neural networks Midjourney, Stable Diffusion, and DALL-E and ways of their implementation in the educational process of students of design specialities. Scientific Bulletin of Mukachevo State University Series "Pedagogy and Psychology" 2023, 9, 36–44. [CrossRef]
- Sánchez-Peralta, L.F.; Bote-Curiel, L.; Picón, A.; Sánchez-Margallo, F.M.; Pagador, J.B. Deep learning to find colorectal polyps in colonoscopy: A systematic literature review. Artif Intell Med 2020, 108, 101923. [CrossRef]
- Salimans, T.; et al. Improved Techniques for Training GANs. In Advances in Neural Information Processing Systems; Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., Garnett, R., Eds.; Curran Associates, Inc., 2016.
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems; Guyon, I., Von Luxburg, U., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc., 2017.
- Wang, P.; et al. Development and validation of a deep-learning algorithm for the detection of polyps during colonoscopy. Nat Biomed Eng 2018, 2, 741–748. [CrossRef]
- Misawa, M.; et al. Artificial Intelligence-Assisted Polyp Detection for Colonoscopy: Initial Experience. Gastroenterology 2018, 154, 2027-2029.e3. [CrossRef]
- Guo, Y.; Bernal, J.; Matuszewski, B.J. Polyp Segmentation with Fully Convolutional Deep Neural Networks—Extended Evaluation Study. J Imaging 2020, 6, 69. [CrossRef]
- Borgli, H.; et al. HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy. Sci Data 2020, 7, 283. [CrossRef]
- Beaumont, R. LAION-5B: A NEW ERA OF OPEN LARGE-SCALE MULTI-MODAL DATASETS. 2022. [Online]. Available: https://laion.ai/blog/laion-5b/.
- Chicco, D.; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 2020, 21, 6. [CrossRef]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit Lett 2006, 27, 861–874. [CrossRef]
- Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; Savarese, S. Generalized Intersection Over Union: A Metric and A Loss for Bounding Box Regression. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2019; pp. 658–666. [CrossRef]
- Powers, D.M.W. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. CoRR 2020, abs/2010.16061.
- Hore, A.; Ziou, D. Image Quality Metrics: PSNR vs. SSIM. In International Conference on Pattern Recognition; IEEE, 2010; pp. 2366–2369. [CrossRef]
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing 2004, 13, 600–612. [CrossRef]
- Taha, A.A.; Hanbury, A. Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. BMC Med Imaging 2015, 15, 29. [CrossRef]
- Ejiga, P.O.; Oluwafemi, O. Text-Guided Synthesis for Colon Cancer Screening. GitHub repository 2024. https://github.com/Ejigsonpeter/Text-Guided-Synthesis-for-Colon-Cancer-Screening.
- Hicks, S.; Storås, A.; Halvorsen, P.; De Lange, T.; Riegler, M.; Thambawita, V. Overview of ImageCLEFmedical 2023 - Medical Visual Question Answering for Gastrointestinal Tract. 2023. [Online]. Available: https://ceur-ws.org/Vol-3497/paper-107.pdf.
- Wang, W.; Tian, J. CP-CHILD records the colonoscopy data. figshare 2020. https://figshare.com/articles/dataset/CP-CHILD_zip/12554042?file=23383508.
- Rahman, M.S. Binary Polyps Classification. 2024. [Online]. Available: https://www.kaggle.com/datasets/mdsahilurrahman71/binary-polyps-classification?resource=download.
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. arXiv (Cornell University) 2015. [CrossRef]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid Scene Parsing Network. arXiv.org 2017. https://arxiv.org/abs/1612.01105.
- Lin, T.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. arXiv (Cornell University) 2016. [CrossRef]
- Chaurasia, A.; Culurciello, E. LinkNet: Exploiting encoder representations for efficient semantic segmentation. IEEE Visual Communications and Image Processing (VCIP) 2017. https://arxiv.org/abs/1707.03718.
- Safari, F.; Savić, I.; Kunze, H.; Ernst, J.; Gillis, D. A Review of AI-based MANET Routing Protocols. IEEE WiMob 2023, 43-50. [CrossRef]
- HuggingFace. Mask Generation. 2024. [Online]. Available: https://huggingface.co/docs/transformers/tasks/mask_generation.


















| Dataset | Model | FID | IS avg | IS std | IS med |
|---|---|---|---|---|---|
| single | CLIP | 0.11 | 1.57 | 0.03 | 1.56 |
| multi | CLIP | 0.11 | 1.57 | 0.03 | 1.56 |
| both | CLIP | 0.12 | 1.57 | 0.03 | 1.56 |
| single | SD | 0.06 | 2.33 | 0.07 | 2.34 |
| multi | SD | 0.06 | 2.33 | 0.07 | 2.34 |
| both | SD | 0.07 | 2.33 | 0.07 | 2.34 |
| single | DB+LoRa | 0.11 | 2.36 | 0.05 | 2.36 |
| multi | DB+LoRa | 0.07 | 2.36 | 0.05 | 2.36 |
| both | DB+LoRa | 0.08 | 2.36 | 0.05 | 2.36 |
| Dataset | Model | G1 | G2 | G3 | G4 | G5 | G6 | G7 | G8 | G9 | G10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| single | CLIP | 1.56 | 1.55 | 1.55 | 1.57 | 1.55 | 1.55 | 1.54 | 1.54 | 1.56 | 1.51 |
| multi | CLIP | 1.56 | 1.55 | 1.55 | 1.57 | 1.55 | 1.55 | 1.54 | 1.54 | 1.56 | 1.51 |
| both | CLIP | 1.56 | 1.55 | 1.55 | 1.57 | 1.55 | 1.55 | 1.54 | 1.54 | 1.56 | 1.51 |
| single | SD | 2.37 | 2.40 | 2.40 | 2.39 | 2.24 | 2.26 | 2.31 | 2.22 | 2.25 | 2.77 |
| multi | SD | 2.37 | 2.40 | 2.40 | 2.39 | 2.24 | 2.26 | 2.31 | 2.22 | 2.25 | 2.77 |
| both | SD | 2.37 | 2.40 | 2.40 | 2.39 | 2.24 | 2.26 | 2.31 | 2.22 | 2.25 | 2.77 |
| single | DB+LoRa | 2.35 | 2.27 | 2.39 | 2.41 | 2.31 | 2.33 | 2.34 | 2.36 | 2.46 | 2.38 |
| multi | DB+LoRa | 2.35 | 2.27 | 2.39 | 2.41 | 2.31 | 2.33 | 2.34 | 2.36 | 2.46 | 2.38 |
| both | DB+LoRa | 2.35 | 2.27 | 2.39 | 2.41 | 2.31 | 2.33 | 2.34 | 2.36 | 2.46 | 2.38 |
| Model | IoU | F1 Score | Precision | Recall | PSNR | SSIM | Dice Coef. |
|---|---|---|---|---|---|---|---|
| UNet [49] | 0.4253 | 0.5599 | 0.7208 | 0.5300 | 382.8118 | 1.0000 | 0.5599 |
| PSPNet [50] | 0.2212 | 0.3451 | 0.7048 | 0.3053 | 381.5486 | 1.0000 | 0.3451 |
| FPN [51] | 0.6373 | 0.7238 | 0.7143 | 0.8716 | 385.8368 | 1.0000 | 0.7238 |
| LinkNet [52] | 0.1863 | 0.3095 | 0.7123 | 0.2359 | 381.7352 | 1.0000 | 0.3095 |
| MANet [53] | 0.0775 | 0.1406 | 0.7359 | 0.0842 | 381.9626 | 1.0000 | 0.1406 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).