Submitted:
01 July 2026
Posted:
01 July 2026
You are already at the latest version
Abstract
Pixel-level annotation remains a major bottleneck for semantic segmentation, motivating methods that synthesize image-label pairs directly from generative models. Prior synthetic dataset generators typically obtain pseudo-labels from cross-attention maps or learned decoders over generative features; however, recent text-to-image (T2I) models increasingly use multimodal diffusion transformers (MM-DiTs), where concept localization is no longer exposed through a single cross-attention pathway but instead distributed across many layers and attention heads. Existing MM-DiT localization methods address this by aggregating saliency across heads, but we observe that this averaging can dilute clean target localizers due to attention head heterogeneity. We introduce HEADHUNTER, a training-free, open-vocabulary segmentation framework that uses aggregate concept saliency as a self-guided proxy to select the single attention head that best localizes a queried textual concept, yielding cleaner segmentation masks. We then use HEADHUNTER to turn target classes into training data automatically: a large language model (LLM) diversifies prompts, an MM-DiT generates images, HEADHUNTER produces pseudo-labels, and a vision language model (VLM) verifies each image-mask pair before acceptance. HEADHUNTER achieves strong zero-shot segmentation performance (81.2 mIoU on PASCAL VOC2012 and 73.2 mIoU on ImageNet-Segmentation), outperforming head aggregation and other interpretability methods. Using only our generated image-label pairs, we train segmentation models which reach 64.9 mIoU on VOC2012 validation, matching or outperforming comparable synthetic dataset generators and showing that the proposed pipeline produces effective dense labels without human intervention.
Keywords:
1. Introduction
- We show that aggregating MM-DiT attention heads discards fine boundary detail, and that selecting a single concept-aligned head yields sharper segmentation masks.
- We introduce HEADHUNTER, a training-free, open-vocabulary MM-DiT segmentation framework that selects this head using aggregate concept saliency as a self-guided proxy, achieving strong zero-shot segmentation performance.
- We propose a fully automatic annotated dataset synthesis pipeline built around HEADHUNTER, producing image-label pairs for target concepts without any human annotation, delivering competitive segmentation performance on the VOC2012 benchmark.
2. Related Work
2.1. Synthetic Dataset Generation for Semantic Segmentation
2.2. Concept Localization in Diffusion Transformers
3. Preliminaries
3.1. Rectified Flow Modelling Framework
3.2. Multimodal Attention Mechanism
3.3. Saliency Maps from Multimodal Attention
4. Method
4.1. Prompt and Concept Generation
4.2. Self-Guided Attention Head Selection
4.3. Mask Generation and Dataset QA
5. Experiments
5.1. Experimental Setup
5.2. Zero-Shot Segmentation
5.3. Synthetic Dataset Evaluation
5.4. Ablation Studies
5.5. Limitations
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| CA | Cross Attention |
| DiT | Diffusion Transformer |
| GT | Ground Truth |
| I2T | Image-to-Text |
| LLM | Large Language Model |
| mIoU | Mean Intersection over Union |
| MM-DiT | Multimodal Diffusion Transformer |
| QA | Quality Assurance |
| SA | Self Attention |
| SAM2 | Segment Anything Model 2 |
| SD | Stable Diffusion |
| T2I | Text-to-Image |
| VLM | Vision Language Model |
| VOC | Visual Object Classes |
Appendix A. Prompt Planner Specifications
| Output | Purpose |
| Prompt | Conditions the MM-DiT T2I model |
| Target Concept | Defines the object queried by HEADHUNTER |
| Proxy Concepts | Provides the concepts used for localization and the proxy mask |
| Expected Target Count | Indicates whether multiple visible instances should be masked |
| Mask Policy | Specifies which visible target pixels should be included |
| QA Requirements | Provides sample-specific checks used by the VLM QA stage |
Appendix B. Multimodal Attention Extraction Settings
| Component | Setting |
| Backbone | FLUX.1-dev |
| Dual-stream DiT blocks | 19 |
| Single-stream DiT blocks | 38 |
| Attention heads per block | 24 |
| Candidate layers | All dual-stream MM-DiT layers |
| Candidate heads | All 24 heads per selected layer |
| Candidate maps per image | 19 × 24 = 456 |
Appendix C. Dataset QA Specifications
| Criterion | Instruction |
| Target Presence | The target object must be visible in the generated image |
| Count Consistency | The number of masked instances must match the target count |
| Target Coverage | The mask must cover the visible target structure |
| Non-Target Exclusion | Background/support/non-target regions must be excluded |
| Major Failures | Reject masks that miss structure/targets or fail the above |
| Minor Failures | Non-structural omissions/boundary issues may pass |
| Output Format | Return JSON only incl. the pass/fail decision with a reason |
Appendix D. Segmentation Model Training Details
| Component | Setting |
| Training Data | QA-accepted synthetic image-label pairs |
| Evaluation Data | PASCAL VOC2012 validation set |
| Classes | 21 VOC classes: 20 foreground classes + background |
| Label Construction | → all non-mask pixels → background |
| Segmenter | DeepLabV3 |
| Backbones | ResNet-50, ResNet-101 |
| Generated Resolution | |
| Training Resolution | |
| Optimizer | AdamW |
| Weight Decay | 10-4 |
| Batch Size | 8 |
| Training Iterations | 5,000 |
| Precision | Mixed precision |
| Learning Rate Schedule | Polynomial decay, power 0.9 |
| ResNet-50 Learning Rates | backbone, classifier |
| ResNet-101 Learning Rates | backbone, classifier |
| Loss | Pixel-wise cross-entropy + foreground Dice + auxiliary classifier cross-entropy |
| Loss Weights | Dice weight 0.5, auxiliary cross-entropy weight 0.4 |
| Augmentations | Random resized crop, horizontal flip, color jitter, Gaussian blur, perspective transform |
References
- Lin, H.; Upchurch, P.; Bala, K. Block Annotation: Better Image Annotation for Semantic Segmentation with Sub-Image Decomposition 2020.
- The Future of Artificial Intelligence in the Face of Data Scarcity. Comput. Mater. Contin. 2025, 84, 1073–1099. [CrossRef]
- Roux, R.L.; Khaksar, S.; Sepehri, M.; Murray, I. A Hybrid Game Engine–Generative AI Framework for Overcoming Data Scarcity in Open-Pit Crack Detection. Mach. Learn. Knowl. Extr. 2026, 8. [Google Scholar] [CrossRef]
- Zhang, Y.; Ling, H.; Gao, J.; Yin, K.; Lafleche, J.-F.; Barriuso, A.; Torralba, A.; Fidler, S. DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort 2021. [PubMed]
- Li, D.; Ling, H.; Kim, S.W.; Kreis, K.; Barriuso, A.; Fidler, S.; Torralba, A. BigDatasetGAN: Synthesizing ImageNet with Pixel-Wise Annotations 2022. [PubMed]
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Networks 2014. [PubMed]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models 2022. [PubMed]
- Wu, W.; Zhao, Y.; Shou, M.Z.; Zhou, H.; Shen, C. DiffuMask: Synthesizing Images with Pixel-Level Annotations for Semantic Segmentation Using Diffusion Models 2024. [PubMed]
- Nguyen, Q.; Vu, T.; Tran, A.; Nguyen, K. Dataset Diffusion: Diffusion-Based Synthetic Dataset Generation for Pixel-Level Semantic Segmentation 2023. [PubMed]
- Wu, W.; Zhao, Y.; Chen, H.; Gu, Y.; Zhao, R.; He, Y.; Zhou, H.; Shou, M.Z.; Shen, C. DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models 2023. [PubMed]
- Xie, J.; Li, W.; Li, X.; Liu, Z.; Ong, Y.S.; Loy, C.C. MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation 2024. [PubMed]
- Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis 2024. [PubMed]
- Black Forest Labs FLUX. 1 2024.
- Helbling, A.; Meral, T.H.S.; Hoover, B.; Yanardag, P.; Chau, D.H. ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features 2025. [PubMed]
- Kim, C.; Shin, H.; Hong, E.; Yoon, H.; Arnab, A.; Seo, P.H.; Hong, S.; Kim, S. Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers 2025. [PubMed]
- Everingham, M.; Van Gool, L.; Williams, C.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef]
- Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; Fei-Fei, L. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, June 2009; pp. 248–255. [Google Scholar]
- Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. LAION-5B: An Open Large-Scale Dataset for Training next Generation Image-Text Models 2022. [PubMed]
- Ahn, J.; Kwak, S. Learning Pixel-Level Semantic Affinity with Image-Level Supervision for Weakly Supervised Semantic Segmentation 2018. [PubMed]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C.L.; Dollár, P. Microsoft COCO: Common Objects in Context 2015.
- Petit, O.; Thome, N.; Rambour, C.; Soler, L. U-Net Transformer: Self and Cross Attention for Medical Image Segmentation 2021. [PubMed]
- Peebles, W.; Xie, S. Scalable Diffusion Models with Transformers 2023. [PubMed]
- Liu, X.; Gong, C.; Liu, Q. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow 2022. [PubMed]
- Qwen Team Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model. Available online: https://qwen.ai/blog?id=qwen3.6-27b (accessed on 21 June 2026).
- Otsu, N. A Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man. Cybern. 1979, 9, 62–66. [Google Scholar] [CrossRef]
- Ravi, N.; Gabeur, V.; Hu, Y.-T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. SAM 2: Segment Anything in Images and Videos 2024. [PubMed]
- Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; et al. Qwen3-VL Technical Report 2025. [CrossRef] [PubMed]
- Chen, L.-C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking Atrous Convolution for Semantic Image Segmentation 2017. [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition 2015. [PubMed]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Int. J. Comput. Vis. 2020, 128, 336–359. [Google Scholar] [CrossRef]
- Tang, R.; Liu, L.; Pandey, A.; Jiang, Z.; Yang, G.; Kumar, K.; Stenetorp, P.; Lin, J.; Ture, F. What the DAAM: Interpreting Stable Diffusion Using Cross Attention 2022. [PubMed]
- Binder, A.; Montavon, G.; Lapuschkin, S.; Müller, K.-R.; Samek, W. Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers; 2016; ISBN 978-3-319-44780-3. [Google Scholar]
- Abnar, S.; Zuidema, W. Quantifying Attention Flow in Transformers. Available online: https://arxiv.org/abs/2005.00928v2 (accessed on 1 July 2026).
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. Available online: https://arxiv.org/abs/2010.11929v2 (accessed on 1 July 2026).
- Gandelsman, Y.; Efros, A.A.; Steinhardt, J. Interpreting CLIP’s Image Representation via Text-Based Decomposition. Available online: https://arxiv.org/abs/2310.05916v4 (accessed on 24 June 2026).
- Chefer, H.; Gur, S.; Wolf, L. Transformer Interpretability Beyond Attention Visualization. Available online: https://arxiv.org/abs/2012.09838v2 (accessed on 1 July 2026).
- Sun, S.; Li, R.; Torr, P.; Gu, X.; Li, S. CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor. Available online: https://arxiv.org/abs/2312.07661v3 (accessed on 1 July 2026).
- Marcos-Manchón, P.; Alcover-Couso, R.; SanMiguel, J.C.; Martínez, J.M. Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion Models. Available online: https://arxiv.org/abs/2403.14291v1 (accessed on 1 July 2026).
- Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; Joulin, A. Emerging Properties in Self-Supervised Vision Transformers; 2021; p. 9640. [Google Scholar]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning Robust Visual Features without Supervision 2024.
- Darcet, T.; Oquab, M.; Mairal, J.; Bojanowski, P. Vision Transformers Need Registers. Available online: https://arxiv.org/abs/2309.16588v2 (accessed on 1 July 2026).
- Zhou, J.; Wei, C.; Wang, H.; Shen, W.; Xie, C.; Yuille, A.; Kong, T. iBOT: Image BERT Pre-Training with Online Tokenizer. Available online: https://arxiv.org/abs/2111.07832v3 (accessed on 1 July 2026).











| Method | Generator | Annotation Mechanism | Training Free |
| DatasetGAN [4] | GAN | Feature decoder | ✗ |
| BigDatasetGAN [5] | GAN | Feature decoder | ✗ |
| DiffuMask [8] | U-Net SD | CA | ✗ |
| Dataset Diffusion [9] | U-Net SD | CA + SA | ✓ |
| DatasetDM [10] | U-Net SD | Latent decoder | ✗ |
| MosaicFusion [11] | U-Net SD | CA | ✓ |
| Ours | MM-DiT | Selected MM-DiT head | ✓ |
| Method | Architecture | ImageNet-Segmentation | VOC2012 Single-Class | ||||
| Acc | mIoU | mAP | Acc | mIoU | mAP | ||
| LRP [32] | CLIP ViT | 51.1 | 32.9 | 55.7 | 48.8 | 31.4 | 52.9 |
| Partial-LRP [32] | CLIP ViT | 76.3 | 57.9 | 84.7 | 71.5 | 51.4 | 84.9 |
| Rollout [33] | CLIP ViT | 73.5 | 55.4 | 84.8 | 69.8 | 51.3 | 85.3 |
| ViT Attention [34] | CLIP ViT | 67.8 | 46.4 | 80.2 | 68.5 | 44.8 | 83.6 |
| GradCAM [30] | CLIP ViT | 64.4 | 40.8 | 71.6 | 70.4 | 44.9 | 76.8 |
| TextSpan [35] | CLIP ViT | 75.2 | 54.5 | 81.6 | 75.0 | 56.2 | 84.8 |
| TransInterp [36] | CLIP ViT | 79.7 | 62.0 | 86.0 | 76.9 | 57.1 | 86.7 |
| CLIPasRNN [37] | CLIP ViT | 74.1 | 58.8 | 84.9 | 61.8 | 41.5 | 76.6 |
| OVAM [38] | SDXL U-Net | 79.4 | 65.0 | 88.1 | 73.5 | 58.1 | 87.9 |
| DINO SA [39] | DINO ViT | 82.0 | 69.4 | 86.1 | 80.7 | 64.3 | 88.9 |
| DINOv2 SA [40] | DINOv2 ViT | 77.4 | 63.1 | 84.2 | 79.7 | 57.6 | 87.3 |
| DINOv2 Reg SA [41] | DINOv2 Reg | 72.0 | 56.3 | 80.8 | 77.2 | 56.6 | 86.4 |
| iBOT SA [42] | iBOT ViT | 76.3 | 61.7 | 82.0 | 75.0 | 55.8 | 85.3 |
| DAAM [31] | SDXL U-Net | 78.5 | 64.6 | 88.8 | 72.8 | 56.0 | 88.3 |
| DAAM [31] | SD2 U-Net | 64.5 | 47.6 | 78.0 | 64.3 | 45.0 | 83.0 |
| Cross Attention [14] | FLUX.1 MM-DiT | 74.9 | 58.9 | 87.2 | 80.4 | 54.8 | 89.1 |
| ConceptAttention [14] | FLUX.1 MM-DiT | 83.1 | 71.0 | 90.5 | 87.9 | 76.5 | 90.2 |
| Ours | FLUX.1 MM-DiT | 84.5 | 73.2 | 85.3 | 91.9 | 81.2 | 91.7 |
| Dataset | Segmenter | Backbone | # Images | mIoU |
| Dataset Diffusion [9] | DeepLabV3 | ResNet-101 | 10k/40k | 63.8/64.8 |
| Ours | DeepLabV3 | ResNet-101 | 10k | 64.9 |
| DiffuMask [8] | Mask2Former | ResNet-50 | 60k | 57.4 |
| Dataset Diffusion [9] | DeepLabV3 | ResNet-50 | 40k | 61.6 |
| Ours | DeepLabV3 | ResNet-50 | 10k | 60.7 |
| Class | ResNet-50 | ResNet-101 | Class | ResNet-50 | ResNet-101 |
| Aeroplane | 83.0 | 83.0 | Dining Table | 9.0 | 10.9 |
| Bicycle | 31.8 | 31.3 | Dog | 74.2 | 77.5 |
| Bird | 83.3 | 89.9 | Horse | 68.9 | 73.8 |
| Boat | 58.4 | 70.0 | Motorbike | 64.5 | 74.5 |
| Bottle | 61.9 | 72.2 | Person | 67.3 | 70.2 |
| Bus | 76.5 | 88.2 | Potted Plant | 38.0 | 43.0 |
| Car | 78.4 | 81.8 | Sheep | 72.5 | 76.7 |
| Cat | 79.5 | 84.5 | Sofa | 36.4 | 39.4 |
| Chair | 23.7 | 16.3 | Train | 74.1 | 78.1 |
| Cow | 77.1 | 80.7 | TV/Monitor | 55.4 | 56.3 |
| Class | DiffuMask [8] | Ours—ResNet-50 | Dataset Diffusion [9] | Ours—ResNet-101 |
| Aeroplane | 80.7 | 83.0 | 81.6 | 83.0 |
| Bird | 86.7 | 83.3 | 73.3 | 89.9 |
| Boat | 56.9 | 58.4 | 62.2 | 70.0 |
| Bus | 81.2 | 76.5 | 85.5 | 88.2 |
| Car | 74.2 | 78.4 | 64.8 | 81.8 |
| Cat | 79.3 | 79.5 | 78.2 | 84.5 |
| Chair | 14.7 | 23.7 | 21.6 | 16.3 |
| Cow | 63.4 | 77.1 | 69.2 | 80.7 |
| Dog | 65.1 | 74.2 | 71.8 | 77.5 |
| Horse | 64.6 | 68.9 | 78.2 | 73.8 |
| Person | 71.0 | 67.3 | 70.8 | 70.2 |
| Sheep | 64.7 | 72.5 | 77.8 | 76.7 |
| Sofa | 26.7 | 36.4 | 41.8 | 39.4 |
| Configuration | Dog | Car | mIoU |
| Full (ours) | 77.5 | 81.8 | 64.9 |
| − Mask refinement | 62.7 (−14.8) | 69.7 (−12.1) | 60.3 (−4.6) |
| − Prompt diversity | 57.4 (−20.1) | 69.2 (−12.6) | 59.7 (−5.2) |
| − QA verification | 66.1 (−11.4) | 68.1 (−13.7) | 59.7 (−5.2) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).