Submitted:
28 January 2026
Posted:
29 January 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction

- We propose Synergistic Multimodal Diffusion Transformer (SyMDit), a novel unified discrete diffusion model that integrates an Adaptive Cross-Modal Transformer (ACMT) with a new Synergistic Attention Module (SAM) and Hierarchical Semantic Visual Tokenization (HSVT) for enhanced visual-language interaction and representation learning.
- We introduce advanced training strategies, including hyper-fine-grained recaptioning of vast image-text datasets using large language models, incorporation of high-quality synthetic data, and selective image region masking during supervised fine-tuning to boost multimodal understanding and generation capabilities.
- We demonstrate that SyMDit achieves superior performance across diverse multimodal generation and understanding benchmarks (Text-to-Image, Image-to-Text, VQA), setting new state-of-the-art results while also offering significant improvements in inference efficiency.
2. Related Work
2.1. Unified Multimodal Diffusion Models
2.2. Large Vision-Language Models and Multimodal Understanding
3. Method
3.1. Overall Architecture of SyMDit

3.2. Adaptive Cross-Modal Transformer (ACMT)
3.2.1. Synergistic Attention Module (SAM)
3.3. Hierarchical Semantic Visual Tokenization (HSVT)
3.4. Context-Aware Text Embedding
3.5. Unified Discrete Diffusion Process
3.6. Lightweight Multitask Decoding Head
4. Experiments
4.1. Experimental Setup
4.1.1. Datasets and Preprocessing
4.1.2. Inference Settings
4.2. Baseline Methods
- DALL-E 3 [2]: A prominent Text-to-Image diffusion model known for its exceptional image quality and compositional understanding. It is specialized for T2I generation.
- Stability Diffusion 3 (SD 3) [1]: Another advanced Text-to-Image diffusion model, lauded for its high-resolution image synthesis and broad creative capabilities. Also specialized for T2I.
- Chameleon: An autoregressive (AR) multimodal model capable of generating both images and text. It uses an AR architecture for both modalities, making it versatile but often slower.
- LLaVA-Next [3]: A leading visual language understanding model, primarily focused on Image-to-Text and VQA tasks, typically based on an autoregressive transformer. It does not perform T2I generation.
- Show-O (512×512): A multimodal model leveraging an autoregressive text generation architecture and a discrete diffusion model for image generation. It demonstrates capabilities across modalities.
- D-DiT (512×512): A model that combines a discrete diffusion framework for text generation with a continuous diffusion model for images, representing a hybrid approach to multimodal tasks.
- Muddit (512×512) [2]: A unified discrete diffusion multimodal model that serves as our closest architectural baseline. It demonstrates the feasibility of using discrete diffusion for multiple multimodal tasks within a single framework.
4.3. Quantitative Results
4.4. Ablation Study
- Synergistic Attention Module (SAM): When SAM is integrated into the SyMDit-Base, the GenEval score improves from 0.60 to 0.63, and VQAv2 accuracy sees a noticeable increase from 66.8% to 67.9%. This demonstrates that the dynamic adjustment of cross-modal attention weights in SAM facilitates more precise visual-textual alignment and deeper semantic interaction, which is crucial for complex multimodal tasks.
- Hierarchical Semantic Visual Tokenization (HSVT): Replacing the standard VQ-VAE with HSVT on the SyMDit-Base further boosts GenEval to 0.64 and MS-COCO CIDEr to 59.5. This indicates that the multi-scale, semantically rich visual tokenization provided by HSVT significantly enhances the model’s ability to capture intricate details and complex structures in images, leading to better visual generation and understanding.
- Context-Aware Text Embedding with<camask>: The inclusion of the <camask> token, which dynamically adapts its semantic representation based on context, shows an improvement in VQAv2 accuracy from 66.8% to 67.5% and MME accuracy from 1089.1 to 1098.3 on SyMDit-Base. This validates the effectiveness of providing richer contextual information during text diffusion denoising, especially for reasoning-heavy tasks.
4.5. Human Evaluation
4.6. Efficiency and Resource Analysis
4.7. Detailed Task-Specific Performance
4.8. Robustness to Complex Instructions
5. Conclusions
References
- Tang, R.; Liu, L.; Pandey, A.; Jiang, Z.; Yang, G.; Kumar, K.; Stenetorp, P.; Lin, J.; Ture, F. What the DAAM: Interpreting Stable Diffusion Using Cross Attention. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics, 2023; Volume 1, pp. 5644–5659. [Google Scholar] [CrossRef]
- Shi, Q.; Bai, J.; Zhao, Z.; Chai, W.; Yu, K.; Wu, J.; Song, S.; Tong, Y.; Li, X.; Li, X.; et al. Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model. arXiv arXiv:2505.23606v3.
- Luo, G.; Darrell, T.; Rohrbach, A. NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; Association for Computational Linguistics; pp. 6801–6817. [Google Scholar] [CrossRef]
- Qi, L.; Wu, J.; Choi, J.M.; Phillips, C.; Sengupta, R.; Goldman, D.B. Over++: Generative Video Compositing for Layer Interaction Effects. arXiv arXiv:2512.19661.
- Gong, B.; Qi, L.; Wu, J.; Fu, Z.; Song, C.; Jacobs, D.W.; Nicholson, J.; Sengupta, R. The Aging Multiverse: Generating Condition-Aware Facial Aging Tree via Training-Free Diffusion. arXiv arXiv:2506.21008.
- Qi, L.; Wu, J.; Gong, B.; Wang, A.N.; Jacobs, D.W.; Sengupta, R. Mytimemachine: Personalized facial age transformation. ACM Transactions on Graphics (TOG) 2025, 44, 1–16. [Google Scholar] [CrossRef]
- Zhang, X.; Li, W.; Zhao, S.; Li, J.; Zhang, L.; Zhang, J. VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning. arXiv arXiv:2506.18564.
- Li, W.; Zhang, X.; Zhao, S.; Zhang, Y.; Li, J.; Zhang, L.; Zhang, J. Q-insight: Understanding image quality via visual reinforcement learning. arXiv arXiv:2503.22679. [CrossRef]
- Xu, Z.; Zhang, X.; Zhou, X.; Zhang, J. AvatarShield: Visual Reinforcement Learning for Human-Centric Video Forgery Detection. arXiv arXiv:2505.15173.
- Wang, J.; Cui, X. Multi-omics Mendelian Randomization Reveals Immunometabolic Signatures of the Gut Microbiota in Optic Neuritis and the Potential Therapeutic Role of Vitamin B6. Molecular Neurobiology 2025, 1–12. [Google Scholar] [CrossRef] [PubMed]
- Xuehao, C.; Dejia, W.; Xiaorong, L. Integration of Immunometabolic Composite Indices and Machine Learning for Diabetic Retinopathy Risk Stratification: Insights from NHANES 2011–2020. Ophthalmology Science 2025, 100854. [Google Scholar] [CrossRef] [PubMed]
- Hui, J.; Cui, X.; Han, Q. Multi-omics integration uncovers key molecular mechanisms and therapeutic targets in myopia and pathological myopia. Asia-Pacific Journal of Ophthalmology 2026, 100277. [Google Scholar] [CrossRef] [PubMed]
- Li, X.L.; Holtzman, A.; Fried, D.; Liang, P.; Eisner, J.; Hashimoto, T.; Zettlemoyer, L.; Lewis, M. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics, 2023; Volume 1, pp. 12286–12312. [Google Scholar] [CrossRef]
- Eichenberg, C.; Black, S.; Weinbach, S.; Parcalabescu, L.; Frank, A. MAGMA – Multimodal Augmentation of Generative Models through Adapter-based Finetuning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2022; Association for Computational Linguistics, 2022; pp. 2416–2428. [Google Scholar] [CrossRef]
- Lin, B.; Ye, Y.; Zhu, B.; Cui, J.; Ning, M.; Jin, P.; Yuan, L. Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024; Association for Computational Linguistics; pp. 5971–5984. [Google Scholar] [CrossRef]
- Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; Zhuang, Y. DiffusionNER: Boundary Diffusion for Named Entity Recognition. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics, 2023; Volume 1, pp. 3875–3890. [Google Scholar] [CrossRef]
- Sun, S.; Chen, Y.C.; Li, L.; Wang, S.; Fang, Y.; Liu, J. LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021; Association for Computational Linguistics; pp. 982–997. [Google Scholar] [CrossRef]
- Changpinyo, S.; Kukliansy, D.; Szpektor, I.; Chen, X.; Ding, N.; Soricut, R. All You May Need for VQA are Image Captions. In Proceedings of the Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022; Association for Computational Linguistics; pp. 1947–1963. [Google Scholar] [CrossRef]
- Yang, J.; Yu, Y.; Niu, D.; Guo, W.; Xu, Y. ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis. Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics 2023, Volume 1, 7617–7630. [Google Scholar] [CrossRef]
- Yang, X.; Feng, S.; Zhang, Y.; Wang, D. Multimodal Sentiment Detection Based on Multi-channel Graph Neural Networks. Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021, Volume 1, 328–339. [Google Scholar] [CrossRef]
- Qin, H.; Song, Y. Reinforced Cross-modal Alignment for Radiology Report Generation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2022; Association for Computational Linguistics, 2022; pp. 448–458. [Google Scholar] [CrossRef]
- Wu, Y.; Zhan, P.; Zhang, Y.; Wang, L.; Xu, Z. Multimodal Fusion with Co-Attention Networks for Fake News Detection. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021; Association for Computational Linguistics, 2021; pp. 2560–2569. [Google Scholar] [CrossRef]
- Yan, H.; Dai, J.; Ji, T.; Qiu, X.; Zhang, Z. A Unified Generative Framework for Aspect-based Sentiment Analysis. Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021, Volume 1, 2416–2429. [Google Scholar] [CrossRef]
- Xu, H.; Yan, M.; Li, C.; Bi, B.; Huang, S.; Xiao, W.; Huang, F. E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning. Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021, Volume 1, 503–513. [Google Scholar] [CrossRef]
- Gui, L.; Wang, B.; Huang, Q.; Hauptmann, A.; Bisk, Y.; Gao, J. KAT: A Knowledge Augmented Transformer for Vision-and-Language. In Proceedings of the Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022; Association for Computational Linguistics; pp. 956–968. [Google Scholar] [CrossRef]
- Yang, J.; Wang, Y.; Yi, R.; Zhu, Y.; Rehman, A.; Zadeh, A.; Poria, S.; Morency, L.P. MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021; Association for Computational Linguistics; pp. 1009–1021. [Google Scholar] [CrossRef]
- Zhang, N.; Chen, M.; Bi, Z.; Liang, X.; Li, L.; Shang, X.; Yin, K.; Tan, C.; Xu, J.; Huang, F.; et al. CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark. In Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics, 2022; Volume 1, pp. 7888–7915. [Google Scholar] [CrossRef]
- Ross, C.; Katz, B.; Barbu, A. Measuring Social Biases in Grounded Vision and Language Embeddings. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics, 2021; pp. 998–1008. [Google Scholar] [CrossRef]
- Artetxe, M.; Bhosale, S.; Goyal, N.; Mihaylov, T.; Ott, M.; Shleifer, S.; Lin, X.V.; Du, J.; Iyer, S.; Pasunuru, R.; et al. Efficient Large Scale Language Modeling with Mixtures of Experts. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022; Association for Computational Linguistics; pp. 11699–11732. [Google Scholar] [CrossRef]
- Le, H.; Pino, J.; Wang, C.; Gu, J.; Schwab, D.; Besacier, L. Lightweight Adapter Tuning for Multilingual Speech Translation. Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021, Volume 2, 817–824. [Google Scholar] [CrossRef]
- Chi, Z.; Huang, S.; Dong, L.; Ma, S.; Zheng, B.; Singhal, S.; Bajaj, P.; Song, X.; Mao, X.L.; Huang, H.; et al. XLM-E: Cross-lingual Language Model Pre-training via ELECTRA. Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics 2022, Volume 1, 6170–6182. [Google Scholar] [CrossRef]


| Model | Text Gen Arch | Image Gen Arch | GenEval ↑ | MS-COCO ↑ | VQAv2 ↑ | MME ↑ | GQA ↑ |
|---|---|---|---|---|---|---|---|
| DALL-E 3 | - | Diffusion | 0.67 | - | - | - | - |
| SD 3 | - | Diffusion | 0.62 | - | - | - | - |
| Chameleon | AR | AR | 0.39 | 18.0 | - | - | - |
| LLaVA-Next | AR | - | - | - | 82.8 | 1575.0 | 65.4 |
| Show-O (512×512) | AR | Discrete Diff. | 0.68 | - | 69.4 | 1097.2 | 58.0 |
| D-DiT (512×512) | Discrete Diff. | Diffusion | 0.65 | 56.2 | 60.1 | 1124.7 | 59.2 |
| Muddit (512×512) | Discrete Diff. | Discrete Diff. | 0.61 | 59.7 | 67.7 | 1104.6 | 57.1 |
| SyMDit (Ours) | Discrete Diff. | Discrete Diff. | 0.69 | 60.5 | 70.1 | 1142.3 | 59.5 |
| Model Variant | Key Components | GenEval ↑ | MS-COCO ↑ | VQAv2 ↑ | MME ↑ |
|---|---|---|---|---|---|
| SyMDit-Base | Base DiT, VQ-VAE, static mask | 0.60 | 58.5 | 66.8 | 1089.1 |
| SyMDit w/ SAM | SyMDit-Base + SAM | 0.63 | 59.2 | 67.9 | 1105.4 |
| SyMDit w/ HSVT | SyMDit-Base + HSVT | 0.64 | 59.5 | 68.2 | 1110.7 |
| SyMDit w/ CAMask | SyMDit-Base + <camask> | 0.62 | 58.9 | 67.5 | 1098.3 |
| SyMDit (Full) | All components | 0.69 | 60.5 | 70.1 | 1142.3 |
| Task | Metric | DALL-E 3 | SD 3 | LLaVA-Next | Muddit | SyMDit (Ours) |
|---|---|---|---|---|---|---|
| T2I | CLIP Score ↑ | 0.32 | 0.30 | - | 0.29 | 0.34 |
| FID ↓ | 7.8 | 8.5 | - | 8.9 | 7.5 | |
| IS ↑ | 120.5 | 115.2 | - | 110.3 | 125.1 | |
| I2T | BLEU-4 ↑ | - | - | - | 35.1 | 36.8 |
| METEOR ↑ | - | - | - | 29.5 | 30.7 | |
| ROUGE-L ↑ | - | - | - | 55.8 | 57.2 | |
| CIDEr ↑ | - | - | - | 59.7 | 60.5 | |
| VQA | VQAv2 (Acc.) ↑ | - | - | 82.8 | 67.7 | 70.1 |
| A-OKVQA (Acc.) ↑ | - | - | 52.3 | 48.9 | 51.1 |
| Task | Metric | DALL-E 3 | SD 3 | Muddit | SyMDit (Ours) |
|---|---|---|---|---|---|
| T2I | Comp. Acc. ↑ | 85.1% | 81.3% | 78.5% | 87.4% |
| F-G Attr. Acc. ↑ | 80.5% | 77.9% | 75.2% | 82.9% | |
| P. Adherence (1-5) ↑ | 4.1 | 3.9 | 3.8 | 4.4 | |
| VQA | M-S Reas. Acc. ↑ | - | - | 55.7% | 58.2% |
| Counterfactual VQA (Acc.) ↑ | - | - | 62.1% | 65.5% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.