Submitted:
26 October 2025
Posted:
28 October 2025
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Research Background and Challenges
1.2. Core Contributions
2. Methodology: MCP Framework Design
2.1. Model Layer: Functional Decoupling
2.2. Controller Layer: Dynamic Routing and RL Policy
2.2.1. Dynamic Routing Algorithm
2.2.2. RL Policy Feedback Loop
2.3. Presenter Layer: Task Adaptation and Interpretability
2.4. Modular LoRA (mLoRA) Integration
3. Experiments
3.1. Experimental Setup
3.1.1. Datasets
3.1.2. Baseline Models
3.2. Main Experimental Results
3.2.1. Multimodal Performance
3.2.2. Interpretability Validation
3.3. Ablation Experiments
3.4. Case Study: Tuberculosis (TB) Diagnosis
4. Discussion and Conclusions
4.1. Theoretical and Practical Significance
4.2. Limitations and Future Work
4.3. Conclusions
Funding
Acknowledgments
References
- Y. He et al., "Layer-adaptive structured pruning guided by latency," Adv. Neural Inf. Process. Syst., vol. 34, pp. 12497–12509, 2021.
- W. Fedus et al., "Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity," Adv. Neural Inf. Process. Syst., vol. 34, pp. 1037–1051, 2021.
- A. Wang et al., "GLUE: A multi-task benchmark and analysis platform for natural language understanding," in Proc. Int. Conf. Learn. Represent., 2019.
- T.-Y. Lin et al., "Microsoft COCO: Common objects in context," in Proc. Eur. Conf. Comput. Vis., 2014, pp. 740–755.
- P. Lu et al., "Learn to explain: Multimodal reasoning via thought chains for science question answering," Adv. Neural Inf. Process. Syst., vol. 36, 2023.
- H. Touvron et al., "LLaMA 2: Open foundation and fine-tuned chat models," arXiv preprint arXiv:2307.09288, 2023.
- OpenAI, "GPT-3.5 technical report," OpenAI Res., 2023.
- D. Lepikhin et al., "Gshard: Scaling giant models with conditional computation and automatic sharding," in Proc. Int. Conf. Mach. Learn., 2022, pp. 12965–12977.
- K. Chua et al., "Deep reinforcement learning in a handful of trials using probabilistic dynamics models," in Proc. Int. Conf. Mach. Learn., 2018, pp. 4754–4765.
- T. P. Lillicrap et al., "Continuous control with deep reinforcement learning," in Proc. Int. Conf. Mach. Learn., 2015, pp. 1871–1880.
- C. Raffel et al., "Exploring the limits of transfer learning with a unified text-to-text transformer," J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, 2020.
- D. Zhou et al., "Automatic chain of thought prompting in large language models," Adv. Neural Inf. Process. Syst., vol. 35, pp. 1788–1801, 2022.
- Z. Chen et al., "Dynamicvit: Efficient vision transformers with dynamic token sparsification," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 13907–13916.
- Zhang et al., "Awq: Activation-aware weight quantization for llm compression and acceleration," arXiv preprint arXiv:2306.00978, 2023.
- Y. Li et al., "Dyhead: Unifying object detection heads with attentions," Adv. Neural Inf. Process. Syst., vol. 34, pp. 22661–22672, 2021.


| Sub-module | Core Function | Params (±Fluctuation) | Latency | Key Technology |
|---|---|---|---|---|
| Reasoning | Logical deduction, knowledge verification | 43.7M (±1.2%) | 8.2ms±0.3ms | Sparse Attention Cluster [1,15] (SAC) |
| Generation | Creative content synthesis | 37.5M (±0.9%) | 6.5ms±0.4ms | Length-Aware Decoding (LAD) [13] |
| Retrieval | Knowledge retrieval, semantic matching | 18.8K (±0.7%) | 4.1ms±0.2ms | Hierarchical Hybrid Index (HHI) [14] |
| Dataset | Metric | LLaMA-2 7B | GPT-3.5 | MCP |
|---|---|---|---|---|
| GLUE (Avg.) | Accuracy | 78% | 85% | 92% |
| COCO | Latency (ms) | 416 | 352 | 212 |
| ScienceQA | Energy (J/task) | 22.6 | 18.3 | 10.3 |
| Metric | Exp. Group vs. Ablation 1 | Exp. Group vs. Ablation 2 |
|---|---|---|
| GLUE Accuracy Gain | +19% | +12% |
| COCO Latency Reduction | -35% | -22% |
| Energy Efficiency Gain | +42% | +29% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).