A Review of Multimodal Large Language Models: Fusion Mechanisms and Capability Evolution
Main Article Content
Keywords
multimodal large language models, fusion mechanisms, multimodal understanding, multimodal generation
Abstract
With the continuous advancement of large language models, MLLMs have gradually emerged as a major research focus in the field of artificial intelligence. Although existing studies have reviewed relevant progress from different perspectives, such as multimodal learning, vision-language models, and domain-specific applications, there is still a lack of systematic analysis of the intrinsic relationships among fusion mechanisms, capability formation, and application expansion in MLLMs. This paper reviews the principal technical paradigms of multimodal fusion, including early fusion, intermediate fusion, late fusion, and hybrid fusion, and compares the structural characteristics and applicable scenarios of different fusion approaches. Building on this analysis, the paper further explores the development of MLLMs from the perspectives of vision-language understanding, multimodal content generation, multimodal interaction and agent-oriented tasks, as well as domain-specific applications, thereby outlining the evolutionary trajectory of their core capabilities and broader application trends.
References
- [1] Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., & Ng, A. Y. (2011). Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning (pp. 689–696).
- [2] Baltrušaitis, T., Ahuja, C., & Morency, L. P. (2019). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443.
- [3] Li, L. H., Yatskar, M., Yin, D., et al. (2019). VisualBERT: A simple and performant baseline for vision and language. arXiv. https://arxiv.org/abs/1908.03557
- [4] Kim, W., Son, B., & Kim, I. (2021). ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the 38th International Conference on Machine Learning (pp. 5583–5594).
- [5] Chameleon Team. (2024). Chameleon: Mixed-modal early-fusion foundation models. arXiv. https://arxiv.org/abs/2405.09818
- [6] Wang, X., Zhang, X., Luo, Z., et al. (2024). Emu3: Next-token prediction is all you need. arXiv. https://arxiv.org/abs/2409.18869
- [7] Alayrac, J. B., Donahue, J., Luc, P., et al. (2022). Flamingo: A visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35, 23716–23736.
- [8] Li, J., Li, D., Savarese, S., et al. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (pp. 19730–19742).
- [9] Tsimpoukelli, M., Menick, J. L., Cabi, S., et al. (2021). Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34, 200–212.
- [10] Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259.
- [11] Atrey, P. K., Hossain, M. A., El Saddik, A., et al. (2010). Multimodal fusion for multimedia analysis: A survey. Multimedia Systems, 16(6), 345–379.
- [12] Mai, S., Hu, H., & Xing, S. (2019). Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 481–492).
- [13] Tsai, Y. H. H., Liang, P. P., Zadeh, A., Morency, L.-P., & Salakhutdinov, R. (2019). Learning factorized multimodal representations. In International Conference on Learning Representations. https://openreview.net/forum?id=rygqqsA9KX
- [14] Yao, Y., Yu, T., Zhang, A., et al. (2024). MiniCPM-V: A GPT-4V level MLLM on your phone. arXiv. https://arxiv.org/abs/2408.01800
- [15] Singh, A., Natarajan, V., Shah, M., et al. (2019). Towards VQA models that can read. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8317–8326).
- [16] Mathew, M., Karatzas, D., & Jawahar, C. V. (2021). DocVQA: A dataset for VQA on document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 2200–2209).
- [17] Han, Y., Zhang, C., Chen, X., et al. (2023). ChartLLaMA: A multimodal LLM for chart understanding and generation. arXiv. https://arxiv.org/abs/2311.16483
- [18] Zhao, H., Cai, Z., Si, S., et al. (2023). MMICL: Empowering vision-language model with multi-modal in-context learning. arXiv. https://arxiv.org/abs/2309.07915
- [19] Fu, C., Dai, Y., Luo, Y., et al. (2025). Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 24108–24118).
- [20] Zhan, J., Dai, J., Ye, J., et al. (2024). AnyGPT: Unified multimodal LLM with discrete sequence modeling. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 9637–9662).
- [21] Wu, S., Fei, H., Qu, L., et al. (2024). NExT-GPT: Any-to-any multimodal LLM. In Proceedings of the 41st International Conference on Machine Learning (pp. 53366–53397).
- [22] Lu, J., Clark, C., Lee, S., et al. (2024). Unified-IO 2: Scaling autoregressive multimodal models with vision, language, audio, and action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 26439–26455).
- [23] Tang, Z., Yang, Z., Khademi, M., et al. (2024). CoDi-2: In-context, interleaved, and interactive any-to-any generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 27425–27434).
- [24] Tong, S., Fan, D., Zhu, J., et al. (2025). MetaMorph: Multimodal understanding and generation via instruction tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 17001–17012).
- [25] Wu, C., Yin, S., Qi, W., et al. (2023). Visual ChatGPT: Talking, drawing and editing with visual foundation models. arXiv. https://arxiv.org/abs/2303.04671
- [26] Liu, S., Cheng, H., Liu, H., et al. (2024). LLaVA-Plus: Learning to use tools for creating multimodal agents. In Proceedings of the European Conference on Computer Vision (pp. 126–142).
- [27] Yang, Z., Li, L., Wang, J., et al. (2023). MM-REACT: Prompting ChatGPT for multimodal reasoning and action. arXiv. https://arxiv.org/abs/2303.11381
- [28] Surís, D., Menon, S., & Vondrick, C. (2023). ViperGPT: Visual inference via Python execution for reasoning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 11888–11898).
- [29] Han, S., Zhang, Q., Yao, Y., et al. (2024). LLM multi-agent systems: Challenges and open problems. arXiv. https://arxiv.org/abs/2402.03578
- [30] Meskó, B. (2023). The impact of multimodal large language models on health care’s future. Journal of Medical Internet Research, 25, e52865.
- [31] Yi, Z., Xiao, T., & Albert, M. V. (2025). A survey on multimodal large language models in radiology for report generation and visual question answering. Information, 16(2), 136.
- [32] Huang, H., Zheng, O., Wang, D., et al. (2023). ChatGPT for shaping the future of dentistry: The potential of multi-modal large language model. International Journal of Oral Science, 15(1), 29.
- [33] Cui, C., Ma, Y., Cao, X., et al. (2024). A survey on multimodal large language models for autonomous driving. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (pp. 958–979).
