Research Article

A Comprehensive Investigation of Advances in Music Understanding and Generation Technologies Based on Large Language Models

Jiayi ZhangSt.George’s School*

* Corresponding author: [email protected]

Abstract

The intersection of artificial intelligence and music has developed rapidly in recent years, driven by advances in deep learning and the increasing availability of multimodal datasets. This review surveys recent progress in music understanding and generation with Artificial Intelligence (AI) through Large Language Models (LLMs) along three lines: agent/controller systems (e.g., AudioGPT, MusicAgent, CoComposer), multimodal fusion/decoders (MuMu-LLaMA, DeepResonance), and symbolic score models (ChatMusician, MuseCoco). The paper summarizes what each does well, such as tool-orchestrated workflows and cross-modal alignment for text-/image-/video-to-music. Next, the paper introduces the common, fixable challenges occurring in each category, such as limited long-form coherence, uneven controllability, and dataset bias. Some key challenges include orchestration reliability and dependence on pretrained decoders. The paper then proposes some short term remedies such as including multi-agent planning, longer-context modelling, and broader, well-labeled pretraining data, before predicting that the field of music based Large Language Models is moving toward hybrid systems that will integrate themselves within real workflows in the near future. This mapping creates a concise summary of advances in musical understanding of Large Language Models.

Keywords: Large-Language-Model; Artificial Intelligence; Agent-based; Multimodal; Symbolic
Published: November 5, 2025
DOI: 10.54254/2753-7064/2025.NS29129
Volume: CHR Vol.91
pp. 165-170
Download PDF

References

  1. Huang, R., Li, M., Yang, D., Shi, J., Chang, X., Ye, Z., ... & Watanabe, S. (2024). Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 21, pp. 23802-23804).
  2. Liu, S., Hussain, A. S., Wu, Q., Sun, C., & Shan, Y. (2024). Mumu-llama: Multi-modal music understanding and generation via large language models. arXiv preprint arXiv: 2412.06660, 3(5), 6.
  3. Yu, D., Song, K., Lu, P., He, T., Tan, X., Ye, W., ... & Bian, J. (2023). Musicagent: An ai agent for music understanding and generation with large language models. arXiv preprint arXiv: 2310.11954.
  4. Xing, P., Plaat, A., & van Stein, N. (2025). CoComposer: LLM Multi-agent Collaborative Music Composition. arXiv preprint arXiv: 2509.00132.
  5. Mao, Z., Zhao, M., Wu, Q., Wakaki, H., & Mitsufuji, Y. (2025). Deepresonance: Enhancing multimodal music understanding via music-centric multi-way instruction tuning. arXiv preprint arXiv: 2502.12623.
  6. Yuan, R., Lin, H., Wang, Y., Tian, Z., Wu, S., Shen, T., ... & Guo, Y. (2024). Chatmusician: Understanding and generating music intrinsically with llm. arXiv preprint arXiv: 2402.16153.
  7. Shu, Y., Xu, H., Zhou, Z., Hengel, A. V. D., & Liu, L. (2024). MuseBarControl: Enhancing fine-grained control in symbolic music generation through pre-training and counterfactual loss. arXiv preprint arXiv: 2407.04331.
  8. Lu, P., Xu, X., Kang, C., Yu, B., Xing, C., Tan, X., & Bian, J. (2023). Musecoco: Generating symbolic music from text. arXiv preprint arXiv: 2306.00110.
  9. Liu, S., Hussain, A. S., Sun, C., & Shan, Y. (2024). Music understanding llama: Advancing text-to-music generation with question answering and captioning. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 286-290). IEEE.
  10. Brachman, M., El-Ashry, A., Dugan, C., & Geyer, W. (2025). Current and Future Use of Large Language Models for Knowledge Work. arXiv preprint arXiv: 2503.16774.