Artificial Intelligence-based Music Evaluation: Progress, Challenges and Prospects
Abstract
Music is among the most expressive and subjective of art forms, which makes its evaluation uniquely difficult. Traditional rule-based methods that rely on objective features such as pitch, rhythm, and audio quality often fail to capture the richness of human perception, particularly when it comes to subjective qualities like creativity, emotion, and cultural context. With the rise of Artificial Intelligence, researchers have begun exploring new approaches for music evaluation that better align with human judgment. This paper surveys the three major strands of work in this emerging field: human-grounded datasets for preference learning, embedding- and distribution-based metrics, and learned predictors and foundation-model evaluators. Each category is examined in detail, with representative works introduced alongside their respective strengths and limitations. The paper then compares these approaches, identifies challenges common across the field, and discusses possible future directions. By analyzing existing research and future prospects, this paper highlights the potential of Artificial Intelligence (AI) to transform music evaluation into a more reliable, inclusive, and scalable process.
References
- Fitch, W. (2013) Musical protolanguage: Darwin’s theory of language evolution revisited. Frontiers in Psychology, 4, 1-15.
- Wojnar, A. (1963) An analysis and synthesis procedure for feedback FM systems. IRE Transactions on Audio, 11, 54-62.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł. and Polosukhin, I. (2017) Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998-6008.
- Rombach, R., Blattmann, A., Lorenz, D., Esser, P. and Ommer, B. (2022) High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10684-10695.
- Yao, J., Zhang, K., Li, M., Chen, H., Wang, T., Liu, X. and Zhao, Y. (2025) SongEval: A benchmark dataset for song aesthetics evaluation. arXiv preprint arXiv: 2505.10793. Retrieved from https: //arxiv.org/abs/2505.10793
- Wang, S., Bao, Z. and E, J. (2021) Armor: A benchmark for meta-evaluation of artificial music. Proceedings of the 29th ACM International Conference on Multimedia, 3182-3190.
- Kim, Y., Park, J., Choi, H., Lee, J., Chen, J. and Nam, J. (2025) Music Arena: Live evaluation for text-to-music. arXiv preprint arXiv: 2507.20900. Retrieved from https: //arxiv.org/abs/2507.20900
- Kilgour, K., Zuluaga, M., Roblek, D. and Sharifi, M. (2018) Fréchet audio distance: A metric for evaluating music enhancement algorithms. arXiv preprint arXiv: 1812.08466. Retrieved from https: //arxiv.org/abs/1812.08466
- Huang, Y., Li, J., Xu, K., Chen, S., Wang, R. and Zhang, Y. (2025) Aligning text-to-music evaluation with human preferences. arXiv preprint arXiv: 2503.16669. Retrieved from https: //arxiv.org/abs/2503.16669
- Gui, A., Li, T., Sun, Q., Wang, Y. and Yang, X. (2024) Adapting Fréchet audio distance for generative music evaluation. ICASSP 2024-IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1061-1065. IEEE.
- Lo, C., Fu, S., Huang, W., Wang, X., Yamagishi, J., Yu, C. and Tsao, Y. (2019) MOSNet: Deep learning based objective assessment for voice conversion. arXiv preprint arXiv: 1904.08352. Retrieved from https: //arxiv.org/abs/1904.08352
- Reddy, C.K.A., Gopal, V. and Cutler, R. (2021) DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors. ICASSP 2021-IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6493-6497. IEEE.
- Mittag, G., Naderi, B., Möller, S., Ribeiro, F. and Cutler, R. (2021) NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets. arXiv preprint arXiv: 2104.09494. Retrieved from https: //arxiv.org/abs/2104.09494
- Deshmukh, S., Park, H., Li, C., Wang, Y., Liu, J., Ma, C. and Li, H. (2024) PAM: Prompting audio-language models for audio quality assessment. arXiv preprint arXiv: 2402.00282. Retrieved from https: //arxiv.org/abs/2402.00282