Research Article

Continuous Valence–Arousal Regression for Music Emotion Estimation Using Machine Learning with OpenL3 Embeddings

Jingyi LyuCommunication University of China*

* Corresponding author: [email protected]

Abstract

Music Emotion Recognition (MER) aims to model the mapping between acoustic features and emotional representations. MER has important value in applications such as music recommendation, automatic accompaniment, and automatic music generation, yet remains challenging due to strong subjectivity and complex temporal dynamics. Based on the public MediaEval Database for Emotional Analysis of Music (DEAM), this study employs OpenL3 pre-trained audio embeddings as a unified feature representation and performs frame-level feature extraction with temporal alignment to 2 Hz continuous Valence-Arousal (VA) annotations. On this basis, three lightweight regression models, including a Multilayer Perception (MLP), a Bidirectional Long Short-term Memory network (BiLSTM), and a Transformer Encoder are constructed to VA regression. Model performance is evaluated using Root Square Error (RMSE) and the Pearson Correlation Coefficient (PCC). Experimental results show that the MLP achieves the best overall performance on the test set, with lower RMSE and higher correlation than the BiLSTM and Transformer Encoder. These results demonstrate that pre-trained audio representations enable stable and efficient continuous music emotion regression with lightweight machine learning models.

Keywords: Music emotion recognition; machine learning; valence; arousal; OpenL3
Published: February 2, 2026
DOI: 10.54254/2753-7064/2026.HT31555
Volume: CHR Vol.102
pp. 149-155
Download PDF

References

  1. Han, D., Kong, Y., Han, J. and Wang, G. (2022) A survey of music emotion recognition. Frontiers of Computer Science, 16, 166335.
  2. Jiang, X., Zhang, Y., Lin, G. and Yu, L. (2024) Music emotion recognition based on deep learning: A review. IEEE Access, 12, 157716–157745.
  3. Russell, J.A. (1980) A circumplex model of affect. Journal of Personality and Social Psychology, 39, 1161–1178.
  4. Schaab, L. and Kruspe, A. (2024) Joint sentiment analysis of lyrics and audio in music. arXiv preprint, arXiv: 2405.01988.
  5. Soleymani, M., Aljanaki, A. and Yang, Y.-H. (2018) DEAM: MediaEval database for emotional analysis in music. University of Geneva and Academia Sinica, Switzerland.
  6. Cramer, A., Wu, H.-H., Salamon, J. and Bello, J.P. (2019) Look, listen, and learn more: Design choices for deep audio embeddings. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 3852–3856.
  7. Koh, E. and Dubnov, S. (2021) Comparison and analysis of deep audio embeddings for music emotion recognition. arXiv preprint, arXiv: 2104.06517.
  8. Kang, J. and Herremans, D. (2025) Towards unified music emotion recognition across dimensional and categorical models. arXiv preprint, arXiv: 2502.03979.
  9. Liyanarachchi, R., Joshi, A. and Meijering, E. (2025) A survey on multimodal music emotion recognition. arXiv preprint, arXiv: 2504.18799.
  10. Yang, Y.-H., Lin, Y.-C., Su, Y.-F. and Chen, H.H. (2008) A regression approach to music emotion recognition. IEEE Transactions on Audio, Speech and Language Processing, 16, 448–457.
  11. Chaki, S., Doshi, P., Bhattacharya, S. and Patnaik, P. (2020) Explaining perceived emotion predictions in music: An attentive approach. Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 150–156.
  12. Chai, T. and Draxler, R.R. (2014) Root mean square error (RMSE) or mean absolute error (MAE)? — Arguments against avoiding RMSE in the literature. Geoscientific Model Development, 7, 1247–1250.
  13. Pearson, K. (1895) Notes on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London, 58, 240–242.