Cross-Cultural Comparative Analysis of Nostalgic Imagery in Classical Chinese and Western Literature via Large Language Model Semantic Embeddings
Abstract
The imagery of nostalgia is a common emotional prototype of human beings, but its semantic expression in Chinese and foreign classical literature shows significant cultural differences. This study constructs a nostalgic imagery semantic embedding framework based on the multilingual large language model, and conducts high-dimensional semantic coding of nostalgia images in Chinese classical poetry, modern nostalgia prose and English, German and Russian classical literature corpus through the BGE M3-Embedding and XLM-RoBERTa models. The study uses UMAP dimension reduction and HDBSCAN clustering methods to analyze the semantic spatial structure of images, and designs two types of indicators, cross-cultural semantic distance and imagery commensurability index, to quantify the similarities and differences between Chinese and foreign nostalgic images. On the corpus covering 2,847 image samples, the image classification accuracy rate reached 89.3%, and the cross-cultural alignment F1 value was 0.84. The study found that core images such as "moon", "homecoming road" and "home" have high cross-cultural commensurability (cosine similarity ≥ 0.80), while images such as "plum blossom", "hometown dialect" and "sangzi" show obvious cultural specificity. This research provides a quantifiable and reproducible methodological framework for the comparison of cross-cultural literature in the field of digital humanities.
References
- Bode K. The Equivalence of "Close" and "Distant" Reading; or, Toward a New Object for Data-Rich Literary History. Modern Language Quarterly, 2017, 78(1): 77-106. https: //doi.org/10.1215/00267929-3699787
- Chen J, Xiao S, Zhang P, Luo K, Lian D, Liu Z. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv preprint arXiv: 2402.03216, 2024. https: //doi.org/10.48550/arXiv.2402.03216
- Bommasani R, Hudson D A, Adeli E, et al. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv: 2108.07258, 2021. https: //doi.org/10.48550/arXiv.2108.07258
- Piper A. Novel Devotions: Conversional Reading, Computational Modeling, and the Modern Novel. New Literary History, 2015, 46(1): 63-98. https: //doi.org/10.1353/nlh.2015.0008
- Hershcovich D, Frank S, Lent H, et al. Challenges and Strategies in Cross-Cultural NLP. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022: 6997-7013. https: //doi.org/10.18653/v1/2022.acl-long.482
- Cao Y, Zhou L, Lee S, et al. Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study. In: Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP), 2023: 53-67. https: //doi.org/10.48550/arXiv.2303.17466
- Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019: 4171-4186. https: //doi.org/10.18653/v1/N19-1423
- Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, Grave E, Ott M, Zettlemoyer L, Stoyanov V. Unsupervised Cross-lingual Representation Learning at Scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020: 8440-8451. https: //doi.org/10.18653/v1/2020.acl-main.747
- Pires T, Schlinger E, Garrette D. How multilingual is Multilingual BERT? In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019: 4996-5001. https: //doi.org/10.18653/v1/P19-1493
- Wang L, Yang N, Huang X, Yang L, Majumder R, Wei F. Multilingual E5 Text Embeddings: A Technical Report. arXiv preprint arXiv: 2402.05672, 2024. https: //doi.org/10.48550/arXiv.2402.05672
- Sun Y, Wang S, Feng S, et al. ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation. arXiv preprint arXiv: 2107.02137, 2021. https: //doi.org/10.48550/arXiv.2107.02137
- Zeng A, Liu X, Du Z, et al. GLM-130B: An Open Bilingual Pre-trained Model. In: The Eleventh International Conference on Learning Representations, 2023. https: //doi.org/10.48550/arXiv.2210.02414
- Liu P, Yuan W, Fu J, Jiang Z, Hayashi H, Neubig G. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 2023, 55(9): 1-35. https: //doi.org/10.1145/3560815
- Reimers N, Gurevych I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 2019: 3982-3992. https: //doi.org/10.18653/v1/D19-1410
- Brown T, Mann B, Ryder N, et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 2020, 33: 1877-1901. https: //doi.org/10.48550/arXiv.2005.14165
- McInnes L, Healy J, Saul N, Großberger L. UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software, 2018, 3(29): 861. https: //doi.org/10.21105/joss.00861