Reference Fidelity in AI-Generated Short-Form Video Across Four Prompt Packages
Abstract
Generative video systems can produce short clips from textual and visual instructions, yet their ability to preserve the content of a human reference remains uncertain. This exploratory study examines 160 videos generated with Doubao-Seedance-2.0 from 40 human-made short-form videos drawn from Douyin and Xiaohongshu. The references were divided evenly among humor, informational, demonstration, and storytelling content. Qwen3.8-Max analyzed each reference and produced four prompt packages labeled vague, vivid, keyword, and image-based. The fourth package also contained timestamped reference frames, frame descriptions, identity constraints, and sequence instructions, so it is treated as a multimodal reference condition rather than a pure image condition. Each prompt was used once under the same default settings, and no output was edited after generation. Metadata inspection covered the complete corpus. All 40 groups received a standardized frame review, and eight groups were examined at three temporal positions. Outputs usually retained the source topic and central action. Fidelity was stronger for one-subject situations than for multi-step procedures, factual labels, comic timing, or longer causal sequences. The multimodal package appeared closer to source composition in several cases, yet the one-shot design and unequal prompt information prevent a causal ranking. The study therefore evaluates reference fidelity rather than virality or prompt efficacy.
References
- Grzenkowicz, M., & Wildfeuer, J. (2025). Addressing TikTok's multimodal complexity: A multi-level annotation scheme for the audio-visual design of short video content.Digital Scholarship in the Humanities, 40(4), 1143-1166. https://doi.org/10.1093/llc/fqaf047
- Lubart, T. (2005). How can computers be partners in the creative process: Classification and commentary on the special issue.International Journal of Human-Computer Studies, 63(4-5), 365-369. https://doi.org/10.1016/j.ijhcs.2005.04.002
- Rezwana, J., & Maher, M. L. (2023). Designing creative AI partners with COFI: A framework for modeling interaction in human-AI co-creative systems.ACM Transactions on Computer-Human Interaction, 30(5), Article 67, 1-28. https://doi.org/10.1145/3519026
- Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B., & Yang, Q. (2023). Why Johnny can't prompt: How non-AI experts try and fail to design LLM prompts. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Article 437, pp. 1-21). https://doi.org/10.1145/3544548.3581388
- Berger, J., & Milkman, K. L. (2012). What makes online content viral?Journal of Marketing Research, 49(2), 192-205. https://doi.org/10.1509/jmr.10.0353
- Broxton, T., Interian, Y., Vaver, J., & Wattenhofer, M. (2013). Catching a viral video.Journal of Intelligent Information Systems, 40(2), 241-259. https://doi.org/10.1007/s10844-011-0191-2
- France, S. L., Vaghefi, M. S., & Zhao, H. (2016). Characterizing viral videos: Methodology and applications.Electronic Commerce Research and Applications, 19, 19-32. https://doi.org/10.1016/j.elerap.2016.07.002
- Jiang, L., Miao, Y., Yang, Y., Lan, Z., & Hauptmann, A. G. (2014). Viral video style: A closer look at viral videos on YouTube. InProceedings of the International Conference on Multimedia Retrieval(Glasgow, United Kingdom, pp. 193-200). https://doi.org/10.1145/2578726.2578754
- Ling, C., Blackburn, J., De Cristofaro, E., & Stringhini, G. (2022). Slapping cats, bopping heads, and Oreo shakes: Understanding indicators of virality in TikTok short videos. InProceedings of the 14th ACM Web Science Conference(pp. 164-173). https://doi.org/10.1145/3501247.3531551
- Chen, Z., He, Q., Mao, Z., Chung, H.-M., & Maharjan, S. (2019). A study on the characteristics of Douyin short videos and implications for edge caching. InProceedings of the ACM Turing Celebration Conference China(pp. 1-6). https://doi.org/10.1145/3321408.3323082