The verbal component of video game text as a resource for retraining large language models
DOI:
https://doi.org/10.33910/2686-830X-2026-8-1-6-13Keywords:
video game text, verbal component, fine-tuning, machine learning, large language modelAbstract
The study was conducted in the field of digital linguistics, whose objectives include studying the parameters of natural language processing. The problems of large language models with understanding communicative context are related to fundamental limitations of their nature, among the absence of genuine emotional intelligence, the complexity of interpreting multimodal signals, and difficulties with contextual understanding of situations. One way to overcome these limitations is to use a retraining method on specialised data sets. The article examines the verbal component of polycode video game text as a most important component of multimodal discourse, which determines the perception of events, characters and the nature of the player’s interaction with the virtual game world. Particular attention is paid to the specific characteristics of the linguistic material in video games: the complexity of the lexical, syntactic and stylistic organisation of game dialogues, emotional intensity, functional informativeness, and contextual variability. It is shown that game dialogues possess built-in annotations of emotions and pragmatic functions, which makes them a valuable resource for natural language processing tasks, including sentiment analysis, speech act recognition, and training of empathetic language models. Using fragments from the verbal component of the video game Disco Elysium, we demonstrate the implementation of emotional and contextual speech parameters of characters in the English and Russian versions. The importance of video game texts as a source of linguistic data for research in the field of corpus technologies and machine learning is noted. It is concluded that the verbal component of video games is not only an artistic tool for narrative organisation, but also a promising basis for the development of interactive interfaces and context-sensitive language models.
References
СПИСОК ЛИТЕРАТУРЫ
Кац, Н. Г., Рубцова, А. В., Аитов, В. Ф. (2024) Иноязычный мультимодальный текст как основа разработки учебных материалов для цифровой образовательной среды вуза. Вестник Томского государственного университета, № 504, с. 156–163. https://doi.org/10.17223/15617793/504/17
Кожемякин, Е. А., Иванова, Л. В., Куприянова, А. В. (2024) Прагматический аспект мультимодальной журналистской коммуникации: взаимодействие реципиентов с цифровым дата-текстом. Вестник Московского университета. Серия 10: Журналистика, т. 49, № 3, с. 89–121.
Финогеева, А. А. (2021) Мультимодальность в современном университетском интернет-дискурсе. Russian Linguistic Bulletin, № 4 (28), с. 162–165. https://doi.org/10.18454/RULB.2021.28.4.39
Blum, A., Mitchell, T. (1998) Combining Labeled and Unlabeled Data with Co-Training. In: COLT’ 98: Proceedings of the eleventh annual conference on Computational learning theory. New York: Association for Computing Machinery Publ., pp. 92–100. https://doi.org/10.1145/279943.279962
Callison-Burch, C., Dredze, M. (2010) Creating Speech and Language Data with Amazon’s Mechanical Turk. In: Proceedings of the NAACL HLT 2009 Workshop on Creating Speech and Language Data with Amazon Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 1–12.
Finin, T., Murnane, W., Karandikar, A. et al. (2010) Annotating Named Entities in Twitter Data with Crowdsourcing. In: Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 80–88.
Hämäläinen, M., Alnajjar, K., Poibeau, T. (2022) Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. arXiv. [Online]. Available at: https://doi.org/10.48550/arXiv.2212.02168 (дата обращения 20.09.2025).
Juraska, J., Bowden, K., Walker, M. (2019) ViGGO: A Video Game Corpus for Data-To-Text Generation in Open- Domain Conversation. In: Proceedings of the 12th International Conference on Natural Language Generation. Tokyo: Association for Computational Linguistics Publ., pp. 164–172. https://doi.org/10.18653/v1/W19-8623
Nangia, N., Bowman, S. R. (2019) Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics Publ., pp. 4566–4575. https://doi.org/10.18653/v1/P19-1449
Pavlick, E., Post, M., Irvine, A. et al. (2014) The Language Demographics of Amazon Mechanical Turk. Transactions of the Association for Computational Linguistics, vol. 2, pp. 79–92 https://doi.org/10.1162/tacl_a_00167
Rennick, S., Roberts, S. G. (2023) The Video Game Dialogue Corpus. Corpora, vol. 19, no. 1, pp. 93–106. https://doi.org/10.3366/cor.2024.0299
Zhu, X. (2005) Semi-Supervised Learning Literature Survey. Madison: University of Wisconsin Publ., 60 p.
SOURCES
Disco Elysium. (2025) [Online]. Available at: https://store.steampowered.com/app/632470/Disco_Elysium__The_Final_Cut/ (accessed 20.09.2025).
Disco Elysium Wiki. (2025) [Online]. Available at: https://discoelysium.fandom.com/wiki/Disco_Elysium_Wiki (accessed 20.09.2025)
Disco Reader. (2025) [Online]. Available at: https://disco-reader.gitlab.io/disco-reader/#/ (accessed 20.09.2025)
REFERENCES
Blum, A., Mitchell, T. (1998) Combining Labeled and Unlabeled Data with Co-Training. In: COLT’ 98: Proceedings of the eleventh annual conference on Computational learning theory. New York: Association for Computing Machinery Publ., pp. 92–100. https://doi.org/10.1145/279943.279962 (In English)
Callison-Burch, C., Dredze, M. (2010) Creating Speech and Language Data with Amazon’s Mechanical Turk. In: Proceedings of the NAACL HLT 2009 Workshop on Creating Speech and Language Data with Amazon Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 1–12. (In English)
Finin, T., Murnane, W., Karandikar, A. et al. (2010) Annotating Named Entities in Twitter Data with Crowdsourcing. In: Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 80–88. (In English)
Finogeeva, A. A. (2021) Multimodality in the modern internet discourse. Russian Linguistic Bulletin, vol. 4 (28), pp. 162–165. https://doi.org/10.18454/RULB.2021.28.4.39 (In Russian)
Hämäläinen, M., Alnajjar, K., Poibeau, T. (2022) Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. arXiv. [Online]. Available at: https://doi.org/10.48550/arXiv.2212.02168 (дата обращения 20.09.2025). (In English)
Juraska, J., Bowden, K., Walker, M. (2019) ViGGO: A Video Game Corpus for Data-To-Text Generation in Open- Domain Conversation. In: Proceedings of the 12th International Conference on Natural Language Generation. Tokyo: Association for Computational Linguistics Publ., pp. 164–172. https://doi.org/10.18653/v1/W19-8623 (In English)
Kats, N. G., Rubtsova, A. V., Aitov, V. F. (2024) Multimodal text as the basis of instructional material design in an academic digital environment. Tomsk State University Journal, vol. 504, pp. 156–163. https://doi.org/10.17223/15617793/504/17 (In Russian)
Kozhemyakin, E. A., Ivanova, L. V., Kupriyanova, A. V. (2024) Pragmatic aspect of multimodal journalistic communication: interaction of recipients with digital data-text. Vestnik Moskovskogo universiteta. Seriya 10. Zhurnalistika, vol. 49, no. 3, pp. 89–121. (In Russian)
Nangia, N., Bowman, S. R. (2019) Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics Publ., pp. 4566–4575. https://doi.org/10.18653/v1/P19-1449 (In English)
Pavlick, E., Post, M., Irvine, A. et al. (2014) The Language Demographics of Amazon Mechanical Turk. Transactions of the Association for Computational Linguistics, vol. 2, pp. 79–92 https://doi.org/10.1162/tacl_a_00167 (In English)
Rennick, S., Roberts, S. G. (2023) The Video Game Dialogue Corpus. Corpora, vol. 19, no. 1, pp. 93–106. https://doi.org/10.3366/cor.2024.0299 (In English)
Zhu, X. (2005) Semi-Supervised Learning Literature Survey. Madison: University of Wisconsin Publ., 60 p. (In English)
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Olga Yu. Kustova, Gleb A. Poltavets

This work is licensed under a Creative Commons Attribution 4.0 International License.
The work is provided under the terms of the Public Offer and of Creative Commons public license Creative Commons Attribution 4.0 International (CC BY 4.0).
This license permits an unlimited number of users to copy and redistribute the material in any medium or format, and to remix, transform, and build upon the material for any purpose, including commercial use.
This license retains copyright for the authors but allows others to freely distribute, use, and adapt the work, on the mandatory condition that appropriate credit is given. Users must provide a correct link to the original publication in our journal, cite the authors' names, and indicate if any changes were made.
Copyright remains with the authors. The CC BY 4.0 license does not transfer rights to third parties but rather grants users prior permission for use, provided the attribution condition is met. Any use of the work will be governed by the terms of this license.





