The verbal component of video game text as a resource for retraining large language models

Authors

DOI:

https://doi.org/10.33910/2686-830X-2026-8-1-6-13

Keywords:

video game text, verbal component, fine-tuning, machine learning, large language model

Abstract

The study was conducted in the field of digital linguistics, whose objectives include studying the parameters of natural language processing. The problems of large language models with understanding communicative context are related to fundamental limitations of their nature, among the absence of genuine emotional intelligence, the complexity of interpreting multimodal signals, and difficulties with contextual understanding of situations. One way to overcome these limitations is to use a retraining method on specialised data sets. The article examines the verbal component of polycode video game text as a most important component of multimodal discourse, which determines the perception of events, characters and the nature of the player’s interaction with the virtual game world. Particular attention is paid to the specific characteristics of the linguistic material in video games: the complexity of the lexical, syntactic and stylistic organisation of game dialogues, emotional intensity, functional informativeness, and contextual variability. It is shown that game dialogues possess built-in annotations of emotions and pragmatic functions, which makes them a valuable resource for natural language processing tasks, including sentiment analysis, speech act recognition, and training of empathetic language models. Using fragments from the verbal component of the video game Disco Elysium, we demonstrate the implementation of emotional and contextual speech parameters of characters in the English and Russian versions. The importance of video game texts as a source of linguistic data for research in the field of corpus technologies and machine learning is noted. It is concluded that the verbal component of video games is not only an artistic tool for narrative organisation, but also a promising basis for the development of interactive interfaces and context-sensitive language models.

References

СПИСОК ЛИТЕРАТУРЫ

Кац, Н. Г., Рубцова, А. В., Аитов, В. Ф. (2024) Иноязычный мультимодальный текст как основа разработки учебных материалов для цифровой образовательной среды вуза. Вестник Томского государственного университета, № 504, с. 156–163. https://doi.org/10.17223/15617793/504/17

Кожемякин, Е. А., Иванова, Л. В., Куприянова, А. В. (2024) Прагматический аспект мультимодальной журналистской коммуникации: взаимодействие реципиентов с цифровым дата-текстом. Вестник Московского университета. Серия 10: Журналистика, т. 49, № 3, с. 89–121.

Финогеева, А. А. (2021) Мультимодальность в современном университетском интернет-дискурсе. Russian Linguistic Bulletin, № 4 (28), с. 162–165. https://doi.org/10.18454/RULB.2021.28.4.39

Blum, A., Mitchell, T. (1998) Combining Labeled and Unlabeled Data with Co-Training. In: COLT’ 98: Proceedings of the eleventh annual conference on Computational learning theory. New York: Association for Computing Machinery Publ., pp. 92–100. https://doi.org/10.1145/279943.279962

Callison-Burch, C., Dredze, M. (2010) Creating Speech and Language Data with Amazon’s Mechanical Turk. In: Proceedings of the NAACL HLT 2009 Workshop on Creating Speech and Language Data with Amazon Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 1–12.

Finin, T., Murnane, W., Karandikar, A. et al. (2010) Annotating Named Entities in Twitter Data with Crowdsourcing. In: Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 80–88.

Hämäläinen, M., Alnajjar, K., Poibeau, T. (2022) Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. arXiv. [Online]. Available at: https://doi.org/10.48550/arXiv.2212.02168 (дата обращения 20.09.2025).

Juraska, J., Bowden, K., Walker, M. (2019) ViGGO: A Video Game Corpus for Data-To-Text Generation in Open- Domain Conversation. In: Proceedings of the 12th International Conference on Natural Language Generation. Tokyo: Association for Computational Linguistics Publ., pp. 164–172. https://doi.org/10.18653/v1/W19-8623

Nangia, N., Bowman, S. R. (2019) Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics Publ., pp. 4566–4575. https://doi.org/10.18653/v1/P19-1449

Pavlick, E., Post, M., Irvine, A. et al. (2014) The Language Demographics of Amazon Mechanical Turk. Transactions of the Association for Computational Linguistics, vol. 2, pp. 79–92 https://doi.org/10.1162/tacl_a_00167

Rennick, S., Roberts, S. G. (2023) The Video Game Dialogue Corpus. Corpora, vol. 19, no. 1, pp. 93–106. https://doi.org/10.3366/cor.2024.0299

Zhu, X. (2005) Semi-Supervised Learning Literature Survey. Madison: University of Wisconsin Publ., 60 p.

SOURCES

Disco Elysium. (2025) [Online]. Available at: https://store.steampowered.com/app/632470/Disco_Elysium__The_Final_Cut/ (accessed 20.09.2025).

Disco Elysium Wiki. (2025) [Online]. Available at: https://discoelysium.fandom.com/wiki/Disco_Elysium_Wiki (accessed 20.09.2025)

Disco Reader. (2025) [Online]. Available at: https://disco-reader.gitlab.io/disco-reader/#/ (accessed 20.09.2025)

REFERENCES

Blum, A., Mitchell, T. (1998) Combining Labeled and Unlabeled Data with Co-Training. In: COLT’ 98: Proceedings of the eleventh annual conference on Computational learning theory. New York: Association for Computing Machinery Publ., pp. 92–100. https://doi.org/10.1145/279943.279962 (In English)

Callison-Burch, C., Dredze, M. (2010) Creating Speech and Language Data with Amazon’s Mechanical Turk. In: Proceedings of the NAACL HLT 2009 Workshop on Creating Speech and Language Data with Amazon Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 1–12. (In English)

Finin, T., Murnane, W., Karandikar, A. et al. (2010) Annotating Named Entities in Twitter Data with Crowdsourcing. In: Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk. Los Angeles: Association for Computational Linguistics Publ., pp. 80–88. (In English)

Finogeeva, A. A. (2021) Multimodality in the modern internet discourse. Russian Linguistic Bulletin, vol. 4 (28), pp. 162–165. https://doi.org/10.18454/RULB.2021.28.4.39 (In Russian)

Hämäläinen, M., Alnajjar, K., Poibeau, T. (2022) Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. arXiv. [Online]. Available at: https://doi.org/10.48550/arXiv.2212.02168 (дата обращения 20.09.2025). (In English)

Juraska, J., Bowden, K., Walker, M. (2019) ViGGO: A Video Game Corpus for Data-To-Text Generation in Open- Domain Conversation. In: Proceedings of the 12th International Conference on Natural Language Generation. Tokyo: Association for Computational Linguistics Publ., pp. 164–172. https://doi.org/10.18653/v1/W19-8623 (In English)

Kats, N. G., Rubtsova, A. V., Aitov, V. F. (2024) Multimodal text as the basis of instructional material design in an academic digital environment. Tomsk State University Journal, vol. 504, pp. 156–163. https://doi.org/10.17223/15617793/504/17 (In Russian)

Kozhemyakin, E. A., Ivanova, L. V., Kupriyanova, A. V. (2024) Pragmatic aspect of multimodal journalistic communication: interaction of recipients with digital data-text. Vestnik Moskovskogo universiteta. Seriya 10. Zhurnalistika, vol. 49, no. 3, pp. 89–121. (In Russian)

Nangia, N., Bowman, S. R. (2019) Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics Publ., pp. 4566–4575. https://doi.org/10.18653/v1/P19-1449 (In English)

Pavlick, E., Post, M., Irvine, A. et al. (2014) The Language Demographics of Amazon Mechanical Turk. Transactions of the Association for Computational Linguistics, vol. 2, pp. 79–92 https://doi.org/10.1162/tacl_a_00167 (In English)

Rennick, S., Roberts, S. G. (2023) The Video Game Dialogue Corpus. Corpora, vol. 19, no. 1, pp. 93–106. https://doi.org/10.3366/cor.2024.0299 (In English)

Zhu, X. (2005) Semi-Supervised Learning Literature Survey. Madison: University of Wisconsin Publ., 60 p. (In English)

Published

2026-05-16

Issue

Section

Linguistics and Interdisciplinary Research in Language and Discourse

Similar Articles

1-10 of 56

You may also start an advanced similarity search for this article.