<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">vestrea</journal-id><journal-title-group><journal-title xml:lang="ru">Вестник Российского экономического университета имени Г. В. Плеханова</journal-title><trans-title-group xml:lang="en"><trans-title>Vestnik of the Plekhanov Russian University of Economics</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2413-2829</issn><issn pub-type="epub">2587-9251</issn><publisher><publisher-name>Plekhanov Russian University of Economics</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.21686/2413-2829-2026-1-72-78</article-id><article-id custom-type="elpub" pub-id-type="custom">vestrea-2593</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>УПРАВЛЕНИЕ ИННОВАЦИЯМИ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>INNOVATION MANAGEMENT</subject></subj-group></article-categories><title-group><article-title>Интеграция мультимодальных данных в генерацию текстовых описаний: методы, вызовы и перспективы</article-title><trans-title-group xml:lang="en"><trans-title>Integration of Multi-Modal Data into Generation of Text Descriptions: Methods, Challenges and Prospects</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Чиняков</surname><given-names>Н. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Chinyakov</surname><given-names>N. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Никита Александрович Чиняков – аспирант кафедры информатики РЭУ им. Г. В. Плеханова </p><p>109992, Москва, Стремянный пер., д. 36</p></bio><bio xml:lang="en"><p>Nikita A. Chinyakov, Post-Graduate Student of the Department for Informatics of the PRUE.</p><p>36 Stremyanny Lane, Moscow, 109992</p></bio><email xlink:type="simple">arinachinyakova99@mail.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Российский экономический университет имени Г. В. Плеханова</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Plekhanov Russian University of Economics</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>30</day><month>01</month><year>2026</year></pub-date><volume>0</volume><issue>1</issue><fpage>72</fpage><lpage>78</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Чиняков Н.А., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Чиняков Н.А.</copyright-holder><copyright-holder xml:lang="en">Chinyakov N.A.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://vest.rea.ru/jour/article/view/2593">https://vest.rea.ru/jour/article/view/2593</self-uri><abstract><p>Современные системы искусственного интеллекта все чаще используют мультимодальные данные, комбинируя визуальную, текстовую и аудиальную информацию для решения сложных задач. Одной из ключевых областей применения таких систем является генерация текстовых описаний на основе изображений и видео. Интеграция мультимодальных данных позволяет повысить точность и выразительность создаваемых текстов, обеспечивая более полное и осмысленное представление содержимого. В статье рассматриваются современные методы интеграции мультимодальных данных в генерацию текстовых описаний, анализируются ключевые вызовы, с которыми сталкиваются исследователи, а также обсуждаются перспективные направления развития этой области. Особое внимание уделяется использованию сверточных нейронных сетей (CNN) и трансформеров для обработки визуальной информации, а также механизмов внимания и моделей последовательной генерации текста. Исследуются подходы к фьюжну данных из разных модальностей, включая раннее и позднее объединение признаков, а также мультимодальные модели, обученные на больших корпусах данных. Несмотря на значительный прогресс, интеграция мультимодальных данных сопровождается рядом вызовов, включая проблему синхронизации информации, сложности в интерпретации и контексте, ограничения в обучающих данных и др. Обсуждаются перспективные направления развития. Полученные результаты могут быть полезны для разработчиков систем компьютерного зрения, обработки естественного языка и мультимодального машинного обучения, а также для создания интеллектуальных приложений в области автоматической аннотации изображений, видеосуммаризации и человеко-машинного взаимодействия.</p></abstract><trans-abstract xml:lang="en"><p>Current systems of AI more and more often use multi-modal data by combining visual, text and audio information to resolve complicated problems. One key sphere of such system application is description generation on the basis of images and video. Integration of multi-modal data provides an opportunity to improve accuracy and expressiveness of texts being created, which gives more complete and sensible representation of the content. The article studies current methods of multi-modal data integration into generation of text descriptions, analyzes key challenges that face researchers and discusses promising trends in this field of development. Special attention is paid to application of convolution neuron nets (CNN) and transformers to process visual information, as well as attention mechanisms and models of successive text generation. The article researches approaches to data fusion formed by different modalities, including earlier and later combination of signs and multi-modal models trained on big blocks of data. In spite of serious progress, integration of multi-modal data provokes a number of challenges, including information synchronization, problems in interpretation and context, restrictions in learning data and others. Promising lines in development are being discussed. Obtained results can be used by developers of computer vision systems, in processing natural language and multi-modal machine learning and for elaboration of intellectual applications in the field of automatic image abstracts, video-summarizing and man-machine interaction.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>задача детекции</kwd><kwd>сверточные нейронные сети</kwd><kwd>мультимодальные модели</kwd><kwd>искусственный интеллект</kwd></kwd-group><kwd-group xml:lang="en"><kwd>detection tasks</kwd><kwd>convolution neuron nets</kwd><kwd>multi-modal models</kwd><kwd>AI</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Baltrusaitis T., Ahuja C., Morency L.-P. Multimodal Machine Learning: A Survey and Taxonomy // IEEE Transactions on Pattern Analysis and Machine Intelligence. – 2019. – Vol. 41 (2). – P. 423–443.</mixed-citation><mixed-citation xml:lang="en">Baltrusaitis T., Ahuja C., Morency L.-P. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, Vol. 41 (2), pp. 423–443.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Buolamwini J., Gebru T. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification //Proceedings of the 1st Conference on Fairness, Accountability and Transparency. – URL: https://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a.pdf?utm_source=chatgpt.com</mixed-citation><mixed-citation xml:lang="en">Buolamwini J., Gebru T. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of the 1st Conference on Fairness, Accountability and Transparency. Available at: https://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a.pdf?utm_source=chatgpt.com</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Esteva A., Kuprel B., Novoa R. A. Dermatologist-Level Classification of Skin Cancer with Deep Neural Networks. – URL: https://www.researchgate.net/publication/312890808_Dermatologist-level_classification_of_skin_cancer_with_deep_neural_networks</mixed-citation><mixed-citation xml:lang="en">Esteva A., Kuprel B., Novoa R. A. Dermatologist-Level Classification of Skin Cancer with Deep Neural Networks. Available at: https://www.researchgate.net/publication/312890808_Dermatologist-level_classification_of_skin_cancer_with_deep_neural_networks</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Goodfellow I., Pouget-Abadie J., Mirza M. Generative Adversarial Nets. – URL: https://www.researchgate.net/publication/263012109_Generative_Adversarial_Networks</mixed-citation><mixed-citation xml:lang="en">Goodfellow I., Pouget-Abadie J., Mirza M. Generative Adversarial Nets. Available at: https://www.researchgate.net/publication/263012109_Generative_Adversarial_Networks</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Hochreiter S., Schmidhuber J. Long Short-Term Memory //Neural Computation. – 1997. – Vol. 9 (8). – P. 1735–1780.</mixed-citation><mixed-citation xml:lang="en">Hochreiter S., Schmidhuber J. Long Short-Term Memory. Neural Computation, 1997, Vol. 9 (8), pp. 1735–1780.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Jobin A., Ienca M., Andorno R. The Global Landscape of AI Ethics Guidelines //Nature Machine Intelligence. – URL: https://www.nature.com/articles/s42256-019-0088-2?utm_source=chatgpt.com</mixed-citation><mixed-citation xml:lang="en">Jobin A., Ienca M., Andorno R. The Global Landscape of AI Ethics Guidelines. Nature Machine Intelligence. Available at: https://www.nature.com/articles/s42256-019-0088-2?utm_source=chatgpt.com</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Kizilcec R. F., Piech C., Schneider E. F. Deconstructing Disengagement: Analyzing Learner Subpopulations in Massive Open Online Courses. – URL: https://www.researchgate.net/publication/260265661_Deconstructing_Disengagement_Analyzing_Learner_Subpopulations_in_Massive_Open_Online_Courses</mixed-citation><mixed-citation xml:lang="en">Kizilcec R. F., Piech C., Schneider E. F. Deconstructing Disengagement: Analyzing Learner Subpopulations in Massive Open Online Courses. Available at: https://www.researchgate.net/publication/260265661_Deconstructing_Disengagement_Analyzing_Learner_Subpopulations_in_Massive_Open_Online_Courses</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">McCormack J., Gifford T., Hutchings P. Autonomy, Authenticity and the Role of the Artist in the Age of AI. – URL: https://www.researchgate.net/publication/331562062_Autonomy_Authenticity_Authorship_and_Intention_in_computer_generated_art</mixed-citation><mixed-citation xml:lang="en">McCormack J., Gifford T., Hutchings P. Autonomy, Authenticity and the Role of the Artist in the Age of AI. Available at: https://www.researchgate.net/publication/331562062_Autonomy_Authenticity_Authorship_and_Intention_in_computer_generated_art</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Nguyen H., Wang Y., Zhang J. Multimodal Sentiment Analysis: A Survey on Methods and Applications //IEEE Transactions on Affective Computing. – URL: https://arxiv.org/abs/2305.07611?utm_source=chatgpt.com</mixed-citation><mixed-citation xml:lang="en">Nguyen H., Wang Y., Zhang J. Multimodal Sentiment Analysis: A Survey on Methods and Applications. IEEE Transactions on Affective Computing. Available at: https://arxiv.org/abs/2305.07611?utm_source=chatgpt.com</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Rihem F. Image Captioning Using Multimodal Deep Learning Approach //Computers, Materials &amp; Continua. – 2024. – Vol. 81 (3). – P. 3951–3968.</mixed-citation><mixed-citation xml:lang="en">Rihem F. Image Captioning Using Multimodal Deep Learning Approach. Computers, Materials &amp; Continua, 2024, Vol. 81 (3), pp. 3951–3968.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Stojkoska B. R., Avramova A. P., Chatzimisios P. Application of Wireless Sensor Networks for Indoor Temperature Regulation. – URL: https://arxiv.org/abs/1606.07386</mixed-citation><mixed-citation xml:lang="en">Stojkoska B. R., Avramova A. P., Chatzimisios P. Application of Wireless Sensor Networks for Indoor Temperature Regulation. Available at: https://arxiv.org/abs/1606.07386</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Sutskever I., Vinyals O., Le Q. V. Sequence to Sequence Learning with Neural Networks. – URL: https://arxiv.org/abs/1409.3215</mixed-citation><mixed-citation xml:lang="en">Sutskever I., Vinyals O., Le Q. V. Sequence to Sequence Learning with Neural Networks. Available at: https://arxiv.org/abs/1409.3215</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Gomez A. N., Kaiser L., Polosukhin I. Attention is All You Need. – URL: https://arxiv.org/abs/1706.03762</mixed-citation><mixed-citation xml:lang="en">Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Gomez A. N., Kaiser L., Polosukhin I. Attention is All You Need. Available at: https://arxiv.org/abs/1706.03762</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Vinyals O., Toshev A., Bengio S., Erhan D. Show and Tell: A Neural Image Caption Generator. – URL: https://arxiv.org/abs/1411.4555</mixed-citation><mixed-citation xml:lang="en">Vinyals O., Toshev A., Bengio S., Erhan D. Show and Tell: A Neural Image Caption Generator. Available at: https://arxiv.org/as/1411.4555</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Zadeh A. B., Liang P. P., Poria S., Cambria E., Morency L-P. Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. – URL: https://aclanthology.org/P18-1208/</mixed-citation><mixed-citation xml:lang="en">Zadeh A. B., Liang P. P., Poria S., Cambria E., Morency L-P. Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. Available at: https://aclanthology.org/P18-1208/</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Zhang Y., Liu F., Wang H., Hu Z. Multimodal Learning for Medical Image Analysis: A Survey //Medical Image Analysis. – 2023. – N 85. – P. 102759.</mixed-citation><mixed-citation xml:lang="en">Zhang Y., Liu F., Wang H., Hu Z. Multimodal Learning for Medical Image Analysis: A Survey. Medical Image Analysis, 2023, No. 85, p. 102759.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
