METHOD FOR INTELLIGENT ANALYSIS OF UKRAINIAN-LANGUAGE MESSAGES FOR PROPAGANDA DETECTION
DOI:
https://doi.org/10.31891/2219-9365-2026-87-53Keywords:
propaganda detection, intelligent text analysis, natural language processing, sentiment analysis, Ukrainian-language messages, classificationAbstract
This paper proposes and experimentally validates a method for binary propaganda detection in short Ukrainian language messages based on combining lexical statistical features with tonal and emotional indicators. The study addresses short text classification, where lexical sparsity, informal writing, and rapid topic shifts limit purely lexical models. The proposed approach represents each message using a TF–IDF model configured for short texts with n–gram features, and extends this representation with a compact vector of sentiment–related numeric indicators computed from a Ukrainian sentiment lexicon and expressive markers. The tonal block quantifies polarity and emotional intensity using lexicon–based positive and negative scores normalized by message length, the share of emotionally marked words, and additional indicators of expressive style such as normalized exclamation usage and capitalization ratio. The final feature space is formed by fusing the sparse TF–IDF vector with the dense tonal vector and by applying controlled weighting of feature groups to isolate the contribution of tonal information. Classification is performed with a linear Passive–Aggressive model, which is suitable for high–dimensional sparse representations and fast training. The evaluation follows a controlled comparison of two pipelines that differ only in the presence of the tonal feature block, TF–IDF only versus TF–IDF plus tonal and emotional indicators. Performance is assessed using accuracy, precision, recall, and F1–score, supported by confusion matrices and reproducible reporting. The results show that adding tonal features increases recall by about 2.7 percent on average while keeping precision around 80 percent, and yields a modest improvement in overall accuracy of about 1.4 percent and in F1 of about 1.5 percent. In practical testing scenarios, the method achieves roughly 83 percent accuracy and about 86 percent recall, supporting its applicability for monitoring information flows and flagging potentially manipulative content for subsequent expert review.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Сергій Мостовий, Андрій Северін

This work is licensed under a Creative Commons Attribution 4.0 International License.


