QUANTITATIVE ASSESSMENT OF FEATURE DICTIONARY INFORMATIVENESS FOR CLASSIFICATION OF SOURCE CODE WRITTEN BY HUMANS AND GENERATIVE ARTIFICIAL INTELLIGENCE MODELS
DOI:
https://doi.org/10.31891/2219-9365-2026-87-52Keywords:
eature dictionary informativeness, source code authorship attribution, generative AI code detection, classifier model tournament, intelligent monitoringAbstract
The article addresses the task of quantitatively assessing the informativeness of a feature dictionary for source code classification under the widespread use of generative artificial intelligence (AI) models such as ChatGPT and Claude Code. Existing approaches evaluate dictionary quality through aggregated metrics — best-model accuracy or mean accuracy across the tournament of classifier models — which characterise either a single specific architecture or an averaged value but do not reflect the distribution of model quality across the entire tournament. To overcome this limitation, a novel informativeness metric is proposed: the number of classifier models in the tournament that achieve 100% classification accuracy at the window level and at the class level. This metric is positioned as a complement to traditional metrics and characterises not a single model but the distribution of quality across the entire set of tournament architectures. The proposed metric is grounded in the input-data-array (IDA) formation methodology of the S.V. Holub research school, originally developed for Ukrainian text classification and extended in this work to source code classification. Experimentally, on a dataset of 12 authors (10 humans and 2 AI models) across 4 programming languages (Java, JavaScript, TypeScript, Python), 22 series were investigated in three scenarios: within-language (4 series), cross-language (10 series), and Human/AI discrimination (8 series). A tournament of 39 classifier model architectures including Dense, CNN, LSTM, GRU, BiLSTM, Attention, BERT, and the Group Method of Data Handling (GMDH) was used. It is shown that the addition of symbolic cross-links between words to the baseline feature dictionary increases the number of models with 100% accuracy at the class level by +3.23 models per series on average. The largest gain is observed in the within-language scenario (+6.25 models) and in the Human/AI discrimination task (+4.38 models). The proposed metric provides a more detailed characterisation of feature dictionary informativeness than aggregated mean accuracy, which is particularly important for selecting effective feature sets in source code authorship attribution.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Руслан Немов

This work is licensed under a Creative Commons Attribution 4.0 International License.


