STOCHASTIC ATTACK–DEFENSE CONFRONTATION MODEL FOR WEB APPLICATION SECURITY TESTING BASED ON REINFORCEMENT LEARNING

Authors

DOI:

https://doi.org/10.31891/2219-9365-2026-87-14

Keywords:

web application security testing, reinforcement learning, Markov game, time to compromise, two-agent model, stochastic modeling, cybersecurity, machine learning, cyberattack, threat, artificial intelligence

Abstract

The paper proposes a stochastic attack–defense confrontation model for web application security testing based on reinforcement learning. The relevance of the work stems from the growing complexity of modern web applications and the adaptive nature of cyber threats, for which static attack graphs and one-sided testing models fail to account for active defense responses and the temporal dynamics of compromise. An analysis of existing approaches showed that most models consider only the attacking or only the defending side, do not use time-to-compromise as the main optimization criterion, and generate excessively large state spaces that hinder the analysis of adaptive scenarios. Unlike one-sided models, the approach is implemented as a two-agent Markov game, in which the actions of the testing agent and the defense agent are considered simultaneously in a shared state space with a two-sided reward function oriented toward time-to-compromise (TTC), while the policy is trained using a risk-sensitive reinforcement learning method that directs the search toward short critical trajectories. The model uses an aggregated representation of system components as an aggregation of an attack graph built with the Meta Attack Language, together with a stochastic initialization of initial conditions, which reduces the dimensionality of the state space and ensures the computational tractability of the analysis. Empirical evaluation on the ToN-IoT and CIC-IDS2017 datasets showed that the proposed model reduces the average time-to-compromise by 34.7–37.7 % compared to classical attack graphs and decreases the dimensionality of the actually visited state space by a factor of 8–12, which accelerates scenario analysis and improves scalability. The active behavior of the defense makes it possible to evaluate not only the fact but also the temporal characteristics of compromise. The results can be used as a formal basis for adaptive web application security testing tools in environments with partial observability and variable defense parameters.

Published

2026-09-10

How to Cite

PRYTULA А., & KUPERSHTEIN Л. (2026). STOCHASTIC ATTACK–DEFENSE CONFRONTATION MODEL FOR WEB APPLICATION SECURITY TESTING BASED ON REINFORCEMENT LEARNING. MEASURING AND COMPUTING DEVICES IN TECHNOLOGICAL PROCESSES, (3), 126–134. https://doi.org/10.31891/2219-9365-2026-87-14