SESSION MEMORY AND CONTEXT COMPRESSION FOR LOW-LATENCY RESPONSE GENERATION IN MESSAGING APPLICATIONS

Authors

DOI:

https://doi.org/10.31891/2219-9365-2026-87-6

Keywords:

large language models, messaging applications, session memory, context compression, hybrid retrieval, reranking

Abstract

The generation of responses in messaging applications under long dialogs and strict response-time requirements is analyzed. An architecture is proposed in which session memory forms a compact, incrementally updated representation of the conversation with citation anchors; context compression selects salient units under a fixed token budget; hybrid retrieval (BM25 and vector representations with messenger-specific signals) followed by reranking refines candidates; and the generation module provides streaming output with inference accelerations (efficient attention, speculative decoding, weight quantization, key–value caching). A methodology for factuality control with automated consistency checking and standardized inline citation is developed. Experimental evaluation is conducted on public retrieval-augmented and dialogue corpora complemented with anonymized fragments of real conversations; retrieval quality (nDCG@k, Recall@k), response usefulness and factuality (human side-by-side judgments and automatic metrics), as well as timing characteristics (median and 95th-percentile end-to-end latency with stage breakdown) are reported. The impact of context-compression degree and of hybrid-retrieval and reranking parameters on source consistency is investigated, with comparisons to baselines without session memory, without compression, and without reranking. Latency reductions are demonstrated while preserving or improving quality; the contribution of individual components is established via ablation, and practical recommendations are provided for different response-time budgets (mobile and server settings). The scientific novelty lies in unifying session memory with controlled context compression, hybrid retrieval, and reranking within a single system augmented with engineering accelerations to ensure low latency without loss of factuality; the practical significance is a robust methodology suitable for real-world deployment under privacy and reproducibility requirements.

Published

2026-09-10

How to Cite

MOCHERNIUK Ю., & MOKRYTSKYI А. (2026). SESSION MEMORY AND CONTEXT COMPRESSION FOR LOW-LATENCY RESPONSE GENERATION IN MESSAGING APPLICATIONS. MEASURING AND COMPUTING DEVICES IN TECHNOLOGICAL PROCESSES, (3), 53–60. https://doi.org/10.31891/2219-9365-2026-87-6