SESSION MEMORY AND CONTEXT COMPRESSION FOR LOW-LATENCY RESPONSE GENERATION IN MESSAGING APPLICATIONS
DOI:
https://doi.org/10.31891/2219-9365-2026-87-6Keywords:
large language models, messaging applications, session memory, context compression, hybrid retrieval, rerankingAbstract
The generation of responses in messaging applications under long dialogs and strict response-time requirements is analyzed. An architecture is proposed in which session memory forms a compact, incrementally updated representation of the conversation with citation anchors; context compression selects salient units under a fixed token budget; hybrid retrieval (BM25 and vector representations with messenger-specific signals) followed by reranking refines candidates; and the generation module provides streaming output with inference accelerations (efficient attention, speculative decoding, weight quantization, key–value caching). A methodology for factuality control with automated consistency checking and standardized inline citation is developed. Experimental evaluation is conducted on public retrieval-augmented and dialogue corpora complemented with anonymized fragments of real conversations; retrieval quality (nDCG@k, Recall@k), response usefulness and factuality (human side-by-side judgments and automatic metrics), as well as timing characteristics (median and 95th-percentile end-to-end latency with stage breakdown) are reported. The impact of context-compression degree and of hybrid-retrieval and reranking parameters on source consistency is investigated, with comparisons to baselines without session memory, without compression, and without reranking. Latency reductions are demonstrated while preserving or improving quality; the contribution of individual components is established via ablation, and practical recommendations are provided for different response-time budgets (mobile and server settings). The scientific novelty lies in unifying session memory with controlled context compression, hybrid retrieval, and reranking within a single system augmented with engineering accelerations to ensure low latency without loss of factuality; the practical significance is a robust methodology suitable for real-world deployment under privacy and reproducibility requirements.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Мочернюк Юрій Миколайович, Мокрицький Андрій Анатолійович

This work is licensed under a Creative Commons Attribution 4.0 International License.


