Tag
This paper introduces MORSE, a compression-aware method for evidence-preserving context ordering in large language models that improves evidence retention and downstream QA performance.
This paper studies hidden states in long horizon language model agents, revealing that memory compression and recall needs are encoded before actions. It proposes the PaMER framework to reduce context consumption while maintaining task performance through state-guided compression and evidence retrieval.
The paper proposes a Context-to-Answer-Aligned Memory Compression (CMC) framework that compresses long input contexts into compact memory embeddings to reduce LLM inference costs without modifying decoder weights, achieving significant performance and efficiency gains.
This paper empirically analyzes cost savings in context-compression gateways for multi-turn coding agents, revealing that tool-schema filtering provides fixed token savings, while content compression saves quadratically but can be offset by recalls, offering actionable insights for cost optimization.
Jev introduced on-the-fly compression for AI agents, which scores each tool call to retain important content and delete irrelevant data directly, eliminating the need for model-based summarization and making context compression faster and more lightweight.
In 72 hours, the Jev ecosystem expanded from 46 to 160 projects, focusing on context compression, platform integrations, and new domains like financial trading, with debates on its novelty and implementation.
This paper introduces a paired-history audit method to evaluate update sufficiency in context compression for AI memory systems, revealing failures where memories answer correctly now but cannot handle future updates.
This paper proposes a Tri-Metric Router, a deterministic framework for adaptive routing among inference pipelines to address the Compression Paradox in long-context RAG on commodity GPUs, achieving zero OOM failures and improved performance.
This paper investigates the trade-off between context compression and citation attribution in retrieval-augmented generation, evaluating multiple compression methods and revealing significant gaps between answer quality and attribution accuracy.
REVA is a framework that mines historical attention traces to create reusable evidence views for context-efficient RAG serving, improving generation quality while reducing compression overhead and latency.
FlexComp is a method-agnostic framework that decouples compression ratios from training and deployment, allowing a single model to compress LLM context at any ratio using Matryoshka-style training and per-input budget selection for efficient inference.
LatentPress introduces a method to compress conversational and document context into continuous memory tokens, enabling frozen decoders to read directly without text reconstruction, achieving higher compression ratios and improved performance on long-context tasks.
This paper proposes telemetry-informed adaptive compression for edge-based RAG systems, showing experimental evidence that intermediate compression can reduce GPU energy by up to 53.2% with negligible quality loss.
A user shares their setup using Ling 3.0 Tiny as an auxiliary model for Hermes (Qwen 3.8 27B) to handle simple tasks like context compression and summarization, improving speed and efficiency without quality loss.
This article provides a detailed introduction to Pi's coding agent's context compression mechanism, analyzing the reasons for context expansion, compression trigger conditions, and the cost of cache invalidation, aiming to optimize AI performance in long conversations.
This paper investigates how epistemic stance (qualifiers, attributions) survives memory compression in AI agent memory systems. It finds that making the stance explicit as a labelled field improves retention significantly, while merely lengthening the text does not.
This empirical study investigates how recurrent context compression affects long-horizon agent behavior, showing that compression can weaken recent interaction influence and cause instability. The authors introduce TRACE, a verifier-guided framework that improves compression reliability and performance on AppWorld.
SeDeM is a selective decompression framework that stores long-context hidden states in a compact memory bank and decompresses only query-relevant blocks for decoder conditioning, improving QA accuracy and efficiency over compression baselines.
SALT is an open-source tool that compresses long documents into a fixed-size plain-text prompt for LLMs, using a keyword trie to avoid theme collapse and efficiently select informative sentences under a token budget.
This paper presents AutoWorldBuilder, a multi-agent LLM system for automated fictional worldbuilding that addresses context explosion, creative diversity, and quality assurance through hierarchical context compression, DAG-based scheduling, and iterative review, achieving 95% success rate and generating 56–103 self-consistent concepts per world.