context-compression

Tag

Cards List
#context-compression

MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression

arXiv cs.CL ↗ · 5d ago Cached

This paper introduces MORSE, a compression-aware method for evidence-preserving context ordering in large language models that improves evidence retention and downstream QA performance.

0 favorites 0 likes
#context-compression

Memory Control Signals Emerge Before Action in Long Horizon Agents

arXiv cs.AI ↗ · 5d ago Cached

This paper studies hidden states in long horizon language model agents, revealing that memory compression and recall needs are encoded before actions. It proposes the PaMER framework to reduce context consumption while maintaining task performance through state-guided compression and evidence retrieval.

0 favorites 0 likes
#context-compression

Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

arXiv cs.CL ↗ · 6d ago Cached

The paper proposes a Context-to-Answer-Aligned Memory Compression (CMC) framework that compresses long input contexts into compact memory embeddings to reduce LLM inference costs without modifying decoder weights, achieving significant performance and efficiency gains.

0 favorites 0 likes
#context-compression

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

arXiv cs.CL ↗ · 2026-09-22 Cached

This paper empirically analyzes cost savings in context-compression gateways for multi-turn coding agents, revealing that tool-schema filtering provides fixed token savings, while content compression saves quadratically but can be offset by recalls, offering actionable insights for cost optimization.

0 favorites 0 likes
#context-compression

@FinanceYF5: It's 2026 already, so why are Agents' context compressions still relying on summarization prompts? Jev found a more dir…

X AI KOLs Following ↗ · 2026-09-20

Jev introduced on-the-fly compression for AI agents, which scores each tool call to retain important content and delete irrelevant data directly, eliminating the need for model-based summarization and making context compression faster and more lightweight.

0 favorites 0 likes
#context-compression

@yibie: Jev Ecosystem 72 Hours: From 46 to 160 Jev turned "judgment" into a primitive that's cheap enough to call directly in c…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

In 72 hours, the Jev ecosystem expanded from 46 to 160 projects, focusing on context compression, platform integrations, and new domains like financial trading, with debates on its novelty and implementation.

0 favorites 0 likes
#context-compression

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper introduces a paired-history audit method to evaluate update sufficiency in context compression for AI memory systems, revealing failures where memories answer correctly now but cannot handle future updates.

0 favorites 0 likes
#context-compression

Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs

arXiv cs.LG ↗ · 2026-09-17 Cached

This paper proposes a Tri-Metric Router, a deterministic framework for adaptive routing among inference pipelines to address the Compression Paradox in long-context RAG on commodity GPUs, achieving zero OOM failures and improved performance.

0 favorites 0 likes
#context-compression

The Attribution-Compression Frontier in Retrieval-Augmented Generation

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper investigates the trade-off between context compression and citation attribution in retrieval-augmented generation, evaluating multiple compression methods and revealing significant gaps between answer quality and attribution accuracy.

0 favorites 0 likes
#context-compression

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

arXiv cs.LG ↗ · 2026-09-11 Cached

REVA is a framework that mines historical attention traces to create reusable evidence views for context-efficient RAG serving, improving generation quality while reducing compression overhead and latency.

0 favorites 0 likes
#context-compression

FlexComp: One Model for Every Ratio in Context Compression

arXiv cs.CL ↗ · 2026-09-11 Cached

FlexComp is a method-agnostic framework that decouples compression ratios from training and deployment, allowing a single model to compress LLM context at any ratio using Matryoshka-style training and per-input budget selection for efficient inference.

0 favorites 0 likes
#context-compression

LatentPress: Context Compression Beyond Text and Vision

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

LatentPress introduces a method to compress conversational and document context into continuous memory tokens, enabling frozen decoders to read directly without text reconstruction, achieving higher compression ratios and improved performance on long-context tasks.

0 favorites 0 likes
#context-compression

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

arXiv cs.AI ↗ · 2026-08-21 Cached

This paper proposes telemetry-informed adaptive compression for edge-based RAG systems, showing experimental evidence that intermediate compression can reduce GPU energy by up to 53.2% with negligible quality loss.

0 favorites 0 likes
#context-compression

Ling 3.0 Tiny makes an amazing auxillery model for Hermes (Qwen 3.8 27B as the primary model)

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user shares their setup using Ling 3.0 Tiny as an auxiliary model for Hermes (Qwen 3.8 27B) to handle simple tasks like context compression and summarization, improving speed and efficiency without quality loss.

0 favorites 0 likes
#context-compression

@LanLance24: Pi has released a blog post detailing their compression mechanism design, with in-depth analysis of how context expands, how compression is triggered, and what costs caching incurs.

X AI KOLs Timeline ↗ · 2026-08-17 Cached

This article provides a detailed introduction to Pi's coding agent's context compression mechanism, analyzing the reasons for context expansion, compression trigger conditions, and the cost of cache invalidation, aiming to optimize AI performance in long conversations.

0 favorites 0 likes
#context-compression

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper investigates how epistemic stance (qualifiers, attributions) survives memory compression in AI agent memory systems. It finds that making the stance explicit as a labelled field improves retention significantly, while merely lengthening the text does not.

0 favorites 0 likes
#context-compression

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

arXiv cs.LG ↗ · 2026-08-10 Cached

This empirical study investigates how recurrent context compression affects long-horizon agent behavior, showing that compression can weaken recent interaction influence and cause instability. The authors introduce TRACE, a verifier-guided framework that improves compression reliability and performance on AppWorld.

0 favorites 0 likes
#context-compression

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

arXiv cs.CL ↗ · 2026-08-04 Cached

SeDeM is a selective decompression framework that stores long-context hidden states in a compact memory bank and decompresses only query-relevant blocks for decoder conditioning, improving QA accuracy and efficiency over compression baselines.

0 favorites 0 likes
#context-compression

looking for contributors - trie based memory efficient LLM runner

Reddit r/artificial ↗ · 2026-07-21 Cached

SALT is an open-source tool that compresses long documents into a fixed-size plain-text prompt for LLMs, using a keyword trie to avoid theme collapse and efficiently select informative sentences under a token budget.

0 favorites 0 likes
#context-compression

Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review

arXiv cs.AI ↗ · 2026-07-13 Cached

This paper presents AutoWorldBuilder, a multi-agent LLM system for automated fictional worldbuilding that addresses context explosion, creative diversity, and quality assurance through hierarchical context compression, DAG-based scheduling, and iterative review, achieving 95% success rate and generating 56–103 self-consistent concepts per world.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback