Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories
Summary
The paper introduces a training-free, deterministic pipeline for lexical prompt compression in large language models, featuring empirical Pareto analysis across eleven task categories.
View Cached Full Text
Cached at: 09/15/26, 08:31 AM
# Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories Source: [https://arxiv.org/abs/2609.13154](https://arxiv.org/abs/2609.13154) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
This paper audits extractive prompt compressors across ten languages, revealing that English-trained models exhibit significant performance gaps on non-English text at high compression rates and proposes a translate-then-compress pipeline as a more effective alternative.
Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality
This paper presents a large-scale analysis of prompt lexical sensitivity in large language models, revealing a scaling law for prompt performance stability and introducing an automated Prompt-Refining Agent that reduces performance variance in tasks like code generation.
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
This paper introduces MedTPE, a method for efficient, lossless prompt compression of electronic health records for large language models, significantly reducing token length and inference latency in clinical prediction tasks.
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
OmniPack proposes a training-free token compression framework for omni-modal LLMs, combining structural pre-LLM compression with task-relevant inner-LLM semantic refinement, achieving strong performance-efficiency trade-offs on multiple benchmarks.
Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
This paper proposes using linguistic rules alone as prompt compressors without LM forward passes, achieving performance similar to advanced strategies under light-to-moderate compression.