tool-prune: prune 50+ tool schemas down to candidates in 0.4ms (-92% prompt tokens, zero deps)

Reddit r/LocalLLaMA Tools

Summary

tool-prune is a tool that prunes tool schemas before calling local LLM models, reducing prompt tokens by 92% and avoiding hallucinations with no extra LLM round-trips.

When running local 7B/8B models with tools, 50+ schemas in context degrades attention and causes distractor hallucinations. Using in-band progressive search (tool_search) costs an entire extra LLM generation turn (+2s), and small models frequently forget to call the search tool or get stuck in discovery loops. tool-prune runs pre-flight in the client before the model is called: const topTools = await router.filter(userPrompt); 0 extra LLM round-trips 0.4ms offline via FWHT (pure JS & Python, 23µs with SIMD) -92% prompt tokens Direct dispatch option for deterministic queries in <1ms Zero external dependencies Playground: https://hemanth.github.io/tool-prune/ GitHub: https://github.com/hemanth/tool-prune
Original Article

Similar Articles

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Hugging Face Daily Papers

SWE-Pruner Pro leverages the coding agent's own internal representations to prune long code context, saving up to 39% of tokens while maintaining or improving task performance on multi-turn benchmarks.

Optimizing Korean-Centric LLMs via Token Pruning

arXiv cs.CL

This paper presents a systematic benchmark of token pruning—a compression technique that removes tokens and embeddings for irrelevant languages—applied to Korean-centric LLM tasks. The study evaluates popular multilingual models (Qwen3, Gemma-3, Llama-3, Aya) across different vocabulary configurations and finds that token pruning significantly improves generation stability and reduces memory footprint for domain-specific deployments.

An AI4AI Framework for Visual Token Pruning

Hugging Face Daily Papers

AutoPrune is a training-free framework that uses LLMs to automatically design visual-token pruning policies for multimodal LLMs via a domain-specific language and residual search, achieving high efficiency with minimal performance loss (99% performance retained while removing 94.4% of visual tokens).

Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning

arXiv cs.CL

This paper proposes STOP (SuperTOken for Pruning), a systematic framework for pruning inefficient reasoning paths early in parallel reasoning with Large Reasoning Models. The method achieves superior efficiency and effectiveness across models from 1.5B to 20B parameters, boosting GPT-OSS-20B accuracy on AIME25 from 84% to 90% under fixed compute budgets.