Tag
ToolSearcher is a reinforcement learning framework designed to optimize tool selection for large language models in large-scale tool repositories, using techniques such as category-constrained discrimination and event-level search modeling to improve performance in iterative search and tool composition tasks.
Harness Router is a decision layer for AI coding agents that improves tool selection before execution using Jev for ambiguity and MCTS for multi-step consequences.
The article discusses design patterns for integrating Jev with LLMs in AI agents, specifically how agents handle tool selection and planning when only one tool schema is exposed at a time.
Toollery is a training-free candidate-compression framework that improves scalability and efficiency for LLM agent tool and skill selection through retrieval-based methods.
The author updates on a read-only analytics agent project using Gemini Flash Lite, discusses caching inefficiencies and bilingual routing challenges, and seeks advice on date handling.
The paper introduces Gavel, a method that elicits native skill routing from frozen LLMs via linear projections, enabling efficient tool selection without context overload and outperforming existing pipelines on benchmarks.
A developer describes building a read-only analytics agent for restaurant POS systems using Node.js, TypeScript, and Gemini Flash, seeking advice on tool selection, grounding techniques, and patterns for implementing write actions.
A study measured 16,893 sessions to analyze how AI coding agents like Claude Code, Codex, and Cursor select tools such as databases, finding consistent recommendations across varied contexts.
This article offers a practical framework for evaluating AI agent tools by focusing on key questions about job clarity, access needs, autonomy, maintenance, cost, and exit options, moving beyond feature comparisons.
The author observed that adding more tools to an AI agent decreased its accuracy in selecting the correct tool due to increased classification complexity, and found that using multiple smaller agents with narrower toolsets improved reliability.
This paper presents a case study on skill discovery and routing in a multimodal agent harness, showing that partial in-prompt exposure of skills can create lexical competition that hinders correct selection, linking small-scale in-context retrieval to large-scale approaches.
OpenMed demonstrated using Liquid AI's LFM2.5-VL-3B model to analyze a generated skin image, mapping regions and measuring dimensions locally on a Mac Studio for visual review.
Unused tools in an AI agent's toolset still consume tokens and add noise to tool selection, so agents should load only the tools required for the current task.
This paper introduces 'canary tools' as diagnostic probes to identify specific tool-selection reasoning failures in LLM agents, proposing a six-type taxonomy and evaluating eight models across 8,640 task runs to show capability-tier disparities and robustness.
Ratel is an open-source tool that reduces input tokens by 79% and improves tool selection accuracy for AI agents by loading only needed tools using a BM25 index, instead of all available tools.
DocOCR-Eval proposes an annotation-free framework that uses a correction and ranking strategy to evaluate and select OCR tools without ground truth labels, showing that aggregating multiple multimodal large language models improves alignment with human rankings.
Recent analysis reveals that retrieval-based tool selection for LLM agents caps out at recovering ~23% of failures, while readout-side interventions addressing attention biases recover 59-91% of failures, indicating that the real bottleneck is in the model's output processing rather than input filtering.
Liquid AI demonstrates using LFM2.5-ColBERT-350M as a filter to select only the five most relevant tools from 151 options, reducing latency and improving tool selection accuracy.
LFM2.5-ColBERT-350M is a model that reliably selects the most relevant tools from a set of 151, saving tokens and improving accuracy, ideal for agentic edge models.
This paper investigates over-privileged tool selection in LLM agents, introducing ToolPrivBench to evaluate and mitigate unnecessary use of high-privilege tools. It finds that safety alignment does not ensure least-privilege choices, and proposes a post-training defense that reduces excessive privilege use without sacrificing performance.