Tag
A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.
Jev is a tool that acts as a layer between deterministic software and LLMs, optimized for semantic judgment tasks where possible answers are predefined but input interpretation requires fuzzy logic, such as in support-ticket routing.
This article is an engineering note that re-examines the design of coding agents from first principles, questions the impact of KV cache on current architectures, and proposes new methods for context management and decision-making.
This paper shows that complexity-based routing in large language models biases against non-standard English registers like African American English, due to length-based signals, leading to lower capacity allocation and compounded bias.
The author built an automated system to route LLM choices based on cost and performance data, aiming to optimize AI spending, with plans to open-source the tool.
The article discusses issues with OpenRouter's automatic provider routing, such as inconsistent model behavior across different backends, and suggests using options like provider.only for better control.
This paper proposes algorithms for adaptively routing prompts to LLM experts in an online setting with limited feedback, formulated as a bandit problem to minimize regret and maximize response quality.
This paper formulates the adaptive routing of prompts to large language model experts as a contextual bandit problem with limited feedback, proposing algorithms that achieve sublinear regret and demonstrate efficient learning of high-quality routing strategies.
SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.
IQ Routing is a trajectory-aware LLM routing system designed to reduce the cost of AI agents by optimizing task routing based on trajectory data.
Ramp has launched Router, an AI model routing service that enables companies to access and switch between various large language models via an API, with features for cost optimization and benchmark-based routing.
This paper introduces InflationAgent, a routing system for agentic LLMs that measures token inflation, predicts task difficulty using CoT Branching Entropy, and optimizes model selection to maximize accuracy per cost, achieving higher accuracy with fewer tokens on benchmarks like GSM8K.
A new paper on arXiv introduces an open-source library called LLMRouter with over 16 router implementations and a benchmark xRouteBench, demonstrating that learned routers can outperform fixed-model baselines by 14.6%.
LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.
Describes Iris Ai, a system that routes queries across 8 specialized LLMs on consumer hardware, achieving large-model performance with low memory by keeping only one model active at a time and dynamic model swapping.
The article argues that model routing isn't the real problem to solve—token efficiency is. It advocates for a holistic closed-loop approach combining cheaper defaults, preference-aware routing, and better caching (e.g., Coinbase's 5%→60% cache hit improvement) to maximize useful intelligence per dollar.
Manifest explains why it deprecated its LLM router, arguing that model routing introduces unpredictability, breaks behavior consistency, and that prompt complexity cannot be inferred from the prompt alone, making caching and deliberate model selection more effective for most use cases.
A paper proposing a latency-aware LLM query router that jointly optimizes latency, accuracy, and cost using a lightweight latency estimator, achieving up to 40% improvement in accuracy–cost utility while maintaining comparable latencies.
A new paper proposes VDAR-Router, a difficulty-aware retrieval-based routing framework for LLMs that adaptively selects models based on query difficulty, achieving better cost-performance trade-offs.
Ramp Router uses EWMA for failure rates and Thompson sampling for latency to select the cheapest LLM model and service tier meeting deadlines, achieving 30% cost savings without performance loss.