Tag
This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.
The article discusses economic strategies for enterprises transitioning AI from experimentation to production, focusing on evaluating ownership versus consumption models based on workload demands and cost predictability.
The paper introduces RouteFM, a foundation model for LLM routing that learns reusable routing capabilities via episodic pretraining across heterogeneous environments, allowing a frozen router to adapt to new domains, modalities, and candidate pools through behavioral context alone. RouteFM outperforms the strongest baseline by 2.23 quality points on MMR-Bench with only eight observations per candidate, supporting a 'pretrain once, route anywhere' paradigm.
FlexRouter proposes a coverage-oriented LLM routing framework that uses Determinantal Point Processes to model complementarity among models, maximizing the probability that at least one selected model answers correctly while avoiding redundant selections and fixed budgets.
The paper proposes SaveRouter, a sparse-supervision LLM routing framework that selectively acquires query-model feedback and shares capability information across related queries, cutting supervision costs while maintaining competitive routing quality and reducing break-even deployment volume by 1.9-9.5x.
SeLMRoute introduces an LLM routing framework that separates candidate-independent semantic evidence extraction from performance learning, achieving 72.08% average accuracy on LLMRouterBench across 15 datasets and 20 candidate models, outperforming the strongest fixed candidate (69.23%) and enabling both performance-oriented and cost-aware routing decisions.
A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.
The paper challenges the assumption that rank portability implies feasibility portability in cross-device hardware evaluation, using benchmarks to show high rank correlation does not guarantee safe deployment decisions under joint constraints.
A company with a strong AI focus has implemented daily budget limits for state-of-the-art models, shifting to cheaper models for most tasks, indicating a trend in AI cost management.
The article questions whether smaller quantized models are becoming the preferred choice for local AI applications, emphasizing their balance of VRAM usage, performance, and capability like tool calling.
FedFIbOS proposes a Fisher importance-based method for optimal submodel selection in heterogeneous federated learning, theoretically grounded and achieving about 10% higher accuracy than state-of-the-art methods under non-IID settings.
Eve introduces automatic tool approvals and model selection using TypeSafe AI's Jev, enabling dynamic choice of AI models for different tasks via the AI SDK evaluation API.
The paper introduces inference networks, a graph-based framework for optimizing the use of multiple LLMs in inference, with optimal activation policies that minimize cost while meeting performance targets.
The author built an automated system to route LLM choices based on cost and performance data, aiming to optimize AI spending, with plans to open-source the tool.
A Python and TypeScript library for downloading LLM catalogs and building pipelines to filter and select models based on criteria like price, latency, and benchmarks.
The author shares personal experience showing how LLMs like Qwen 27B drastically reduce programming task time, offering rules of thumb for model selection.
SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.
The author is developing Agent-PGO, a tool that profiles AI agent executions to dynamically substitute cheaper models for less critical tasks while maintaining quality through evaluation benchmarks.
The paper proposes EEG-AS, an algorithm selection framework that enables instance-level selection among multiple EEG foundation models by reconstructing their behaviors, thereby improving neural decoding performance.
This paper introduces a recursion formula for efficiently computing the stochastic complexity of vectors with cluster structure using the Normalized Maximum Likelihood model, reducing time complexity from polynomial to linear.