Tag
FlexRouter 提出一种显式建模模型互补性的 LLM 路由框架,将路由建模为基于覆盖的子集选择问题,并用 Determinantal Point Processes 对模型能力与冗余进行联合建模,同时通过边缘化失败集合的目标函数直接优化答案覆盖率。在 RouterEval 基准上,该方法在域内与域外任务上以更低冗余实现了更高的覆盖率,且推理成本灵活可控。
This paper introduces a taxonomy of 9 model selection algorithms for multi-LLM collaboration, showing that capability-aware selection strategies outperform random or heuristic team assembly by up to 36.1% across math, coding, QA, and reasoning tasks.
The article discusses economic strategies for enterprises transitioning AI from experimentation to production, focusing on evaluating ownership versus consumption models based on workload demands and cost predictability.
The paper introduces RouteFM, a foundation model for LLM routing that learns reusable routing capabilities via episodic pretraining across heterogeneous environments, allowing a frozen router to adapt to new domains, modalities, and candidate pools through behavioral context alone. RouteFM outperforms the strongest baseline by 2.23 quality points on MMR-Bench with only eight observations per candidate, supporting a 'pretrain once, route anywhere' paradigm.
FlexRouter proposes a coverage-oriented LLM routing framework that uses Determinantal Point Processes to model complementarity among models, maximizing the probability that at least one selected model answers correctly while avoiding redundant selections and fixed budgets.
The paper proposes SaveRouter, a sparse-supervision LLM routing framework that selectively acquires query-model feedback and shares capability information across related queries, cutting supervision costs while maintaining competitive routing quality and reducing break-even deployment volume by 1.9-9.5x.
SeLMRoute introduces an LLM routing framework that separates candidate-independent semantic evidence extraction from performance learning, achieving 72.08% average accuracy on LLMRouterBench across 15 datasets and 20 candidate models, outperforming the strongest fixed candidate (69.23%) and enabling both performance-oriented and cost-aware routing decisions.
A study revealed that an LLM router trained to select between models learned task recognition instead of difficulty, causing poor generalization on held-out tasks, but deferral based on the cheap model's output yielded better performance.
The paper challenges the assumption that rank portability implies feasibility portability in cross-device hardware evaluation, using benchmarks to show high rank correlation does not guarantee safe deployment decisions under joint constraints.
A company with a strong AI focus has implemented daily budget limits for state-of-the-art models, shifting to cheaper models for most tasks, indicating a trend in AI cost management.
The article questions whether smaller quantized models are becoming the preferred choice for local AI applications, emphasizing their balance of VRAM usage, performance, and capability like tool calling.
FedFIbOS proposes a Fisher importance-based method for optimal submodel selection in heterogeneous federated learning, theoretically grounded and achieving about 10% higher accuracy than state-of-the-art methods under non-IID settings.
Eve introduces automatic tool approvals and model selection using TypeSafe AI's Jev, enabling dynamic choice of AI models for different tasks via the AI SDK evaluation API.
The paper introduces inference networks, a graph-based framework for optimizing the use of multiple LLMs in inference, with optimal activation policies that minimize cost while meeting performance targets.
The author built an automated system to route LLM choices based on cost and performance data, aiming to optimize AI spending, with plans to open-source the tool.
A Python and TypeScript library for downloading LLM catalogs and building pipelines to filter and select models based on criteria like price, latency, and benchmarks.
The author shares personal experience showing how LLMs like Qwen 27B drastically reduce programming task time, offering rules of thumb for model selection.
SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.
The author is developing Agent-PGO, a tool that profiles AI agent executions to dynamically substitute cheaper models for less critical tasks while maintaining quality through evaluation benchmarks.
The paper proposes EEG-AS, an algorithm selection framework that enables instance-level selection among multiple EEG foundation models by reconstructing their behaviors, thereby improving neural decoding performance.