cost-aware-inference

Tag

Cards List
#cost-aware-inference

Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems

arXiv cs.CL · 2026-08-17 Cached

This paper introduces InflationAgent, a routing system for agentic LLMs that measures token inflation, predicts task difficulty using CoT Branching Entropy, and optimizes model selection to maximize accuracy per cost, achieving higher accuracy with fewer tokens on benchmarks like GSM8K.

0 favorites 0 likes
#cost-aware-inference

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

arXiv cs.AI · 2026-07-29 Cached

This paper introduces value-router, a simulation study for cost-aware routing between cheap heuristics and expensive LLM calls in recommender systems, showing that value-weighted routing improves precision and handles seasonal demand surges with adaptive budgets.

0 favorites 0 likes
← Back to home

Submit Feedback