INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
Summary
INAR-VL proposes a lightweight routing system for edge-cloud vision-language inference that dynamically selects between edge and cloud models based on query complexity, achieving significant latency and energy reductions while preserving near-cloud accuracy.
Similar Articles
Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
Pro-Router introduces a token-aware progressive model routing method for efficient multimodal LLM inference, leveraging adaptive edge-cloud collaboration to improve throughput and reduce costs.
VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval
A new paper proposes VDAR-Router, a difficulty-aware retrieval-based routing framework for LLMs that adaptively selects models based on query difficulty, achieving better cost-performance trade-offs.
LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
LayerRoute introduces an action-conditioned routing interface for Vision-Language-Action policies that dynamically adapts access to VLM layer representations, improving robot manipulation performance with minimal additional parameters.
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
This paper introduces a carbon-aware routing framework for function-calling LLMs in edge-cloud systems, reducing operational carbon emissions by an average of 4× while maintaining cloud-level accuracy.
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
VisualThink-VLA introduces a visual intermediate reasoning framework for vision-language-action policies that preserves spatial precision and dramatically reduces latency compared to text-based reasoning, achieving sub-second inference and state-of-the-art success rates on robot manipulation benchmarks.