INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
Summary
INAR-VL proposes a lightweight routing system for edge-cloud vision-language inference that dynamically selects between edge and cloud models based on query complexity, achieving significant latency and energy reductions while preserving near-cloud accuracy.
Similar Articles
VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval
A new paper proposes VDAR-Router, a difficulty-aware retrieval-based routing framework for LLMs that adaptively selects models based on query difficulty, achieving better cost-performance trade-offs.
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
VisualThink-VLA introduces a visual intermediate reasoning framework for vision-language-action policies that preserves spatial precision and dramatically reduces latency compared to text-based reasoning, achieving sub-second inference and state-of-the-art success rates on robot manipulation benchmarks.
Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
This paper introduces CARGO, a training-free routing framework that uses the local LLM's own inference-time agreement across sampled responses to decide when to offload to a cloud model, enabling controllable collaboration ratios without additional training.
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Proposes Reroute, a training-free plug-in for vision-language models that replaces irreversible visual-token pruning with recoverable routing, allowing tokens to re-enter the pipeline later to improve grounding under aggressive token reduction while maintaining VQA performance.
@songhan_mit: explore VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
MIT researchers developed VLASH, a method enabling vision-language-action models to predict future robot states, doubling speed and reducing lag in tasks like pick-and-place and table tennis.