adaptive-inference

Tag

Cards List
#adaptive-inference

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

arXiv cs.LG · 2026-08-18 Cached

This paper introduces a paired exact-reset evaluation protocol to determine when switching from a medium to a full world-model predictor improves task-specific decision loss, with experiments in PushT and PyBullet showing incremental routing benefits.

0 favorites 0 likes
#adaptive-inference

Looped Language Models Improve Compositional Tool Calling

Hugging Face Daily Papers · 2026-08-17 Cached

Looped language models enhance compositional tool calling by leveraging recurrent computation, improving accuracy on multi-step tasks while adaptive inference optimizes the balance between performance and compute cost. The study suggests these models are promising for reliable agentic systems.

0 favorites 0 likes
#adaptive-inference

When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions

arXiv cs.LG · 2026-05-25 Cached

This paper investigates when chain-of-thought reasoning is beneficial for LLMs, showing that early-stage entropy dynamics reliably indicate reasoning utility, and introduces EDRM, a lightweight, training-free framework that adaptively selects inference strategies to achieve significant token savings while maintaining or improving accuracy.

0 favorites 0 likes
#adaptive-inference

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Hugging Face Daily Papers · 2026-04-16 Cached

AVR is an adaptive visual reasoning framework that dynamically selects optimal reasoning formats to reduce token usage by 50-90% while maintaining accuracy in visual reasoning tasks. The method addresses reasoning path redundancy by decomposing visual reasoning into three cognitive functions and using FS-GRPO training to encourage efficient format selection.

0 favorites 0 likes
#adaptive-inference

Fast and Faithful: Real-Time Verification for Long-Document Retrieval-Augmented Generation Systems

Papers with Code Trending · 2026-03-04 Cached

This paper presents a real-time verification system for retrieval-augmented generation that processes long documents up to 32K tokens, using adaptive inference strategies to balance latency and verification coverage. It provides practical guidance for building reliable RAG systems.

0 favorites 0 likes
← Back to home

Submit Feedback