Tag
TRACE is a training-free framework that optimizes GUI agent efficiency by ranking visual evidence based on utility and diversity, reducing latency and memory usage through adaptive token management and KV contraction.
CoVeR is a training-free spatial token selector that improves 3D reasoning in Vision-Language Models by enforcing exact token budgets and full scene coverage, outperforming prior state-of-the-art methods.
AdaptiveSpec is a training-free per-step speculative decoding method that adaptively adjusts token verification and draft tree shape to enhance LLM inference throughput, improving performance by up to 56% while maintaining high accuracy across benchmarks.
The article highlights the development of TwinDEX, a co-designed robotic interface that enables robots to learn complex dexterity manipulation tasks without on-robot training, signaling a major advancement in robotics.
GeoSPRINT is a training-free framework that detects geometrically redundant steps in diffusion trajectories using hyperplanarity tests to optimize sampling schedules, improving inference efficiency on models like Stable Diffusion v1.5 without retraining.
This paper proposes CAT-OV and CAT-OT, two lightweight, training-free algorithms that adapt step-sizes in Flow Matching sampling based on curvature, improving image quality and reducing generation steps by up to 40%.
HeadWiseKV is a training-free framework that compresses KV caches in hybrid long-context language models, reducing GPU memory usage and extending context lengths while maintaining quality.
IDEEA proposes a training-free, input-dependent steering method for large language models that clusters activations and uses optimal matching to improve truthfulness in TruthfulQA by up to 23.5% over baselines.
GAPS introduces dimension-level gating for conditional activation steering in language models, combining static and dynamic gates to selectively intervene and improve behavior-capability trade-off, with significant gains in toxicity mitigation and concept removal tasks.
EditVid is a unified training-free video editing framework that supports instruction-guided and subject-guided edits using sparse causal memory, token injection, and soft latent blending, achieving high fidelity and outperforming baseline methods in benchmarks.
This paper introduces ClusterAttention, a training-free method to speed up bidirectional attention in transformers by using recursive clustering for block-sparse attention, achieving 2-6x speedups on tabular data and 1.8x on video generation while maintaining high accuracy.
GRAS is a training-free method for reward alignment in discrete diffusion models that reduces variance in guided proposals and adaptively selects particles, achieving state-of-the-art results on DNA and protein design tasks.
The paper proposes a survival-guided length predictor for diffusion language models that speeds up inference by up to 7x on reasoning and code-generation benchmarks without sacrificing accuracy.
This paper introduces PAYN, a plug-and-play token compression method for MLLM-based referring expression segmentation that relies solely on position information, outperforming existing techniques by preserving spatial relational consistency.
ExFold is a unified training-free framework that accelerates MoE model inference by folding excluded expert contributions into retained experts, achieving up to 1.41× speedup while maintaining high quality.
RECAP-Forcing improves long video generation by indexing memory based on appearance novelty instead of recency, preserving key-value caches for new content to maintain consistency and quality without additional training.
ChorusTIC is a training-free foundation model for multivariate time series classification that uses in-context learning to handle heterogeneous channel configurations without target-task updates, demonstrating strong performance on standard benchmarks.
HIRA is a training-free, on-premises retrieval-augmented cascade system for document classification in regulated industries that uses human feedback and on-premises LLMs to improve accuracy while reducing model retraining and LLM calls.
The paper proposes COEC, a training-free compensation framework for structured pruning of large language models that applies orthogonal rotations and calibration to reduce output error and improve accuracy after column removal.
EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction using structured evidence packages, achieving state-of-the-art performance without training.