training-free

Tag

Cards List
#training-free

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

arXiv cs.AI ↗ · 2026-08-20 Cached

SESSE is a training-free framework that decomposes holistic LLM-as-judge evaluations into structured sub-questions, enabling better interpretability and diagnosis of label ambiguity while achieving competitive performance with fine-tuned models.

0 favorites 0 likes
#training-free

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

FlowEvo is a training-free framework that enables large language model agents to co-evolve reusable skills and workflows at inference time, achieving state-of-the-art accuracy and efficiency across benchmarks like ALFWorld, HumanEval, and GSM8K.

0 favorites 0 likes
#training-free

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Hugging Face Daily Papers ↗ · 2026-08-19 Cached

SparsePR is a training-free method that accelerates video transformers by using response-coupled partitioning and probe-fitted residual reconstruction to reduce attention error while maintaining generation quality.

0 favorites 0 likes
#training-free

Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers

arXiv cs.LG ↗ · 2026-08-18 Cached

The paper proposes dLLM-SetScore, a training-free framework using discrete masked-diffusion language models for multi-label text classification, achieving competitive performance with minimal validation data.

0 favorites 0 likes
#training-free

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

arXiv cs.LG ↗ · 2026-08-17 Cached

This paper proposes Activation-Prune-Merge (APM), a training-free framework for cross-scale fusion that improves smaller language models using larger donors without semantic alignment, achieving performance gains on multiple benchmarks.

0 favorites 0 likes
#training-free

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Hugging Face Daily Papers ↗ · 2026-08-13 Cached

Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during LLM agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.

0 favorites 0 likes
#training-free

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper identifies a privacy vulnerability in RL-trained multimodal large reasoning models, which can leak sensitive facts in their reasoning traces even after unlearning, and proposes LEMUR, a training-free inference-time framework that uses entropy dynamics to detect and suppress such leakage.

0 favorites 0 likes
#training-free

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

arXiv cs.CL ↗ · 2026-08-13 Cached

Proposes Ripple-Pivot Search, a training-free decoding method for diffusion large language models that proactively commits mid-entropy pivot positions to reduce uncertainty and accelerate parallel decoding, achieving 4-10x speedup.

0 favorites 0 likes
#training-free

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

arXiv cs.AI ↗ · 2026-08-13 Cached

This paper presents a unified global-to-local paradigm for video anomaly detection, introducing a training-free framework (GtS) and a tool-augmented agentic reasoning method with reinforcement learning, along with a new benchmark VAGU-T and metric JeAUG.

0 favorites 0 likes
#training-free

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

arXiv cs.AI ↗ · 2026-08-13 Cached

CORA-Diff is a training-free method that accelerates diffusion language model inference by using native confidence and persistence signals to accept residual positions early, skipping redundant dense denoising passes while preserving task quality.

0 favorites 0 likes
#training-free

Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models

arXiv cs.CL ↗ · 2026-08-11 Cached

Introduces Archer, a training-free KV caching method for diffusion language models that adaptively reuses cached hidden states to reduce recomputation while preserving rollback capabilities, achieving up to 2.95x speedup and improved generation quality.

0 favorites 0 likes
#training-free

SkillAligner: Treating Retrieved Skills as Adaptable Drafts at Execution Time

arXiv cs.LG ↗ · 2026-08-10 Cached

This paper introduces SkillAligner, a training-free framework that treats retrieved skills as adaptable drafts, jointly adapting them to task requirements, execution environments, and other skills to mitigate skill-execution misfit and improve agent performance.

0 favorites 0 likes
#training-free

KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty

arXiv cs.LG ↗ · 2026-08-10 Cached

KReF introduces a training-free retrieval framework for long-term time-series forecasting that constructs empirical predictive distributions from similar historical lookback-future pairs, achieving strong CRPS performance across multiple benchmarks.

0 favorites 0 likes
#training-free

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

Reddit r/LocalLLaMA ↗ · 2026-08-09 Cached

KLQ is a training-free LLM quantization method that allocates bits per direction based on measured KL divergence, outperforming existing training-free rotation-based methods on W4A4KV4-bit settings for models like Llama 3.2 1B and Qwen 2.5.

0 favorites 0 likes
#training-free

Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

arXiv cs.CL ↗ · 2026-08-07 Cached

The paper introduces Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol that extracts answer-space constraints before solving and checks intermediate/final results against them, improving math problem solving on competition benchmarks.

0 favorites 0 likes
#training-free

ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

arXiv cs.CL ↗ · 2026-08-07 Cached

ConWriter introduces a training-free framework for long-form story generation that maintains narrative consistency through scene-level incremental writing, symbolic state reasoning, and uncertainty-aware risk signals. Evaluated on ConStory-Bench across multiple models and lengths, it aims to prevent consistency errors from propagating in extended contexts.

0 favorites 0 likes
#training-free

Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper introduces elbow-based routing, a training-free inference-time method for MoE models that dynamically adjusts the number of active experts per token by detecting the elbow point in router probability distributions, achieving a 5.3% average latency reduction while maintaining accuracy.

0 favorites 0 likes
#training-free

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models

arXiv cs.LG ↗ · 2026-08-04 Cached

This paper proposes a training-free, uncertainty-aware inference framework for using large language models in operations research. The method uses short lookahead simulations and importance resampling to improve the coherence of mathematical formulations, outperforming standard baselines on OR benchmarks.

0 favorites 0 likes
#training-free

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Hugging Face Daily Papers ↗ · 2026-08-04 Cached

OmniPack proposes a training-free token compression framework for omni-modal LLMs, combining structural pre-LLM compression with task-relevant inner-LLM semantic refinement, achieving strong performance-efficiency trade-offs on multiple benchmarks.

0 favorites 0 likes
#training-free

GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization

Hugging Face Daily Papers ↗ · 2026-08-02 Cached

GaussianSelector is a training-free framework for interactive 3D object selection from sparse views using scribble guidance, operating directly on Gaussian primitives via graph-cut optimization.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback