inference-time

Tag

Cards List
#inference-time

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

arXiv cs.AI · 6d ago Cached

This paper introduces a question-level audit framework distinguishing 'realized' from 'reachable' answers on LLM benchmarks, showing that aggregate score gains often come from producing already-reachable answers rather than expanding true capability, and that random layer routing matches structured search under matched budgets.

0 favorites 0 likes
#inference-time

Inference-Time Policy Alignment for Fair Reinforcement Learning

arXiv cs.LG · 2026-08-04 Cached

This paper proposes an inference-time policy shaping framework to steer pretrained reinforcement learning policies toward welfare-based fairness objectives without retraining, inspired by inference-time alignment in LLMs.

0 favorites 0 likes
#inference-time

Steering Instruction Hierarchies at Inference Time

arXiv cs.CL · 2026-07-30 Cached

Introduces V-Steer, a training-free inference-time method that edits cached value vectors to restore instruction hierarchy in language models, raising primary constraint accuracy from under 18% to 92% on controlled benchmarks with negligible overhead.

0 favorites 0 likes
#inference-time

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models

Hugging Face Daily Papers · 2026-07-20 Cached

Proposes DiFA, a training-free framework that reframes inference-time data prediction refinement as sequential state estimation using Kalman filtering, significantly improving generative fidelity on CIFAR-10 and ImageNet.

0 favorites 0 likes
#inference-time

Motion4Motion: Motion Transfer Across Subjects at Inference

Hugging Face Daily Papers · 2026-07-13 Cached

Motion4Motion is a training-free motion transfer framework that models motion flow from a video instead of relying on skeleton structures, enabling motion transfer across diverse species (e.g., humans to animals) at inference time.

0 favorites 0 likes
#inference-time

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

arXiv cs.LG · 2026-07-07 Cached

Proposes LARA, a framework for safe inference-time alignment that uses Lagrangian dualization to derive an augmented reward from separate reward and cost models, improving the helpfulness-harmlessness tradeoff without retraining.

0 favorites 0 likes
#inference-time

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

arXiv cs.AI · 2026-06-29 Cached

This paper reveals that hallucination in large vision-language models is caused by a dynamic structural misalignment where certain attention heads act as risky mediators, decoupling from visual evidence to lock onto language priors. The authors propose Fox, a training-free causal intervention framework that diagnoses and physically severs these pathological shortcuts, achieving state-of-the-art performance in faithful decoding.

0 favorites 0 likes
#inference-time

Graph-Based Phonetic Error Correction of Noisy ASR

arXiv cs.CL · 2026-06-25 Cached

Proposes G-SPIN, a lightweight framework that combines phonetic graph modeling with contextual language understanding for correcting ASR errors, using a GNN to generate phonetically plausible candidate tokens, an MLM for local scoring, and an LLM for final re-ranking, all operating at inference time.

0 favorites 0 likes
#inference-time

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

Hugging Face Daily Papers · 2026-06-23 Cached

This paper introduces a conversational voice agent system that uses a lightweight on-device 'Talker' model to start responding immediately, then incorporates knowledge from a frontier LLM 'Reasoner' as it becomes available, achieving 7-19x faster time-to-first-response while approaching frontier-level performance on a laptop.

0 favorites 0 likes
#inference-time

ARIADNE: Agnostic Routing for Inference-time Adapter DyNamic sElection

arXiv cs.AI · 2026-06-18 Cached

Proposes ARIADNE, a training-free, adapter-agnostic routing framework that selects the optimal PEFT adapter at inference time by measuring input proximity to adapter-specific centroids in embedding space, recovering 97.44% of upper-bound performance on 23 tasks.

0 favorites 0 likes
#inference-time

From Consumption to Reflection: Designing Human-AI Relations for Stable Reasoning

arXiv cs.AI · 2026-06-11 Cached

This paper introduces Relational Reflective Intelligence (RRI), an inference-time governance layer that uses auditable reasoning loops to stabilize human-AI reasoning, addressing cognitive vulnerabilities shared by humans and LLMs.

0 favorites 0 likes
#inference-time

Calibrating Overconfidence Without Sacrificing Confidence: Probe-Conditioned Head Intervention for LLMs

arXiv cs.LG · 2026-06-10 Cached

The paper introduces Probe-Conditioned Head Intervention (PCHI), an inference-time method for LLMs that selectively reduces overconfidence on wrong answers without significantly reducing confidence on correct ones, by conditionally rescaling attention head outputs when the model is likely wrong but confident.

0 favorites 0 likes
#inference-time

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

Hugging Face Daily Papers · 2026-06-10 Cached

Evoflux uses evolutionary search at inference time to repair failed tool workflows for compact language models, boosting execution feasibility significantly over fine-tuning methods.

0 favorites 0 likes
#inference-time

I built an inference-time epistemic framework that extends coherent LLM threads to 325k–1M tokens. Here's how it works.

Reddit r/artificial · 2026-06-05

An independent researcher introduces Epistemic Lattice Tethering (ELT), an inference-time scaffolding framework that extends coherent LLM threads to 325k–1M tokens by applying epistemic and ontological governance.

0 favorites 0 likes
#inference-time

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

arXiv cs.AI · 2026-06-04 Cached

This paper demonstrates that LLM safety vulnerabilities extend beyond 'shallow safety' (first-token alignment) to any point during generation, showing that short token injections mid-sequence can redirect models toward harmful outputs. The authors propose training on generation trajectories with simulated mid-sequence perturbations to improve robustness.

0 favorites 0 likes
#inference-time

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

arXiv cs.AI · 2026-06-04 Cached

This paper presents the 'Digital Apprentice,' a framework for scalable and safe agentic AI in which autonomy is earned incrementally through observational learning, human authorization, and continuous alignment correction. It introduces ADAPT, an inference-time control plane that operationalizes graduated autonomy tiers and converts human corrections into reusable preference data.

0 favorites 0 likes
#inference-time

Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization

arXiv cs.CL · 2026-06-03 Cached

This paper proposes Dynamic Contextual Orthogonalization (DCO), an inference-time method that reduces hallucinations in large language models by aligning attention head outputs with the context manifold, achieving superior faithfulness on benchmarks with Llama-3 models.

0 favorites 0 likes
#inference-time

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

arXiv cs.AI · 2026-06-02 Cached

Introduces Latent Reward Steering (Lrs), an adaptive inference-time framework that uses sparse autoencoder latent states and a learned reward model to implicitly promote cognitive behaviors like verification and backtracking in reasoning LLMs, improving performance across multiple models and benchmarks.

0 favorites 0 likes
#inference-time

TIGER: Traceable Inference with Graph-Based Evidence Routing for Mitigating Hallucinations in Multimodal Generation

arXiv cs.AI · 2026-06-02 Cached

TIGER is an inference-time framework that mitigates hallucinations in multimodal generation by extracting observation and claim graphs and assigning risk scores to repair unsupported facts. It reduces unsupported content across image-to-text, image+text-to-text, audio-to-text, and video-to-text tasks.

0 favorites 0 likes
#inference-time

Scalable Inference-Time Annealing with Surrogate Likelihood Estimators

Hugging Face Daily Papers · 2026-06-01

SITA (Scalable Inference-Time Annealing) introduces a method for efficiently sampling molecular Boltzmann distributions by retraining flow-based models along a temperature ladder using energy-based surrogate likelihoods, avoiding costly divergence computations. The approach achieves state-of-the-art performance on Alanine Dipeptide and Tripeptide benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback