efficient-decoding

Tag

Cards List
#efficient-decoding

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Hugging Face Daily Papers · 2026-08-28 Cached

Parallel Tube Decoding enables efficient simultaneous spatial and temporal video grounding by eliminating autoregressive dependencies, reducing latency and improving accuracy over standard methods.

0 favorites 0 likes
#efficient-decoding

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

arXiv cs.AI · 2026-08-13 Cached

CORA-Diff is a training-free method that accelerates diffusion language model inference by using native confidence and persistence signals to accept residual positions early, skipping redundant dense denoising passes while preserving task quality.

0 favorites 0 likes
#efficient-decoding

Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models

arXiv cs.CL · 2026-08-11 Cached

Introduces Archer, a training-free KV caching method for diffusion language models that adaptively reuses cached hidden states to reduce recomputation while preserving rollback capabilities, achieving up to 2.95x speedup and improved generation quality.

0 favorites 0 likes
#efficient-decoding

Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding

arXiv cs.AI · 2026-08-03 Cached

This paper proposes GenCDSR, a generative framework for cross-domain sequential recommendation with hybrid tokenization and serial-parallel decoding, achieving improved accuracy and significantly reduced inference latency compared to state-of-the-art baselines.

0 favorites 0 likes
#efficient-decoding

MUGEN: A Unified Framework for Efficient Motion Understanding and Generation

arXiv cs.LG · 2026-07-31 Cached

MUGEN introduces a unified motion-language framework that avoids discrete codebooks and iterative decoding, using a single adaptive-length autoencoder with continuous latent slots and one-shot generation to achieve efficient, high-quality text-to-motion and motion-to-text performance across HumanML3D and SnapMoGen benchmarks.

0 favorites 0 likes
#efficient-decoding

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

arXiv cs.CL · 2026-07-30 Cached

DIRECT is a framework for sequence labeling using large language models that improves domain alignment through Direct Preference Optimization (DPO) after supervised fine-tuning and increases inference efficiency via controlled decoding with template-filling and KV cache reuse.

0 favorites 0 likes
#efficient-decoding

nvidia/Nemotron-Labs-Diffusion-14B

Hugging Face Models Trending · 2026-04-22 Cached

NVIDIA releases Nemotron-Labs-Diffusion, a family of tri-mode language models (3B, 8B, 14B) supporting AR, diffusion, and self-speculation decoding, achieving 2.7x-4x speed-ups over standard AR decoding.

0 favorites 0 likes
← Back to home

Submit Feedback