decoding

Tag

Cards List
#decoding

hLLM: Single Pass Decoding for Generative Reranking

arXiv cs.LG · 5d ago Cached

This paper introduces hLLM, a decoding strategy for generative reranking that uses the Hungarian algorithm to achieve single-pass decoding, resulting in a 64x speed-up while maintaining ranking quality.

0 favorites 0 likes
#decoding

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

arXiv cs.LG · 6d ago Cached

EEG-VID introduces task-guided latent predictive pretraining to improve EEG decoding under session and subject shifts, achieving above-chance target selection in assistive robotics scenarios.

0 favorites 0 likes
#decoding

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

arXiv cs.LG · 6d ago Cached

Introduces Faster Flash Decoding (FFD), a training-free hardware-algorithm co-design framework that accelerates long-context decoding in LLMs by exploiting attention sparsity, achieving up to 11.6x speedup and scaling to 256K context length.

0 favorites 0 likes
#decoding

Firefox 157 will include JPEG XL by default on all platforms

Hacker News Top · 2026-08-25 Cached

Firefox 157 intends to enable JPEG XL decoding by default on all platforms, using the jxl-rs Rust library for improved performance and feature parity with other image formats.

0 favorites 0 likes
#decoding

@BetaTomorrow: #DeepManifoldInterpretation Paper: On the Entropy Calibration of Language Models Author: Steven Cao, Gregory Valiant, a…

X AI KOLs Following · 2026-08-15 Cached

The paper 'On the Entropy Calibration of Language Models' interprets rising entropy as increasing diffusion of accessible pathways in autoregressive generation, proposing that scaling has limited benefits due to heavy-tailed data and suggesting a pathway-aware decoding alternative.

0 favorites 0 likes
#decoding

Why Tiny JPEGs Look Different in Chrome

Hacker News Top · 2026-08-12 Cached

The article explains why tiny JPEGs can look different in Chrome compared to other browsers, due to a JPEG decoding optimization that skips high-frequency DCT coefficients during heavy downscaling.

0 favorites 0 likes
#decoding

Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models

arXiv cs.CL · 2026-08-11 Cached

This paper investigates when classifier-free guidance (CFG) is actually necessary in masked diffusion language models, showing that guidance dependence is prompt-specific and can often be removed without losing constraint satisfaction, leading to a defined 'commitment horizon'.

0 favorites 0 likes
#decoding

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

arXiv cs.CL · 2026-08-10 Cached

This paper critiques existing benchmark contamination mitigation metrics and proposes SA-PPG (Stratified Aggregate of Per-question Probability Gaps) for more reliable evaluation, alongside RailCap, a decoding-time mitigation method that caps greedy fallback tokens to suppress memorization.

0 favorites 0 likes
#decoding

Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models

arXiv cs.CL · 2026-07-31 Cached

This paper introduces LATCH, a training-free candidate-aware early-exit framework for diffusion language models that separates when to stop from where to accelerate, achieving 9.3-17.8x speedups on short-answer tasks and 2-3.3x on long-reasoning tasks with minimal accuracy loss on LLaDA and Dream.

0 favorites 0 likes
#decoding

DeepLook: Deeper Thinking with Lookahead

arXiv cs.AI · 2026-07-28 Cached

DeepLook is a training-free framework that improves LLM reasoning by allocating compute at uncertainty bottlenecks, reducing token generation by 87.3% on average while improving accuracy on competition math benchmarks.

0 favorites 0 likes
#decoding

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

arXiv cs.AI · 2026-07-24 Cached

DecodeShare proposes a method to identify a low-dimensional subspace consistently shared across tasks in LLM decode-time hidden states and shows that disturbing this subspace degrades decision performance more than random or prefill-derived subspaces, with implications for activation steering.

0 favorites 0 likes
#decoding

Decoupling Task-Solving and Output Formatting in LLM Generation

arXiv cs.CL · 2026-07-13 Cached

Introduces Deco-G, a decoding framework that separates format adherence from problem-solving in LLMs, using a Format Estimation Module to ensure compliance without degrading reasoning. Achieves improved accuracy on mathematical reasoning, event extraction, and LLM-as-a-judge tasks.

0 favorites 0 likes
#decoding

Producing Structured Outputs from LLMs with Constrained Sampling

Reddit r/LocalLLaMA · 2026-07-10

Discusses methods for generating structured outputs from large language models using constrained sampling techniques.

0 favorites 0 likes
#decoding

Decoding the obfuscated bash script on a Uniqlo t-shirt

Hacker News Top · 2026-07-08 Cached

A blog post details the discovery and decoding of an obfuscated bash script printed on a Uniqlo t-shirt as part of Akamai's Peace for All campaign, which reveals a hidden Easter egg message when executed.

0 favorites 0 likes
#decoding

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding

arXiv cs.CL · 2026-07-07 Cached

TACG is a training-free decoder for diffusion language models that uses trajectory-aware signals to decide when to commit tokens, improving accuracy and efficiency on code and math benchmarks.

0 favorites 0 likes
#decoding

Prefill vs. decoding and local LLM ROI: is prefill underrated?

Reddit r/LocalLLaMA · 2026-07-06

An analysis comparing prefill and decoding phases in LLM inference, questioning whether prefill is underappreciated in terms of ROI for local LLM deployments.

0 favorites 0 likes
#decoding

@josh_bickett: Fable 5 shows models are insanely good at decoding shorthand. me: "ittiaa real product here for m rn. lsaeawl" Fable 5:…

X AI KOLs Following · 2026-07-06 Cached

Fable 5 demonstrates impressive ability to decode shorthand where users type the first letter of each word, accurately reconstructing the intended message.

0 favorites 0 likes
#decoding

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding

arXiv cs.LG · 2026-07-03 Cached

Set Diffusion introduces a new class of language models that interpolates between autoregressive and diffusion models by factorizing token generation over flexible-position, flexible-length token sets. This enables faster decoding and flexible token ordering, achieving better speed-quality tradeoffs on reasoning, summarization, and unconditional generation tasks.

0 favorites 0 likes
#decoding

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

arXiv cs.CL · 2026-07-02 Cached

Proposes Dual-Confidence Contrastive Decoding (DCCD), a training-free method for retrieval-augmented generation that handles intra-context conflicts in multi-document settings by combining document-level and token-level confidence signals, and introduces the DRQA benchmark for factual-conflict QA.

0 favorites 0 likes
#decoding

Breaking the Bird Barrier: Scientist Decodes Zebra Finch Language

Hacker News Top · 2026-06-30 Cached

Dr. Julie Elie of UC Berkeley won the 2026 Coller-Dolittle Prize for decoding the core vocabulary of zebra finches, identifying 11 distinct calls by combining observation, machine learning, and behavioral experiments. Her work marks a significant step in interspecies communication.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback