decoding-time

Tag

Cards List
#decoding-time

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

arXiv cs.LG · 2026-08-13 Cached

This paper introduces Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates supervised fine-tuning effects via logit-space transport, achieving competitive personalization performance with less than 7% of the computation.

0 favorites 0 likes
#decoding-time

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

arXiv cs.LG · 2026-06-30 Cached

Hard-KV introduces a Cascade Cache hierarchy and Logits Calibration mechanism to resolve the static-dynamic mismatch in head-adaptive KV cache compression, achieving up to 2x throughput improvement in long-context LLM inference.

0 favorites 0 likes
← Back to home

Submit Feedback