system-efficiency

Tag

Cards List
#system-efficiency

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

arXiv cs.LG · 2026-06-30 Cached

Hard-KV introduces a Cascade Cache hierarchy and Logits Calibration mechanism to resolve the static-dynamic mismatch in head-adaptive KV cache compression, achieving up to 2x throughput improvement in long-context LLM inference.

0 favorites 0 likes
#system-efficiency

AESOP: Adversarial Execution-path Selection to Overload Deep Learning Pipelines

arXiv cs.LG · 2026-05-13 Cached

This paper introduces AESOP, a framework for adversarial execution-path selection that significantly inflates FLOPs and latency in deep learning inference pipelines, revealing new efficiency-based vulnerabilities.

0 favorites 0 likes
← Back to home

Submit Feedback