early-exits

Tag

Cards List
#early-exits

Accelerating Large Language Model Inference with Self-Supervised Early Exits

arXiv cs.CL · 2026-07-13 Cached

This paper introduces a self-supervised early exit method for LLMs, allowing computation to stop early at intermediate layers when confidence is high, thereby reducing inference cost. It also presents Dynamic Self-Speculative Decoding (DSSD) which achieves higher token acceptance than existing baselines.

0 favorites 0 likes
← Back to home

Submit Feedback