per-token-routing

Tag

Cards List
#per-token-routing

LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment

arXiv cs.LG · 2026-08-03 Cached

LARA is a method for efficient adaptation that adds low-rank corrections to a frozen model's residual stream instead of modifying weights, matching LoRA's performance while enabling composable behaviors and inference-time steering.

0 favorites 0 likes
← Back to home

Submit Feedback