transformer-pruning

Tag

Cards List
#transformer-pruning

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

arXiv cs.LG · 5d ago Cached

The paper proposes REP-LIE, a resource-efficient pruning method for Transformer models that uses low-rank importance estimation from LoRA gradients to enable pruning during finetuning, with competitive performance on LLaMA-7B and Mistral-7B.

0 favorites 0 likes
#transformer-pruning

CausalGate: Causal Importance Distillation for Transformer Module Pruning

arXiv cs.LG · 2026-07-28 Cached

CausalGate introduces a method that uses causal interventions to measure the importance of transformer sub-layers and distills this into static scalar gates for efficient inference without runtime overhead, outperforming existing pruning and routing methods.

0 favorites 0 likes
← Back to home

Submit Feedback