pythia

Tag

Cards List
#pythia

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

arXiv cs.LG · 2026-08-05 Cached

This paper studies what transfers between transformer models of different sizes in the same family (Pythia), showing that representations align while weights don't, and that conversion works best via initialization rather than direct weight projection.

0 favorites 0 likes
#pythia

Accelerating Large Language Model Inference with Self-Supervised Early Exits

arXiv cs.CL · 2026-07-13 Cached

This paper introduces a self-supervised early exit method for LLMs, allowing computation to stop early at intermediate layers when confidence is high, thereby reducing inference cost. It also presents Dynamic Self-Speculative Decoding (DSSD) which achieves higher token acceptance than existing baselines.

0 favorites 0 likes
#pythia

Finetuned a Early 2023-Era Model on 2 Instruction Following Datasets and it Became Good

Reddit r/LocalLLaMA · 2026-06-12

A finetuned Pythia-6.9B model on two instruction-following datasets for 550 steps becomes capable in 13 languages, showing significant improvement over the base model.

0 favorites 0 likes
#pythia

A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting

arXiv cs.AI · 2026-06-03 Cached

This paper investigates whether direct activation transfer between language models can improve reasoning, using a linear translation layer from Pythia-160M to Pythia-410M. Despite achieving high representational alignment, the transferred activations do not improve multi-hop question answering, yielding a negative result.

0 favorites 0 likes
← Back to home

Submit Feedback