FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
Summary
The paper introduces FocuSFT, a bilevel optimization framework that enhances long-context language model performance by addressing attention dilution through parametric memory. It demonstrates significant improvements in accuracy and context engagement on benchmarks like BABILong and RULER.
View Cached Full Text
Cached at: 05/13/26, 08:12 AM
Paper page - FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
Source: https://huggingface.co/papers/2605.09932
Abstract
Training framework FocuSFT improves long-context language model performance by addressing attention allocation issues through bilevel optimization with parametric memory that focuses attention on semantically relevant content.
Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains limited. We trace this gap to howattention budgetis spent duringsupervised fine-tuning(SFT) on long sequences:positional biasesandattention sinkscause the model to allocate most of its attention to positionally privileged tokens rather than semantically relevant content. This training-timeattention dilution(the starvation of content tokens in the attention distribution) weakens the gradient signal, limiting the model’s ability to learn robust long-context capabilities. We introduce FocuSFT, abilevel optimizationframework that addresses this problem at training time. An inner loop adapts lightweightfast-weight parameterson the training context to form aparametric memorythat concentrates attention on relevant content, and the outer loop performs SFT conditioned on this sharpened representation. Both loops apply bidirectional attention over context tokens while preservingcausal maskingfor responses, reducing the causal asymmetry that gives rise toattention sinksand aligning inner-outer behavior. On BABILong, FocuSFT improves accuracy by up to +14pp across 4K--32K context lengths; on RULER, it raises CWE aggregation from 72.9\% to 81.1\% at 16K; and on GPQA with agentic tool use, it yields a 24\% relative gain in pass@1. Attention analysis shows that FocuSFT reducesattention sink massby 529times and triplescontext engagementduring training. Code: https://github.com/JarvisPei/FocuSFT
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2605\.09932
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.09932 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.09932 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.09932 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models?
LiFT is a longitudinal instruction fine-tuning framework that unifies diverse temporal NLP tasks under a shared instruction schema with curriculum-based training. Evaluated across OLMo, LLaMA, and Qwen models, LiFT consistently outperforms base-model in-context learning, especially on out-of-distribution data and rare change events.
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
This paper presents a novel method for fine-tuning transformer language models with sparse attention to enable efficient long-context inference, often outperforming models trained with exact attention, and introduces an efficient implementation and a new open-source library.
Task Specialization Fine-Tuning for Contextual Reinforcement Learning
The paper introduces Task Specialization Fine-Tuning (TSFT) for contextual reinforcement learning, a framework that uses a simple parametric model and integer linear programming to allocate fine-tuning budgets efficiently, significantly outperforming baselines in task coverage.
Learnability-Informed Fine-Tuning of Diffusion Language Models
We propose LIFT, a learnability-informed fine-tuning algorithm for diffusion language models that aligns training with token difficulty and time step, achieving substantial gains on reasoning benchmarks.
Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
This paper investigates the dissociation between attention-level proxies and behavioral performance in large language models under fine-tuning, revealing that attention sensitivity alone is unreliable for diagnosing in-context learning capabilities.