Form Over Content In Gradient-Based Data Attribution Methods
Summary
This paper resolves a debate on gradient-based data attribution methods for large language models by demonstrating that they primarily track answer format rather than task semantics, challenging their reliability in targeted instruction tuning.
View Cached Full Text
Cached at: 09/18/26, 08:57 AM
# Form Over Content In Gradient-Based Data Attribution Methods
Source: [https://arxiv.org/html/2609.19589](https://arxiv.org/html/2609.19589)
Seokwon JungSohyung KimSeong Joon OhAlice OhAffiliation:KAISTAffiliation:\{sunwoo\.kim, tjrdnjs0313\}@kaist\.ac\.krks000225@gmail\.comEmail:[coallaoh@kaist\.ac\.kr](mailto:)alice\.oh@kaist\.edu
###### Abstract
Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated\. Some interpret it as identifying task\-relevant skills, while other work reports that surface form is the main factor\. We resolve this debate for supervised fine\-tuning examples by varying task and answer format independently\. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task\. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly \(disattenuated cosine near 0\.4\), while same benchmarks rendered with different answer format classes show no alignment \(near 0\.0\)\. We demonstrate that this ordering holds from the earliest pretraining checkpoints through post\-training, and across model scales and families\. We then analyze the released selections of LESS, a gradient\-based data selection method for instruction tuning, and find that each target’s selections over\-represent the target’s own answer format\. Hence, we demonstrate that gradient\-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability\.
## 1Introduction
Various data attribution methods have been developed to improve the analysis of training data for deep learning models such as transformer\-based large language models \(LLMs\)\. Many such methods leverage the loss gradient to measure the influence of training data\[[1](https://arxiv.org/html/2609.19589#bib.bib4),[2](https://arxiv.org/html/2609.19589#bib.bib6),[3](https://arxiv.org/html/2609.19589#bib.bib7),[4](https://arxiv.org/html/2609.19589#bib.bib1),[5](https://arxiv.org/html/2609.19589#bib.bib8)\]\. A representative method upon which many other methods are based on is TracIn, which uses the raw similarity, or dot product, between loss gradients of data points to calculate the influence of a training data point on a test data point\[[5](https://arxiv.org/html/2609.19589#bib.bib8)\]\.
The literature disagrees on why these methods work and what the gradient actually represents\.[Xia et al\. \[4\]](https://arxiv.org/html/2609.19589#bib.bib1)state that the gradient signal “goes beyond surface form cues to identify data that exemplifies the necessary reasoning skills,” while[Wang et al\. \[6\]](https://arxiv.org/html/2609.19589#bib.bib2)report that answer format dominates over the semantic content of the task in the fine\-tuning gradient\. The confusion has persisted because format and semantics are spuriously correlated in data, especially benchmarks, so observation alone cannot separate them; for example, factual knowledge benchmarks tend to share the same multiple\-choice format, so a gradient that matches format and a gradient that matches knowledge make the same predictions\. The confusion is pertinent totargeted instruction tuning, where the gradient signal is used to pick training examples intended to exemplify the capabilities a target requires, based on the assumption that the gradient encodes semantics\[[4](https://arxiv.org/html/2609.19589#bib.bib1)\]\. We resolve the confusion by varying thetask111Note that we use the termstaskandbenchmarkinterchangeably, as the benchmarks used assess unique capabilities\.and theanswer formatof the same questions independently, and measuring each factor’s effect on gradient alignment\.
Our contributions are threefold:
- •We show that answer format, not task content, governs gradient similarity in LLMs, meaning that influence as measured by gradient\-based data attribution methods does not imply task relatedness\.
- •We find that such dominance of format over content is consistent across pre\-training and post\-training steps, as well as across model scales\.
- •We analyze the selection resulting from an existing gradient\-based data selection method, namely LESS\[[4](https://arxiv.org/html/2609.19589#bib.bib1)\], to demonstrate that the resulting selection is skewed towards same\-format samples\.
## 2Related Work
#### Gradient\-based attribution and its semantic reading\.
The representative gradient\-based data attribution method is TracIn, which estimates the influence of a training example on a test example by accumulating dot products of their loss gradients over training checkpoints\[[5](https://arxiv.org/html/2609.19589#bib.bib8)\]\. Later work scales this approach to larger models and corpora\[[3](https://arxiv.org/html/2609.19589#bib.bib7),[1](https://arxiv.org/html/2609.19589#bib.bib4),[2](https://arxiv.org/html/2609.19589#bib.bib6),[7](https://arxiv.org/html/2609.19589#bib.bib3)\]\. LESS applies it to data selection and interprets the signal as identifying data that exemplifies the necessary reasoning skills\[[4](https://arxiv.org/html/2609.19589#bib.bib1)\]\. This semantic reading of gradient geometry is widely shared: gradients are described as task\-sensitive representations of knowledge\[[8](https://arxiv.org/html/2609.19589#bib.bib9)\], and gradient clusters are treated as task experts\[[9](https://arxiv.org/html/2609.19589#bib.bib10)\], task\-agnostic coresets\[[10](https://arxiv.org/html/2609.19589#bib.bib11)\], separable basic abilities\[[11](https://arxiv.org/html/2609.19589#bib.bib13)\], or attributions of reasoning ability\[[12](https://arxiv.org/html/2609.19589#bib.bib12)\]\. The closest supporting evidence, which are influence analyses that find increasingly abstract relationships between influential documents and model generations\[[13](https://arxiv.org/html/2609.19589#bib.bib18),[1](https://arxiv.org/html/2609.19589#bib.bib4)\], concerns a different object: the influence of pretraining documents on model generations, not the gradients of supervised fine\-tuning examples\. In short, there is prior work that assumes that gradient alignment between fine\-tuning examples reflects shared task content, and we test this assumption\.
#### Evidence for surface form\.
Other research suggests the gradient encodes surface format more strongly than content\. At the outcome level, LESS fails to beat random selection at large pool scales\[[14](https://arxiv.org/html/2609.19589#bib.bib5)\], lexical baselines remain competitive with gradient attribution at fact tracing\[[7](https://arxiv.org/html/2609.19589#bib.bib3)\], and gradient similarity fails to indicate task relatedness in multi\-task classification\[[15](https://arxiv.org/html/2609.19589#bib.bib15)\]\. At the signal level, the only direct gradient measurement is confined to a single dataset:[Wang et al\. \[6\]](https://arxiv.org/html/2609.19589#bib.bib2)shuffle one dataset’s labels and find that early fine\-tuning gradients track the format\. Whether format organizes the gradientsbetweendatasets, the comparison that attribution and selection actually perform, remains unmeasured\. The confound surfaces in deployed pipelines as well\. Influence magnitudes track completion length inside LESS itself\[[16](https://arxiv.org/html/2609.19589#bib.bib14)\], gradient\-matched selections of harmful data turn out to be lists, bullet points, and math questions\[[17](https://arxiv.org/html/2609.19589#bib.bib16)\], and influence retrieval for verbalized confidence latches onto surface markers of certainty rather than content\[[18](https://arxiv.org/html/2609.19589#bib.bib17)\]\. We isolate the confound with a controlled design, varying task and answer format independently in the setting attribution methods operate in, where candidate and target come from different datasets\.
Figure 1:Overview of method used to extract pairwise cosine similarities between datasets\.
## 3Method
### 3\.1Preliminaries
#### Per\-example loss gradients\.
Letfθf\_\{\\theta\}be a language model with parametersθ\\theta, and let a supervised examplez=\(x,y\)z=\(x,y\)consist of a promptxxand an answeryy\. The lossℓ\(z,θ\)\\ell\(z;\\theta\)is the mean cross\-entropy of the answer tokens given the prompt, and the per\-example gradient isg\(z,θ\)=∇θℓ\(z,θ\)g\(z;\\theta\)=\\nabla\_\{\\theta\}\\,\\ell\(z;\\theta\)\. Because the loss is masked to the answer span,ggis the update that the example would contribute during instruction tuning, and it is the basic object of gradient\-based attribution\.
#### Gradient\-based influence\.
TracIn\[[5](https://arxiv.org/html/2609.19589#bib.bib8)\]estimates the influence of a training example on a test example as a learning\-rate\-weighted sum of gradient dot products over training checkpointsθ1,…,θK\\theta\_\{1\},\\dots,\\theta\_\{K\}\. LESS\[[4](https://arxiv.org/html/2609.19589#bib.bib1)\]adapts this score to instruction tuning: Notably, gradient alignments are calculated by cosine so that long answers do not dominate through gradient norm\. The training pool is ranked by the aggregated similarity score with a target set across checkpoints, and the top 5% is selected for fine\-tuning\. We analyze the selections this method produces in Section[3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px2)\.
### 3\.2Our Method: Measuring Dataset\-wise Gradient Alignment
We measure gradient alignment between datasets in three steps as shown in Figure[1](https://arxiv.org/html/2609.19589#S2.F1)\.
#### Step 1: Render datasets\.
We create adatasetby rendering the questions of a benchmark into one answer format, keeping the questions themselves fixed\. We cross five benchmarks, TriviaQA\[[19](https://arxiv.org/html/2609.19589#bib.bib21)\], SQuAD\[[20](https://arxiv.org/html/2609.19589#bib.bib22)\], GSM8K\[[21](https://arxiv.org/html/2609.19589#bib.bib23)\], ARC\-Challenge\[[22](https://arxiv.org/html/2609.19589#bib.bib24)\], and CRUXEval\[[23](https://arxiv.org/html/2609.19589#bib.bib25)\], with seven answer formats spanning three structure classes: binary judgment \(True/False, Yes/No\), four\-option selection \(A–D, 1–4, W–Z, a–d\), and structured generation \(JSON\); details in Appendix[C](https://arxiv.org/html/2609.19589#A3)\. The resulting 35 datasets let any two be compared while changing only the task or only the answer format\.
#### Step 2: Compute gradients for each dataset\.
Forn=64n=64examples per dataset we compute the per\-example answer\-span gradientg\(z,θ\)g\(z;\\theta\)of Section[3\.1](https://arxiv.org/html/2609.19589#S3.SS1), taken with respect to the transformer body only; excluding the embedding and unembedding tables ensures that vocabulary overlap between answer tokens cannot drive similarity\. Each gradient is compressed with a seeded CountSketch projection\[[24](https://arxiv.org/html/2609.19589#bib.bib20)\]tok=16,384k=16\{,\}384dimensions, which preserves inner products in expectation\. Each dataset becomes a matrix inℝ64×16,384\\mathbb\{R\}^\{64\\times 16\{,\}384\}\.
#### Step 3: Compute pairwise cosines\.
Theinter\-datasetalignment of datasetsAAandBBis the mean cosine over all cross\-dataset example pairs:
c^\(A,B\)=1\|A\|\|B\|∑z∈A∑z′∈Bcos\(g\(z,θ\),g\(z′,θ\)\)\.\\hat\{c\}\(A,B\)\\;=\\;\\frac\{1\}\{\|A\|\\,\|B\|\}\\sum\_\{z\\in A\}\\ \\sum\_\{z^\{\\prime\}\\in B\}\\cos\\\!\\big\(g\(z;\\theta\),\\,g\(z^\{\\prime\};\\theta\)\\big\)\.\(1\)The same statistic measured within a dataset gives theintra\-datasetcoherencecA=c^\(A,A\)c\_\{A\}=\\hat\{c\}\(A,A\), evaluated over distinct question pairs\. We disattenuate by coherence, as datasets differ in how noisy their example gradients are, so raw values of Equation[1](https://arxiv.org/html/2609.19589#S3.E1)are not comparable across pairs\. If each example gradient is a dataset\-level signalμ\\muplus independent noise, the measured mean is attenuated by exactlycAcB\\sqrt\{c\_\{A\}c\_\{B\}\}, so dividing it out recovers the alignment of the signals themselves\[[25](https://arxiv.org/html/2609.19589#bib.bib19)\]:
𝔼c^\(A,B\)=cos\(μA,μB\)cAcB⟹c~\(A,B\)=c^\(A,B\)cAcB\.\\mathbb\{E\}\\,\\hat\{c\}\(A,B\)\\;=\\;\\cos\(\\mu\_\{A\},\\mu\_\{B\}\)\\,\\sqrt\{c\_\{A\}c\_\{B\}\}\\quad\\Longrightarrow\\quad\\widetilde\{c\}\(A,B\)\\;=\\;\\frac\{\\hat\{c\}\(A,B\)\}\{\\sqrt\{c\_\{A\}c\_\{B\}\}\}\.\(2\)More details on disattenuation may be found in Appendix[B](https://arxiv.org/html/2609.19589#A2)\. As a reference for zero, the 280 dataset pairs that share neither task nor format class have mean−0\.001\-0\.001and standard deviation0\.0080\.008\(maximum0\.030\.03\), so we read values within±0\.02\\pm 0\.02as null\. Note, error bars in figures are±1\\pm 1standard error across benchmark identities, so that they reflect whether the effect is consistent across tasks rather than estimator noise \(Appendix[E](https://arxiv.org/html/2609.19589#A5)\)\.
### 3\.3Experimental Setup
#### Models\.
To test whether the format ordering is a property of fully trained models only, we repeat the measurement at ten OLMo pretraining checkpoints \(1B to 4T tokens\)\[[26](https://arxiv.org/html/2609.19589#bib.bib26)\], at ten Pythia checkpoints\[[27](https://arxiv.org/html/2609.19589#bib.bib27)\], and at the base, SFT, DPO, and Instruct stages of OLMo\. To test whether it is specific to a scale or model family, we repeat it at three OLMo scales \(1B, 7B, 13B\) and on Llama 3\.1\[[28](https://arxiv.org/html/2609.19589#bib.bib28)\]and Qwen3\[[29](https://arxiv.org/html/2609.19589#bib.bib29)\]models, both base and instruction\-tuned\.
#### LESS selections\.
We analyze the selections released by[Xia et al\. \[4\]](https://arxiv.org/html/2609.19589#bib.bib1)at[princeton\-nlp/less\_data](https://huggingface.co/datasets/princeton-nlp/less_data): the top 5% \(13,533 examples\) of a 270K\-example pool for each of three targets \(MMLU\[[30](https://arxiv.org/html/2609.19589#bib.bib30)\], BBH\[[31](https://arxiv.org/html/2609.19589#bib.bib31)\], TydiQA\[[32](https://arxiv.org/html/2609.19589#bib.bib32)\]\) and three seeds\. We label every selected example and the pool by answer format with a rule\-based classifier \(~90% agreement with manual labels on a held\-out sample\), and report enrichment as the share of a format among selected examples divided by its share of the pool\. Refer to Appendix[A](https://arxiv.org/html/2609.19589#A1)for more details on the classifier and Appendix[D](https://arxiv.org/html/2609.19589#A4)for more details on pool composition\.
## 4Results
### 4\.1Answer Format, Not Task, Governs Gradient Alignment
Figure[2\(a\)](https://arxiv.org/html/2609.19589#S4.F2.sf1)shows cross\-dataset cosine similarity while Figure[2\(b\)](https://arxiv.org/html/2609.19589#S4.F2.sf2)shows the same, aggregated for each model\. The heatmap shows that same format class dataset pairs have high alignment, while dataset pairs with different format classes have near zero alignment\. This implies that the gradient encodes more for answer format than task semantics\. Figure[2\(b\)](https://arxiv.org/html/2609.19589#S4.F2.sf2)shows that format dominance in the gradient is consistent across model scale and family\.
\(a\)
\(b\)
Figure 2:Gradient alignment follows answer format, not task\. \(a\) Disattenuated cosine between all 35 datasets, averaged over four model families\. Boxes mark the three answer\-structure classes: alignment is high inside a class whatever the task, and at zero across classes even for the same task\. Cells on the diagonal denote alignment for same answer format, different task pairs, while cells off\-diagonal denote same task, different answer format pairs\. \(b\) The same two contrasts measured per model, spanning 1B to 13B and four families\.
### 4\.2Format Dominance Holds Throughout Pretraining and Post\-Training
We show that surface form dominance in gradient similarity is not an artifact of the amount of model pre\-training or post\-training in Figure[3](https://arxiv.org/html/2609.19589#S4.F3)\. The figure shows that different benchmark, same answer class cosine similarity stays above the same benchmark, different answer class similarity\.
Figure 3:Answer form decides gradient similarity at every pre\-training checkpoint\.
### 4\.3LESS Selects for the Target’s Answer Format
Figure 4:LESS’s released selections, by answer format, for the three target formats\. Each cell is the share of a format among a target’s selections divided by its share of the pool; teal boxes mark each target’s own format\.We analyze the selections that[Xia et al\. \[4\]](https://arxiv.org/html/2609.19589#bib.bib1)released for each target benchmark as shown in Figure[4](https://arxiv.org/html/2609.19589#S4.F4)\. Although the pool contains no data from any of the three target benchmarks, each target’s selections are enriched in the target’s own answer format: letter answers are selected at 3\.7 times their pool share for MMLU, chain\-of\-thought at 2\.4 times for BBH, and short answers at 1\.9 times for TydiQA\. The figure including analysis of other answer formats are in Figure[5](https://arxiv.org/html/2609.19589#A4.F5)\. The pattern is stable across the three released seeds\.
## 5Discussion & Future Work
Our results bear on the reliability and robustness of gradient\-based attribution and selection: their efficacy may partly be an artifact of format coinciding with semantics in current evaluation settings, and may not transfer to pools where the two decouple\. More broadly, researchers should test gradient\-based data attribution and selection methods against format\-controlled baselines before relying on them\. Future work includes interventional studies: For example, whether unifying a pool’s answer format makes selection more content\-driven and fine\-tuning more effective\. Additionally, using learned parameter weightings that separate style from content may remove format dominance\.
## 6Limitations
We assume that task relatedness is a desirable factor in data attribution and especially selection\. However, format\-matched data often genuinely improves benchmark scores in the context of targeted instruction tuning; we falsify the semanticinterpretationbehind gradient\-based selection to question its robustness, not its measured gains\. Additionally, the scope of our experiments may be enlarged: the measurement basis is five English question\-answering benchmarks and seven answer formats in three structure classes, with JSON as the only structured\-generation member\. The answer format classifications are also our choices rather than the result of rigorous analysis\.
## References
- \[1\]R\. Grosse, J\. Bae, C\. Anil, N\. Elhage, A\. Tamkin, A\. Tajdini, B\. Steiner, D\. Li, E\. Durmus, E\. Perez, E\. Hubinger, K\. Lukošiūtė, K\. Nguyen, N\. Joseph, S\. McCandlish, J\. Kaplan, and S\. R\. Bowman\(2023\)Studying large language model generalization with influence functions\.External Links:2308\.03296,[Link](https://arxiv.org/abs/2308.03296)Cited by:[§1](https://arxiv.org/html/2609.19589#S1.p1.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[2\]S\. K\. Choe, H\. Ahn, J\. Bae, K\. Zhao, Y\. Chung, A\. Pratapa, W\. Neiswanger, E\. Strubell, T\. Mitamura, J\. Schneider, E\. Hovy, R\. B\. Grosse, and E\. P\. Xing\(2025\)What is your data worth to GPT? LLM\-scale data valuation with influence functions\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=zPKeJAEo27)Cited by:[§1](https://arxiv.org/html/2609.19589#S1.p1.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[3\]S\. M\. Park, K\. Georgiev, A\. Ilyas, G\. Leclerc, and A\. Madry\(2023\)TRAK: attributing model behavior at scale\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2609.19589#S1.p1.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[4\]M\. Xia, S\. Malladi, S\. Gururangan, S\. Arora, and D\. Chen\(2024\)LESS: selecting influential data for targeted instruction tuning\.InInternational Conference on Machine Learning \(ICML\),Cited by:[3rd item](https://arxiv.org/html/2609.19589#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2609.19589#S1.p1.1),[§1](https://arxiv.org/html/2609.19589#S1.p2.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.19589#S3.SS1.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2609.19589#S4.SS3.p1.1)\.
- \[5\]G\. Pruthi, F\. Liu, S\. Kale, and M\. Sundararajan\(2020\)Estimating training data influence by tracing gradient descent\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 19920–19930\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/e6385d39ec9394f2f3a354d9d2b88eec-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.19589#S1.p1.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.19589#S3.SS1.SSS0.Px2.p1.1)\.
- \[6\]Y\. Wang, S\. Si, D\. Li, M\. Lukasik, F\. Yu, C\. Hsieh, I\. S\. Dhillon, and S\. Kumar\(2024\)Two\-stage LLM fine\-tuning with less specialization and more generalization\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=pCEgna6Qco)Cited by:[§1](https://arxiv.org/html/2609.19589#S1.p2.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[7\]T\. A\. Chang, D\. Rajagopal, T\. Bolukbasi, L\. Dixon, and I\. Tenney\(2025\)Scalable influence and fact tracing for large language model pretraining\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gLa96FlWwn)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[8\]Y\. Zhao, L\. Du, X\. Ding, Y\. Ouyang, H\. Wang, K\. Xiong, J\. Gao, Z\. Sun, D\. Xu, Q\. Yang, D\. Li, B\. Qin, and T\. Liu\(2025\)Beyond similarity: a gradient\-based graph method for instruction tuning data selection\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 24391–24404\.External Links:[Link](https://aclanthology.org/2025.acl-long.1189/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1189),ISBN 979\-8\-89176\-251\-0Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[9\]Y\. Li, V\. R\. Gao, C\. Zhang, and M\. Torkamani\(2025\)Ensembles of low\-rank expert adapters\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=l0gZS0sAlf)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[10\]J\. Zhang, Y\. Qin, R\. Pi, W\. Zhang, R\. Pan, and T\. Zhang\(2025\)TAGCOS: task\-agnostic gradient clustered coreset selection for instruction tuning data\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 4686–4701\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.264/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.264),ISBN 979\-8\-89176\-195\-7Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[11\]B\. Wang, X\. Li, C\. Li, J\. Chi, G\. Niu, and M\. Sugiyama\(2026\)Decomposing the basic abilities of large language models: mitigating cross\-task interference in multi\-task instruct\-tuning\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=FFAHL32Wok)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[12\]S\. Kou, Q\. Tian, H\. Xu, Z\. Zeng, and Z\. Deng\(2025\)Which data attributes stimulate math and code reasoning? an investigation via influence functions\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=b7uniOw0sZ)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[13\]L\. Ruis, M\. Mozes, J\. Bae, S\. R\. Kamalakara, D\. Gnaneshwar, A\. Locatelli, R\. Kirk, T\. Rocktäschel, E\. Grefenstette, and M\. Bartolo\(2025\)Procedural knowledge in pretraining drives reasoning in large language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1hQKHHUsMx)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px1.p1.1)\.
- \[14\]T\. Xia, B\. Yu, K\. Dang, A\. Yang, Y\. Wu, Y\. Tian, Y\. Chang, and J\. Lin\(2025\)Rethinking data selection at scale: random selection is almost all you need\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 2698–2711\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.146/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.146),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[15\]J\. Ni, Z\. Jin, Q\. Wang, M\. Sachan, and M\. Leippold\(2023\)When does aggregating multiple skills with multi\-task learning work? a case study in financial NLP\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 7465–7488\.External Links:[Link](https://aclanthology.org/2023.acl-long.412/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.412)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[16\]Q\. Dai, D\. Zhang, J\. W\. Ma, and H\. Peng\(2025\)Improving influence\-based instruction tuning data selection for balanced learning of diverse capabilities\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 7079–7102\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.373/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.373),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[17\]L\. He, M\. Xia, and P\. Henderson\(2024\)What is in your safe data? identifying benign data that breaks safety\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Hi8jKh4HE9)Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[18\]Y\. Xia, L\. Schoenegger, and B\. Roth\(2026\)Influential training data retrieval for explaining verbalized confidence of llms\.InAdvances in Information Retrieval,R\. Campos, A\. Jatowt, Y\. Lan, M\. Aliannejadi, C\. Bauer, S\. MacAvaney, A\. Anand, Z\. Ren, S\. Verberne, N\. Bai, and M\. Mansoury \(Eds\.\),Cham,pp\. 529–547\.External Links:ISBN 978\-3\-032\-21289\-4Cited by:[§2](https://arxiv.org/html/2609.19589#S2.SS0.SSS0.Px2.p1.1)\.
- \[19\]M\. Joshi, E\. Choi, D\. Weld, and L\. Zettlemoyer\(2017\)TriviaQA: a large scale distantly supervised challenge dataset for reading comprehension\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),R\. Barzilay and M\. Kan \(Eds\.\),Vancouver, Canada,pp\. 1601–1611\.External Links:[Link](https://aclanthology.org/P17-1147/),[Document](https://dx.doi.org/10.18653/v1/P17-1147)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px1.p1.1)\.
- \[20\]P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. Liang\(2016\)SQuAD: 100,000\+ questions for machine comprehension of text\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,J\. Su, K\. Duh, and X\. Carreras \(Eds\.\),Austin, Texas,pp\. 2383–2392\.External Links:[Link](https://aclanthology.org/D16-1264/),[Document](https://dx.doi.org/10.18653/v1/D16-1264)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px1.p1.1)\.
- \[21\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px1.p1.1)\.
- \[22\]P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord\(2018\)Think you have solved question answering? try arc, the ai2 reasoning challenge\.External Links:1803\.05457,[Link](https://arxiv.org/abs/1803.05457)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px1.p1.1)\.
- \[23\]A\. Gu, B\. Roziere, H\. J\. Leather, A\. Solar\-Lezama, G\. Synnaeve, and S\. Wang\(2024\)CRUXEval: a benchmark for code reasoning, understanding and execution\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 16568–16621\.External Links:[Link](https://proceedings.mlr.press/v235/gu24c.html)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px1.p1.1)\.
- \[24\]K\. Weinberger, A\. Dasgupta, J\. Langford, A\. Smola, and J\. Attenberg\(2009\)Feature hashing for large scale multitask learning\.InProceedings of the 26th Annual International Conference on Machine Learning,ICML ’09,New York, NY, USA,pp\. 1113–1120\.External Links:ISBN 9781605585161,[Link](https://doi.org/10.1145/1553374.1553516),[Document](https://dx.doi.org/10.1145/1553374.1553516)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px2.p1.1)\.
- \[25\]C\. Spearman\(1904\)The proof and measurement of association between two things\.The American Journal of Psychology15\(1\),pp\. 72–101\.External Links:ISSN 00029556,[Link](http://www.jstor.org/stable/1412159)Cited by:[§3\.2](https://arxiv.org/html/2609.19589#S3.SS2.SSS0.Px3.p1.2)\.
- \[26\]T\. OLMo, P\. Walsh, L\. Soldaini, D\. Groeneveld, K\. Lo, S\. Arora, A\. Bhagia, Y\. Gu, S\. Huang, M\. Jordan, N\. Lambert, D\. Schwenk, O\. Tafjord, T\. Anderson, D\. Atkinson, F\. Brahman, C\. Clark, P\. Dasigi, N\. Dziri, A\. Ettinger, M\. Guerquin, D\. Heineman, H\. Ivison, P\. W\. Koh, J\. Liu, S\. Malik, W\. Merrill, L\. J\. V\. Miranda, J\. Morrison, T\. Murray, C\. Nam, J\. Poznanski, V\. Pyatkin, A\. Rangapur, M\. Schmitz, S\. Skjonsberg, D\. Wadden, C\. Wilhelm, M\. Wilson, L\. Zettlemoyer, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi\(2025\)2 olmo 2 furious\.External Links:2501\.00656,[Link](https://arxiv.org/abs/2501.00656)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px1.p1.1)\.
- \[27\]S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. Van Der Wal\(2023\)Pythia: a suite for analyzing large language models across training and scaling\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 2397–2430\.External Links:[Link](https://proceedings.mlr.press/v202/biderman23a.html)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px1.p1.1)\.
- \[28\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px1.p1.1)\.
- \[29\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu\(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px1.p1.1)\.
- \[30\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt\(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px2.p1.1)\.
- \[31\]M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. Le, E\. Chi, D\. Zhou, and J\. Wei\(2023\)Challenging BIG\-bench tasks and whether chain\-of\-thought can solve them\.InFindings of the Association for Computational Linguistics: ACL 2023,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 13003–13051\.External Links:[Link](https://aclanthology.org/2023.findings-acl.824/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.824)Cited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px2.p1.1)\.
- \[32\]J\. H\. Clark, E\. Choi, M\. Collins, D\. Garrette, T\. Kwiatkowski, V\. Nikolaev, and J\. Palomaki\(2020\)TyDi qa: a benchmark for information\-seeking question answering in typologically diverse languages\.Transactions of the Association for Computational Linguistics8,pp\. 454–470\.External Links:ISSN 2307\-387X,[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00317),[Link](https://doi.org/10.1162/tacl_a_00317),https://direct\.mit\.edu/tacl/article\-pdf/doi/10\.1162/tacl\_a\_00317/1923348/tacl\_a\_00317\.pdfCited by:[§3\.3](https://arxiv.org/html/2609.19589#S3.SS3.SSS0.Px2.p1.1)\.
- \[33\]S\. Longpre, L\. Hou, T\. Vu, A\. Webson, H\. W\. Chung, Y\. Tay, D\. Zhou, Q\. V\. Le, B\. Zoph, J\. Wei, and A\. Roberts\(2023\)The flan collection: designing data and methods for effective instruction tuning\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 22631–22648\.External Links:[Link](https://proceedings.mlr.press/v202/longpre23a.html)Cited by:[Table 3](https://arxiv.org/html/2609.19589#A4.T3)\.
- \[34\]A\. Köpf, Y\. Kilcher, D\. von Rütte, S\. Anagnostidis, Z\. R\. Tam, K\. Stevens, A\. Barhoum, D\. Nguyen, O\. Stanley, R\. Nagyfi, S\. ES, S\. Suri, D\. Glushkov, A\. Dantuluri, A\. Maguire, C\. Schuhmann, H\. Nguyen, and A\. Mattick\(2023\)OpenAssistant conversations \- democratizing large language model alignment\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 47669–47681\.External Links:[Document](https://dx.doi.org/10.52202/075280-2064),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/949f0f8f32267d297c2d4e3ee10a2e7e-Paper-Datasets_and_Benchmarks.pdf)Cited by:[Table 3](https://arxiv.org/html/2609.19589#A4.T3)\.
- \[35\]Databricks\(2023\)Free dolly: introducing the world’s first truly open instruction\-tuned llm\.GitHub\.Note:Blog postExternal Links:[Link](https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm)Cited by:[Table 3](https://arxiv.org/html/2609.19589#A4.T3)\.
## NeurIPS Paper Checklist
1. 1\.Claims\.Answer: Yes\. The abstract and introduction state the three contributions, each matched by a results subsection\.
2. 2\.Limitations\.Answer: Yes\. See the Limitations section\.
3. 3\.Theory, assumptions and proofs\.Answer: Yes\. The one identity used \(Eq\. 2\) states its assumption in the text; the derivation is in Appendix[B](https://arxiv.org/html/2609.19589#A2)\.
4. 4\.Experimental result reproducibility\.Answer: Yes\. All models are open\-weight, all datasets public, and the Method and Setup sections specify the estimator \(examples per dataset, projection, parameter scope, masking\); classifier rules are in Appendix[A](https://arxiv.org/html/2609.19589#A1)\.
5. 5\.Open access to data and code\.Answer: No\. All datasets and models used are publicly available and cited; our measurement code is not yet released\.
6. 6\.Experimental setting/details\.Answer: Yes\. See Method and Experimental Setup\.
7. 7\.Error bars\.Answer: Yes\. Error bars are±1\\pm 1standard error over benchmark groups, defined in Experimental Setup; the LESS analysis reports all three released seeds\.
8. 8\.Compute resources\.Answer: Yes\. See Appendix[F](https://arxiv.org/html/2609.19589#A6)\(Compute\)\.
9. 9\.Code of ethics\.Answer: Yes\.
10. 10\.Broader impacts\.Answer: Yes\. The work analyzes existing methods; we foresee no direct negative societal impact, and a clearer picture of what drives data selection may improve curation practice\.
11. 11\.Safeguards\.Answer: NA\. No models or high\-risk assets are released\.
12. 12\.Licenses for existing assets\.Answer: Yes\. All datasets and models are cited and used under their released licenses and terms\.
13. 13\.New assets\.Answer: NA\. No new assets are released\.
14. 14\.Crowdsourcing and human subjects\.Answer: NA\.
15. 15\.IRB approvals\.Answer: NA\.
16. 16\.LLM usage\.Answer: NA\. LLMs are not a component of the methodology; the format classifier is rule\-based\.
## Appendix APost\-hoc Classifier Details
Each example is labeled by rules applied to its assistant turn \(and its user turn, to detect offered options\), in priority order: \(1\) a bare yes/no/true/false answer isbinary; \(2\) an answer longer than 25 words containing reasoning markers or a trailing “the answer is \(X\)” ischain of thought, checked before the letter rule so that rationales ending in a letter are not misrouted; \(3\) a bare option letter, with options present in the prompt, ischoice; \(4\) a bare number isnumeric; \(5\) an answer that parses as JSON isJSON; \(6\) an answer of at most 12 words isshort answer; \(7\) anything else islong generation\. Numeric is merged into short answer in all reported results, as both are short constrained spans\. On 30 randomly sampled examples labeled by hand, the classifier agreed on 27 \(90%\)\. Pool base rates are computed from 5,000 sampled examples per source, weighted by the true source sizes\.
## Appendix BDisattenuation Identity
Assume each unit\-normalized example gradient of datasetAAdecomposes asgi=cAμ^A\+1−cAηig\_\{i\}=\\sqrt\{c\_\{A\}\}\\,\\hat\{\\mu\}\_\{A\}\+\\sqrt\{1\-c\_\{A\}\}\\,\\eta\_\{i\}, whereμ^A\\hat\{\\mu\}\_\{A\}is the unit dataset\-level direction and the noise termsηi\\eta\_\{i\}have mean zero and are independent across examples and datasets\. Then for examplesi≠ji\\neq jwithinAA,𝔼⟨gi,gj⟩=cA\\mathbb\{E\}\\langle g\_\{i\},g\_\{j\}\\rangle=c\_\{A\}, socAc\_\{A\}is by construction the expected within\-dataset cosine of Equation[1](https://arxiv.org/html/2609.19589#S3.E1)\. Across datasets, all noise cross\-terms vanish in expectation, leaving𝔼⟨giA,gjB⟩=cAcBcos\(μA,μB\)\\mathbb\{E\}\\langle g^\{A\}\_\{i\},g^\{B\}\_\{j\}\\rangle=\\sqrt\{c\_\{A\}c\_\{B\}\}\\,\\cos\(\\mu\_\{A\},\\mu\_\{B\}\), which is Equation[2](https://arxiv.org/html/2609.19589#S3.E2); dividing the measured mean bycAcB\\sqrt\{c\_\{A\}c\_\{B\}\}therefore recoverscos\(μA,μB\)\\cos\(\\mu\_\{A\},\\mu\_\{B\}\)\.
## Appendix CBenchmarks and Answer Formats
Table[1](https://arxiv.org/html/2609.19589#A3.T1)lists the five benchmarks and Table[2](https://arxiv.org/html/2609.19589#A3.T2)the seven answer formats and their structure classes\.
Table 1:The five benchmarks we cross with the answer formats of Table[2](https://arxiv.org/html/2609.19589#A3.T2); each crossing of answer format with a benchmark serves as adataset, totaling 35 datasets\. For the four\-option renderings, ARC supplies its own hand\-written distractors and GSM8K uses near\-miss arithmetic values; the rest draw distractors from other answers in the same dataset, so the options are always in\-distribution\.Table 2:The seven answer formats we cross with each benchmark, grouped by the structure of the answer the model must produce\.
## Appendix DThe LESS Selection Pool
Table[3](https://arxiv.org/html/2609.19589#A4.T3)gives the composition of LESS’s 270,679\-example selection pool: the four source datasets, their sizes, and the answer\-format shares our classifier assigns within each source\. The bottom row is the source\-weighted pool base rate, the denominator of every enrichment in Figure[4](https://arxiv.org/html/2609.19589#S4.F4)\. Note that our classifier labels only 6% of the CoT source as chain\-of\-thought format, since rationales without explicit reasoning markers fall under long generation\. Relabeling every CoT\-source example as chain\-of\-thought instead gives BBH 1\.7×\\timeson its own format \(versus 2\.4×\\timesunder the rule\-based labels\), leaves the diagonal of Figure[4](https://arxiv.org/html/2609.19589#S4.F4)the maximum of every row, and changes no other conclusion\.
Table 3:Composition of the LESS selection pool\[[33](https://arxiv.org/html/2609.19589#bib.bib33),[34](https://arxiv.org/html/2609.19589#bib.bib34),[35](https://arxiv.org/html/2609.19589#bib.bib35)\]\. Format shares are estimated from 5,000 sampled examples per source and weighted by true source sizes\. Formats are labeled by the rule\-based classifier of Appendix[A](https://arxiv.org/html/2609.19589#A1); rationales without explicit reasoning markers fall under long generation, so the chain\-of\-thought base rate is conservative\.Figure 5:Enrichment of LESS’s selections over all six answer formats; the left block is Figure[4](https://arxiv.org/html/2609.19589#S4.F4)\. Formats outside each target’s structure class are depleted \(for MMLU, long generation at 0\.2×\\timesand JSON at 0\.04×\\times\)\.
## Appendix EAggregation and Error Bars
The trajectory and per\-model figures reduce the full grid to two contrasts, and their error bars are designed to answer one question: Does the format effect depend on which benchmark was used? Each contrast is therefore aggregated in two stages, first averaging within a grouping unit defined by benchmark identity, then reporting the mean and±1\\pm 1standard error across those units\.
#### Same class, different task\.
This contrast compares datasets from two different benchmarks that share a format class, so the grouping unit is a benchmark pair\. With five benchmarks there are ten pairs\. Within each pair we average the disattenuated cosine over every same\-class format combination \(for the pair GSM8K–ARC, for example, GSM8K rendered as A–D against ARC rendered as 1–4, GSM8K as True/False against ARC as Yes/No, and so on\), yielding one value per pair; the plotted point is the mean over the ten pairs and the error bar their standard error\.
#### Different class, same task\.
This contrast compares a benchmark with itself rendered in different format classes, so no benchmark pair is involved and the grouping unit is the single benchmark\. With five benchmarks there are five units\. Within each we average over every cross\-class format combination \(for GSM8K, A–D against True/False, JSON against Yes/No, and so on\), and the plotted point is the mean and standard error over the five\.
Grouping by benchmark identity is deliberate\. Each cell already averages64×6464\\times 64example pairs, so a standard error over example pairs would be very small and would measure only estimator precision\. The standard error over benchmarks instead measures whether the finding generalizes across task content, which is the uncertainty relevant to our claim; the ten\-versus\-five asymmetry simply reflects how many benchmark identities each contrast involves\.
## Appendix FCompute
All gradient measurements ran on NVIDIA RTX A6000 GPUs \(48 GB\), using at most three concurrently\. Measuring the full grid for one 7B\-scale checkpoint takes about one GPU\-hour; the 29 reported checkpoints total roughly 30 GPU\-hours\. The analysis of LESS’s released selections runs on CPU only\.Similar Articles
Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution
Introduces PRIG, a gradient attribution method that localizes prompt ambiguity in large language models by training a linear probe to distinguish clear from ambiguous prompts and attributing the probe score to token representations in the residual stream, achieving strong performance on synthetic and human-written benchmarks.
DataDignity: Training Data Attribution for Large Language Models
This paper introduces DataDignity, a framework and benchmark (FakeWiki) for pinpoint provenance, aiming to identify the specific training data sources that support an LLM's response. It proposes ScoringModel and SteerFuse methods to improve attribution accuracy over standard retrieval baselines.
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
Introduces MultAttnAttrib, a training-free method for multimodal attribution in long document QA, along with the MultAttrEval benchmark. It outperforms prompting-based methods and matches frontier models like GPT-5.4.
When Attribution Patching Lies: Diagnosis and a Second-Order Correction
This paper diagnoses systematic errors in attribution patching, a gradient-based approximation used for causal localization in language models, and proposes a second-order correction using Hessian-vector products that improves reliability with minimal additional computational cost.
The Attribution Contract: Feature Attribution for Generative Language Models
This paper introduces the Attribution Contract, a specification for feature-attribution claims in generative language models, addressing ambiguities in what constitutes a feature and how attribution methods should be evaluated. It uses autoregressive and diffusion models as case studies to show when attribution is informative or misleading.