Scaling Laws for Behavioral Foundation Models over User Event Sequences

arXiv cs.LG Papers

Summary

This paper studies scaling laws for behavioral foundation models trained on sequences of user actions, finding that a small event embedder is compute-optimal and that the evaluation metric itself influences the optimal compute allocation.

arXiv:2606.05257v1 Announce Type: new Abstract: Foundation models are increasingly trained on sequences of user actions in recommendation, payments, fraud, and commerce, but these models still lack the kind of compute calibration that scaling laws provide for language models. We study a common two-part behavioral-model architecture: a feature-based event embedder maps each multi-modal item to a vector, and a decoder-only transformer predicts the next event from the resulting sequence. Across roughly 600 runs on real interaction data, spanning $10^{15}$-$10^{19}$ training FLOPs, we jointly vary four deployment-relevant axes: the two-part parameter split, critical batch size, model/data allocation, and the number of sampled negatives used after freezing the embedder. A small embedder ($s^{\star}\!\approx\!2\%$ of parameters) is compute-optimal at every budget we test because embedder parameters are both more expensive per step and exposed to far more repeated items than contextualizer parameters. Compute-optimal training is data-heavy relative to text at low compute, but its $D/N$ ratio moves toward the Chinchilla heuristic as compute increases. The sampled training objective and deployed ranking metrics disagree in ways that themselves scale: critical batch size, optimal negative count after freezing, and the agreement between loss and ranking quality all shift with compute and with the chosen evaluation metric. For negative sampling, larger budgets increasingly prefer more negatives; by $10^{19}$ FLOPs the active constraint is candidate-axis memory rather than FLOPs. In behavioral foundation models, the evaluation metric is therefore part of the scaling law: changing it can change the compute-optimal recipe.
Original Article
View Cached Full Text

Cached at: 06/05/26, 08:09 AM

# Scaling Laws for Behavioral Foundation Models over User Event Sequences
Source: [https://arxiv.org/html/2606.05257](https://arxiv.org/html/2606.05257)
###### Abstract

Foundation models are increasingly trained on sequences of user actions in recommendation, payments, fraud, and commerce, but these models still lack the kind of compute calibration that scaling laws provide for language models\. We study a common two\-part behavioral\-model architecture: a feature\-based event embedder maps each multi\-modal item to a vector, and a decoder\-only transformer predicts the next event from the resulting sequence\. Across roughly 600 runs on real interaction data, spanning101510^\{15\}–101910^\{19\}training FLOPs, we jointly vary four deployment\-relevant axes: the two\-part parameter split, critical batch size, model/data allocation, and the number of sampled negatives used after freezing the embedder\. A small embedder \(s⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%of parameters\) is compute\-optimal at every budget we test because embedder parameters are both more expensive per step and exposed to far more repeated items than contextualizer parameters\. Compute\-optimal training is data\-heavy relative to text at low compute, but itsD/ND/Nratio moves toward the Chinchilla heuristic as compute increases\. The sampled training objective and deployed ranking metrics disagree in ways that themselves scale: critical batch size, optimal negative count after freezing, and the agreement between loss and ranking quality all shift with compute and with the chosen evaluation metric\. For negative sampling, larger budgets increasingly prefer more negatives; by101910^\{19\}FLOPs the active constraint is candidate\-axis memory rather than FLOPs\. In behavioral foundation models, the evaluation metric is therefore part of the scaling law: changing it can change the compute\-optimal recipe\.

###### Contents

1. [1Introduction](https://arxiv.org/html/2606.05257#S1)
2. [2Experimental Setup](https://arxiv.org/html/2606.05257#S2)
3. [3Scaling the Event Embedder](https://arxiv.org/html/2606.05257#S3)1. [3\.1Embedder Share](https://arxiv.org/html/2606.05257#S3.SS1) 2. [3\.2Depth as a Secondary Knob](https://arxiv.org/html/2606.05257#S3.SS2)
4. [4Critical Batch Size Across Metrics](https://arxiv.org/html/2606.05257#S4)
5. [5Model/Data Allocation Across Compute Budgets](https://arxiv.org/html/2606.05257#S5)
6. [6Scaling the Negative Candidate Pool](https://arxiv.org/html/2606.05257#S6)
7. [7Cross\-Metric and Cross\-Regime Evaluation](https://arxiv.org/html/2606.05257#S7)
8. [8Related Work](https://arxiv.org/html/2606.05257#S8)
9. [9Discussion](https://arxiv.org/html/2606.05257#S9)
10. [10Conclusion](https://arxiv.org/html/2606.05257#S10)
11. [References](https://arxiv.org/html/2606.05257#bib)
12. [AMetric Definitions](https://arxiv.org/html/2606.05257#A1)
13. [BAdditional Architecture Results](https://arxiv.org/html/2606.05257#A2)
14. [CPhase 2: Per\-Metric Trajectories](https://arxiv.org/html/2606.05257#A3)
15. [DPhase 3 Train\-Surrogate Allocation](https://arxiv.org/html/2606.05257#A4)
16. [EPhase 4 Metric\-Stratified Sampling Efficiency](https://arxiv.org/html/2606.05257#A5)
17. [FMaximal Update Parameterization](https://arxiv.org/html/2606.05257#A6)
18. [GCross\-Metric Details](https://arxiv.org/html/2606.05257#A7)
19. [HContext\-Length Scoring Robustness \(full slice matrices\)](https://arxiv.org/html/2606.05257#A8)

## 1Introduction

Scaling laws turned large language model development from a sequence of ad hoc training runs into a quantitative allocation problem: given a compute budget, how many parameters should be trained on how many tokens, at what batch size, and with which optimizer recipe? The same question is now appearing outside text\. Industrial systems increasingly train foundation models over sequences of human actions: recommendations, purchases, payments, financial events, workforce activity, and other behavioral traces\[[5](https://arxiv.org/html/2606.05257#bib.bib5),[6](https://arxiv.org/html/2606.05257#bib.bib6),[13](https://arxiv.org/html/2606.05257#bib.bib13),[14](https://arxiv.org/html/2606.05257#bib.bib14),[16](https://arxiv.org/html/2606.05257#bib.bib16),[17](https://arxiv.org/html/2606.05257#bib.bib17),[18](https://arxiv.org/html/2606.05257#bib.bib18),[22](https://arxiv.org/html/2606.05257#bib.bib22)\]\. These models are large enough that scaling\-law mistakes are expensive, yet the field has little Chinchilla\-style guidance for them\.

Behavioral foundation models differ from text LMs in a way that matters for scaling\. A token in a language model is usually an opaque vocabulary id\. An event in a behavioral model is feature\-rich: a product or transaction can have text, category metadata, visual features, price, timestamp, payment channel, and other structured fields\. Catalogues also change over time\. Modern systems therefore tend to use a two\-part architecture: a feature\-based*event embedder*maps raw item features into a dense representation, and a*contextualizer*, typically a decoder\-only transformer, models sequences of those event embeddings\.

This architecture creates scaling questions that do not appear in ordinary text LMs\. The embedder is suitably trained end\-to\-end\[[19](https://arxiv.org/html/2606.05257#bib.bib19)\]on the sequential objective, but re\-embedding every candidate item during a full softmax is prohibitively expensive for million\-item catalogues\. Prior work addresses this with a two\-stage recipe: train the embedder and contextualizer jointly with an in\-batch sampled softmax, then freeze and cache the embedder and continue training only the contextualizer with a larger sampled candidate pool\[[19](https://arxiv.org/html/2606.05257#bib.bib19),[22](https://arxiv.org/html/2606.05257#bib.bib22)\]\. The resulting system has at least four coupled knobs: how much capacity belongs in the embedder, which batch size is data\-efficient, how compute should be split between model size and data, and how many negatives should be sampled after the embedder is frozen\.

We calibrate those knobs on a single behavioral\-model stack and a single real interaction corpus\. The study spans approximately 600 runs andC∈\[1015,1019\]C\\\!\\in\\\!\[10^\{15\},10^\{19\}\]FLOPs, with a shared evaluation pipeline across all experiments\. The goal is not to propose a new architecture, but to answer a more basic question: if a practitioner is already training this now\-standard event\-embedder→\\rightarrowtransformer stack, what scaling laws should guide the next run?

Our main findings are:

- •The compute\-optimal event embedder is small\.Across four decades of compute, two\-term iso\-FLOP fits place the optimal embedder share in a narrow band arounds⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%of parameters\. The optimum is explained by two asymmetries: embedder parameters are touched many more times per item, and popular items are repeated hundreds to thousands of times while contextualizer windows rarely repeat\.
- •Behavioral scaling is initially data\-heavy but moves toward the Chinchilla heuristic\.The compute\-optimalD/ND/Nratio decreases from roughly340340at101510^\{15\}FLOPs to roughly3636at101910^\{19\}FLOPs, approaching the text\-LM rule of thumb at larger budgets\. The fitted model\-size exponent isN⋆∝C0\.617±0\.025N^\{\\star\}\\\!\\propto\\\!C^\{0\.617\\pm 0\.025\}under validation loss, with similar exponents under the headline ranking metrics\.
- •The evaluation metric is part of the scaling law\.Cross\-budget exponents are relatively stable across metrics, but the actionable recipe is not\. Critical batch size, optimal negative count after freezing, and the agreement between loss and ranking quality all depend on compute, evaluation regime, and target metric\. In particular, the sampled\-softmax loss used during training is not always a reliable proxy for full\-catalogue ranking quality\.
- •Negative sampling shifts from a compute question to a memory question at scale\.At smaller budgets, smooth fits place useful negative counts in the low hundreds of thousands and the optimum is metric\-dependent\. AtC=1019C\\\!=\\\!10^\{19\}, every headline metric is still improving at the largestKKwe train, so the active constraint becomes candidate\-axis memory rather than available FLOPs\.

Table[1](https://arxiv.org/html/2606.05257#S1.T1)summarizes the experimental map\. The remainder of the paper follows the table: Section[2](https://arxiv.org/html/2606.05257#S2)defines the model, training, compute accounting, and evaluation protocol; Sections[3](https://arxiv.org/html/2606.05257#S3)–[6](https://arxiv.org/html/2606.05257#S6)calibrate the four knobs; Section[7](https://arxiv.org/html/2606.05257#S7)explains why metric choice changes the recipe; and Section[9](https://arxiv.org/html/2606.05257#S9)distills the practical schedule and limitations\.

Table 1:Experimental map\.Each axis is swept on the same two\-part behavioral\-model stack\. Labels retain the original phase numbers for cross\-reference with the appendix, but the main text treats them as four scaling\-law questions\.
## 2Experimental Setup

#### Model\.

All experiments use the same two\-part next\-event prediction architecture\. The feature\-based embedder consumes the raw fields of each catalogue item and produces an event embedding at hidden sizehh\. The contextualizer is a decoder\-only transformer that consumes a sequence ofLseq=256L\_\{\\mathrm\{seq\}\}\\\!=\\\!256event embeddings and predicts the next event with a sampled softmax\. The parameter countNNexcludes vocabulary parameters and is decomposed into embedder parameterspep\_\{e\}and contextualizer parameterspp; the embedder share iss=pe/Ns\\\!=\\\!p\_\{e\}/N\.

#### Two\-stage training\.

Following the deployed recipe\[[19](https://arxiv.org/html/2606.05257#bib.bib19),[22](https://arxiv.org/html/2606.05257#bib.bib22)\], all runs use two stages\. In*Stage 1*the embedder and contextualizer are trained jointly on the next\-event objective with an*in\-batch*sampled softmax: the positive is the true next item and the negatives are the other targets in the same global batch, so the candidate pool is the≤B​Lseq\\leq\\\!BL\_\{\\mathrm\{seq\}\}unique items present in the batch\. In*Stage 2*the trained embedder is frozen and its event embeddings are cached, and only the contextualizer continues training, now scoring each position against a larger pool ofKKuniformly sampled extra negatives on top of the in\-batch candidates\. Stage 1 sets architecture, batch size, and the\(N,D\)\(N,D\)allocation; Stage 2 isolates the negative\-pool sizeKKat a fixed Stage 1 backbone\. This split is what makes re\-embedding a million\-item catalogue tractable at serving time and defines the two evaluation regimes below\.

#### Data\.

All sweeps train on an anonymized real\-world retail interaction corpus that combines offline and online consumer activity \(product searches, views, clicks and purchases\), chunked into sequences ofLseq=256L\_\{\\mathrm\{seq\}\}\\\!=\\\!256events per training example\. Each event is multi\-modal: a free\-text description, categorical fields \(e\.g\. event type, merchant, device, etc\), numerical fields \(e\.g\. price, timestamp, etc\), and optional visual features\. The catalogue contains on the order of10810^\{8\}unique actions, and training consumes on the order of10910^\{9\}event tokens\. Item popularity is strongly heavy\-tailed: a small head of frequent actions is observed many times within a single training run while much of the long tail is observed only once or twice\. This asymmetry drives the embedder/contextualizer compute tradeoff \(§[3](https://arxiv.org/html/2606.05257#S3)\)\.

#### Training recipe\.

We train with AdamW, weight decay0\.10\.1,bf16mixed precision, and fully sharded data parallelism\. The architecture, allocation, and sampling experiments use cosine learning\-rate decay with55–10%10\\%linear warm\-up\. The batch\-size experiment uses a constant learning rate so that “updates to target” measures optimization efficiency rather than schedule shape\. Learning rate is selected per cell from the training loss, not from validation metrics\.

#### Compute accounting\.

We report training compute in standardized buckets and targetC≈6​N​DC\\\!\\approx\\\!6NDfollowing the language\-model scaling\-law convention, whereD=T​B​LseqD\\\!=\\\!TBL\_\{\\mathrm\{seq\}\}is the number of event tokens consumed byTToptimizer steps at global batch sizeBB\. The experiment generator uses the finer Kaplan\-style per\-step formula

Fstep=6​B​Lseq​\(t​pe\+p\+3​B​Lseq​h\),F\_\{\\text\{step\}\}\\;=\\;6\\,B\\,L\_\{\\mathrm\{seq\}\}\\,\\bigl\(t\\,p\_\{e\}\+p\+3\\,B\\,L\_\{\\mathrm\{seq\}\}\\,h\\bigr\),\(1\)wheret=24t\\\!=\\\!24is the embedder context length in events\. The hidden sizehhis not a global constant: it is set per architectural cell and determines both the event\-embedding dimension used by the contextualizer and the dimension at which in\-batch contrastive scoring is performed\. In the embedder\-share sweep, changing the target shares=pe/Ns\\\!=\\\!p\_\{e\}/Nchanges bothhhand the embedder parameter countpep\_\{e\}; the contextualizer parameter countppthen follows from the target totalNNat the fixedD/ND/Nratio\. In the depth sweep,hhandpep\_\{e\}are held fixed while contextualizer depth changespp, so the embedder contributes the same6​B​Lseq​t​pe6BL\_\{\\mathrm\{seq\}\}tp\_\{e\}compute tax across depth cells\. The term3​B​Lseq​h3BL\_\{\\mathrm\{seq\}\}his the in\-batch contrastive scoring term after candidate embeddings are all\-gathered across data\-parallel ranks\. We also useNeff≡p\+t​peN\_\{\\text\{eff\}\}\\\!\\equiv\\\!p\+tp\_\{e\}when interpreting the embedder/contextualizer tradeoff\.

#### Evaluation\.

The evaluation protocol intentionally separates the two regimes used by the training system\. Before the embedder is frozen, checkpoints are evaluated against the unique target items in each validation batch, matching the in\-batch negative distribution used during training\. After freezing the embedder, Stage 2 checkpoints are evaluated against the full cached deployed product catalogue, matching the deployed retrieval setting\. We report cross\-entropy, perplexity, recall@kk, NDCG@kk\[[23](https://arxiv.org/html/2606.05257#bib.bib23)\], MRR@10\[[24](https://arxiv.org/html/2606.05257#bib.bib24)\], coverage@kk\[[26](https://arxiv.org/html/2606.05257#bib.bib26)\], and predictive entropy; the standard ranking metrics followManning et al\. \[[25](https://arxiv.org/html/2606.05257#bib.bib25)\], and Appendix[A](https://arxiv.org/html/2606.05257#A1)gives the exact formulas as we compute them\. Throughout, “training loss” means the sampled\-softmax objective optimized by the model, while “validation loss” means the corresponding held\-out cross\-entropy under the evaluation regime being used\.

## 3Scaling the Event Embedder

### 3\.1Embedder Share

The first question is architectural: at fixed compute, how much capacity should belong to the feature\-based event embedder rather than the transformer contextualizer? The answer is stable and surprisingly small\. Across four compute budgets, the best share is about2%2\\%of parameters\. This is not just an empirical accident: the embedder is more expensive per parameter and sees a more repetitive effective data distribution than the contextualizer\. The scale is also a useful point of comparison to multimodal foundation models, where a comparatively small modality encoder feeds a larger language model: LLaVA, Flamingo and BLIP\-2 use vision encoders on the order of a few to tens of percent of total parameters\[[8](https://arxiv.org/html/2606.05257#bib.bib8),[9](https://arxiv.org/html/2606.05257#bib.bib9),[10](https://arxiv.org/html/2606.05257#bib.bib10)\]\. Those systems have not, to our knowledge, been calibrated with the same encoder\-share scaling\-law sweep; checking whether similar share laws hold for multimodal encoders is natural future work\.

Setup\.We hold the data\-to\-parameter ratioD/N≈15D/N\\\!\\approx\\\!15\(Chinchilla\) fixed across the sweep, so the width study asks*where on the embedder/contextualizer split*the loss bottoms out given that we are Chinchilla\-matched\. This complements the model/data allocation sweep of §[5](https://arxiv.org/html/2606.05257#S5), which instead asks where along the iso\-FLOP frontier to sit at eachCC\. Each\(C,s\)\(C,s\)cell solves the Kaplan formula \([1](https://arxiv.org/html/2606.05257#S2.E1)\) for\(N,D\)\(N,D\)at that ratio and sizes the joint embedding dimhh\(which also setspep\_\{e\}via the fixed embedder architecture\) to hit the target shares=pe/Ns\\\!=\\\!p\_\{e\}/N; the realizedN/DN/Dproxy stays essentially flat acrossssat each budget \(Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px2)\), so the sweep slides along the fixed\-ratio line rather than drifting off it\. We sweep embedder sharessfrom0%0\\%to50%50\\%\(finer spacing below6%6\\%\) at four compute budgets and three learning rates per cell, select the best LR on training loss without using the validation set for that choice, and fit each per\-budget share\-loss curve with the two\-term iso\-FLOP form

sign⋅y​\(s\)=a​sα⏟contextualizerstarvation\+b​s−β⏟embedderstarvation\+E,\\mathrm\{sign\}\\cdot y\(s\)\\;=\\;\\underbrace\{a\\,s^\{\\alpha\}\}\_\{\\begin\{subarray\}\{c\}\\text\{contextualizer\}\\\\ \\text\{starvation\}\\end\{subarray\}\}\\;\+\\;\\underbrace\{b\\,s^\{\-\\beta\}\}\_\{\\begin\{subarray\}\{c\}\\text\{embedder\}\\\\ \\text\{starvation\}\\end\{subarray\}\}\\;\+\\;E,\(2\)with closed\-form optimums⋆=\(b​β/\(a​α\)\)1/\(α\+β\)s^\{\\star\}\\\!=\\\!\(b\\beta/\(a\\alpha\)\)^\{1/\(\\alpha\+\\beta\)\}andsign=\+1\\mathrm\{sign\}\\\!=\\\!\+1for smaller\-better metrics\. The smallest swept cell carries a nominal target share of0%0\\%, but a minimum embedder width \(the text encoder is never removed\) means it realizess≈0\.5%s\\\!\\approx\\\!0\.5\\%; we therefore fit and plot it at its realizedsfloor≈0\.5%s\_\{\\text\{floor\}\}\\\!\\approx\\\!0\.5\\%rather than at the nominal0%0\\%\.

Figure[1](https://arxiv.org/html/2606.05257#S3.F1)summarizes every headline metric vs\. target embedder sharess, with the per\-\(metric, budget\) two\-term fits of \([2](https://arxiv.org/html/2606.05257#S3.E2)\) overlaid and the analytic optimums⋆s^\{\\star\}marked per panel; aval\_loss\-only zoom on the small\-share band is Figure[10](https://arxiv.org/html/2606.05257#A2.F10)\. The Kaplan FLOP\-share cross\-check \(same cells re\-plotted vs\. the realized embedder\-side FLOP sharefffrom \([1](https://arxiv.org/html/2606.05257#S2.E1)\)\) and the justification for parameterizing inssrather thanffare in Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px2)\. The reason is thatsslinearly trades off the two parameter pools, whileffis a nonlinear, cell\-dependent pushforward that breaks the two\-term fit’s conditioning\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/share_sweep_all_metrics.png)Figure 1:Width sweep: every headline eval metric vs\. target embedder parameter sharess\.One curve per compute budget; solid lines are the per\-\(metric, budget\) two\-term starvation fits of \([2](https://arxiv.org/html/2606.05257#S3.E2)\)\.⋆\\star= analytics⋆s^\{\\star\}inside the swept ranges∈\[1%,50%\]s\\\!\\in\\\!\[1\\%,50\\%\];\+\+= boundary case \(s⋆s^\{\\star\}extrapolated outside the sweep, almost alwayscoverage@10at largeCC\)\. The Kaplan\-FLOP\-share view of the same cells is Appendix Figure[11](https://arxiv.org/html/2606.05257#A2.F11); aval\_loss\-only zoom on the small\-share band is Figure[10](https://arxiv.org/html/2606.05257#A2.F10)\.Findings\.The share–loss relationship from0%0\\%to50%50\\%is clean: loss is essentially flat in the small\-share bands∈\[2,6\]%s\\\!\\in\\\!\[2,6\]\\%\(Δ≤0\.06\\Delta\\\!\\leq\\\!0\.06nat at every budget\); monotonically increasing above6%6\\%, where going from6%6\\%to27%27\\%costs∼0\.40\\sim\\\!0\.40nat atC=1015C\\\!=\\\!10^\{15\}, falling to∼0\.09\\sim\\\!0\.09nat atC=1018C\\\!=\\\!10^\{18\}; and shallow below2%2\\%\. The two\-term fitL​\(s\)=E\+a​sα\+b​s−βL\(s\)\\\!=\\\!E\\\!\+\\\!a\\,s^\{\\alpha\}\\\!\+\\\!b\\,s^\{\-\\beta\}collapses every budget into a single closed\-form optimums⋆∈\[1\.1%,3\.7%\]s^\{\\star\}\\\!\\in\\\!\[1\.1\\%,\\,3\.7\\%\]:*the optimal embedder share is constant ats⋆≈2s^\{\\star\}\\\!\\approx\\\!2–3%3\\%across all four budgets*\(val\_loss fitss⋆∝C\+0\.07s^\{\\star\}\\\!\\propto\\\!C^\{\+0\.07\}, slope indistinguishable from zero given a0\.5%0\.5\\%swept\-share resolution\)\. Repeating the analysis for every other headline eval metric agrees within1​σ1\\sigma:s⋆∝C±0\.10s^\{\\star\}\\\!\\propto\\\!C^\{\\pm 0\.10\}across val\_entropy,recall@10,NDCG@10,NDCG@100, all clustering inside the\[1%,6%\]\[1\\%,\\,6\\%\]band \(Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px2), Figure[2](https://arxiv.org/html/2606.05257#S3.F2)\);recall@100drifts upward most, withs⋆∈\[1\.4%,4\.1%\]s^\{\\star\}\\\!\\in\\\!\[1\.4\\%,\\,4\.1\\%\]and ans⋆∝C\+0\.11s^\{\\star\}\\\!\\propto\\\!C^\{\+0\.11\}trend, since a100100\-deep recommendation list tolerates more capacity in the catalogue representation\.

Table 2:Analytics⋆s^\{\\star\}from the two\-term starvation fit \([2](https://arxiv.org/html/2606.05257#S3.E2)\), per budget and per metric\.All cells fall inside the swept ranges∈\[1%,50%\]s\\\!\\in\\\!\[1\\%,50\\%\]\(bold\) except for three boundary cases \(italic “boundary”\), where the fit’s closed\-form optimum falls outside the sweep because the curve is flat enough that the analytic minimum is not bracketed\.β\\betais the embedder\-starvation exponent \(penalty fors→0s\\\!\\to\\\!0\);α\\alphais the contextualizer\-starvation exponent \(penalty for over\-allocation to the embedder\)\. Both exponents are reported only forval\_loss; full per\-metric exponents are in Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px2)\. Figure[2](https://arxiv.org/html/2606.05257#S3.F2)is the visual summary of thes⋆s^\{\\star\}column across budgets and metrics; the Kaplan\-FLOP\-share cross\-check is Appendix Figure[11](https://arxiv.org/html/2606.05257#A2.F11)\.Analytics⋆s^\{\\star\}\(%\) \(discrete\-grid argmax in parentheses\)val\_loss exponentsBudgetval\_lossrecall@10NDCG@10coverage@10α\\alphaβ\\beta101510^\{15\}1\.11\.1\(22\)1\.11\.1\(22\)*bound\.*\(22\)1\.31\.3\(0\)0\.400\.401\.031\.03101610^\{16\}3\.73\.7\(22\)3\.03\.0\(22\)2\.82\.8\(22\)3\.53\.5\(0\)0\.940\.940\.260\.26101710^\{17\}2\.02\.0\(22\)2\.12\.1\(22\)1\.91\.9\(22\)*bound\.*\(0\)0\.750\.751\.651\.65101810^\{18\}2\.42\.4\(55\)2\.32\.3\(22\)1\.81\.8\(22\)*bound\.*\(0\)0\.480\.481\.061\.06![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/share_sweep_sstar_vs_C.png)Figure 2:s⋆​\(C\)s^\{\\star\}\(C\)summary, per metric\.Visual companion to Table[2](https://arxiv.org/html/2606.05257#S3.T2): stars are the analytics⋆s^\{\\star\}from the per\-\(metric, budget\) two\-term fit when it lands inside the swept ranges∈\[1%,50%\]s\\\!\\in\\\!\[1\\%,50\\%\]; faded\+\+markers are boundary\- extrapolateds⋆s^\{\\star\}and are excluded from the dashed power\-law fits\. Every loss / ranking metric fitss⋆∝Cρs^\{\\star\}\\\!\\propto\\\!C^\{\\rho\}with\|ρ\|≤0\.11\|\\rho\|\\\!\\leq\\\!0\.11, effectively flat over four decades of compute\.recall@100drifts upward most strongly \(ρ=\+0\.11\\rho\\\!=\\\!\+0\.11\),NDCG@10downward most strongly \(ρ=−0\.10\\rho\\\!=\\\!\-0\.10\); the gray band is the practitioner\-recipes∈\[1%,6%\]s\\\!\\in\\\!\[1\\%,6\\%\]that contains every interiors⋆s^\{\\star\}in Table[2](https://arxiv.org/html/2606.05257#S3.T2)\.Why the embedder is so small\.Two asymmetries between the embedder and the contextualizer push the compute\-optimal split into the small\-ssband\.*\(i\) Compute asymmetry*: the contextualizer runs once perB×LB\\\!\\times\\\!Lbatch, while the text encoder inside the embedder runs onB​LB\\,Lshort sequences per step\. Each embedder parameter is therefore touched∼t\\sim\\\!ttimes more often per item than each contextualizer parameter, which is exactly theNeff=p\+t​peN\_\{\\text\{eff\}\}\\\!=\\\!p\\\!\+\\\!t\\,p\_\{e\}correction of §[2](https://arxiv.org/html/2606.05257#S2)\.*\(ii\) Effective\-epoch asymmetry*: the contextualizer effectively never sees the sameLL\-event window twice in our sub\-one\-epoch schedule, while the embedder sees every popular item hundreds to thousands of times\. The embedder is thus much more exposed to memorization than the contextualizer under the same global regularization, and thes⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%recommendation should be read as the answer for the current \(essentially unregularized\) recipe\. Adding embedder\-specific regularization may shifts⋆s^\{\\star\}upward; we treat this as future work\.

### 3\.2Depth as a Secondary Knob

The depth sweep is a complementary check on the width result: once embedder share is already small, how should the remaining capacity split between a shallow text encoder and a deeper contextualizer at fixed compute? We tradeLtextL\_\{\\text\{text\}\}againstLctxL\_\{\\text\{ctx\}\}on an iso\-FLOP depth\-sum constraint \(Lctx\+LtextL\_\{\\text\{ctx\}\}\\\!\+\\\!L\_\{\\text\{text\}\}fixed per budget\) and pick the best learning rate per\(Lctx,Ltext\)\(L\_\{\\text\{ctx\}\},L\_\{\\text\{text\}\}\)cell on training loss, matching the width\-sweep protocol\. Per\-cell grids and per\-metric curves are in Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px5)\.

Findings\.We extract per\-budget optima as the analytic minimum of an iso\-FLOP parabola inlog⁡s\\log s, following the model/data allocation sweep’s treatment of\(N,D\)\(N,D\)allocation \(§[5](https://arxiv.org/html/2606.05257#S5)\)\. The parabolic fit recoverss⋆=1\.82%s^\{\\star\}\\\!=\\\!1\.82\\%atC=1016C\\\!=\\\!10^\{16\}ands⋆=1\.80%s^\{\\star\}\\\!=\\\!1\.80\\%atC=1017C\\\!=\\\!10^\{17\}, both within0\.2%0\.2\\%of the embedder\-share recommendations⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%\. AtC=1018C\\\!=\\\!10^\{18\}the parabolics⋆=0\.80%s^\{\\star\}\\\!=\\\!0\.80\\%, but the grid saturates atLctx=32L\_\{\\text\{ctx\}\}\\\!=\\\!32so the true minimum may sit slightly deeper\. AtC=1015C\\\!=\\\!10^\{15\}the val\_loss landscape acrossLctx∈\[1,12\]L\_\{\\text\{ctx\}\}\\\!\\in\\\!\[1,12\]is too flat \(≤0\.21\\leq\\\!0\.21nats end\-to\-end\) for the parabola to resolve an interior minimum \(Appendix Table[10](https://arxiv.org/html/2606.05257#A2.T10), Figure[13](https://arxiv.org/html/2606.05257#A2.F13)\); the underlying depth\-sum sweep coversLctx⋆∈\{12,6,4,32\}L\_\{\\text\{ctx\}\}^\{\\star\}\\\!\\in\\\!\\\{12,6,4,32\\\}across the four budgets \(Appendix Table[9](https://arxiv.org/html/2606.05257#A2.T9)\), but the small\-ssvalley is flat enough that pickings=2%s\\\!=\\\!2\\%instead of the per\-budget optimum costs≤0\.01\\leq\\\!0\.01nats under the parabolic projection atC∈\{1016,1017,1018\}C\\\!\\in\\\!\\\{10^\{16\},10^\{17\},10^\{18\}\\\}, and ranking metrics at the small\-sscells agree withval\_loss\(Appendix Figure[14](https://arxiv.org/html/2606.05257#A2.F14)\)\. We therefore treat depth as a second\-order knob: any configuration in the small\-ssvalley is acceptable onces≈2%s\\\!\\approx\\\!2\\%is set\. We keeps⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%from the width sweep as the primary architectural recommendation\. AtC=1018C\\\!=\\\!10^\{18\}deeper extensions off the iso\-FLOP diagonal do not beat the depth\-sum\-3434winner on held\-out metrics \(Appendix Table[12](https://arxiv.org/html/2606.05257#A2.T12)\)\.

## 4Critical Batch Size Across Metrics

Batch size is usually treated as an optimization detail, but in a scaling\-law recipe it determines how much data efficiency is traded for hardware throughput\. We therefore ask how large the global batch can grow before extra examples stop buying proportional progress\. At the fixed architecture selected above, the answer depends on the metric: loss\-like metrics and recall have a critical batch near∼570\\sim\\\!570, while top\-weighted ranking metrics saturate at roughly one third of that value\.

Setup\.We train atB∈\{64,128,256,512,1024,2048\}B\\\!\\in\\\!\\\{64,128,256,512,1024,2048\\\}with square\-root LR scaling and a constant LR schedule \(so updates\-to\-target reflects optimization efficiency, not schedule shape\), with periodic held\-out evaluation during training\. We then read offSm​\(B\)S\_\{m\}\(B\), the smallest update at which the EWMA\-smoothed metric first crosses a per\-metric iso\-targetTmT\_\{m\}, and fit the Kaplan/McCandlish\[[4](https://arxiv.org/html/2606.05257#bib.bib4)\]critical\-batch model per metric:

Sm​\(B\)=Smmin​\(1\+Bcrit\(m\)B\)\.S\_\{m\}\(B\)\\;=\\;S^\{\\min\}\_\{m\}\\\!\\left\(1\+\\frac\{B\_\{\\mathrm\{crit\}\}^\{\(m\)\}\}\{B\}\\right\)\.\(3\)Smoothing, iso\-target derivation, and the per\-metric trajectories thatSm​\(B\)S\_\{m\}\(B\)is read off of are in Appendix[C](https://arxiv.org/html/2606.05257#A3)\.

Table 3:Per\-metric critical batch size\.Updates\-to\-targetSm​\(B\)S\_\{m\}\(B\)on the EWMA\-smoothed trajectory at the per\-metric iso\-targetTmT\_\{m\}, and the Kaplan fit \([3](https://arxiv.org/html/2606.05257#S4.E3)\)\. Validation loss andrecall@10share aBcritB\_\{\\mathrm\{crit\}\}near∼550\\sim\\\!550; position\-weighted ranking metrics sit substantially lower \(∼200\\sim\\\!200–275275\)\.Sm​\(2048\)≥Sm​\(1024\)S\_\{m\}\(2048\)\\\!\\geq\\\!S\_\{m\}\(1024\)for every metric exceptval\_entropy\.∗val\_entropydecreases monotonically inBB\(S=7650S\\\!=\\\!7650atB=64B\\\!=\\\!64to300300atB=2048B\\\!=\\\!2048\) without a Kaplan plateau;BcritB\_\{\\mathrm\{crit\}\}is poorly constrained and we omit a point estimate \(likely above20482048\)\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/bcrit_per_metric_v3.png)Figure 3:Per\-metric Kaplan fitsof \([3](https://arxiv.org/html/2606.05257#S4.E3)\); dotted lines markBcrit\(m\)B\_\{\\mathrm\{crit\}\}^\{\(m\)\}andSmminS^\{\\min\}\_\{m\}per panel\.val\_lossandrecall@10coincide \(Bcrit≈574B\_\{\\mathrm\{crit\}\}\\\!\\approx\\\!574and544544\);NDCG@10andMRR@10sit notably lower;val\_entropyhas no plateau within ourBBrange \(Tab\.[3](https://arxiv.org/html/2606.05257#S4.T3)footnote\)\. EWMA\-smoothed trajectories and iso\-target crossings are in Appendix[C](https://arxiv.org/html/2606.05257#A3)\.Findings\.*\(i\) Loss and recall share a knee\.*val\_lossandrecall@10have essentially the sameBcrit≈570B\_\{\\mathrm\{crit\}\}\\\!\\approx\\\!570; the batch\-size knee for cross\-entropy and for the dominant retrieval metric coincide\.*\(ii\) Position\-weighted ranking saturates earlier\.*NDCG@10andMRR@10sit atBcrit≈200B\_\{\\mathrm\{crit\}\}\\\!\\approx\\\!200–275275, andval\_entropyhas no plateau within the swept range\. This strict ordering matches a simple sensitivity argument: if a downstream metricM=f​\(L\)M\\\!=\\\!f\(L\)has small\|f′​\(L\)\|\|f^\{\\prime\}\(L\)\|in the operating band, the gradient\-noise reduction that largerBBbuys translates into a proportionally smaller improvement inMMand the Kaplan curvature regime is reached at a smallerBB\(top\-weighted ranking saturates quickly once the head is approximately correct; entropy is at the high\-sensitivity extreme\)\.*\(iii\)B=2048B\\\!=\\\!2048sits just past the val\-loss knee\.*Sm​\(2048\)≥Sm​\(1024\)S\_\{m\}\(2048\)\\\!\\geq\\\!S\_\{m\}\(1024\)for every non\-entropy metric, the predicted data\-efficiency penalty\. This is the largest single\-node batch we measured without GPU\-efficiency loss, so we recommend it for throughput\-limited production training \(§[9](https://arxiv.org/html/2606.05257#S9)\); the smaller batches used in the model/data allocation sweep are picked for iso\-FLOP feasibility on the allocation grid \(§[5](https://arxiv.org/html/2606.05257#S5)\)\.

## 5Model/Data Allocation Across Compute Budgets

We next ask the Chinchilla question for behavioral models: given a fixed compute budget, should we train a larger model on fewer event tokens, or a smaller model on more data? The optimum is more data\-heavy than the text\-LM rule at small budgets, but it moves steadily toward the Chinchilla heuristic as compute grows\. Across metrics, the fitted model\-size exponents are close, even when the exact best cell at a particular budget changes\.

Setup\.We train4545contextualizers on the primary fixed\-architecture grid\(h,Lctx\)∈\{\(128,4\),\(192,4\),\(256,6\),\(320,6\),\(384,6\),\(512,8\)\}\(h,L\_\{\\text\{ctx\}\}\)\\\!\\in\\\!\\\{\(128,4\),\(192,4\),\(256,6\),\(320,6\),\(384,6\),\(512,8\)\\\}plus six additional larger anchors atC=1019C\\\!=\\\!10^\{19\}\(h∈\{768,…,1408\}h\\\!\\in\\\!\\\{768,\\dots,1408\\\},Lctx∈\[12,20\]L\_\{\\text\{ctx\}\}\\\!\\in\\\!\[12,20\]\), all with cosine learning\-rate decay and every\(h,Lctx,LR\)\(h,L\_\{\\text\{ctx\}\},\\mathrm\{LR\}\)cell trained from scratch\. Every cell uses embedder shares=2%s\\\!=\\\!2\\%\(§[3\.1](https://arxiv.org/html/2606.05257#S3.SS1)\)\. The global training batch is budget\-scaled \(B∈\{64,128,256,512,512\}B\\\!\\in\\\!\\\{64,128,256,512,512\\\}atC∈\{1015,1016,1017,1018,1019\}C\\\!\\in\\\!\\\{10^\{15\},10^\{16\},10^\{17\},10^\{18\},10^\{19\}\\\}so every iso\-FLOP cell runs with feasible GPU throughput; the larger\-BBcase for throughput\-oriented training is in §[4](https://arxiv.org/html/2606.05257#S4)\. We takeval\_losson the held\-out set \(batch\-local pool, fixedeval\_batchper budget\) as the primary objective, since it aligns with the headline ranking metrics under the same protocol and is less optimizer\-noisy than the tail\-100100training surrogate \(the train\-surrogate parallel allocation table is Appendix[D](https://arxiv.org/html/2606.05257#A4)\)\. Table[4](https://arxiv.org/html/2606.05257#S5.T4)lists, per budget, theval\_loss\-best grid cell alongside the Hoffmann Approach 2 parabolic optimum inlog⁡N\\log Nand the closest swept\(h/Lctx\)\(h/L\_\{\\text\{ctx\}\}\)anchor\.

Table 4:Val\-loss\-optimal allocation per budget\.*Grid winner:*N⋆N^\{\\star\},D⋆D^\{\\star\}, andD/ND/Nat the run minimizingval\_losson the iso\-FLOP sweep, with observed val\_loss, R@10, and NDCG@10\.*Parabolic \(Approach 2\):*analytic minimum of the quadratic\-in\-log⁡N\\log Nfit per budget \(Fig\.[4](https://arxiv.org/html/2606.05257#S5.F4),val\_losspanel\);D⋆=C/\(6​N⋆\)D^\{\\star\}\\\!=\\\!C/\(6N^\{\\star\}\)\. The negative\-sampling sweep sizes contextualizers from the grid\-winnerN⋆N^\{\\star\}column\. Train\-surrogate minima differ slightly \(Appendix[D](https://arxiv.org/html/2606.05257#A4)\)\.Power\-law fits \(val optima, parabolic\)\.FittingN⋆​\(C\)=a​CbN^\{\\star\}\(C\)\\\!=\\\!aC^\{b\}andD⋆​\(C\)=a′​Cb′D^\{\\star\}\(C\)\\\!=\\\!a^\{\\prime\}C^\{b^\{\\prime\}\}in log\-space on the five per\-budget parabolic minima \(Fig\.[4](https://arxiv.org/html/2606.05257#S5.F4),val\_losspanel; Hoffmann Approach 2\) gives

N⋆​\(C\)\\displaystyle N^\{\\star\}\(C\)=3\.35×10−4​C0\.617±0\.025,\\displaystyle\\;=\\;3\.35\\\!\\times\\\!10^\{\-4\}\\,C^\{0\.617\\pm 0\.025\},\(4\)D⋆​\(C\)\\displaystyle D^\{\\star\}\(C\)=4\.97×102​C0\.383±0\.025,\\displaystyle\\;=\\;4\.97\\\!\\times\\\!10^\{2\}\\,C^\{0\.383\\pm 0\.025\},\(5\)withb\+b′=1b\\\!\+\\\!b^\{\\prime\}\\\!=\\\!1exactly by Approach 2 construction\. The allocation is moderately parameter\-heavy \(bN\>bDb\_\{N\}\\\!\>\\\!b\_\{D\}\); the train\-surrogate parallel \(bN=0\.612±0\.024b\_\{N\}\\\!=\\\!0\.612\\\!\\pm\\\!0\.024\) lands within0\.0050\.005of the val\-surrogate exponent \(Appendix[D](https://arxiv.org/html/2606.05257#A4)\)\.

Data\-heavy relative to text LMs, narrowing toward Chinchilla\.Val\-optimal points giveD⋆/N⋆D^\{\\star\}/N^\{\\star\}that decreases monotonically inCC:344→265→222→110→36344\\\!\\to\\\!265\\\!\\to\\\!222\\\!\\to\\\!110\\\!\\to\\\!36acrossC=1015→1019C\\\!=\\\!10^\{15\}\\\!\\to\\\!10^\{19\}\. Over the lower four budgets we sit nearly an order of magnitude above the text\-LM Chinchilla heuristicD/N≈20D/N\\\!\\approx\\\!20; theC=1019C\\\!=\\\!10^\{19\}anchor brings the optimum to within∼2×\\sim\\\!2\\\!\\timesof Chinchilla and extrapolates the trajectory toward the text\-LM heuristic at production scale\.

Per\-metric allocation laws are metric\-robust\.RefittingN⋆∝CaNN^\{\\star\}\\\!\\propto\\\!C^\{a\_\{N\}\}on each metric’s analytic parabola optimum \(Table[5](https://arxiv.org/html/2606.05257#S5.T5)\) lands all five headline metrics in the tight bandaN∈\[0\.57,0\.66\]a\_\{N\}\\\!\\in\\\!\[0\.57,\\,0\.66\]\(0\.617±0\.0250\.617\\\!\\pm\\\!0\.025val\_loss,0\.616±0\.0290\.616\\\!\\pm\\\!0\.029recall@10,0\.586±0\.0330\.586\\\!\\pm\\\!0\.033NDCG@10,0\.574±0\.0790\.574\\\!\\pm\\\!0\.079coverage@10\); theval\_loss\-vs\-NDCG@10separationΔ​aN≈0\.03\\Delta a\_\{N\}\\\!\\approx\\\!0\.03is inside its own slope uncertainty\. The per\-budget winners do still disagree across metrics on which\(h,L\)\(h,L\)cell to ship: the scaling law is shared, but the recipe is not\. §[7](https://arxiv.org/html/2606.05257#S7)consolidates this\. Figure[4](https://arxiv.org/html/2606.05257#S5.F4)plots the iso\-FLOP curves and parabolic fits per metric; Figure[5](https://arxiv.org/html/2606.05257#S5.F5)shows the implied frontiers in\(N⋆,D⋆,D/N\)\(N^\{\\star\},D^\{\\star\},D/N\)\-space\.

Table 5:Hoffmann\-style exponents under each eval metric\.Per\-budget optima are taken as the analytic minimum \(or maximum, for ranking metrics\) of the iso\-FLOP parabola inlog⁡N\\log N\(Fig\.[4](https://arxiv.org/html/2606.05257#S5.F4)\) on the primary grid plus sixC=1019C\\\!=\\\!10^\{19\}anchors\. Slope errors are the 1\-sigma OLS uncertainties of the five\-budget log\-log power\-law fit;aN\+aDa\_\{N\}\\\!\+\\\!a\_\{D\}is exact at11by construction sinceD⋆=C/\(6​N⋆\)D^\{\\star\}\\\!=\\\!C/\(6N^\{\\star\}\)\. Cell columns list the discrete\(h/Lctx\)\(h/L\_\{\\text\{ctx\}\}\)closest to each parabola optimum\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/chinchilla_isoflop_grid.png)Figure 4:Iso\-FLOP curves per metric, with parabolic fits\.Each panel scatters one headline metric vs\.NNalong each compute budget and overlays the quadratic\-in\-log⁡N\\log Nfit per budget;⋆\\starsits at the parabola’s analytic optimum \(minimum for losses/entropy, maximum for ranking metrics\)\. The plus\-marker variant is used at\(coverage@10,C=1018\)\(\\textsc\{coverage@10\},C\\\!=\\\!10^\{18\}\)where the coverage curve is essentially flat \(a≈0a\\\!\\approx\\\!0\) and the parabola maximum falls just outside the sweptNNrange\. These analytic optima feed the\(h/Lctx\)\(h/L\_\{\\text\{ctx\}\}\)closest\-cell column of Tab\.[5](https://arxiv.org/html/2606.05257#S5.T5)and the per\-metric power\-law fits in Fig\.[5](https://arxiv.org/html/2606.05257#S5.F5)\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/chinchilla_frontiers.png)Figure 5:Chinchilla frontier per metric\.Left:N⋆​\(C\)N^\{\\star\}\(C\); middle:D⋆​\(C\)D^\{\\star\}\(C\); right:D⋆/N⋆D^\{\\star\}/N^\{\\star\}\. Switching the target fromval\_loss/recall@10toNDCG@10pushes the optimum toward smaller models trained on more tokens at every budget\.
## 6Scaling the Negative Candidate Pool

After the embedder is frozen, the cached catalogue makes it possible to train against many more sampled negatives\. This creates a new scaling knob: how large should the extra negative poolKKbe at fixed compute? The useful range is broad but not arbitrary\. Smooth iso\-FLOP fits place most metric\-specific optima in the low10510^\{5\}–10610^\{6\}range; at the largest budget, the limiting constraint is no longer FLOPs but candidate\-axis memory\.

Setup\.AllK⋆K^\{\\star\}claims below are read from*full\-catalogue*evaluation \(the deployed full\-catalogue candidate set; §[2](https://arxiv.org/html/2606.05257#S2)\), not from the training sampled\-softmax whose candidate countVsoftmax=16,384\+KV\_\{\\text\{softmax\}\}\\\!=\\\!16\{,\}384\\\!\+\\\!Kchanges withKK\. We sweepK∈\{0,16​k,32​k,64​k,131​k,262​k,524​k,1​M,2​M\}K\\\!\\in\\\!\\\{0,16\\mathrm\{k\},32\\mathrm\{k\},64\\mathrm\{k\},131\\mathrm\{k\},262\\mathrm\{k\},524\\mathrm\{k\},1\\mathrm\{M\},2\\mathrm\{M\}\\\}at five compute budgets \(102102training runs,123123full\-catalogue evals\), with the contextualizer sized at the model/data grid\-winnerN⋆N^\{\\star\}for each budget \(Table[4](https://arxiv.org/html/2606.05257#S5.T4), grid\-winner column\)\. We model the iso\-FLOP curve as a sum of two opposing terms plus an irreducible floor,

y​\(K\)=a​Kα⏟starvation\+b​K−β⏟sampling bias\+E,y\(K\)\\;=\\;\\underbrace\{a\\,K^\{\\alpha\}\}\_\{\\text\{starvation\}\}\\;\+\\;\\underbrace\{b\\,K^\{\-\\beta\}\}\_\{\\text\{sampling bias\}\}\\;\+\\;E,\(6\)with closed\-form optimumK⋆=\(b​β/\(a​α\)\)1/\(α\+β\)K^\{\\star\}\\\!=\\\!\(b\\beta/\(a\\alpha\)\)^\{1/\(\\alpha\+\\beta\)\}\(the starvation term is the linear\-in\-KKper\-step cost shrinking the available step countTTat fixedCC; the sampling\-bias term is the∝1/K\\propto\\\!1/Kvariance of the sampled\-softmax partition\-function estimator\)\. In\-batch vs extra\-negative channel decomposition and per\-metric bias\-exponent values are in Appendix[E](https://arxiv.org/html/2606.05257#A5)\.

Table 6:AnalyticK⋆K^\{\\star\}from the fitted starvation/bias model \([6](https://arxiv.org/html/2606.05257#S6.E6)\), per budget and per criterion\.The eval curves are nearly flat inlog⁡K\\log K, so the analytic optima carry wide CIs; even so,K⋆K^\{\\star\}for every metric clusters around∼105\\sim\\\!10^\{5\}–10610^\{6\}\. Bold cells are inside the swept rangeK∈\[16​k,2​M\]K\\\!\\in\\\!\[16\\mathrm\{k\},2\\mathrm\{M\}\]; italic “boundary” cells indicate the extrapolatedK⋆K^\{\\star\}falls outside the sweep so we report the right edge; values in parentheses are the discrete\-grid winners\.β\\betais the per\-metric bias\-decay exponent \(sampling\-bias regime; see §[6](https://arxiv.org/html/2606.05257#S6)\);α\\alphais the \(shared\) starvation exponent\.†AtC=1015C\\\!=\\\!10^\{15\}the recall@10 fit hitsβ=2\\beta\\\!=\\\!2\(upper bound\) because the swept\-range curve is monotone; no interiorK⋆K^\{\\star\}\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/k_sweep_all_metrics.png)Figure 6:Negative\-samplingKK\-sweep, per headline metric, with the per\-\(metric,CC\) starvation/bias fit \([6](https://arxiv.org/html/2606.05257#S6.E6)\) overlaid\.⋆\\star: analyticK⋆K^\{\\star\}inside the swept rangeK∈\[16​k,2​M\]K\\\!\\in\\\!\[16\\mathrm\{k\},2\\mathrm\{M\}\];\+\+:K⋆K^\{\\star\}extrapolated outside the sweep \(boundary cases in Table[6](https://arxiv.org/html/2606.05257#S6.T6);1313of2525metric×\\timesbudget cells\)\. Three things are visible directly off the parabolas: \(a\) atC≤1017C\\\!\\leq\\\!10^\{17\}the ranking panels are flat enough acrossK∈\[105,106\]K\\\!\\in\\\!\[10^\{5\},10^\{6\}\]that the discrete\-grid argmax moves nearly a decade between adjacent cells on noise alone, motivating the smooth\-fit estimator of \(i\) below; \(b\)val\_lossandval\_entropykeep falling all the way toK=2K\\\!=\\\!2M atC∈\{1016,1018,1019\}C\\\!\\in\\\!\\\{10^\{16\},10^\{18\},10^\{19\}\\\}, the right\-edge boundary cases of Table[6](https://arxiv.org/html/2606.05257#S6.T6); and \(c\) atC=1019C\\\!=\\\!10^\{19\}every panel is still trending atK=2K\\\!=\\\!2M, consistent with the memory\-bound flip in finding \(iii\)\. The compactK⋆​\(C\)K^\{\\star\}\(C\)summary on the interior cells is Figure[17](https://arxiv.org/html/2606.05257#A5.F17)in Appendix[E](https://arxiv.org/html/2606.05257#A5)\.Findings\.The main result is that the “right”KKis metric\-dependent\. On the discrete grid, full\-catalogueval\_lossandrecall@10can select very different points: atC=1017C\\\!=\\\!10^\{17\}the grid winners are11M and6565k negatives, respectively\. Reading from the smooth fit makes the disagreement much smaller \(400400k vs\.250250k\), but it does not remove it\. Loss\-like metrics still prefer more negatives than ranking metrics whenever both have interior optima \(Table[6](https://arxiv.org/html/2606.05257#S6.T6)\)\.

The actionable range is narrower than the raw grid suggests\. Across the interior fits, the ranking\-metric optima lie inK⋆∈\[125​k,870​k\]K^\{\\star\}\\\!\\in\\\!\[125\\text\{k\},\\,870\\text\{k\}\]\. The fitted cross\-budget slope is weak \(K⋆∝C0\.08K^\{\\star\}\\\!\\propto\\\!C^\{0\.08\}–C0\.15C^\{0\.15\}\), and many curves are shallow inlog⁡K\\log K, so we treat the band rather than the slope as the practitioner summary\. The fitted bias exponent also separates loss from ranking:β≈1\.2\\beta\\\!\\approx\\\!1\.2forval\_lossversusβ∈\[0\.39,0\.56\]\\beta\\\!\\in\\\!\[0\.39,0\.56\]for ranking metrics \(Appendix[E](https://arxiv.org/html/2606.05257#A5)\)\.

At the largest budget,KKstops looking compute\-limited\. AtC=1019C\\\!=\\\!10^\{19\}, every headline metric is still improving at the largest value we trained \(K=2K\\\!=\\\!2M\), and the analytic optimum falls beyond the swept range for every metric\. The bottleneck is then the candidate\-axis softmax\-logit memory footprint, which grows linearly withKK, not the available FLOPs\. Pushing beyond this regime requires candidate\-axis checkpointing, candidate sharding, or a sampled/hierarchical approximation to the partition function\.

The resulting recipe is simple: useKKin the low hundreds of thousands when memory is not binding, cap near10610^\{6\}aroundC=1018C\\\!=\\\!10^\{18\}, and treatC≥1019C\\\!\\geq\\\!10^\{19\}as a memory\-engineering problem rather than a pure compute\-allocation problem \(Table[7](https://arxiv.org/html/2606.05257#S9.T7)\)\.

## 7Cross\-Metric and Cross\-Regime Evaluation

The experiments above point to a single methodological lesson: in behavioral foundation models, the evaluation metric is part of the scaling law\. Changing it can change the compute\-optimal recipe\. The optimizer sees a sampled\-softmax loss, while the deployed system serves a full\-catalogue ranking metric\. Those quantities are often correlated, but they do not always choose the same batch size, architecture cell, or negative\-sampling recipe\. This section gathers the evidence across the study: metric\-specific critical batch sizes, metric\-specificK⋆K^\{\\star\}, a compute\-dependent loss–ranking sign flip, and the asymmetry between batch\-local and full\-catalogue evaluation\. The analysis pools the408408validation evaluations across the architecture, allocation, and sampling experiments\.

Within\-stage correlations are tight\.Within either stage,val\_loss,val\_ppl,val\_entropyand the ranking metrics \(recall@kk,NDCG@kkfork∈\{1,5,10,20,50,100\}k\\\!\\in\\\!\\\{1,5,10,20,50,100\\\}\) are mutually correlated at\|ρS\|≥0\.94\|\\rho\_\{S\}\|\\\!\\geq\\\!0\.94\(Figure[7](https://arxiv.org/html/2606.05257#S7.F7)\)\. Coverage is the only metric that decouples meaningfully, and the loss–coverage link weakens monotonically with scale \(full table in Appendix[G](https://arxiv.org/html/2606.05257#A7)\): once catalogue coverage saturates,*which*cells happen to spread the head distribution furthest is essentially independent of which cells minimize loss\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/correlation_heatmap.png)\(a\)Stage 1 architecture/allocation sweeps \(n=285n\\\!=\\\!285\)\.
![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/correlation_heatmap_sampled.png)\(b\)Negative\-sampling sweep after freezing \(n=123n\\\!=\\\!123\)\.

Figure 7:Spearman rank correlations between headline metrics\.Within either stage the loss/perplexity/entropy/ranking metrics are essentially one quantity; coverage is the only metric that decouples meaningfully\.High correlations are not the whole story\.Two metrics withρS=−0\.99\\rho\_\{S\}\\\!=\\\!\-0\.99can still disagree about which cell wins each budget when the leaderboard is tightly bunched\. The architectural takeaways \(§[3](https://arxiv.org/html/2606.05257#S3)\) are metric\-robust: the optimal embedder share and depth coincide under loss, recall and NDCG\. The per\-budget winners and within\-regime correlations on every other axis are not\. The four axes \(a\)–\(d\) below itemize where\.

#### \(a\) Per\-metric critical batch size \(§[4](https://arxiv.org/html/2606.05257#S4)\)\.

The Kaplan fit \([3](https://arxiv.org/html/2606.05257#S4.E3)\) returnsBcrit≈200B\_\{\\mathrm\{crit\}\}\\\!\\approx\\\!200–275275for position\-weighted ranking \(NDCG@10,MRR@10\),∼570\\sim\\\!570forval\_lossandrecall@10, and no plateau within the swept range forval\_entropy: same model, same iso\-target machinery, four different knees \(Table[3](https://arxiv.org/html/2606.05257#S4.T3)\)\. A downstream metricM=f​\(L\)M\\\!=\\\!f\(L\)with small\|f′​\(L\)\|\|f^\{\\prime\}\(L\)\|in the operating band reaches its Kaplan curvature regime at smallerBB, which is the source of the spread \(§[4](https://arxiv.org/html/2606.05257#S4), finding \(ii\)\)\.

#### \(b\) Per\-metric Stage 2K⋆K^\{\\star\}\(§[6](https://arxiv.org/html/2606.05257#S6)\)\.

The analytic minima of the iso\-CCfit \([6](https://arxiv.org/html/2606.05257#S6.E6)\) disagree across metrics at every budget: atC=1017C\\\!=\\\!10^\{17\},Kloss⋆=400K^\{\\star\}\_\{\\text\{loss\}\}\\\!=\\\!400k vs\.Krecall@10⋆=250K^\{\\star\}\_\{\\text\{recall@10\}\}\\\!=\\\!250k, and across the four interior budgets every ranking\-metricK⋆K^\{\\star\}lies in\[125​k,870​k\]\[125\\mathrm\{k\},\\,870\\mathrm\{k\}\]\. The loss landscape inlog⁡K\\log Kis shallow at the larger budgets, so the analytic optima carry wide confidence intervals, but the loss\-vs\-ranking ordering is preserved at every budget: loss prefers more negatives than the ranking metrics where both have interior optima\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/loss_vs_ranking_kstar.png)Figure 8:Full\-catalogueval\_lossvs\.recall@10\.At fixedCC, each column is one iso\-FLOP budget; curves use the same deployed\-catalogue eval pool \(only trainingKKdiffers\)\. Red:val\_loss\(↓\\downarrow\); green:recall@10\(↑\\uparrow\); stars mark the swept\-grid argmax per metric\.
#### \(c\) Stage 2 loss\-vs\-ranking rank correlation flips sign with compute\.

In the Stage 2KK\-sweep, the Spearman correlation betweenval\_loss\(smaller\-better\) andrecall@10\(larger\-better\) runs fromρS=\+0\.93\\rho\_\{S\}\\\!=\\\!\+0\.93atC=1015C\\\!=\\\!10^\{15\}\(misaligned\) toρS=−1\.00\\rho\_\{S\}\\\!=\\\!\-1\.00atC=1018C\\\!=\\\!10^\{18\}\(perfectly aligned\) \(Table[17](https://arxiv.org/html/2606.05257#A7.T17)\)\. By contrast the Stage 1 loss\-vs\-ranking link is locked atρS≈−0\.98\\rho\_\{S\}\\\!\\approx\\\!\-0\.98across all budgets \(Table[16](https://arxiv.org/html/2606.05257#A7.T16)\)\. Figure[8](https://arxiv.org/html/2606.05257#S7.F8)shows the same effect directly on the full\-catalogue curves: the loss and recall winners can separate at low compute and realign at larger budgets\. The sign flip arises because the source of variation differs: Stage 1 varies*architecture*\(bigger model→\\toboth lower loss and higher recall, locked\), while Stage 2 variesKK\. More negatives sharpen the conditional but can eventually flatten or worsen ranking\. The alignment atC=1018C\\\!=\\\!10^\{18\}confirms that once popularity bias is removed \(large enoughKKfor the sampling distribution to approach uniform\), the two regimes agree on the ranking ofKK\-cells\.

Why training loss and ranking can disagree: mechanism for \(c\) and \(d\)\.At each step the model is scored against a candidate setℬ⊂𝒱\\mathcal\{B\}\\\!\\subset\\\!\\mathcal\{V\}drawn from a sampling distributionqqand minimizes the sampled cross\-entropy

ℒ​\(h,ℬ\)=−log⁡exp⁡⟨h,eθ​\(y\+\)⟩∑x∈ℬexp⁡⟨h,eθ​\(x\)⟩\.\\mathcal\{L\}\(h,\\mathcal\{B\}\)\\;=\\;\-\\log\\frac\{\\exp\\langle h,e\_\{\\theta\}\(y^\{\+\}\)\\rangle\}\{\\sum\_\{x\\in\\mathcal\{B\}\}\\exp\\langle h,e\_\{\\theta\}\(x\)\\rangle\}\.By a standard importance\-weighting argument\[[12](https://arxiv.org/html/2606.05257#bib.bib12)\], this is an unbiased estimator of the full softmax only if each candidate logit is corrected by−log⁡q​\(x\)\-\\log q\(x\); without that correction the unique minimizer satisfiesfθ⋆​\(y∣h\)∝p​\(y∣h\)/q​\(y\)f\_\{\\theta\}^\{\\star\}\(y\\\!\\mid\\\!h\)\\\!\\propto\\\!p\(y\\\!\\mid\\\!h\)/q\(y\)\. In Stage 1 each candidate is itself a target, soq∝pq\\\!\\propto\\\!pand the loss minimizer is uniform on the catalogue \(absolute popularity is unidentifiable from the in\-batch objective; only within\-batch rank order is learned\)\. In Stage 2 withKKuniform extras,qqapproaches uniform asK→∞K\\\!\\to\\\!\\inftyand the minimizer approaches the true conditional\. Eval metrics score against the full catalogue and depend on the marginalfθf\_\{\\theta\}, not just its within\-batch ranks; the1/q1/qcorrection the model never had to learn shows up directly in the loss\-to\-NDCG mapping, mechanistically producing both the sign\-flip of \(c\) and the cross\-regime asymmetry of \(d\)\.

#### \(d\) Batch\-local and full\-catalogue evaluation disagree mainly on loss\.

Batch\-local evaluation and full\-catalogue evaluation use different candidate sets, so their losses need not rank checkpoints the same way\. This is most visible in the Stage 2KK\-sweep, where changingKKalso changes the sampling distribution behind the sampled\-softmax objective\. In paired re\-evaluations of the same checkpoints, batch\-localval\_lossis therefore a poor proxy for full\-catalogueval\_loss\(ρS=−0\.95\\rho\_\{S\}\\\!=\\\!\-0\.95atB=512B\\\!=\\\!512,C=1019C\\\!=\\\!10^\{19\}\)\. Ranking metrics behave differently: batch\-local ranking metrics, especiallyNDCG@10andMRR@10atB=512B\\\!=\\\!512, remain strongly correlated with full\-catalogue ranking metrics \(ρS=\+0\.90\\rho\_\{S\}\\\!=\\\!\+0\.90and\+0\.95\+0\.95respectively; Figure[9](https://arxiv.org/html/2606.05257#S7.F9)\)\. The Phase 3 iso\-FLOP architecture grid at the same budget tells a slightly different story: the six cells cluster tightly enough that the cross\-regime ranking signal is dominated by noise onNDCG@10\(ρS=\+0\.26\\rho\_\{S\}\\\!=\\\!\+0\.26\) andMRR@10\(ρS=\+0\.03\\rho\_\{S\}\\\!=\\\!\+0\.03\), whilecoverage@10stays the most stable proxy \(ρS=\+0\.90\\rho\_\{S\}\\\!=\\\!\+0\.90,n=6n\\\!=\\\!6\) andval\_lossactually correlates*positively*\(ρS=\+0\.66\\rho\_\{S\}\\\!=\\\!\+0\.66\)\. So batch\-local ranking metrics are a good proxy for full\-catalogue ranking metrics when the cells span a real quality gap \(theKK\-sweep\), but lose discriminative power when the cells are quality\-equivalent \(the iso\-FLOP arch grid\)\. We therefore use batch\-local ranking metrics as a practical proxy for comparing architecture/allocation cells, with the caveat that the marginal ranking gap between near\-tied cells is unreliable across regimes\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/local_vs_fullvocab_scatter.png)Figure 9:Batch\-local vs\. full\-catalogue evaluation\.Each point is one checkpoint scored under both regimes\. Red squares mark Phase 4K=0K\\\!=\\\!0only\. The first two rows are Stage 2KK\-sweep strata \(B=64B\\\!=\\\!64,C=1015C\\\!=\\\!10^\{15\}andB=512B\\\!=\\\!512,C=1019C\\\!=\\\!10^\{19\}\)\. The third row is the Phase 3 iso\-FLOP architecture grid atC=1019C\\\!=\\\!10^\{19\},B=512B\\\!=\\\!512\(n=6n\\\!=\\\!6cells re\-evaluated under full\-catalogue\)\. Column 2 is the cross\-metric panel: batch\-localval\_lossvs\. full\-catalogueNDCG@10; the other columns match metrics on both axes\. In the Phase 4KK\-sweep, loss is anti\-correlated across regimes while ranking metrics stay strongly positively correlated\. In the Phase 3 arch grid, the cells cluster tightly within each regime, soNDCG@10andMRR@10carry essentially no cross\-regime ranking signal whereascoverage@10andval\_lossremain positively correlated\.
#### \(e\) Ranking stability across scoring\-history lengths depends on the sweep\.

For each evaluated checkpoint we log metrics at scoring\-history lengthsc​l∈\{3,5,10,20,50,100\}cl\\\!\\in\\\!\\\{3,5,10,20,50,100\\\}\. For each metric and compute budget, we then ask whether cells are ranked similarly at different scoring\-history lengths: for example, do the same checkpoints win atc​l=3cl\\\!=\\\!3andc​l=100cl\\\!=\\\!100? Architecture/allocation sweeps are stable under this test: ranking metrics have worst\-case pairwise Spearman correlationsρmin≥0\.93\\rho\_\{\\min\}\\\!\\geq\\\!0\.93atC≤1018C\\\!\\leq\\\!10^\{18\}\. Stage 2KKsweeps are less stable at smaller budgets: forC≤1017C\\\!\\leq\\\!10^\{17\}, the worst\-case correlations across scoring\-history lengths drop toρmin∈\[0\.26,0\.73\]\\rho\_\{\\min\}\\\!\\in\\\!\[0\.26,0\.73\], meaning that the bestKKcan depend on whether evaluation emphasizes short\- or long\-history queries\. ByC≥1018C\\\!\\geq\\\!10^\{18\}, these Stage 2 rankings realign \(ρmin≥0\.87\\rho\_\{\\min\}\\\!\\geq\\\!0\.87\)\. Full 7×\\times7 slice\-pair matrices and the worst\-case\-ρ\\rho\-vs\-CCsummary are in Appendix[H](https://arxiv.org/html/2606.05257#A8)\(Figures[22](https://arxiv.org/html/2606.05257#A8.F22),[23](https://arxiv.org/html/2606.05257#A8.F23),[24](https://arxiv.org/html/2606.05257#A8.F24)\)\.

Cross\-budget exponents are metric\-robust\.Despite \(a\)–\(d\) above, the cross\-budget Hoffmann*exponents*agree across metrics to within slope uncertainty \(§[5](https://arxiv.org/html/2606.05257#S5), Table[5](https://arxiv.org/html/2606.05257#S5.T5)\): every metric lies inaN∈\[0\.57,0\.66\]a\_\{N\}\\\!\\in\\\!\[0\.57,\\,0\.66\], with theval\_loss\-vs\-NDCG@10separationΔ​aN≈0\.03\\Delta a\_\{N\}\\\!\\approx\\\!0\.03\. So the disagreement is in the per\-budget recipe, namely which\(h,L\)\(h,L\)cell or whichKKto ship, not in the scaling law itself\.

Pick the target before fitting the scaling law\.How much metric choice matters varies by axis\. In the model/data allocation sweep, the per\-metric Hoffmann exponents agree within slope uncertainty \(Δ​aN≈0\.04\\Delta a\_\{N\}\\\!\\approx\\\!0\.04\) and per\-budget winners differ by≤15%\\leq\\\!15\\%inN⋆N^\{\\star\}: the scaling law transfers across metrics, the per\-budget recipe approximately so\. In the negative\-sampling sweep, the per\-budgetK⋆K^\{\\star\}spreads∼1\.5\\sim\\\!1\.5–3×3\\\!\\timesacross metrics at the same budget, the cross\-budget exponent itself spreads \(K⋆∝C0\.08K^\{\\star\}\\\!\\propto\\\!C^\{0\.08\}–C0\.15C^\{0\.15\}\), and the bias\-decay exponentβ\\betasplits cleanly between loss\-like and ranking metrics \(βloss≈1\.2\\beta\_\{\\text\{loss\}\}\\\!\\approx\\\!1\.2vs\.βranking∈\[0\.39,0\.56\]\\beta\_\{\\text\{ranking\}\}\\\!\\in\\\!\[0\.39,\\,0\.56\]\): the scaling law itself does not transfer\. The strong within\-stage correlations of Figure[7](https://arxiv.org/html/2606.05257#S7.F7)do not save the practitioner from this asymmetry: two metrics that move together can still disagree on the per\-budgetK⋆K^\{\\star\}\. Selecting the deployed target metric*before*fitting a scaling law, not after, is the single practical recommendation we would carry to other behavioral foundation models, and matters most for the negative\-sampling axis\.

## 8Related Work

Scaling laws for language models\.The empirical\-scaling\-laws program of Kaplan et al\.\[[1](https://arxiv.org/html/2606.05257#bib.bib1)\]and Hoffmann et al\.\[[2](https://arxiv.org/html/2606.05257#bib.bib2)\]established that loss is well\-described by power laws in compute, parameters and tokens, and that the compute\-optimal allocation lies nearD/N≈20D/N\\\!\\approx\\\!20for text language modeling\. McCandlish et al\.\[[4](https://arxiv.org/html/2606.05257#bib.bib4)\]introduced the critical batch size as an orthogonal axis, and Yang and Hu\[[3](https://arxiv.org/html/2606.05257#bib.bib3)\]reframed initialization and LR transfer\. Our setup transposes this program to a different objective \(sampled in\-batch contrastive loss\), a different architecture \(two\-part feature embedder→\\rightarrowcontextualizer\), and an evaluation regime in which the training loss and the deployed ranking metric do not coincide \(§[7](https://arxiv.org/html/2606.05257#S7)\)\.

Recommender systems and feature\-rich foundation models\.The two\-tower factorization has its roots in the YouTube deep recommender\[[11](https://arxiv.org/html/2606.05257#bib.bib11)\]\. Recent industrial systems generalize the recipe to dynamic catalogues with feature\-based encoders, including Visa TREASURE\[[14](https://arxiv.org/html/2606.05257#bib.bib14)\], TransactionGPT\[[15](https://arxiv.org/html/2606.05257#bib.bib15)\], Stripe PFM\[[16](https://arxiv.org/html/2606.05257#bib.bib16)\], Revolut PRAGMA\[[17](https://arxiv.org/html/2606.05257#bib.bib17)\], and J\.P\. Morgan TradeFM\[[18](https://arxiv.org/html/2606.05257#bib.bib18)\]\. Generative recommenders such as HSTU\[[5](https://arxiv.org/html/2606.05257#bib.bib5)\]and Wukong\[[6](https://arxiv.org/html/2606.05257#bib.bib6)\]demonstrate favorable single\-axis scaling inNN\. Ardalani et al\.\[[7](https://arxiv.org/html/2606.05257#bib.bib7)\]report scaling laws for DLRM\-style hybrids and Netflix’s foundation model\[[13](https://arxiv.org/html/2606.05257#bib.bib13)\]reports that scaling\-up monotonically improves quality without quantitative laws\. None of these works jointly varies architecture, batch,\(N,D\)\(N,D\)allocation and negative sampling on a single stack, none reports per\-metric power\-law exponents, and none quantifies the embedder/contextualizer share as a scaling axis\.

Behavioral foundation models\.The BehaviorGPT line\[[19](https://arxiv.org/html/2606.05257#bib.bib19),[20](https://arxiv.org/html/2606.05257#bib.bib20),[21](https://arxiv.org/html/2606.05257#bib.bib21)\]and its generalization to “Large Behavioral Models”\[[22](https://arxiv.org/html/2606.05257#bib.bib22)\]argue that the embedder must be trained end\-to\-end on the sequential task and frame action\-sequence modeling as its own foundation\-model paradigm, complementary to language modeling\. Within that paradigm, the present paper supplies the joint scaling\-law calibration that the scaling\-law guidance of Table[7](https://arxiv.org/html/2606.05257#S9.T7)relies on\.

## 9Discussion

The four sweeps give a practical recipe for the current two\-part behavioral modeling stack and its two\-stage training recipe\. The robust recommendation is not a single number so much as an ordering of decisions\. First choose the deployment metric; then set a small event embedder, choose the batch size based on the data\-efficiency/throughput tradeoff, allocate compute betweenNNandDD, and finally tune the frozen\-embedder negative pool under the full\-catalogue ranking metric\.

Table 7:Scaling\-law guidance distilled from the four sweeps\.Depth is omitted: the depth sweep showsLctxL\_\{\\text\{ctx\}\}at the val minimum varies by budget while inducedssstays in the width\-sweep band \(Appendix[B](https://arxiv.org/html/2606.05257#A2.SS0.SSS0.Px5)\)\.#### Maximal Update Parameterization: a negative result\.

We swept Maximal Update Parameterization \(MuP\)\[[3](https://arxiv.org/html/2606.05257#bib.bib3)\]across four model sizes \(∼10\\sim\\\!10M to∼500\\sim\\\!500M total trainable parameters\) and learning rates in\[10−5,5⋅10−2\]\[10^\{\-5\},\\,5\\\!\\cdot\\\!10^\{\-2\}\], against our default truncated\-normal–style initialization\. MuP delivers most of its core LR\-transferability promise: the MuP\-optimum LR span across widths is0\.300\.30decades vs\.0\.700\.70for Default, and every MuP optimum is a verified local minimum\. Default nevertheless reaches a strictly lower training loss at every scale by0\.680\.68–0\.920\.92nats \(Appendix[F](https://arxiv.org/html/2606.05257#A6), Table[14](https://arxiv.org/html/2606.05257#A6.T14)\)\. The pattern holds for every other MuP variant we tried, including per\-layer MuP and a FLOP\-budget\-matched variant; we therefore retain the default initialization and pay the modest cost of a per\-phase LR sweep instead\.

#### Limitations: Stage 1 evaluation is batch\-local\.

The architecture and allocation sweeps score checkpoints against the unique target embeddings in each validation batch \(∼5\\sim\\\!5–1010k items\), not against the full catalogue\. Eval batch size is fixed within each compute budget, so within\-budget rankings are apples\-to\-apples; this is what every architecture/allocation takeaway depends on\. Absolute losses across budgets are not directly comparable\. The mechanism argument for why batch\-local rankings should still transfer to the deployed full\-catalogue metric on the architecture axis is axis \(d\) of §[7](https://arxiv.org/html/2606.05257#S7); Figure[9](https://arxiv.org/html/2606.05257#S7.F9)sharpens it on identical checkpoints from Stage 2KK\-sweep cells\. Remaining open work is to extend full\-catalogue re\-evaluation across the full architecture/allocation grid and additional Stage 2 checkpoints at intermediate eval batch sizes\.

#### Limitations: Architecture coverage\.

The width sweep does not vary hidden size at fixed depth; the depth sweep does not vary text\-encoder share\. A full55D grid \(share×\\timeswidth×\\timesdepth×\\timesembedder\-depth×\\timesLR\) is the natural next step\. Thes⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%result is also for the essentially\-unregularized recipe we trained with; embedder regularization may shifts⋆s^\{\\star\}upward \(§[3\.1](https://arxiv.org/html/2606.05257#S3.SS1)\)\.

#### Limitations: Context length\.

All sweeps fix the*training*contextualizer sequence length atLseq=256L\_\{\\mathrm\{seq\}\}\\\!=\\\!256event tokens per example\. Every eval batch is additionally stratified by history position so each metric is also logged atc​l∈\{3,5,10,20,50,100\}cl\\\!\\in\\\!\\\{3,5,10,20,50,100\\\}events of context \(§[2](https://arxiv.org/html/2606.05257#S2)\), which axis \(e\) of §[7](https://arxiv.org/html/2606.05257#S7)\(Figure[22](https://arxiv.org/html/2606.05257#A8.F22)\) uses to bound the*scoring*half of the question: Stage 1 architectural winners are context\-length\-robust \(recall@10,NDCG@10andNDCG@100allρmin≥0\.93\\rho\_\{\\min\}\\\!\\geq\\\!0\.93atC≤1018C\\\!\\leq\\\!10^\{18\}\) so architecture/allocation winners transfer to shorter or longer serving contexts; Stage 2K⋆K^\{\\star\}winners are not atC≤1017C\\\!\\leq\\\!10^\{17\}\(rankingρmin∈\[0\.26,0\.73\]\\rho\_\{\\min\}\\\!\\in\\\!\[0\.26,0\.73\]\), realigning atC≥1018C\\\!\\geq\\\!10^\{18\}on the same compute threshold as axis \(c\)\. What none of this measures is the*training*half: howN⋆N^\{\\star\},D⋆D^\{\\star\},s⋆s^\{\\star\}andK⋆K^\{\\star\}shift when models are trained from scratch at smaller or largerLseqL\_\{\\mathrm\{seq\}\}, how the in\-batch3​B​Lseq​h3\\,BL\_\{\\mathrm\{seq\}\}hcontrastive cost in \([1](https://arxiv.org/html/2606.05257#S2.E1)\) scales withLseqL\_\{\\mathrm\{seq\}\}, and whether the iso\-FLOP trade\-offs continue to be captured by the Kaplan formula once attention’sL⋅Lseq2L\\\!\\cdot\\\!L\_\{\\mathrm\{seq\}\}^\{2\}term becomes non\-negligible\. That sweep is a planned future axis we did not vary\.

#### Limitations: Compute range\.

Our primary model/data allocation budgets spanC∈\[1015,1019\]C\\\!\\in\\\!\[10^\{15\},10^\{19\}\]FLOPs\. Exploring larger budgets remains open work\.

## 10Conclusion

Behavioral foundation models need their own scaling laws\. We study the now\-common two\-part stack, a feature\-based event embedder feeding a transformer contextualizer, under the two\-stage recipe in which the embedder is first trained jointly, then frozen while the contextualizer is trained with extra negatives\. Across this setting, the compute\-optimal event embedder is small: roughly two percent of parameters across the budgets we test\. The reason is structural: embedder parameters are more expensive per step and see a far more repetitive effective data distribution than contextualizer parameters\.

The compute allocation law is also distinctive\. Behavioral models are strongly data\-heavy at small budgets, withD/ND/Nfar above the text\-LM Chinchilla heuristic, but the optimum moves toward the language\-model regime as compute increases\. The recipe around that allocation is metric\-dependent: critical batch size changes with the target metric, the useful number of negatives after freezing changes with both compute and metric, and atC=1019C\\\!=\\\!10^\{19\}the negative\-sampling axis becomes limited by candidate\-axis memory rather than FLOPs\.

Finally, evaluation regime matters\. Batch\-local loss is not a reliable proxy for full\-catalogue loss, while batch\-local ranking metrics are a more practical proxy for comparing architecture/allocation cells\. For practitioners, the most portable lesson is therefore simple: choose the metric the system will serve, then fit the scaling law to that metric\. In behavioral foundation models, the evaluation metric is part of the scaling law because changing it can change the compute\-optimal recipe\.

## Acknowledgements

We thank Adam Fredriksson, Alexander Junco Hagberg, Alexandros Lemonaris, Erik Guander, Gabriel Melin, Gonçalo Marques, Jens Palmborg, Marcel Rød, Nicolas Sanchez, Simon Granström, and Tom Boustedt for their support and contributions\.

## References

- Kaplan et al\. \[2020\]J\. Kaplan et al\.Scaling laws for neural language models\.*arXiv:2001\.08361*, 2020\.
- Hoffmann et al\. \[2022\]J\. Hoffmann et al\.Training compute\-optimal large language models \(Chinchilla\)\.*arXiv:2203\.15556*, 2022\.
- Yang and Hu \[2021\]G\. Yang and E\. Hu\.Tensor Programs IV: Feature learning in infinite\-width neural networks \(MuP\)\.In*ICML*, 2021\.
- McCandlish et al\. \[2018\]S\. McCandlish, J\. Kaplan, D\. Amodei, and the OpenAI Dota Team\.An empirical model of large\-batch training\.*arXiv:1812\.06162*, 2018\.
- Zhai et al\. \[2024\]J\. Zhai et al\.Actions speak louder than words: Trillion\-parameter sequential transducers for generative recommendations \(HSTU\)\.In*ICML*, 2024\.
- Zhang et al\. \[2024\]B\. Zhang et al\.Wukong: Towards a scaling law for large\-scale recommendation\.*arXiv:2403\.02545*, 2024\.
- Ardalani et al\. \[2022\]N\. Ardalani et al\.Understanding scaling laws for recommendation models\.*arXiv:2208\.08489*, 2022\.
- Liu et al\. \[2023\]H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee\.Visual instruction tuning\.In*NeurIPS*, 2023\.
- Alayrac et al\. \[2022\]J\.\-B\. Alayrac et al\.Flamingo: a visual language model for few\-shot learning\.In*NeurIPS*, 2022\.
- Li et al\. \[2023\]J\. Li, D\. Li, S\. Savarese, and S\. Hoi\.BLIP\-2: Bootstrapping language\-image pre\-training with frozen image encoders and large language models\.In*ICML*, 2023\.
- Covington et al\. \[2016\]P\. Covington, J\. Adams, and E\. Sargin\.Deep neural networks for YouTube recommendations\.In*ACM RecSys*, 2016\.
- Bengio and Senécal \[2008\]Y\. Bengio and J\.\-S\. Senécal\.Adaptive importance sampling to accelerate training of a neural probabilistic language model\.*IEEE Transactions on Neural Networks*, 19\(4\):713–722, 2008\.
- Netflix \[2025\]Netflix Technology Blog\.Foundation Model for Personalized Recommendation\.Mar\. 2025\.[https://netflixtechblog\.com/foundation\-model\-for\-personalized\-recommendation\-1a0bd8e02d39](https://netflixtechblog.com/foundation-model-for-personalized-recommendation-1a0bd8e02d39)
- Yeh et al\. \[2025\]C\.\-C\. M\. Yeh, U\. S\. Saini, X\. Dai, X\. Fan, S\. Jain et al\.TREASURE: A transformer\-based foundation model for high\-volume transaction understanding \(Visa Payment Foundation Model\)\.*arXiv:2511\.19693*, 2025\.
- Dou et al\. \[2025\]Y\. Dou, Z\. Jiang, T\. Zhang, M\. Hu, Z\. Xu, Y\. Chen et al\.TransactionGPT\.*arXiv:2511\.08939*, 2025\.
- Stripe \[2025\]G\. Kedia and the Stripe Machine Learning Team\.Stripe’s Payments Foundation Model\.Stripe Sessions / Stripe Engineering, May 2025\.
- Iashin et al\. \[2026\]V\. Iashin et al\.PRAGMA: Revolut foundation model\.*arXiv:2604\.08649*, 2026\.
- Kawawa\-Beaudan et al\. \[2026\]M\. Kawawa\-Beaudan, D\. Borrajo, M\. Veloso et al\.TradeFM: A generative foundation model for trade\-flow and market microstructure\.*arXiv:2602\.23784*, 2026\.
- Brüel Gabrielsson et al\. \[2025a\]R\. Brüel Gabrielsson et al\.A foundation model for consumption, transactions, and actions: The inception of BehaviorGPT\.Unbox AI Research, 2025\.
- Brüel Gabrielsson and Gupta \[2025a\]R\. Brüel Gabrielsson and V\. Gupta\.BehaviorGPT at work: A foundation model for workforce actions and dynamics\.Unbox AI Research, 2025\.
- Brüel Gabrielsson and Gupta \[2025b\]R\. Brüel Gabrielsson and V\. Gupta\.BehaviorGPT for visual art: A foundation model for aesthetics\.Unbox AI Research, 2025\.
- Brüel Gabrielsson et al\. \[2026\]R\. Brüel Gabrielsson et al\.Large behavioral models: A foundation\-model paradigm for human actions\.Unbox AI Research, 2026\.
- Järvelin and Kekäläinen \[2002\]K\. Järvelin and J\. Kekäläinen\.Cumulated gain\-based evaluation of IR techniques\.*ACM Transactions on Information Systems*, 20\(4\):422–446, 2002\.
- Voorhees \[1999\]E\. M\. Voorhees\.The TREC\-8 question answering track report\.In*Proceedings of TREC\-8*, 1999\.
- Manning et al\. \[2008\]C\. D\. Manning, P\. Raghavan, and H\. Schütze\.*Introduction to Information Retrieval*\.Cambridge University Press, 2008\.
- Herlocker et al\. \[2004\]J\. L\. Herlocker, J\. A\. Konstan, L\. G\. Terveen, and J\. T\. Riedl\.Evaluating collaborative filtering recommender systems\.*ACM Transactions on Information Systems*, 22\(1\):5–53, 2004\.

## Appendix AMetric Definitions

#### Notation and candidate set\.

Every metric scores each query position against a*candidate set*𝒞\\mathcal\{C\}and ranks its items by the dot\-product scorezq,j=⟨hq,ej⟩z\_\{q,j\}\\\!=\\\!\\langle h\_\{q\},e\_\{j\}\\rangle, wherehqh\_\{q\}is the contextualizer’s output \(query\) embedding at positionqqandeje\_\{j\}is the cached embedding of candidatej∈𝒞j\\\!\\in\\\!\\mathcal\{C\}\. We writeNNfor the number of scored positions,C=\|𝒞\|C\\\!=\\\!\|\\mathcal\{C\}\|for the candidate\-set size,KKfor the cutoff \(the “@KK”\), and𝟙​\[⋅\]\\mathbbm\{1\}\[\\cdot\]for the indicator\. Each query has exactly one relevant item, its true next event;rqr\_\{q\}is that item’s11\-indexed rank in the descending score order over𝒞\\mathcal\{C\}\(sorq=1r\_\{q\}\\\!=\\\!1is a top hit\)\.

*The choice of𝒞\\mathcal\{C\}is the central batch\-local vs\. global\-catalogue distinction in the paper*, and it governs*every*metric below, not just coverage:

- •Batch\-local \(Stage 1\)\.Before the embedder is frozen,𝒞\\mathcal\{C\}is the set of*unique target items in the current validation batch*\(C≲B​LseqC\\\!\\lesssim\\\!BL\_\{\\mathrm\{seq\}\}\), matching the in\-batch sampled softmax used in training\.
- •Global catalogue \(Stage 2\)\.After freezing,𝒞\\mathcal\{C\}is the*full cached deployed catalogue*\(C∼108C\\\!\\sim\\\!10^\{8\}\), matching the deployed retrieval setting\. For tractable repeated evaluation we score against a fixed∼13\.6\\sim\\\!13\.6M\-item subset of this catalogue, still∼1500×\\sim\\\!1500\\timeslarger than the batch\-local pool\.

Because𝒞\\mathcal\{C\}fixes both the ranking pool \(hencerqr\_\{q\}\) and the softmax normalizer \(hencepqp\_\{q\}below\), absolute loss, entropy, and ranking scores are comparable only*within*a stage; this is exactly why §[7](https://arxiv.org/html/2606.05257#S7)reports loss–ranking correlations per stage\.

#### Ranking metrics\[[25](https://arxiv.org/html/2606.05257#bib.bib25)\]\.

With a single relevant item and binary relevance \(rq=1r\_\{q\}\\\!=\\\!1best\):

recall@​K\\displaystyle\\text\{recall@\}K=1N​∑q=1N𝟙​\[rq≤K\],\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{q=1\}^\{N\}\\mathbbm\{1\}\[\\,r\_\{q\}\\\!\\leq\\\!K\\,\],NDCG@​K\\displaystyle\\text\{NDCG@\}K=1N​∑q=1N𝟙​\[rq≤K\]log2⁡\(rq\+1\),\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{q=1\}^\{N\}\\frac\{\\mathbbm\{1\}\[\\,r\_\{q\}\\\!\\leq\\\!K\\,\]\}\{\\log\_\{2\}\(r\_\{q\}\+1\)\},MRR@​K\\displaystyle\\text\{MRR@\}K=1N​∑q=1N𝟙​\[rq≤K\]rq\.\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{q=1\}^\{N\}\\frac\{\\mathbbm\{1\}\[\\,r\_\{q\}\\\!\\leq\\\!K\\,\]\}\{r\_\{q\}\}\.\(7\)recall@KK\(here equal to hit@KK\) counts how often the true item lands in the topKK; NDCG@KK\[[23](https://arxiv.org/html/2606.05257#bib.bib23)\]and MRR@KK\[[24](https://arxiv.org/html/2606.05257#bib.bib24)\]additionally reward placing it near the top\. The ideal DCG is11\(one relevant item\), so NDCG@KKis just that item’s rank discount1/log2⁡\(rq\+1\)1/\\log\_\{2\}\(r\_\{q\}\+1\)\.

#### Coverage\.

coverage@KKis the catalogue fraction the model actually surfaces in its top\-KKpredictions, a recommender diversity measure\[[26](https://arxiv.org/html/2606.05257#bib.bib26)\]\. It has two protocols; we report thebatch\-localone, because the global mask saturates toward11once enough batches are pooled \(§[4](https://arxiv.org/html/2606.05257#S4)\):

coverage@​K=1Beval​∑b=1Beval\|\{distinct top\-​K​items in batch​b\}\|C⏟batch\-local \(reported\),\|⋃b\{top\-​K​items in batch​b\}\|C⏟global mask,\\underbrace\{\\text\{coverage@\}K=\\frac\{1\}\{B\_\{\\text\{eval\}\}\}\\sum\_\{b=1\}^\{B\_\{\\text\{eval\}\}\}\\frac\{\\bigl\|\\\{\\text\{distinct top\-\}K\\text\{ items in batch \}b\\\}\\bigr\|\}\{C\}\}\_\{\\text\{batch\-local \(reported\)\}\}\\;,\\qquad\\underbrace\{\\frac\{\\bigl\|\\bigcup\_\{b\}\\\{\\text\{top\-\}K\\text\{ items in batch \}b\\\}\\bigr\|\}\{C\}\}\_\{\\text\{global mask\}\},\(8\)whereBevalB\_\{\\text\{eval\}\}is the number of evaluation batches\.

#### Loss and predictive uncertainty\.

Letpqp\_\{q\}be the softmax over𝒞\\mathcal\{C\},pq,j=ezq,j/∑j′∈𝒞ezq,j′p\_\{q,j\}\\\!=\\\!e^\{z\_\{q,j\}\}\\big/\\sum\_\{j^\{\\prime\}\\in\\mathcal\{C\}\}e^\{z\_\{q,j^\{\\prime\}\}\}\. We report the sampled\-softmax cross\-entropy, its exponential, and the mean predictive entropy \(in nats\):

CE=1N​∑q=1N\(−log⁡pq​\(targetq\)\),perplexity=eCE,entropy=1N​∑q=1N\(−∑j∈𝒞pq,j​log⁡pq,j\)\.\\text\{CE\}=\\frac\{1\}\{N\}\\sum\_\{q=1\}^\{N\}\\\!\\bigl\(\-\\log p\_\{q\}\(\\text\{target\}\_\{q\}\)\\bigr\),\\quad\\text\{perplexity\}=e^\{\\text\{CE\}\},\\quad\\text\{entropy\}=\\frac\{1\}\{N\}\\sum\_\{q=1\}^\{N\}\\\!\\Bigl\(\-\\\!\\\!\\sum\_\{j\\in\\mathcal\{C\}\}p\_\{q,j\}\\log p\_\{q,j\}\\Bigr\)\.\(9\)

## Appendix BAdditional Architecture Results

#### Per\-budget iso\-FLOPs accounting\.

Table[8](https://arxiv.org/html/2606.05257#A2.T8)writes the Kaplan per\-step formula \([1](https://arxiv.org/html/2606.05257#S2.E1)\) out cell by cell at the near\-optimal reference shares=6%s\\\!=\\\!6\\%we use for the iso\-FLOPs schedule \(the upper edge of the flats∈\[2,6\]%s\\\!\\in\\\!\[2,6\]\\%band; the recommendation remainss⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%, §[3\.1](https://arxiv.org/html/2606.05257#S3.SS1)\)\. All four budgets land within0\.01%0\.01\\%of their nominal target\.

Table 8:Iso\-FLOPs accounting at the near\-optimal reference shares=6%s\\\!=\\\!6\\%\.D/ND/Nstays at the Chinchilla∼15\\sim\\\!15that motivated the schedule\.
#### Validation loss vs\. parameter share\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/share_sweep_two_stage.png)Figure 10:Embedder\-share sweep \(val\_lossonly\)\.*Left:*validation loss vs\. embedder share overs∈\[0,50\]%s\\\!\\in\\\!\[0,50\]\\%at four compute budgets\. Curves are monotone increasing inssat every budget overs∈\[6,50\]%s\\\!\\in\\\!\[6,50\]\\%\.*Right:*zoom ons∈\[0,6\]%s\\\!\\in\\\!\[0,6\]\\%around the per\-budget optimum\. Solid lines are per\-budget two\-term starvation fitsL​\(s\)=E\+a​sα\+b​s−βL\(s\)\\\!=\\\!E\\\!\+\\\!a\\,s^\{\\alpha\}\\\!\+\\\!b\\,s^\{\-\\beta\}; stars mark the closed\-form analytic optimums⋆=\(b​β/\(a​α\)\)1/\(α\+β\)s^\{\\star\}\\\!=\\\!\(b\\beta/\(a\\alpha\)\)^\{1/\(\\alpha\+\\beta\)\}\(Table[2](https://arxiv.org/html/2606.05257#S3.T2)\)\. The analytic optima cluster ats⋆∈\[1\.1%,3\.7%\]s^\{\\star\}\\\!\\in\\\!\[1\.1\\%,\\,3\.7\\%\]across all four budgets \(vs\. a discrete\-grid argmax that bounces betweens=2%s\\\!=\\\!2\\%ands=5%s\\\!=\\\!5\\%\), withs⋆∝C\+0\.07s^\{\\star\}\\\!\\propto\\\!C^\{\+0\.07\}, effectively flat in compute\. The Kaplan\-FLOP\-share cross\-check on the same cells is Figure[11](https://arxiv.org/html/2606.05257#A2.F11)\.
#### Kaplan embedder\-side FLOP share \(cross\-check\)\.

The width–share grid is controlled by a target*parameter*embedder sharess, which need not coincide with the fraction of the Kaplan per\-step inner sum from \([1](https://arxiv.org/html/2606.05257#S2.E1)\) spent on the text embedder, the non\-text embedder stack, and the in\-batch contrastive3​v​h3\\,v\\,hterm \(numerator\) vs\. the contextualizer term \(denominator\)\. Figure[11](https://arxiv.org/html/2606.05257#A2.F11)re\-plots the same per\-metric, train\-best\-LR cells as the main\-text Figure[1](https://arxiv.org/html/2606.05257#S3.F1)with that Kaplan fractionffon the horizontal axis; the two\-term template \([2](https://arxiv.org/html/2606.05257#S3.E2)\) is still overlaid for visual continuity, but we reads⋆s^\{\\star\}only from thess\-axis fit \(Table[2](https://arxiv.org/html/2606.05257#S3.T2)\) because theff\-axis refit is poorly conditioned\. Hereffis a nonlinear, cell\-dependent pushforward ofss\(the3​B​s​h3Bshterm varies with the joint embedding dimhhthat the width search picks per cell\), the swept points cover only the bandf∈\[0\.1,0\.5\]f\\\!\\in\\\!\[0\.1,0\.5\]vs\.s∈\[0\.5%,50%\]s\\\!\\in\\\!\[0\.5\\%,50\\%\], and the two\-term form gets pushed to its parameter bounds\. The picture is nonetheless useful as a sanity check: the small\-embedder optimum is not a parameter\-counting artifact\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/share_sweep_all_metrics_compute_share.png)Figure 11:Width sweep vs\. Kaplan embedder\-side FLOP shareff\.Same cells as Figure[1](https://arxiv.org/html/2606.05257#S3.F1); horizontal axis is the Kaplan fractionfffrom \([1](https://arxiv.org/html/2606.05257#S2.E1)\)\.⋆\\star= analytic optimum inside the swept range;\+\+= boundary extrapolation\. Reporteds⋆s^\{\\star\}values use thess\-axis fit \(Table[2](https://arxiv.org/html/2606.05257#S3.T2)\)\.
#### N/DN/Dproxy across the width–share sweep\.

The iso\-D/ND/Nframing of §[3\.1](https://arxiv.org/html/2606.05257#S3.SS1)is verified empirically here\. For every train\-best cell we plot a microbatch proxy for tokens seen against total embedder\-plus\-contextualizer parameters \(iso\-FLOP accounting ats=6%s\\\!=\\\!6\\%in Table[8](https://arxiv.org/html/2606.05257#A2.T8)\)\. Figure[12](https://arxiv.org/html/2606.05257#A2.F12)\(a\) shows thatNemb\+ctx/DproxyN\_\{\\mathrm\{emb\+ctx\}\}/D\_\{\\mathrm\{proxy\}\}is nearly flat in target embedder sharessat each budget, so the share sweep does slide along an approximately constant ratio rather than along a largeN/DN/Dmove that would confound the embedder/contextualizer tradeoff with a Chinchilla allocation move\. Figure[12](https://arxiv.org/html/2606.05257#A2.F12)\(b\) plots validation loss against the same ratio; the gray dashed line is an OLS fit whoseR2R^\{2\}stays low, while loss varies much more systematically withssin Figure[10](https://arxiv.org/html/2606.05257#A2.F10), confirming that quality is driven by the embedder–contextualizer split, not by the small residualN/DN/Dmovement\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/p1w_n_over_d_vs_target_share_paper_appendix_mono.png)\(a\)Nemb\+ctx/DproxyN\_\{\\mathrm\{emb\+ctx\}\}/D\_\{\\mathrm\{proxy\}\}vs\. target embedder share \(single color; train\-best LR per cell\)\.
![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/p1w_val_loss_vs_n_over_d_proxy_paper.png)\(b\)Validation loss vs\.Nemb\+ctx/DproxyN\_\{\\mathrm\{emb\+ctx\}\}/D\_\{\\mathrm\{proxy\}\}with a linear OLS overlay; per\-panelR2R^\{2\}is small relative to the sharpss\-dependence in the main text\.

Figure 12:N/DN/Dproxy diagnostics on the width sweep\.Varyingssmoves compute between embedder and contextualizer at nearly fixed training\-to\-parameter ratio:Nemb\+ctx/DproxyN\_\{\\mathrm\{emb\+ctx\}\}/D\_\{\\mathrm\{proxy\}\}is flat inss\(a\), while validation loss tracksssmuch more strongly than this residual ratio \(b\)\.
#### Depth sweep \(per\-cell val grid and per\-metric eval\)\.

Table[9](https://arxiv.org/html/2606.05257#A2.T9)lists the val\-loss winner per budget; Table[10](https://arxiv.org/html/2606.05257#A2.T10)re\-reads the same grid through the parabolic estimator inlog⁡s\\log sand Figure[13](https://arxiv.org/html/2606.05257#A2.F13)shows the fits visually overlaid on the discrete cells; Table[11](https://arxiv.org/html/2606.05257#A2.T11)is the full per\-cell grid; Figure[14](https://arxiv.org/html/2606.05257#A2.F14)plots every headline metric vs\. embedder sharessand Kaplan FLOP shareff\.

Table 9:Val\-loss\-optimal depth per budget \(appendix detail\)\.ssis the induced embedder share; ranking metrics are at the same checkpoint\. Val\-optimalLctxL\_\{\\text\{ctx\}\}varies;ssstays in the width\-sweep band\.Table 10:Parabolic re\-reading of the Phase 1D val\_loss grid\.For each budget we fitval​\_​loss=a​\(log10⁡s\)2\+b​log10⁡s\+c\\mathrm\{val\\\_loss\}\\\!=\\\!a\(\\log\_\{10\}s\)^\{2\}\\\!\+\\\!b\\log\_\{10\}s\\\!\+\\\!con the depth\-sum\-diagonal cells and report the analytic minimums⋆s^\{\\star\}, the val\_loss there, and the projected val\_loss penalty ats=2%s\\\!=\\\!2\\%\. “flat” marks budgets where the swept range is too flat \(≤0\.21\\leq\\\!0\.21nats end\-to\-end atC=1015C\\\!=\\\!10^\{15\}\) for the parabola to resolve a sharp interior minimum\. Per\-cell penalties versus the parabolic minimum are inparabolic\_depth\_fit\_cells\.csvalongside the table CSV\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/depth_val_loss_parabolic_fits.png)Figure 13:Depth\-sweep parabolic fits behind Table[10](https://arxiv.org/html/2606.05257#A2.T10)\.Per\-budget panels ofval\_lossagainst embedder sharess\(log scale\)\. Filled circles are the discrete depth\-sum\-diagonal cells \(annotated withLctxL\_\{\\mathrm\{ctx\}\}\); the dashed curve is the parabolic fit inlog10⁡s\\log\_\{10\}s; the star marks the analytic minimums⋆s^\{\\star\}; the hollow circle marks the value of the parabola ats=2%s\\\!=\\\!2\\%\(the Phase 1W recommendation\) with the gapΔs=2%\\Delta\_\{s=2\\%\}annotated below\. The gray band is the Phase 1Ws∈\[1\.5%,2\.5%\]s\\\!\\in\\\!\[1\.5\\%,\\,2\.5\\%\]recommendation\. AtC=1015C\\\!=\\\!10^\{15\}the swept range is too flat \(≤0\.21\\leq\\\!0\.21nats end\-to\-end\) for the parabola to resolve a sharp interior minimum; atC=1018C\\\!=\\\!10^\{18\}the depth grid saturates atLctx=32L\_\{\\mathrm\{ctx\}\}\\\!=\\\!32so the fit is monotone\-trending and the analytics⋆s^\{\\star\}should be read as an upper bound on the true value\. AtC=1016,1017C\\\!=\\\!10^\{16\},\\,10^\{17\}the parabolic minima \(s⋆=1\.82%,1\.80%s^\{\\star\}\\\!=\\\!1\.82\\%,\\,1\.80\\%\) land essentially on top of the Phase 1W recommendation of2%2\\%and the projected loss penalty ats=2%s\\\!=\\\!2\\%is≤10−3\\leq\\\!10^\{\-3\}nat\. The noisy discreteLctx⋆L\_\{\\mathrm\{ctx\}\}^\{\\star\}jumps of Table[9](https://arxiv.org/html/2606.05257#A2.T9)are an artifact of the discrete\-grid argmax, not of an underlying shift in the optimum\.Table 11:Held\-outval\_lossper \(budget, contextualizer depth\)\.“–” = outside swept grid; bold = best atC=1018C\\\!=\\\!10^\{18\}among depth\-sum\-3434cells \(†/‡= deeper off\-diagonal extensions\)\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/depth_eval_metrics_two_views_grid_compute_share.png)Figure 14:Depth sweep: per\-metric eval vs\.ssand Kaplan FLOP shareff\.Stars mark per\-budget optima\. Gray band: width\-sweeps⋆≈2%s^\{\\star\}\\\!\\approx\\\!2\\%\.
#### Held\-out evaluation atC=1018C\\\!=\\\!10^\{18\}\.

Table[12](https://arxiv.org/html/2606.05257#A2.T12)reports validation metrics on the depth\-sum\-3434diagonal\.L=32L\\\!=\\\!32wins every metric; deeper off\-diagonal extensions \(L=40,56L\\\!=\\\!40,56\) do not improve held\-out quality\.

Table 12:Held\-out evaluation atC=1018C\\\!=\\\!10^\{18\}\(depth\-sum\-3434diagonal and off\-diagonal extensions\)\.

## Appendix CPhase 2: Per\-Metric Trajectories

Figure[15](https://arxiv.org/html/2606.05257#A3.F15)shows the EWMA\-smoothed metric trajectories vs\. optimizer updates for eachB∈\{64,128,256,512,1024,2048\}B\\\!\\in\\\!\\\{64,128,256,512,1024,2048\\\}at the Phase 1 winner \(C=1017C\\\!=\\\!10^\{17\},s≈2%s\\\!\\approx\\\!2\\%; §[4](https://arxiv.org/html/2606.05257#S4)\)\. Dashed horizontal lines are the per\-metric iso\-targetsTmT\_\{m\}; the first crossing of each curve definesSm​\(B\)S\_\{m\}\(B\)in Tab\.[3](https://arxiv.org/html/2606.05257#S4.T3)\. The Kaplan fits of those updates\-to\-target points are Fig\.[3](https://arxiv.org/html/2606.05257#S4.F3)in the main text\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/metric_trajectories_v3.png)Figure 15:Per\-metric trajectories vs\. batch size\.One curve perBB; dashed line: iso\-targetTmT\_\{m\}\.
## Appendix DPhase 3 Train\-Surrogate Allocation

The main\-text Phase 3 fits useval\_lossas the primary objective \(§[5](https://arxiv.org/html/2606.05257#S5), Table[4](https://arxiv.org/html/2606.05257#S5.T4)\)\. For completeness we record the parallel tail\-100100train\-surrogate analysis on the same architecture grid, the classic Chinchilla diagnostic\. Train and val parabolic minima need not coincide: they diverge notably atC=1018C\\\!=\\\!10^\{18\}\(39\.539\.5M train vs\.19\.419\.4M val\) but land nearh1152\_L16on both surrogates atC=1019C\\\!=\\\!10^\{19\}\. Parabolic smoothing also moves the train\-surrogate per\-budget winner relative to its own discrete\-cell choice in Table[13](https://arxiv.org/html/2606.05257#A4.T13):0\.920\.92M→706\\\!\\to\\\!706k at101510^\{15\},31\.931\.9M→39\.5\\\!\\to\\\!39\.5M at101810^\{18\},251251M→203\\\!\\to\\\!203M at101910^\{19\}\.

Table 13:Train\-surrogate Chinchilla allocation\.N⋆N^\{\\star\}andD⋆D^\{\\star\}minimize the tail\-100100training loss;L⋆L^\{\\star\}is the fitted irreducible\-loss intercept trajectory \(Eq\.[12](https://arxiv.org/html/2606.05257#A4.E12)\)\.FittingN⋆​\(C\)=a​CbN^\{\\star\}\(C\)\\\!=\\\!aC^\{b\}in log\-space on the five parabolic train\-surrogate minima gives

Ntrain⋆​\(C\)\\displaystyle N^\{\\star\}\_\{\\text\{train\}\}\(C\)=4\.06×10−4​C0\.612±0\.024,\\displaystyle\\;=\\;4\.06\\\!\\times\\\!10^\{\-4\}\\,C^\{0\.612\\pm 0\.024\},\(10\)Dtrain⋆​\(C\)\\displaystyle D^\{\\star\}\_\{\\text\{train\}\}\(C\)=4\.11×102​C0\.388±0\.024,\\displaystyle\\;=\\;4\.11\\\!\\times\\\!10^\{2\}\\,C^\{0\.388\\pm 0\.024\},\(11\)Ltrain⋆​\(C\)\\displaystyle L^\{\\star\}\_\{\\text\{train\}\}\(C\)=2\.485\+1\.50×103​C−0\.210,\\displaystyle\\;=\\;2\.485\+1\.50\\\!\\times\\\!10^\{3\}\\,C^\{\-0\.210\},\(12\)withb\+b′=1b\\\!\+\\\!b^\{\\prime\}\\\!=\\\!1exact by Approach 2 construction\. For comparison, the discrete\-winner refit on the five rows of Table[13](https://arxiv.org/html/2606.05257#A4.T13)gives slightly steeper exponents \(bN=0\.609±0\.052b\_\{N\}\\\!=\\\!0\.609\\\!\\pm\\\!0\.052,bD=0\.439±0\.028b\_\{D\}\\\!=\\\!0\.439\\\!\\pm\\\!0\.028, summing to1\.0481\.048\), the same∼5%\\sim\\\!5\\%inflation that shows up on the val surrogate in §[5](https://arxiv.org/html/2606.05257#S5)\. The parametric three\-term fit

Ltrain​\(N,D\)=2\.570\+2\.46×103N0\.696\+8\.08×102D0\.384L\_\{\\text\{train\}\}\(N,D\)\\;=\\;2\.570\+\\frac\{2\.46\\\!\\times\\\!10^\{3\}\}\{N^\{0\.696\}\}\+\\frac\{8\.08\\\!\\times\\\!10^\{2\}\}\{D^\{0\.384\}\}\(13\)achieves RMSE=0\.084\\\!=\\\!0\.084nats across the4848merged training points\.

## Appendix EPhase 4 Metric\-Stratified Sampling Efficiency

#### Training\-loss decomposition \(not comparable acrossKKas quality\)\.

Figure[16](https://arxiv.org/html/2606.05257#A5.F16)plots the two channels the Stage 2 trainer logs atC=1017C\\\!=\\\!10^\{17\}: in\-batch CE \(fixed\|ℬbatch\|≈16,384\|\\mathcal\{B\}\_\{\\text\{batch\}\}\|\\\!\\approx\\\!16\{,\}384candidates\) and extra CE over theKKsampled catalog negatives, averaged for optimization\. These curves are*not*full\-catalogue evaluation: the extra term is defined over a candidate set that grows withKK, so the combined training loss is not an apples\-to\-apples quality metric across the sweep\. We include the panel only to show why the implemented objective has two opposing pieces; allK⋆K^\{\\star\}numbers in §[6](https://arxiv.org/html/2606.05257#S6)come from full\-catalogue evaluation\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/loss_components_vs_k_1e17.png)Figure 16:Training\-loss channels atC=1017C\\\!=\\\!10^\{17\}\(not full\-catalogue evaluation\)\.Blue dashed: in\-batch CE; red dashed: extra\-negative CE \(only forK\>0K\\\!\>\\\!0\); black: their average, which SGD minimizes\. The extra channel rises withKKby construction; the in\-batch channel moves with the checkpoint trained at eachKK\.Figure[17](https://arxiv.org/html/2606.05257#A5.F17)summarizes the interior analytic optima of Figure[6](https://arxiv.org/html/2606.05257#S6.F6)on a singleK⋆K^\{\\star\}\-vs\-CCpanel, and is the figure the practical recipe band of §[6](https://arxiv.org/html/2606.05257#S6)is read off of\. Three caveats on reading it\.*First*,1313of2525metric×\\timesbudget cells are boundary \(Table[6](https://arxiv.org/html/2606.05257#S6.T6)\), so the dashed power\-law slopes are fit on onlyn=2n\\\!=\\\!2–44interior cells per metric and should be read as a within\-band summary rather than as tightly\-identified scaling exponents\. This is why finding \(ii\) in §[6](https://arxiv.org/html/2606.05257#S6)reports the bandK⋆∈\[125​k,870​k\]K^\{\\star\}\\\!\\in\\\!\[125\\mathrm\{k\},\\,870\\mathrm\{k\}\]alongside the slope\.*Second*, the bias exponentβ\\betaof \([6](https://arxiv.org/html/2606.05257#S6.E6)\) partitions the metrics into three behaviors that are already visible in the parabolas of Figure[6](https://arxiv.org/html/2606.05257#S6.F6): bias\-dominated \(val\_loss,val\_entropy;β∼1\.2\\beta\\\!\\sim\\\!1\.2and∼0\.2\\sim\\\!0\.2–0\.60\.6, still falling atK=2K\\\!=\\\!2M\), saturating ranking metrics \(recall@10,NDCG@10,MRR@10;β∼0\.4\\beta\\\!\\sim\\\!0\.4–0\.60\.6, interior peak at every budget≥1016\\geq\\\!10^\{16\}\), and popularity\-collapse diversity \(coverage@10, monotonically decreasing inKKat everyCC; the fit returnsβ\>0\\beta\\\!\>\\\!0but with the bias term acting as a positive penalty, so we recordcoverage@10as a boundary case on the larger\-better side\)\.*Third*, every metric atC=1019C\\\!=\\\!10^\{19\}is boundary \(Table[6](https://arxiv.org/html/2606.05257#S6.T6)\), because all panels at that budget are still trending atK=2K\\\!=\\\!2M; finding \(iii\) of §[6](https://arxiv.org/html/2606.05257#S6)attributes this to the binding constraint onKKflipping from compute to memory\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/k_sweep_kstar_vs_C.png)Figure 17:AnalyticK⋆​\(C\)K^\{\\star\}\(C\)from \([6](https://arxiv.org/html/2606.05257#S6.E6)\), per metric \(interior\-cell summary of Figure[6](https://arxiv.org/html/2606.05257#S6.F6)\)\.Filled markers: interior analytic optimum from the per\-\(metric,CC\) parabola; faded\+\+: boundary cases \(K⋆K^\{\\star\}extrapolated outside the swept range\), excluded from the dashed power\-law fit\. The fit is onn=2n\\\!=\\\!2–44interior cells per metric \(1313of2525cells are boundary; see Table[6](https://arxiv.org/html/2606.05257#S6.T6)\), so the dashed slopesK⋆∝C0\.09−C0\.15K^\{\\star\}\\\!\\propto\\\!C^\{0\.09\}\\\!\\\!\-\\\!\\\!C^\{0\.15\}summarize the interior cluster but should not be read as tightly\-fit scaling exponents\. The primary takeaway is the band, with rankingK⋆K^\{\\star\}falling in\[125​k,870​k\]\[125\\mathrm\{k\},\\,870\\mathrm\{k\}\]acrossC∈\[1016,1018\]C\\\!\\in\\\!\[10^\{16\},10^\{18\}\], not the slope\.
#### Surrogacy stratified by\(C,metric\)\(C,\\text\{metric\}\)\.

Figure[18](https://arxiv.org/html/2606.05257#A5.F18)stratifies the Stage 2 surrogacy by\(compute budget, eval metric\)\(\\text\{compute budget, eval metric\}\)\. The training\-time signal here is the*in\-batch*cross\-entropy—theKK\-comparable channel of the Stage 2 objective\. We deliberately do*not*use the optimized combined loss: its extra\-negative channel grows withKKby construction \(Fig\.[16](https://arxiv.org/html/2606.05257#A5.F16)\), so the combined loss carries a mechanicalKK\-trend and is not comparable across the sweep as a quality measure\. The high\-compute rows are then unambiguous: atC=1018C\\\!=\\\!10^\{18\}and101910^\{19\}the in\-batch loss is almost perfectly*anti*\-correlated with every deployed metric—rankingρS≈\+0\.97\\rho\_\{S\}\\\!\\approx\\\!\+0\.97to\+1\.00\+1\.00, andval\_loss,val\_entropy,coverage@10ρS≈−0\.98\\rho\_\{S\}\\\!\\approx\\\!\-0\.98to−1\.00\-1\.00—on a large underlying spread \(Δ​y≈3\\Delta y\\\!\\approx\\\!3–11%11\\%\)\. Both rows share one mechanism: theKKthat minimizes in\-batch loss isK=0K\\\!=\\\!0\(with no extra negatives the contextualizer over\-fits the easy in\-batch task\), yet every full\-catalogue metric improves monotonically out to the largest sampledK=2\.1K\\\!=\\\!2\.1M, so the cheap in\-batch signal points to exactly the wrong end of theKKaxis\. This is in fact what we should expect from the form of the objective: the optimized loss is the average of the batch\-local and extra\-candidate cross\-entropies, and asKKgrows the extra\-candidate term—scored against an ever\-larger negative pool—increasingly dominates the gradient\. The optimizer therefore trades away in\-batch fit, so the batch\-local loss drifts*up*withKKeven as the model gets better at the full\-catalogue task that the deployed metrics reward\. AtC≤1017C\\\!\\leq\\\!10^\{17\}the in\-batch loss instead carries little reliable signal: the ranking metrics are nearly flat acrossKK\(Δ​y≲2%\\Delta y\\\!\\lesssim\\\!2\\%, faded cells\) and the residualρS\\rho\_\{S\}is weak and mixed in sign\. In short, the cheap training\-time loss is at best uninformative and at worst actively misleading for choosingKK\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/train_vs_eval_phase4_K_surrogacy.png)Figure 18:Negative\-sampling efficiency surrogacy\.Cell color is SpearmanρS\\rho\_\{S\}between theKK\-comparable*in\-batch*training loss and the eval metric across theKKaxis \(the optimized combined objective folds in aKK\-dependent extra\-negative channel and is not comparable acrossKK; cf\. Fig\.[16](https://arxiv.org/html/2606.05257#A5.F16)\)\. Cells where the underlyingKK\-sweep is essentially flat are faded\. AtC=1018C\\\!=\\\!10^\{18\}–101910^\{19\}the in\-batch loss is near\-perfectly*anti*\-correlated with every deployed metric \(\|ρS\|≈0\.97\|\\rho\_\{S\}\|\\\!\\approx\\\!0\.97–1\.001\.00\): theKKthat minimizes it \(K=0K\\\!=\\\!0\) is the worstKKfor full\-catalogue performance\. At smaller budgets the eval metrics are nearly flat acrossKKand the signal is weak\.

## Appendix FMaximal Update Parameterization

We swept four model sizes \(tiny/small/medium/large\), spanning∼10\\sim\\\!10M to∼500\\sim\\\!500M*total*trainable parameters with the embedder and contextualizer scaled together, and learning rates in\[10−5,5⋅10−2\]\[10^\{\-5\},\\,5\\\!\\cdot\\\!10^\{\-2\}\]with two initialization strategies:Default\(our usual truncated\-normal–style init, as in the rest of this work\) andMuP\(output and hidden weights LR\-scaled by1/fan​\_​in1/\\mathrm\{fan\\\_in\}, embeddings and biases unscaled\)\. We use88MuP learning rates spanning\{10−4,…,5⋅10−2\}\\\{10^\{\-4\},\\dots,5\\\!\\cdot\\\!10^\{\-2\}\\\}so that the MuP optimum is bracketed from above at every size, i\.e\. verified as a local minimum rather than as the right edge of the swept range\.

Table 14:MuP halves the LR drift but does not improve loss\.Training loss per \(size, init\) at the optimal LR; the MuP optima are verified local minima \(bracketed from above by strictly\-higher LRs that are strictly worse\)\. Our default initialization wins by0\.680\.68–0\.920\.92nats at every scale\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/all_losses_vs_lr_mupv5_stdv5.png)\(a\)Loss vs\. LR for every \(size, init\) cell\.
![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/mup_lr_stability_mupv5_stdv5.png)\(b\)Optimal LR vs\. model size\.

Figure 19:MuP halves the LR drift but does not improve loss\.Solid: MuP, dashed: Default\. MuP optima land at10−210^\{\-2\}attiny,smallandlarge, and at5⋅10−35\\\!\\cdot\\\!10^\{\-3\}atmedium\(a0\.300\.30\-decade band\) versus0\.700\.70decades for Default\. Default initialization sits below MuP at every model size\.Findings\.MuP delivers most of its core promise: the optimum LR is exactly10−210^\{\-2\}at three of the four sizes \(the MuP\-optimum span is0\.300\.30decades against0\.700\.70for Default\)\. Every MuP optimum is a verified local minimum: at every size,2⋅10−22\\\!\\cdot\\\!10^\{\-2\}and5⋅10−25\\\!\\cdot\\\!10^\{\-2\}produce strictly worse loss \(Figure[19\(a\)](https://arxiv.org/html/2606.05257#A6.F19.sf1)\)\. But Default reaches a lower training loss at every size by0\.680\.68–0\.920\.92nats \(Table[14](https://arxiv.org/html/2606.05257#A6.T14)\)\. The same pattern holds for every other MuP variant we tried, including a per\-layer MuP and a FLOP\-budget\-matched variant; we therefore retain our default truncated\-normal–style initialization \(Table[7](https://arxiv.org/html/2606.05257#S9.T7)\) and pay the modest cost of a per\-phase LR sweep instead\.

## Appendix GCross\-Metric Details

#### How well does training loss track eval metrics?

Within Stage 1 the sampled\-softmax training loss correlates with the headline eval metrics \(excludingMRR@10andcoverage@10\) at\|ρS\|≥0\.99\|\\rho\_\{S\}\|\\\!\\geq\\\!0\.99\(Table[15](https://arxiv.org/html/2606.05257#A7.T15)\)\. This is partly tautological: Stage 1 evaluates against a batch\-local pool, i\.e\. the same construction the optimizer sees, and at our budgets the train/val gap is small\. Stage 2 \(Phase 4\) is the genuine batch\-local\-vs\-full\-catalogue comparison; using theKK\-comparable in\-batch cross\-entropy, pooled across the whole sweep it tracks the eval metrics only moderately \(\|ρS\|∈\[0\.51,0\.90\]\|\\rho\_\{S\}\|\\\!\\in\\\!\[0\.51,0\.90\]\)\. This pooled figure is, however, dominated by the compute axis \(largerCClowers in\-batch loss and raises every metric\) and*masks*the per\-budget behavior alongKK: at fixed high compute the in\-batch loss*anti*\-tracks the deployed metrics almost perfectly \(Fig\.[18](https://arxiv.org/html/2606.05257#A5.F18)\), so it is not a usable surrogate for choosingKK\.

Table 15:How well batch\-local training loss tracks validation eval metrics\.Stage 1 medians are tautologically high \(same in\-batch construction\)\. Stage 2 uses theKK\-comparable in\-batch cross\-entropy channel and pools all\(K,C\)\(K,C\)cells; this pooled value is dominated by the compute axis and should be read together with the per\-budget Fig\.[18](https://arxiv.org/html/2606.05257#A5.F18), which shows the in\-batch loss*anti*\-tracking the deployed metrics at fixed high compute\.
#### Stratifying by compute budget\.

The loss–ranking link stays essentially flat atρS≈−0\.98\\rho\_\{S\}\\\!\\approx\\\!\-0\.98across the full compute range, whereas the loss–coverage link weakens monotonically and*collapses*atC=1018C\\\!=\\\!10^\{18\}\(Table[16](https://arxiv.org/html/2606.05257#A7.T16), Figure[20](https://arxiv.org/html/2606.05257#A7.F20)\): once the model is good enough that catalogue coverage saturates, which cells happen to spread the head distribution out furthest is essentially independent of which cells minimize loss\.

Table 16:Per\-budget Spearman correlationsbetween headline metrics \(Stage 1 pool\)\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/correlation_heatmap_by_budget.png)Figure 20:Correlation matrices stratified by budget\.The loss/perplexity/entropy/ranking block stays saturated at everyCC, but the coverage rows and columns visibly fade with scale\.
#### Stage 2 per\-budget correlations\.

Table[17](https://arxiv.org/html/2606.05257#A7.T17)is the Phase 4 counterpart to Table[16](https://arxiv.org/html/2606.05257#A7.T16): per\-budget Spearman correlations on the best\-LR\-per\-cellKK\-sweep \(8–9KK\-cells per budget, full\-catalogue evaluation\)\. Three contrasts with the Stage 1 table are worth flagging\.*\(i\)*The loss–ranking link is*not*locked:val\_loss–recall@10swings from\+0\.93\+0\.93atC=1015C\\\!=\\\!10^\{15\}\(loss and recall move in the*same*direction acrossKK, so pushingKKup reduces both\) through near\-zero atC∈\{1016,1017\}C\\\!\\in\\\!\\\{10^\{16\},10^\{17\}\\\}to−1\.00\-1\.00atC=1018C\\\!=\\\!10^\{18\}\(perfect alignment in the expected direction\)\.*\(ii\)*The loss–coverage sign*flips*:−0\.97\-0\.97–−0\.22\-0\.22in Stage 1 \(bigger architectures get both lower loss and higher coverage\) versus\+0\.60\+0\.60–\+1\.00\+1\.00in Stage 2 \(largerKKsharpens predictions, which lowers loss but*narrows*coverage\)\.*\(iii\)*The Stage 2nnis small \(8–9 per budget\) so the exact magnitudes are noisy, but the qualitative pattern is robust: theKK\-axis decouples the four loss\-like and ranking\-like metrics in a budget\-dependent way\. The mechanism is the same one analysed mathematically in §[7](https://arxiv.org/html/2606.05257#S7)\(importance\-reweighting\): the in\-batch contrastive loss does not learn absolute popularity, so the loss minimizer and the full\-catalogue ranking maximizer need not coincide, and the gap shrinks as more sampled negatives pushqqtoward uniform\.

Table 17:Per\-budget Spearman correlations \(Stage 2, Phase 4KK\-sweep\)\.Best\-LR\-per\-cell;nnis the number ofKK\-cells available at that budget\.*Both*the “loss” here and every ranking, coverage and entropy metric arefull\-catalogue evaluationquantities, computed against the full∼13\.6\\sim\\\!13\.6M\-item Stage\-2 catalogue: “loss” is the full\-catalogueval\_loss, not the training\-time loss\. These are therefore eval\-metric vs\. eval\-metric correlations acrossKK, distinct from the train\-loss surrogacy of Fig\.[18](https://arxiv.org/html/2606.05257#A5.F18)\. Compare with the Stage 1 table \([16](https://arxiv.org/html/2606.05257#A7.T16)\): loss–ranking is no longer locked atρS≈−0\.98\\rho\_\{S\}\\\!\\approx\\\!\-0\.98\(it swings from\+0\.93\+0\.93to−1\.00\-1\.00acrossCC\), and the loss–coverage sign flips relative to Stage 1\.
#### Why absolute loss is incomparable across stages\.

Stage 1 evaluates on a batch\-local pool of∼5\\sim\\\!5–1010k items, Stage 2 on the full deployed catalogue\. The cross\-entropy of a uniform predictor differs by several nats just from this candidate\-pool denominator, which by itself accounts for most of the observed∼7\\sim\\\!7\-nat gap in Figure[21](https://arxiv.org/html/2606.05257#A7.F21)\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/regime_contrast_loss_vs_recall.png)Figure 21:Same recall@10, very different absolute val\_losses depending on stage\.Blue: Stage 1 \(batch\-local pool\); red: Stage 2 \(full catalogue\)\. Within either cloud\|ρS\|≥0\.97\|\\rho\_\{S\}\|\\\!\\geq\\\!0\.97; across clouds the link breaks because the partition function is computed over a∼1500×\\sim\\\!1500\\\!\\timeslarger set\.

## Appendix HContext\-Length Scoring Robustness \(full slice matrices\)

This appendix backs axis \(e\) of §[7](https://arxiv.org/html/2606.05257#S7)\. The eval pipeline stratifies every batch by*scoring*history position and logs every headline metric atc​l∈\{3,5,10,20,50,100\}cl\\\!\\in\\\!\\\{3,5,10,20,50,100\\\}events of preceding context alongside the aggregateval/all\. The question is whether the cell ranking \(architectural cells for Stage 1,KK\-cells for Stage 2\) depends on which slice we score on\. All models are trained at the fullLseq=256L\_\{\\mathrm\{seq\}\}\\\!=\\\!256context length, so this measures the scoring\-position half of the context\-length sensitivity, not the training\-context half \(§[9](https://arxiv.org/html/2606.05257#S9)\)\.

For each \(regime, metric, budget\) we compute the Spearmanρ\\rhobetween every pair of scoring\-slice columns across the cells at that budget\. Figure[22](https://arxiv.org/html/2606.05257#A8.F22)reports the worst\-case off\-diagonalρ\\rhoper panel; Figures[23](https://arxiv.org/html/2606.05257#A8.F23)–[24](https://arxiv.org/html/2606.05257#A8.F24)break the same data out as full 7×\\times7 heatmaps for each \(regime, metric, budget\)\.

Stage 1 \(batch\-local; Phase 1W\+1D\+3 architecture sweeps,n=19n\\\!=\\\!19–3333cells per budget\)\.The headline ranking metrics \(recall@10,NDCG@10,NDCG@100\) all holdρmin≥0\.93\\rho\_\{\\min\}\\\!\\geq\\\!0\.93atC≤1018C\\\!\\leq\\\!10^\{18\};val\_lossandval\_entropysit atρmin≥0\.94\\rho\_\{\\min\}\\\!\\geq\\\!0\.94forC≤1017C\\\!\\leq\\\!10^\{17\}and soften toρmin≈0\.77\\rho\_\{\\min\}\\\!\\approx\\\!0\.77atC=1018​–​1019C\\\!=\\\!10^\{18\}\\text\{\-\-\}10^\{19\}where the val\-loss landscape itself is flatter \(§[3\.2](https://arxiv.org/html/2606.05257#S3.SS2)\);MRR@10dips occasionally to≈0\.83\\approx\\\!0\.83on small\-nncells, andcoverage@10decouples earlier as in Table[16](https://arxiv.org/html/2606.05257#A7.T16)\. Either way, Phase\-1/3 architectural winners transfer under shorter\- or longer\-history scoring; the architectural argmax is essentially independent of scoring position\.

Stage 2 \(full\-catalogue; Phase 4KK\-sweep,n=9n\\\!=\\\!9–1111cells per budget\)\.Ranking metrics showρmin∈\[0\.26,0\.73\]\\rho\_\{\\min\}\\\!\\in\\\!\[0\.26,\\,0\.73\]atC≤1017C\\\!\\leq\\\!10^\{17\}and only realign toρmin≥0\.87\\rho\_\{\\min\}\\\!\\geq\\\!0\.87atC≥1018C\\\!\\geq\\\!10^\{18\}\. The geometry visible in Figure[24](https://arxiv.org/html/2606.05257#A8.F24)is that the short\-context slices \(ctx\_3,ctx\_5\) and the long\-context tail \(ctx\_100,val/all\) only weakly rank\-agree; the middle slices \(c​l∈\[10,50\]cl\\\!\\in\\\!\[10,50\]\) sit in between\. This is consistent with the importance\-weighting mechanism of §[7](https://arxiv.org/html/2606.05257#S7): short\-history queries are popularity\-dominated \(little informative context to condition on\) so theKKthat minimises their conditional weights the1/q1/qsampling bias differently from a long\-history query; onceKKis large enough to pushqqtoward uniform the two converge, which is exactly theC≥1018C\\\!\\geq\\\!10^\{18\}realignment\. The practical corollary is that theK⋆K^\{\\star\}band in Table[6](https://arxiv.org/html/2606.05257#S6.T6)should be read at the deployed serving context length: atC=1015​–​1017C\\\!=\\\!10^\{15\}\\text\{\-\-\}10^\{17\}a model picked onrecall@10pooled overval/allis not guaranteed to be the best model for a cold\-start user\. Note the two stages move in opposite directions with compute: Stage 2 slice agreement*rises*with budget \(theC≥1018C\\\!\\geq\\\!10^\{18\}realignment above\), whereas Stage 1 stays high throughout with no upward trend and itsval\_loss/val\_entropyagreement even softens slightly at the top budgets\.

![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/context_length_minrho_vs_C.png)Figure 22:Context\-length scoring robustness: worst\-case cell\-ranking agreement across history\-position slices, per regime\.Per \(regime, metric, budget\) we compute the Spearmanρ\\rhobetween every pair of scoring slices\{\\\{ctx\_3,ctx\_5,ctx\_10,ctx\_20,ctx\_50,ctx\_100,val/all\}\\\}and plot the minimum off\-diagonal pair\.*Left:*batch\-local Stage 1 evals \(architecture sweep\): every headline ranking metric staysρmin≥0\.93\\rho\_\{\\min\}\\\!\\geq\\\!0\.93atC≤1018C\\\!\\leq\\\!10^\{18\}\.*Right:*full\-catalogue Stage 2 evals \(KK\-sweep\): ranking metricsρmin∈\[0\.26,0\.73\]\\rho\_\{\\min\}\\\!\\in\\\!\[0\.26,0\.73\]atC≤1017C\\\!\\leq\\\!10^\{17\}, realigning toρmin≥0\.87\\rho\_\{\\min\}\\\!\\geq\\\!0\.87atC≥1018C\\\!\\geq\\\!10^\{18\}on the same compute threshold at which \(c\)’s loss\-vs\-ranking correlation flips to perfectly aligned\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/context_length_heatmap_batchlocal.png)Figure 23:Per\-budget Spearmanρ\\rhobetween context\-length scoring slices, batch\-local evals \(Phase 1W\+1D\+3 architecture sweep\)\.Each row is a metric, each column a budget; the per\-budget cell\-countnnis in the column header\. Every cell is the Spearman rank correlation between two scoring\-slice columns across the architectural cells at that budget\.![Refer to caption](https://arxiv.org/html/2606.05257v1/figures/context_length_heatmap_fullvocab.png)Figure 24:Per\-budget Spearmanρ\\rhobetween context\-length scoring slices, full\-catalogue evals \(Phase 4KK\-sweep\)\.Same axes as Figure[23](https://arxiv.org/html/2606.05257#A8.F23)\. Therecall@10,NDCG@10andMRR@10panels atC∈\{1015,1016,1017\}C\\\!\\in\\\!\\\{10^\{15\},10^\{16\},10^\{17\}\\\}are the cells driving theρmin∈\[0\.26,0\.73\]\\rho\_\{\\min\}\\\!\\in\\\!\[0\.26,0\.73\]summary of Figure[22](https://arxiv.org/html/2606.05257#A8.F22)\.

Similar Articles

Scaling laws for neural language models

OpenAI Blog

Foundational empirical study demonstrating power-law scaling relationships between language model performance and model size, dataset size, and compute budget, with implications for optimal training allocation and sample efficiency.

Scaling Laws, Carefully (25 minute read)

TLDR AI

A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.

Prescriptive Scaling Laws for Data Constrained Training

Hugging Face Daily Papers

A modified scaling law accounting for data repetition effects provides compute-optimal training strategies for data-constrained scenarios, showing that beyond a point further repetition is counterproductive and compute is better spent on model capacity.