Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
Summary
This monograph develops a unified account of training and inference dynamics in Power Law Decoder Representation language models (PLDR-LLMs), covering exact identities, renormalization, and experimental findings.
View Cached Full Text
Cached at: 09/29/26, 08:15 AM
Paper page - Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
Source: https://huggingface.co/papers/2609.34130
Abstract
ThismonographdevelopsaunifiedaccountoftrainingandinferenceinPowerLawDecoderRepresentationlanguagemodels(PLDR-LLMs).Exactfiniteworkidentitiesdecomposechangesintheabsoluteenergyoftherow-centeredlearnedmapintoparametercontributions,signedinteractions,andnumericalobservationdefects.Positiveaffineblockingretainsrestartsattherow-constantface,whiletheaugmentedAdamWstatesuppliesthecompletedynamicaldescription.Predictiverenormalizationactsonthecompleteconditionaltraininglawforasinglepassoverdistinctcorpustargetblocks,retainingoptimizermemory,remainingdata,schedule,andnumericalpolicy.Autonomousreductionsrequireclosure;approximatereductionscarrysuccessorandemissionerrors.Finite-populationcovariance,matchedphysicalclocks,matrixfluxes,andsignedtemporalenergyconnectrowdynamicstomodel-wideobservations.Absoluterowcollapse,relativerowconcentration,operatorstabilization,andpredictiveaccuracyaredistinguished.Experimentsrevealobserverandoptimizerdependence,rejectthetestedautonomousrow-statecandidates,andsupportfiniteconditionalpredictionandstate-specificoperatorreduction.Independentsingle-passfamiliesexhibitmovingfinitefluctuationregionswithoutestablishingathermodynamiccriticalclass.Conditionalsymmetry,headlimits,covarianceflows,andreadouterrorbudgetsspecifyassumptionsneededtotransferscalinglawstoinference.Thetheoryseparatesexactidentities,conditionaldynamicalclaims,andfiniteempiricalfindings,withproofs,selectedformalchecks,andcompactnumericalevidence.
View arXiv pageView PDFProject pageGitHub0Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34130 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34130 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34130 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
This paper introduces Layer-wise Representation Dynamics (LRD), a framework with three measurement families to analyze how hidden states change across layers in language models. Applied to 31 models on 30 MTEB tasks, LRD reveals architectural differences and enables label-free model selection and inference-time layer pruning.
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
This paper investigates when chain-of-thought reasoning is beneficial for LLMs, showing that early-stage entropy dynamics reliably indicate reasoning utility, and introduces EDRM, a lightweight, training-free framework that adaptively selects inference strategies to achieve significant token savings while maintaining or improving accuracy.
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
This paper introduces ScaleLogic, a framework demonstrating that RL training compute scales as a power law with reasoning depth in LLMs. It highlights that logical expressiveness is key to improving downstream transfer and training efficiency.
Learning to Refine Hidden States for Reliable LLM Reasoning
Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.
GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training
This paper analyzes post-training weight updates in LLMs using singular value decomposition, identifying geometric components that drive performance gains, with insights suggesting that reshaping singular values is less critical than rotating and routing changes.