@snowboat84: I've been pondering this for years: the relationship between statistical mechanics and AI. Statistical mechanics, using a statistical approach to molecular dynamics, reproduces the elegant fundamental theorems of thermodynamics, especially the beautiful relationships between macroscopic quantities like entropy, free energy, and of course temperature and pressure. The question is, does AI have these thermodynamic mac...

X AI KOLs Timeline Papers

Summary

This tweet explores the relationship between statistical mechanics and artificial intelligence, citing a paper that proposes a thermodynamic theory for machine learning systems, introducing concepts like temperature, entropy, and energy, and treating the training process as a phase transition.

I've been pondering this for years: the relationship between statistical mechanics and AI. Statistical mechanics, using a statistical approach to molecular dynamics, reproduces the elegant fundamental theorems of thermodynamics, especially the beautiful relationships between macroscopic quantities like entropy, free energy, and of course temperature and pressure. The question is, does AI have these thermodynamic macroscopic quantities? If AI cannot be expressed in terms of macroscopic quantities, then it cannot fundamentally be studied using the methods of statistical mechanics. A few years ago, I did some thinking and proposed reintroducing the concept of temperature in machine learning systems, viewing the training process as a change in the state of a heat engine. https://arxiv.org/abs/2404.13218
Original Article
View Cached Full Text

Cached at: 06/26/26, 10:11 AM

For years I have been pondering this question: the relationship between statistical mechanics and AI. Statistical mechanics, through the statistical method of molecular dynamics, reproduces the beautiful fundamental theorems of thermodynamics, especially the relationships between some extremely elegant macroscopic quantities in thermodynamics, such as entropy, free energy, and of course temperature, pressure, etc. The question is, does AI possess these thermodynamic macroscopic quantities? If AI cannot be represented by macroscopic quantities, then it is essentially impossible to study it using the methods of statistical mechanics. I did some thinking a few years ago, re-proposing the concept of temperature in machine learning systems, and viewing the training process as a state change process of a heat engine. https://arxiv.org/abs/2404.13218 — # On the Temperature of Machine Learning Systems Source: https://arxiv.org/html/2404.13218 ###### Abstract We develop a thermodynamic theory for machine learning (ML) systems. Similar to physical thermodynamic systems which are characterized by energy and entropy, ML systems possess these characteristics as well. This comparison inspire us to integrate the concept of temperature into ML systems grounded in the fundamental principles of thermodynamics, and establish a basic thermodynamic framework for machine learning systems with non-Boltzmann distributions. We introduce the concept of states within a ML system, identify two typical types of state, and interpret model training and refresh as a process of state phase transition. We consider that the initial potential energy of a ML system is described by the model’s loss functions, and the energy adheres to the principle of minimum potential energy. For a variety of energy forms and parameter initialization methods, we derive the temperature of systems during the phase transition both analytically and asymptotically, highlighting temperature as a vital indicator of system data distribution and ML training complexity. Moreover, we perceive deep neural networks as complex heat engines with both global temperature and local temperatures in each layer. The concept of work efficiency is introduced within neural networks, which mainly depends on the neural activation functions. We then classify neural networks based on their work efficiency, and describe neural networks as two types of heat engines. Keywords: machine learning system, thermodynamics, temperature, entropy, energy, phase transition, heat engine, work efficiency ## 1Introduction From the perspective of information theory, data carries entropy. The concept of entropy originated from thermodynamics and statistical mechanics, where it describes the disorder or randomness in a physical system. Claude Shannon later extended the idea of entropy to information theory to measure the uncertainty of random variables[1 (https://arxiv.org/html/2404.13218v1#bib.bib1)], while Norbert Wiener also discussed entropy in the context of cybernetics, especially differential entropy[2 (https://arxiv.org/html/2404.13218v1#bib.bib2)]. The employment of entropy in machine learning is an adaptation from information theory. For example, cross-entropy and information gain are used for splitting nodes in decision trees and random forests. In unsupervised learning, entropy can be used to evaluate the quality of clusters. Overall, the usage of entropy in data systems and machine learning is fundamentally rooted in the principles of thermodynamics and information theory, demonstrating a diverse and interdisciplinary application of the concept. On the other hand, a physical system has energy, as well as entropy. If the concept of entropy can be introduced into a data system, does data also have energy? In the field of machine learning, there is a category of models known as energy-based models (EBMs)[3 (https://arxiv.org/html/2404.13218v1#bib.bib3),4 (https://arxiv.org/html/2404.13218v1#bib.bib4)]. The origins of these EBMs can be traced back to the Ising model in statistical physics[5 (https://arxiv.org/html/2404.13218v1#bib.bib5),6 (https://arxiv.org/html/2404.13218v1#bib.bib6)]and the Amari-Hopfield network[7 (https://arxiv.org/html/2404.13218v1#bib.bib7),8 (https://arxiv.org/html/2404.13218v1#bib.bib8)]. The Boltzmann Machines (BMs) were proposed as stochastic recurrent neural networks[9 (https://arxiv.org/html/2404.13218v1#bib.bib9)], inspired by the Ising model as well as spin-glass model in physics[10 (https://arxiv.org/html/2404.13218v1#bib.bib10)]. To simplify the training process and improve computational efficiency, the Restricted Boltzmann Machines (RBMs) were later developed[11 (https://arxiv.org/html/2404.13218v1#bib.bib11),12 (https://arxiv.org/html/2404.13218v1#bib.bib12)]. The RBMs introduced a restriction that the neurons must form a bipartite graph, which significantly improved the training efficiency. Since the advent of RBMs, a variety of methods and applications have been proposed under the umbrella of EBMs, contributing to the evolution and expansion of this field[13 (https://arxiv.org/html/2404.13218v1#bib.bib13),14 (https://arxiv.org/html/2404.13218v1#bib.bib14),15 (https://arxiv.org/html/2404.13218v1#bib.bib15),16 (https://arxiv.org/html/2404.13218v1#bib.bib16),17 (https://arxiv.org/html/2404.13218v1#bib.bib17),18 (https://arxiv.org/html/2404.13218v1#bib.bib18),19 (https://arxiv.org/html/2404.13218v1#bib.bib19)]. The fundamental concept of an EBM is to define an energy function that satisfiesEμ⁢(x)=−log⁡pμ⁢(x)subscriptEμxsubscriptpμxE_{\mu}(x)=-\log p_{\mu}(x)italic_E start_POSTSUBSCRIPT italic_μ end_POSTSUBSCRIPT ( italic_x ) = - roman_log italic_p start_POSTSUBSCRIPT italic_μ end_POSTSUBSCRIPT ( italic_x ), orpμ⁢(x)=exp⁡[−Eθ⁢(x)/Zθ]subscriptpμxsubscriptEθxsubscriptZθp_{\mu}(x)=\exp[-E_{\theta}(x)/Z_{\theta}]italic_p start_POSTSUBSCRIPT italic_μ end_POSTSUBSCRIPT ( italic_x ) = roman_exp [ - italic_E start_POSTSUBSCRIPT italic_θ end_POSTSUBSCRIPT ( italic_x ) / italic_Z start_POSTSUBSCRIPT italic_θ end_POSTSUBSCRIPT ], whereEμ⁢(x)subscriptEμxE_{\mu}(x)italic_E start_POSTSUBSCRIPT italic_μ end_POSTSUBSCRIPT ( italic_x )is the energy function with parameter setμμ\muitalic_μ, andZθsubscriptZθZ_{\theta}italic_Z start_POSTSUBSCRIPT italic_θ end_POSTSUBSCRIPTis the partition function as the normalizing constant. This relationship between probability and energy aligns with the principles of statistical physics, and we can optimize either the loss function or the energy function to train the models. Of course, the so-called “energy of data” is not physical energy in the real world. Instead, it is an analogy drawn between data systems and the real physical world, serving as an indicator to describe the properties of data and machine learning systems associated with the data. However, just like the introduction of entropy in information theory, we can consider the energy in data systems as a kind of generalized energy. Building on this concept, it leads us to consider:if a machine learning (ML) system itself has energy, and incorporates the concept of entropy, the machine learning (ML) system can essentially be analogized to a thermodynamic system.This raises an important question:could we define temperature-like quantities to characterize the properties of a ML system?The concept of temperature already exists in machine learning as a scaling parameter used to control the randomness of predictions made by models. For example, in BMs and RBMs, the temperature parameter appears in the Boltzmann distribution, with higher temperatures leading to more uniform distributions over states[20 (https://arxiv.org/html/2404.13218v1#bib.bib20),21 (https://arxiv.org/html/2404.13218v1#bib.bib21),22 (https://arxiv.org/html/2404.13218v1#bib.bib22)]. Similarly, the temperature parameter in softmax classifiers for multi-class classification is applied to the logits before the softmax function, with higher temperatures giving more similar probabilities to all classes[23 (https://arxiv.org/html/2404.13218v1#bib.bib23),24 (https://arxiv.org/html/2404.13218v1#bib.bib24),25 (https://arxiv.org/html/2404.13218v1#bib.bib25)]. Temperature can be used to control the creativity of a generative model, where higher temperatures will make more novel and unexpected predictions more likely[26 (https://arxiv.org/html/2404.13218v1#bib.bib26),27 (https://arxiv.org/html/2404.13218v1#bib.bib27)]. However, the existing concept of temperature in machine learning is merely a single model parameter, which cannot be derived from first principles, nor can it reflect the overall thermodynamic properties of the ML system.Overall, although the concepts of energy, entropy, and temperature exist in the field of machine learning, no one has yet unified the three concepts together, nor viewed ML systems as thermodynamic systems from first principles. This paper systematically proposes a theory of thermodynamics and statistical mechanics in machine learning, with a particular focus on discussing the concept of temperature for ML systems. We can gain inspiration from the thermodynamic potentials in the real physical world. The thermodynamic potentials are fundamental concepts that describes the energy characteristics of a thermodynamic system. They are scalar quantities that provide information about the system state, and are used to understand how the system will respond to changes in temperature, pressure, and volume. There are several thermodynamic potentials, including internal energy (UUUitalic_U), Helmholtz free energy (FFFitalic_F), enthalpy (HHHitalic_H) and Gibbs free energy (GGGitalic_G). Correspondingly, we have a set of four equations known as the fundamental thermodynamic relations to describe these thermodynamic potentials, which are essential in understanding the behavior of thermodynamic systems[28 (https://arxiv.org/html/2404.13218v1#bib.bib28)] d⁢UdU\displaystyle dUitalic_d italic_U=\displaystyle==T⁢d⁢S−P⁢d⁢V+∑iμi⁢d⁢NiTdSPdVsubscriptisubscriptμidsubscriptNi\displaystyle TdS-PdV+\sum_{i}\mu_{i}dN_{i}italic_T italic_d italic_S - italic_P italic_d italic_V + ∑ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_μ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_d italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT(1.1)d⁢FdF\displaystyle dFitalic_d italic_F=\displaystyle==−S⁢d⁢T−P⁢d⁢V+∑iμi⁢d⁢NiSdTPdVsubscriptisubscriptμidsubscriptNi\displaystyle-SdT-PdV+\sum_{i}\mu_{i}dN_{i}- italic_S italic_d italic_T - italic_P italic_d italic_V + ∑ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_μ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_d italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT(1.2)d⁢HdH\displaystyle dHitalic_d italic_H=\displaystyle==T⁢d⁢S+V⁢d⁢P+∑iμi⁢d⁢NiTdSVdPsubscriptisubscriptμidsubscriptNi\displaystyle TdS+VdP+\sum_{i}\mu_{i}dN_{i}italic_T italic_d italic_S + italic_V italic_d italic_P + ∑ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_μ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_d italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT(1.3)d⁢GdG\displaystyle dGitalic_d italic_G=\displaystyle==−S⁢d⁢T+V⁢d⁢P+∑iμi⁢d⁢Ni,SdTVdPsubscriptisubscriptμidsubscriptNi\displaystyle-SdT+VdP+\sum_{i}\mu_{i}dN_{i},- italic_S italic_d italic_T + italic_V italic_d italic_P + ∑ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_μ start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT italic_d italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT ,(1.4)whereSSSitalic_S,TTTitalic_T,PPPitalic_P,VVVitalic_Vare entropy, temperature, pressure and volume of the system respectively. For fixed number of particles, volume or pressure, we have the following equations of state for temperature: T=(∂U∂S)V,{Ni}=(∂H∂S)P,{Ni}.TsubscriptUSVsubscriptNisubscriptHSPsubscriptNiT=\left(\frac{\partial U}{\partial S}\right)_{V,\{N_{i}\}}=\left(\frac{% \partial H}{\partial S}\right)_{P,\{N_{i}\}}.italic_T = ( divide start_ARG ∂ italic_U end_ARG start_ARG ∂ italic_S end_ARG ) start_POSTSUBSCRIPT italic_V , { italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT } end_POSTSUBSCRIPT = ( divide start_ARG ∂ italic_H end_ARG start_ARG ∂ italic_S end_ARG ) start_POSTSUBSCRIPT italic_P , { italic_N start_POSTSUBSCRIPT italic_i end_POSTSUBSCRIPT } end_POSTSUBSCRIPT .(1.5) Refer to captionFigure 1:A machine learning (ML) system includes the initial model design, setting of initial parameters, importing data for training, importing new data for prediction tasks, and the process of keeping the model refreshed with new data. From a physics perspective, such a system with a series of steps is analogous to a heat engine. We can examine the various temperatures of the system during these processes, as well as the changes in energy and entropy.On the other hand, a ML system has energy and entropy, but lacks some other physical quantities such as volume and pressure. We can use a method similar to equation (1.5 (https://arxiv.org/html/2404.13218v1#S1.E5)) to calculate the ML system temperature. Consider a system transitioning from one state to another, where the change in energy isΔ⁢EΔE\Delta Eroman_Δ italic_E, and the change in entropy isΔ⁢SΔS\Delta Sroman_Δ italic_S, the for an equilibrium system, we can useT=Δ⁢E/Δ⁢STΔEΔST=\Delta E/\Delta Sitalic_T = roman_Δ italic_E / roman_Δ italic_Sto calculate the temperature of the system. Note that for a machine learning system, we should not simply view it as a static data processor that goes from raw data input to prediction output. Instead, we should see it as a dynamic process, which includes a series of processes including the initial design of the model architecture, parameter initialization, optimization of the loss function and parameter tuning, and keep refreshing the model dynamically due to data shifts. Therefore, the temperature of a ML system must also be used to describe the entire process above. Figure1 (https://arxiv.org/html/2404.13218v1#S1.F1)shows such a system with multiple processes, and we can explore the various system temperature in different steps. This paper is organized as follows. In Section2 (https://arxiv.org/html/2404.13218v1#S2), we develop the general theory to build a thermodynamic framework for ML systems. Three pivotal elements define an ML system: the data flow, model structure with its parameters, and system energy. In particular, we introduce two states of an ML system, corresponding to the stages of parameter initialization and data shifting respectively (Section2.1 (https://arxiv.org/html/2404.13218v1#S2.SS1)). The ML training process can be viewed as an isothermal phase transition process, and the two states can be unified into a global picture (AppendixC.1 (https://arxiv.org/html/2404.13218v1#A3.E1)). In Section2.2 (https://arxiv.org/html/2404.13218v1#S2.SS2), we assign new physical meaning to the model loss function, viewing it as the internal potential energy of a ML system, which follows the principle of minimum potential energy. We emphasize that the meaning of energy in our theory differs from energy in energy-based models in literature. We then review system entropy in Section2.3 (https://arxiv.org/html/2404.13218v1#S2.SS3)and discuss the relation between discrete and differential entropy under dimension collapse scenarios (AppendixB (https://arxiv.org/html/2404.13218v1#A2)). The comparison of the ML system from an ML perspective versus a thermodynamic perspective is presented in Section2.5 (https://arxiv.org/html/2404.13218v1#S2.SS5). Next, we derive the temperature of various ML systems based on different system energies and parameter initialization methods. In Section3 (https://arxiv.org/html/2404.13218v1#S3), we develop the temperature theory in ML systems with a linear regression model and mean square error (MSE) as the internal energy, while parameters are initialized by normal distribution (Section3.1 (https://arxiv.org/html/2404.13218v1#S3.SS1)), uniform distribution (Section3.2 (https://arxiv.org/html/2404.13218v1#S3.SS2)), and mixed distribution (Section3.3 (https://arxiv.org/html/2404.13218v1#S3.SS3)).

Similar Articles

@snowboat84: This is the second part of the "When Physics Meets AI" series. The role of physics in AI can be divided into four layers: (1) The first layer is the bottommost, providing the computational skeleton—energy, entropy, and free energy are embedded into AI's training objectives. (2) The second layer is the middle layer, where physics shapes the network architecture—Hopfield's Ising energy function, CNN's translational symmetry, and renormalization group correspond to the hierarchical structure of deep networks.

X AI KOLs Timeline

This article explores the four layers of physics' role in AI, from the bottom computational skeleton to the methodological layer, arguing that physics' methodology is migrating from the natural world to the AI domain.

@snowboat84: Several years ago, dissipative systems and nonlinear complex systems were extremely popular in academic and cultural circles. To fully review dissipative systems, one must start with non-dissipative thermodynamics. The second law of thermodynamics (entropy law) states that everything should move towards chaos and stillness. But life grows, forests succeed, and even the large models in data centers are constantly "learning" order. …

X AI KOLs Timeline

This is a popular science article of over 25,000 characters, starting from the origin of entropy, reviewing the development of dissipative system theory, and exploring a three-level analysis of whether AI belongs to dissipative systems (hardware level, training level, static model).

@snowboat84: To add a supplementary note, regarding the phenomena emerging from AI—scaling laws, emergence, double descent, representation geometry—the papers discussing them are already numerous. But there is a big problem: they are all thinking in the way of computer scientists, not physicists. What is a computer sci…

X AI KOLs Timeline

The author comments that current AI research overuses the thinking style of computer science and lacks a physics-based approach, proposing the need to establish an ideal system like 'Cyber Space' to lay a theoretical foundation.

@snowboat84: https://x.com/snowboat84/status/2062686432335184321

X AI KOLs Timeline

This article explores the deep connections between physics and deep learning, analyzes the isomorphism of phenomena such as Scaling Law and emergence with concepts like critical scaling laws and phase transitions in physics, and reviews the current status and prospects of applying physical methodologies in AI.

@freeman1266: Software engineering methodology must shift from the traditional 'state perspective' to a dynamical system perspective. The core view advocates that 'attractor logic takes precedence over governance tools (Harness)', that is, first define the structural invariants that the system should converge to in the long term, rather than merely focusing on local constraints and verification. AI, as a high-frequency and directionless perturbing...

X AI KOLs Timeline

The article proposes that software engineering methodology should shift from a state perspective to a dynamical system perspective, emphasizing that attractor logic takes precedence over governance tools. In the AI era, it is necessary to explicitly model state space, attractors, trajectories, and controls to address architectural drift caused by AI as a high-frequency perturbation source.