How Complexity Contributes to Learning Opacity in Machine Learning

arXiv cs.LG Papers

Summary

This paper analyzes why machine learning, particularly neural networks, remains opaque in its learning process by framing it as a complex dynamical system, identifying three key properties that contribute to learning opacity, and arguing that some sources may be irreducible.

arXiv:2606.24953v1 Announce Type: new Abstract: Machine learning (ML) algorithms are known to be opaque. We do not know the reasons for their predictions. The learning process leading to the prediction function is also opaque. We do not fully understand the time evolution of the weight values of neural nets (NN) and related dynamical phenomena. While prediction opacity is widely studied, learning opacity remains largely underexplored. This article studies learning opacity trough the lens of complex dynamical systems. We argue that NN learning is essentially a complex system and that learning opacity is due to dynamical complexity and the epistemological challenges that arise from it. We identify three key properties of training complexity -- sensitivity to weight initialization, feedback in gradient based optimization, and sensitivity to the training data -- and show how each contributes to learning opacity. As these properties are fundamental to the learning process damping or eliminating them would fundamentally alter how ML systems learn. Some sources of opacity in ML may hence be irreducible.
Original Article
View Cached Full Text

Cached at: 06/25/26, 05:07 AM

# How Complexity Contributes to Learning Opacity in Machine Learning†
Source: [https://arxiv.org/html/2606.24953](https://arxiv.org/html/2606.24953)
Joachim SteinHeidelberger Akademie der WissenschaftenUniversität Tübingen, Cluster of Excellence: Machine Learning – New Perspectives for ScienceEric RaidlUniversität Tübingen, Cluster of Excellence: Machine Learning – New Perspectives for ScienceHeidelberger Akademie der Wissenschaften

###### Abstract

Machine learning \(ML\) algorithms are known to be opaque\. We do not know the reasons for their predictions\. The learning process leading to the prediction function is also opaque\. We do not fully understand the time evolution of the weight values of neural nets \(NN\) and related dynamical phenomena\. While prediction opacity is widely studied, learning opacity remains largely underexplored\. This article studies learning opacity trough the lens of complex dynamical systems\. We argue that NN learning is essentially a complex system and that learning opacity is due to dynamical complexity and the epistemological challenges that arise from it\. We identify three key properties of training complexity – sensitivity to weight initialization, feedback in gradient based optimization, and sensitivity to the training data – and show how each contributes to learning opacity\. As these properties are fundamental to the learning process damping or eliminating them would fundamentally alter how ML systems learn\. Some sources of opacity in ML may hence be irreducible\.

††footnotetext:†\\daggerThis work has been funded by the Deutsche Forschungsgemeinschaft \(EXC number 2064/1, project no\. 390727645\) and the WIN program of the Heidelberg Academy of Sciences and Humanities, financed by the Ministry of Science, Research and the Arts of the State of Baden\-Württemberg\. For helpful discussion and feedback we thank Sarah Bopp, Timo Freiesleben, Miriam Klopotek, Jan\-Willem Romeijn and his group, Hans Rott, Tom Sterkenburg, Max Weinmann, Sebastian Zezulka and the participants of the SAS25\-Uncertainty conference in Stuttgart\.## 1\. Introduction

Machine learning, in particular neural networks \(NNs\), is a powerful automation of induction and learning in general\. Given a vast amount of data, machine learning algorithms are able to learn a reliable prediction function\.111In machine learning research, the learned prediction function is often referred to as ‘the model’\. However, to avoid confusion with scientific models used to represent systems, we will use the termprediction function\.However, the learned prediction function as well as the learning mechanism themselves remain relatively incomprehensible to us — they are termed ‘opaque’\[[5](https://arxiv.org/html/2606.24953#bib.bib188),[8](https://arxiv.org/html/2606.24953#bib.bib124),[45](https://arxiv.org/html/2606.24953#bib.bib119),[10](https://arxiv.org/html/2606.24953#bib.bib187),[48](https://arxiv.org/html/2606.24953#bib.bib118)\]\.

On the one hand we do not know what real\-world patterns are encoded in the prediction function\. Call thisprediction opacity\. Prediction opacity raises epistemic and ethical challenges in contexts like policy making\[[52](https://arxiv.org/html/2606.24953#bib.bib87)\]or scientific research\[[53](https://arxiv.org/html/2606.24953#bib.bib90)\]\. The need to overcome prediction opacity has lead to the field of ‘eXplainable Artifical Intelligence’ \(XAI\)\[\[, cf\.\]\]bordt2025position, molnar2025, DearXAICommu\. Researchers in this field develop interpretable machine learning methods and provide post\-hoc explanation techniques for established machine learning methods\. Existing philosophical contributions to XAI investigate, among other things, the reason for prediction opacity\[[49](https://arxiv.org/html/2606.24953#bib.bib148)\], analyse ways to counteract it\[[10](https://arxiv.org/html/2606.24953#bib.bib187)\]and how using opaque NN changes the way we gain understanding in science\[[5](https://arxiv.org/html/2606.24953#bib.bib188)\]\.

On the other hand, we do not understandhowthe NN has learned\. We take this to mean that we do not understand the dynamics of the learning process and related phenomena\. Call thislearning opacity\. Learning opacity is also a troublesome phenomenon for even the designers of NNs, who have full access to the learning process, might not understand how the network learns\. This is reflected in the great effort of NN researchers to explain phenomena of NN training like benign\-overfitting\[[2](https://arxiv.org/html/2606.24953#bib.bib135)\], implicit regularization\[[41](https://arxiv.org/html/2606.24953#bib.bib8),[7](https://arxiv.org/html/2606.24953#bib.bib13)\], saturation\[[14](https://arxiv.org/html/2606.24953#bib.bib107)\], grokking\[[42](https://arxiv.org/html/2606.24953#bib.bib97)\]or optimization at the edge of stability\[[9](https://arxiv.org/html/2606.24953#bib.bib20)\]\. These phenomena are surprising giving the orthodox view on optimization, and understanding them would help researchers gain better insights into the peculiarities of the optimization process in NNs\. Furthermore, since the prediction function is a direct product of the learning process, illuminating learning opacity most likely also advances the analysis of prediction opacity\.

Understanding the sources of learning opacity is also a necessary condition to develop means to counteract it, allowing researchers to deepen their understanding of the NN’s learning process\. This would have practical implications for debugging, improving training efficiency and real\-world deployment\. However, unlike in the case of prediction opacity, a thorough conceptual analysis of the sources for learning opacity is still lacking\.

In this article, we provide this conceptual analysis of learning opacity in NNs using insights from complexity science\. That even full access to the learning mechanism does not resolve learning opacity shows that it is not merely a matter of limited access\. Rather, learning opacity might be due to more intrinsic properties of the optimization heuristics\[\[, cf\.\]\]Burrell, Creel\_2020,Sullivan\_2022\. We explore this idea in more depth\. We clarify the nature of learning opacity in NN and how complexity may be a major contributing factor to it\. Drawing on complexity analysis from the sciences, we identify three characteristic features of complex dynamics — sensitive dependence on initial conditions, feedback and context sensitivity — and show how each limits understanding of systems in the natural sciences\. We demonstrate that these properties are also present in NN training, and examine how each of them contributes to learning opacity\.

The article is structured as follows\. Section[2](https://arxiv.org/html/2606.24953#S2)outlines the NN learning process and discusses learning opacity\. Section[3](https://arxiv.org/html/2606.24953#S3)introduces the notion of complexity from the natural sciences, argues that it provides a conceptual lens for investigating learning opacity, and identifies three central properties of complex systems: sensitive dependence on initial conditions, feedback, and context sensitivity\. It also introduces dynamical modelling, a central heuristic for understanding dynamical processes\. The remainder of the paper examines these three properties, how they limit researchers’ understanding of natural complex systems, and how they contribute to learning opacity in machine learning\. Section[4](https://arxiv.org/html/2606.24953#S4)focuses on sensitive dependence on initial conditions, Section[5](https://arxiv.org/html/2606.24953#S5)on feedback, and Section[6](https://arxiv.org/html/2606.24953#S6)on context sensitivity\.

## 2\. Neural Networks

Machine learning \(ML\) refers to the process of using a computer algorithm to learn patterns from data\. In this process, researchers provide data to a learning algorithm that aims to identify a function that encodes these patterns and is thus capable of predicting future observations, theprediction function\.222For an introduction to the theory of ML, see\[[44](https://arxiv.org/html/2606.24953#bib.bib112)\]\.\[[15](https://arxiv.org/html/2606.24953#bib.bib193)\]is the standard textbook on NN learning\.\[[46](https://arxiv.org/html/2606.24953#bib.bib9)\]and\[[47](https://arxiv.org/html/2606.24953#bib.bib17)\]provide a philosophical discussion of the foundations of ML\.In this work, we focus on supervised ML\.

In supervised ML, adata point\(x\(i\),y\(i\)\)∈𝒳×𝒴\(x^\{\(i\)\},y^\{\(i\)\}\)\\in\\mathcal\{X\}\\times\\mathcal\{Y\}consisting of an instancex\(i\)x^\{\(i\)\}from someinstance space𝒳\\mathcal\{X\}, e\.g\. a picture of an animal, paired with its labely\(i\)y^\{\(i\)\}from thelabel space𝒴\\mathcal\{Y\}, e\.g\. the name of the animal\. The process is termed supervised because the learning algorithm has access to the instance’s true label during training\.

The prediction function is selected from a predetermined class of functionsf:𝒳→𝒴f:\\mathcal\{X\}\\rightarrow\\mathcal\{Y\}, known as thehypothesis class\. In the case of NN learning, the hypothesis class consists of NNs\. A NN is a mathematical function that can be represented as a graph\(V,E\)\(V,E\), whereVVdenotes the set of nodes, theneurons, andEEthe set of edges transmitting information from the output of one neuron to the input of another\. This work focuses on feed\-forward networks, in which the graph is directed and acyclic, so that all edges point in the same direction and no cycles occur\. Accordingly, a feed\-forward NN is organized as a layered graph\.

A neuron’s outputted information is transformed by the neuron’sactivation functionσ\\sigma, which acts as a threshold function\. This transformed information is furtherweightedby a factorwi:E→ℝw\_\{i\}:E\\rightarrow\\mathbb\{R\}that determines how much information is transmitted along a given edge from one neuron to the next\. The vectorwwconsists of all weight factorswiw\_\{i\}\.333In this paper we usewiw\_\{i\}to denote a weight variable or its specific value\. Accordingly,wwmay denote vector containing variables or scalars\. The intended meaning ofwiw\_\{i\}andwwshould be clear from the context\.A neuron is defined as the function obtained by applying the activation function to the weighted sum of its input information plus a constant term, thebias\.444To simplify the discussion in this paper, we ignore bias terms in the following and focus solely on the weights of the NN\.A specific neuron value is the output of this function for a given input\. A NN,fwf\_\{w\}, is defined as therepeated composition of neurons\. The number of neurons, their distribution within the network, and the activation function are specified by thearchitectureof the network\. The architecture defines a set of neural networks \{fwf\_\{w\}\} that share the same structural properties but differ in their specific weight and bias values\. This set constitutes the hypothesis class, that is, the space of possible prediction functions that may be selected during training\. A concrete NN from this set is identified by specifying its weight parameters\. The architecture is ahyperparameter\. In general hyperparameters are parameters that are needed to define a NN and the learning process, but are neither the weights nor the biases\. Hyperparameters for the learning process are, for example, thelearning rate, thebatch sizeor theweight initialization strategy\.

The predictive success of a NN is quantified by thelossfunction:ℒ:\{fw\}×\(𝒳×𝒴\)→ℝ\\mathcal\{L\}:\\\{f\_\{w\}\\\}\\times\(\\mathcal\{X\}\\times\\mathcal\{Y\}\)\\rightarrow\\mathbb\{R\}\.555Common loss functions are the 0\-1 loss,ℒ​\[fw​\(x\(i\)\),y\(i\)\]=1\\mathcal\{L\}\[f\_\{w\}\(x^\{\(i\)\}\),y^\{\(i\)\}\]=1iffw​\(x\(i\)\)≠y\(i\)f\_\{w\}\(x^\{\(i\)\}\)\\not=y^\{\(i\)\}and 0 otherwise and the squared lossℒ​\[fw​\(x\(i\)\),y\(i\)\]=\(y\(i\)−fw​\(x\(i\)\)\)2\\mathcal\{L\}\[f\_\{w\}\(x^\{\(i\)\}\),y^\{\(i\)\}\]=\(y^\{\(i\)\}\-f\_\{w\}\(x^\{\(i\)\}\)\)^\{2\}\.The goal of the learning process is to find a NN that reliably predicts the true labels for \(unseen\) input instances from some instance space\. To achieve this goal the NN is trained on a finite subset of the whole instance space, thetraining data\. The training process searches for the NN with the lowestempirical risk, that is, the network whose predictions deviate least, on average, from the true labels of the training instances\. The deviation is measured by the chosen loss function\. To identify the NN that minimizes empirical risk on the training data, researchers employ alearning algorithmthat iteratively updates the network’s initial weights in a way that successively reduces the loss on the training data\. The most common learning algorithms are based on the heuristic ofgradient descent\.666We will introduce the heuristic of gradient descent learning in section[5\.1](https://arxiv.org/html/2606.24953#S5.SS1)\.The final step involves testing the resulting prediction function on thetest set, a separate set of data points that were not included in the training set\. If the prediction function performs well on this test set, it is accepted and usually used for deployment\. Otherwise, researchers adjust the hyperparameters and repeat the training process, potentially with a modified hypothesis class\.

### 2\.1\. Learning Opacity

The opacity of the NN learning process has already been noted by some authors\[[35](https://arxiv.org/html/2606.24953#bib.bib92),[5](https://arxiv.org/html/2606.24953#bib.bib188),[45](https://arxiv.org/html/2606.24953#bib.bib119)\]\. For example,\[[35](https://arxiv.org/html/2606.24953#bib.bib92)\]talks about a lack of ‘algorithmic transparency’ in NN and thereby means that researchers do not understand the ‘heuristic optimization procedures’ for NNs \(p\. 40\)\.\[[5](https://arxiv.org/html/2606.24953#bib.bib188)\]identifies a type of opacity that ‘concerns the way in which a DNN automatically alters the instantiated function in response to data’ \(p\. 59\)\. Similarly,\[[45](https://arxiv.org/html/2606.24953#bib.bib119)\]defines ‘training\-opacity’ as the inability of expert humans to say, upon inspection, how the parameters of the DNN were induced as a result of its training data \(p\. 225\)\.

To delineate our notion of the opacity of the NN learning process from the others, we call itlearning opacity\. We take learning opacity to refer to the lack of understanding of the dynamical phenomena of the NN’sweight dynamicsthat appear during the learning process\. By weight dynamics we mean the time evolution of the NN’s weight values during the training process\. To keep our discussion as broad as possible we do not restrict the notion of learning opacity to a specific meaning of ‘dynamical phenomena’ or ‘understanding’\.777The concept of understanding has been extensively debated in philosophy, and there is no universally accepted definition\. For an overview, see\[[1](https://arxiv.org/html/2606.24953#bib.bib19),[16](https://arxiv.org/html/2606.24953#bib.bib18)\]\.Dynamical phenomena in NN weight dynamics may include fine\-grained behaviours, such as the evolution of precise weight values, or more coarse\-grained behaviours, such as the convergence speed of training, the type of minimum reached, or aforementioned phenomena like saturation or benign overfitting\. By not committing to a single definition of understanding we can illustrate the various epistemological challenges that complexity poses for the study of dynamical phenomena of the NN learning process and how these challenges limit different ways of understanding the learning process, thus contributing to learning opacity\.888\[[45](https://arxiv.org/html/2606.24953#bib.bib119)\]conducts a similar project as he discusses possible reasons for training\-opacity\. However, his project differs from ours in two important respects: first, he focuses solely on the lack of understanding of how the parameters of a NN are induced by its training data, rather than on phenomena of the learning process in general and second, he does not address how complexity contributes to this lack of understanding\.However, before doing so, we first need to clarify what we mean bycomplexity\.

## 3\. Complexity in the Natural Sciences

Complex systems in the natural sciences are generally considered to be constituted by many elements that interact in a sophisticated way which results in some kind of self\-organized structure\[\[, cf\.\]\]WiesnerLadyman, Wiesner\_Ladyman\_paper, Sandra\_Mitchell, Mitchell\_Complexity\_A\_Guided\_Tour\.999The term ‘complexity’, when left unspecified, is used to denote the concept of complexity as it is understood in the natural sciences\.Such structures display some regularity, e\.g\. symmetry, periodicity or some form of pattern\. Roughly speaking, the structure is self\-organized in case it arises autonomously out of the interactions among the parts and not from some central control\. The self\-organized structure in complex systems is a form of emergent behaviour\[\[, cf\.\]\]Anderson\_1972\_Moreisdiffernet\.

Complex systems closely resemble NN systems\. NNs are defined as the repeated composition of elements \(the neurons\) that interact with each other \(transmitting information\)\. During training, the strength of the information flow between nodes \(the weights\) changes in a sophisticated manner, eventually resulting in the optimal weight configuration \(the prediction function\), a kind of self\-organized structure\.

The resemblance between NNs and complex systems is useful to investigate learning opacity because it directs attention to the dynamic processes and how complex features of the dynamic process limit understanding\.101010It is less clear whether complexity concepts from computer science, such as computational complexity\[[18](https://arxiv.org/html/2606.24953#bib.bib51)\]or Kolmogorov complexity\[[29](https://arxiv.org/html/2606.24953#bib.bib49)\], whose systems of study are Turing machines, are suited for the study of learning opacity in NN\.Complexity research in the natural sciences examines not only the complexity of the organized results but also the complexity of the dynamic processes generating them\. Having the means to examine the complexity of the dynamic processes is crucial for our research question, since we are interested in the complexity of the NN’s learning process, i\.e\. the dynamic process that generates the prediction function\. Furthermore, there are reflections on how the complexity of the dynamical processes limits our understanding of the system\. These provide the conceptual foundation for investigating how complexity contributes to learning opacity in NN learning\.

Yet, what exactly qualifies as ‘sophisticated’ interaction and what kind of self\-organized structure is required for a system to satisfy the above characterization of a complex system remains a matter of debate\. The first point of contention concerns which qualitative properties are regarded as being necessary and/or sufficient for a system to count as complex\[\[, cf\.\]\]WiesnerLadyman, HOOKER\_Introduction\. Given a specific qualitative property, the second point of contention concerns how it should be quantified\[\[, cf\.\]\] Wiesner\_Ladyman\_paper\. The multiplicity of qualitative and quantitative approaches to complexity reflects both the diversity of systems considered complex and the differing interests of the scientific disciplines studying them\. Nevertheless, three qualitative properties appear to be central to most accounts of a dynamic processes’ complexity\.

A property that often appears in connection to complexity is nonlinearity and the associated property of sensitive dependence on initial conditions \(SI\)\[[4](https://arxiv.org/html/2606.24953#bib.bib169),[23](https://arxiv.org/html/2606.24953#bib.bib85)\]\.111111SI is a defining property of another class of systems, namely chaotic systems\. It should be noted that, although many complex systems display SI, chaos and complexity are distinct concepts\[\[\]\]Zuchowski\_Disentangeling\.The second property is feedback \(F\)\[\[, cf\.\]\]WiesnerLadyman, Sandra\_Mitchell\. Lastly, complex systems are often considered to be context\-sensitive \(CS\)\[[40](https://arxiv.org/html/2606.24953#bib.bib100)\]\.

To elucidate how these properties affect researchers’ understanding of dynamical phenomena, we now introduce one of the most common methods used to understand the dynamical behaviour of natural systems: dynamical modelling\.

### 3\.1\. Dynamical Modelling

Dynamical modelling\[\[, cf\.\]\]BISHOP\_Epistemology, Berger\_1998 is a form of mathematical modelling in which natural systems are represented and analysed through comparison with abstract mathematical entities known asdynamical systems\.

A dynamical system is a mathematical model that describes how the states of a target system evolve over time\. What counts as the relevant states of the target system depends on the researcher’s explanatory goals\. The most accurate and fine\-grained approach characterises the evolution of the target system in terms of the states of its components, the system’smicro\-states\. Thestate spaceof a dynamical system comprises all possible states the system could, in principle, occupy\.

The time\-evolution of the target system is modelled via atime evolution functionwhich is typically derived fromdynamical equationsthat specify how the system’s state changes over time\. These equations take the form of difference equations in the discrete\-time case, or differential equations in the continuous\-time case\. Together with theinitial conditions, that describe the system’s state at the start of its evolution, and theboundary conditions, that describe the constraints on the system, the dynamical equations uniquely determine the time\-evolution function\. Atrajectoryrepresents a particular path that the system follows through the state space as it evolves over time\.

We here propose to view the NN’s weight dynamics during training as a dynamical system\. The micro\-states correspond to the weight values, the time evolution is governed by the learning algorithm, and the external conditions, the context, are given by the training data \(see Table[1](https://arxiv.org/html/2606.24953#S3.T1)\)\[\[, cf\.\]section 4\.1\]Levin\_Dis\. The dynamical systems perspective on NN learning provides a fruitful framework for analysing how SI, F, and CS contribute to learning opacity, which will be addressed in the remainder of the paper\. For each property, we first discuss its characterization in the natural sciences and assess whether it is also present in NN learning\. We then examine its impact on the understanding of complex systems via dynamical modelling, and whether it contributes to learning opacity in a comparable manner\.

Complex systemsNN learningMicro\-statesStates of constituentsWeight valuesState spaceAll possible states of constituentsAll possible weight valuesInitial statesInitial states of constituentsInitial weight valuesContextEnvironmentTraining dataTime evolutionDifferential or difference equations governing micro\-state evolutionWeight update algorithmTable 1:Analogy between complex systems and neural network learning understood as dynamical systems\.

## 4\. Sensitive Dependence on Initial Conditions

Sensitive dependence on initial conditions\(SI\) means that small changes in initial states might be amplified over time and lead to significantly different system behaviour\.

In dynamical systems the phenomenon of SI is closely associated to nonlinearity\[[4](https://arxiv.org/html/2606.24953#bib.bib169)\]\. A functionffislinearif it satisfies both of the following conditions:121212The first condition is calledsuperpositionand the secondhomogeneity\.

f​\(x\+y\)=f​\(x\)\+f​\(y\)\\displaystyle f\(x\+y\)=f\(x\)\+f\(y\)f​\(c​x\)=c​f​\(x\),wherecis a constant\.\\displaystyle f\(cx\)=cf\(x\),\\text\{where $c$ is a constant\.\}Anonlinearfunction violates at least one of these conditions\. In the context of dynamical systems, a nonlinear system is a system whose time\-evolution is nonlinear\. Unlike linear functions, a change in the input of a nonlinear function does not necessarily produce a proportional change in the output\. So, in a nonlinear dynamical system even small changes in the input may be amplified over time, leading to drastically different outcomes\.

SI is commonly understood as the exponential growth of errors signified by a positive global Lyapunov exponent\. A positive global Lyapunov exponent measures the average exponential rate of divergence of infinitesimal perturbations of initial conditions in a dynamical system over infinite times\. However, the global Lyapunov exponent is criticised as an adequate measure of SI\[[4](https://arxiv.org/html/2606.24953#bib.bib169)\]\. A positive Lyapunov exponent only measures theaverageeffect ofinfinitesimalperturbations for aninfinitetime period\. Finite deviations are not guaranteed to exhibit average growth rates for a finite time at the rate predicted by the Lyapunov exponent\. Hence, a positive Lyapunov exponent might not adequately capture a possible rapid growth offinitedeviations infinitetime\. We therefore don’t consider a positive Lyapunov exponent as a necessary condition for SI\. Instead, we adopt a qualitative understanding, according to which SI means that small changes in initial states might eventually lead to significantly different system behaviour\.

An example of a system exhibiting SI is a ball placed on the peak of a gable roof\. Whether the ball rolls down the right hand side of the roof or the left hand side can be determined by very small displacements of its initial position\. The first thorough scientific examination of a system that is sensitive to small perturbations in its initial conditions was done by Edward N\. Lorenz discussing atmospheric convection phenomena\[[36](https://arxiv.org/html/2606.24953#bib.bib40)\]\. A famous one\-dimensional map in the interval\[0,1\]\[0,1\]that is used to study SI is the logistic mapxn\+1=r​xn​\(1−xn\)x\_\{n\+1\}=rx\_\{n\}\(1\-x\_\{n\}\)popularized by the biologist Robert May\[[38](https://arxiv.org/html/2606.24953#bib.bib140)\]\. Population biologists examine the logistic map to model the effect of the population sizexnx\_\{n\}in generationnnon the next generation’s population sizexn\+1x\_\{n\+1\}, whererris interpreted as the population’s growth rate\. Small differences in the initial conditions of the logistic map already have a huge effect after a small number of iterations\. Forr=4r=4and the initial valuex0=0\.51x\_\{0\}=0\.51the map outputsx10=0\.53049x\_\{10\}=0\.53049after 10 iterations\. However, starting with initial valuex0=0\.52x\_\{0\}=0\.52yieldsx10=0\.99577x\_\{10\}=0\.99577\.

### 4\.1\. Weight Initialization Sensitivity in NN Learning

A NN’s learning process displays SI in the sense that its behaviour strongly depends on the strategy used to initialize the weights\.

In the natural sciences, an initial condition is the state of a system at the beginning of the time period under investigation\. Models used to study complex systems frequently exhibit SI, since complex natural systems are often best represented by models with nonlinear dynamics\. NNs also exhibit nonlinearities, namely their nonlinear activation functions\. In NNs, however, it is not the precise weight values at initialization that matter most, but rather theweight initialization strategy\. The general approach to weight initialization is to select a probability distribution and draw the initial weight values as samples from it\. Specific strategies differ in their choice of distribution and in how the parameters of that distribution are set\. Several results demonstrate that nonlinear activation functions interact with the initialization strategy in ways that affect convergence speed and generalization performance\[\[, cf\.\]\]He\_initialization, understanding\_Glorot, Initial\_Conditions\_Generalization\_Performance\.

\[[14](https://arxiv.org/html/2606.24953#bib.bib107)\]show that for certain activation functions the initialization strategy strongly influences the number of neurons that are saturated at initialization and during training\. A neuron is said to besaturated, when its input lies in a region of the activation function where the activation function’s output changes very little with respect to changes of the neuron’s value\. Saturated neurons impact learning behaviour because the magnitude of a weight’s update step is proportional to the sensitivity of the activation function to changes in the neurons’ value that contain that weight\. Thus, weights involved in saturated neurons may update only very slowly\. If many neurons are saturated, the training process may converge extremely slowly or fail to converge at all\.\[[19](https://arxiv.org/html/2606.24953#bib.bib134)\]provide direct empirical evidence that weight initialization strategies can decisively influence convergence speed\. They even show that in one classification task a NN initialized with one weight initialization strategy fails to converge, while the same network initialized with another weight initialization strategy converges\. As for generalization performance,\[[14](https://arxiv.org/html/2606.24953#bib.bib107)\]demonstrate that there is at least one activation function, the hyperbolic tangent,131313The hyperbolic tangent activation function, tanh\(xx\), has the formex−e−xex\+e−x\\frac\{e^\{x\}\-e^\{\-x\}\}\{e^\{x\}\+e^\{\-x\}\}\.for which the initialization strategy influences the test error\. For this activation function the difference in test error between different initialization strategy can rise up to 12%\.

Arguably, whether these findings show that the training process of a NN is sensitive to initial conditions, meaning that even a small change in the initialization strategy can lead to a significant difference in weight dynamics, depends on what counts as a small change in the initialization strategy and what counts as a significant difference in weight dynamics\.

The weight initialization strategies compared in\[[14](https://arxiv.org/html/2606.24953#bib.bib107)\]are both uniform distributions that only differ in their parameters\. Likewise, the initialization strategies examined in\[[19](https://arxiv.org/html/2606.24953#bib.bib134)\]are Gaussian distributions that also differ in just one parameter\. We think that a difference only in the parameters of the probability distribution used to sample the initial weights can be reasonably considered to be a small change in the initialization strategy\. Regarding the effects on weight dynamics, a difference between converging behaviour and non converging behaviour constitutes a significant effect\. A difference of about 12% in test error indicates that the respective weight dynamics ended up in quite different minima\.141414For comparison: the ImageNet Large Scale Visual Recognition Challenge\[[43](https://arxiv.org/html/2606.24953#bib.bib4)\], the standard benchmark for image recognition, saw a total improvement of 13\.7%, from AlexNet’s 16% error rate in 2012 to SENet’s 2\.3% in 2017, when the contest ended\. A 12% drop thus represents nearly the entire progress made by the field over that 5 year period\.

### 4\.2\. Sensitive Dependence on Initial Conditions and the Limits of Representability

In natural complex systems, SI makes it extremely difficult to construct an adequate representation of a system’s dynamics, often forcing researchers to rely on empirical investigation and statistical inference\. This, in turn, challenges our understanding of the system by introducing the limitations inherent in statistical analysis\.

In the dynamical modelling approach, the model used to represent the dynamics of the target system is the dynamical system consisting of the time\-evolution function and the state space\. If the time\-evolution function of the dynamical system exhibits SI, even small variations in the initial states can eventually lead to drastically different outcomes\. However, it is generally impossible to measure the states of a natural system with arbitrary precision\[\[, cf\.\]\]Crutchfield1994\. Consequently, the rapid amplification of errors in the real initial conditions of the target system imposes severe limitations on our ability to use the dynamical system to predict the target system’s evolution in time\.

The challenges of using the dynamical system to predict the target system’s evolution makes it difficult to confirm the dynamical model, discrete or continuous, such that the researchers are confident that the model adequately captures the target system’s dynamics\. This is because SI poses serious challenges to the common strategy of piecemeal confirmation\.151515Other methods for model confirmation are available, but encounter problems similar to those faced by the piecemeal improvement strategy when applied to models exhibiting SI\[\[, cf\.\]\]BISHOP\_Epistemology, Koperski\_model\.There are two strategies for confirming a model in a piecemeal fashion\. One approach is to increase the accuracy of the initial and boundary conditions while keeping the model itself fixed\[[32](https://arxiv.org/html/2606.24953#bib.bib145)\]\. Alternatively, researchers may refine the model incrementally, e\.g\. by de\-idealizing it, while keeping the initial and boundary conditions constant\[[54](https://arxiv.org/html/2606.24953#bib.bib150)\]\.

Assuming that the model correctly represents the target system, a successive increase in the accuracy of the initial conditions should lead to a monotonic convergence of the model’s predictions toward the actual behaviour of the target system\. Similarly, in the second strategy, if the initial model captures the target system’s mechanisms to some extent, successive refinements should produce increasingly accurate predictions\. While monotonic convergence is taken as evidence for the initial model’s adequacy in representing the target system’s mechanisms, a failure to converge suggests that the initial model may not adequately capture those mechanisms\.

The monotonic convergence approach to confirmation, however, faces great difficulties if the model deploys SI\[[30](https://arxiv.org/html/2606.24953#bib.bib146),[4](https://arxiv.org/html/2606.24953#bib.bib169)\]\. In the first case, an increase of the accuracy of the initial conditions of the model means that the initial conditions are slightly changed\. This might lead to an extreme change of the predicted trajectory due to SI\. It is not guaranteed that the resulting prediction is closer to the target system behaviour, even when the new, improved initial conditions more closely approximate the actual initial conditions of the target system\. Moreover, the amplification of measurement error and researchers’ inability for accurate measurements blocks the second convergence strategy, fixing the data and improving the model\. The problem is thateven ifresearchers had the correct \(non\-linear\) time\-evolution function, they would never be able to know whether it is the correct time\-evolution function if it deploys SI, since they would never be able to have infinitely precise initial conditions\. However, if the dynamical model with the correct time\-evolution function failed to predict the target system’s true behaviour, then neither will any of its refined versions, and it may be the case that researchers never attain sufficient confidence in the adequacy of any of the models\.

When researchers are unable to study a target system’s dynamics using an adequate dynamical model, they are compelled to investigate it experimentally, that is, through direct observation\. Reliance on direct observation marks a shift toward a mode of reasoning grounded in statistical inference\. Researchers examine the natural system’s dynamics by conducting repeated experiments, reporting the results and drawing general conclusion from the collected data in an inductive fashion\. Arguably, the frequentist approach remains the most widely used method for inductive inference in scientific practice\. Yet, it is very challenging to actually implement in practice the theoretical assumptions required for a valid inference from frequentist methods\. Moreover, even if the researchers correctly apply frequentist methods they sometimes misinterpret the results, which leads to mistaken conclusions from the empirical data\. All this contributes to the fragile nature of hypotheses about the system’s micro\-states dynamics when they are based on empirical data and inductive inferences\[\[, cf\.\]\]Sprenger\_Frequentism\_vs\_Bayesianism, Schneider\_Problems\_of\_Statistics\.

### 4\.3\. Weight Initialization Sensitivity and Opacity

Sensitivity to the weight initialization strategy does not contribute to learning opacity by complicating the construction of an adequate representation of the target system and thus does not force a shift to empirical studies\. Rather, it limits the researcher’s pragmatic understanding of the learning process\.

Unlike in complexity science, researchers in ML have full access to the target system\. The learning algorithm itself provides a fully accurate representation of the weight dynamics during training and as long as researchers studying the NN learning process have access to the learning algorithm, the hyperparameter settings, the training data and the source code they have access to a complete representation of the learning dynamics\.161616This scenario can be considered to be what Bordt and co\-authors call a ‘fully transparent scenario’\[[6](https://arxiv.org/html/2606.24953#bib.bib14)\]\.

Nevertheless, sensitivity to weight initialization affects the pragmatic understanding researchers have of the learning process\. A scientist has pragmatic understanding of a target system insofar as she is able to recognise qualitatively characteristic consequences of the assumptions of the target system’s representation without exact calculations and can successfully intervene on and control the target phenomenon\[[33](https://arxiv.org/html/2606.24953#bib.bib32),[34](https://arxiv.org/html/2606.24953#bib.bib31)\]\.171717The idea that understanding has to do with anticipating qualitative characteristics is inspired by the notion of ‘intelligibility’\[\[, cf\.\]\]understanding\_scientific\_understanding, deRegtDieks2005\.Arguably, because the learning algorithm itself can be considered as a fully accurate representation of the target system in the context of NN learning, one possesses pragmatic understanding of a NN’s learning behaviours if one can design the NN learning process in such a way that the NN exhibits desired qualitatively characteristic learning behaviours\. However, it is typically very difficult to anticipate even qualitatively the consequences of a given weight initialization strategy for the resulting learning dynamics prior to training\. Consequently, there exist very few principled guidelines for selecting weight initialization strategies that yield a desired learning behaviour and determining the ‘correct’ weight initialization is often a matter of guessing\[\[\]p\. 1\]weight\_initialization\_guess\.181818This observation holds for hyperparameter settings in general, as NN modellers often need to first guess the best setting and then fine\-tune it through trial and error\.

## 5\. Feedback

Feedback\(F\) means that there are mutual dependencies between different parts of the system\. Yet, there are different ways of conceiving the parts involved in such dependencies: feedback may occur from past interactions to current element interactions\[[31](https://arxiv.org/html/2606.24953#bib.bib202)\], between elements and their interactions\[[50](https://arxiv.org/html/2606.24953#bib.bib203),[21](https://arxiv.org/html/2606.24953#bib.bib174)\], or between micro\-level properties and system\-wide properties\[[11](https://arxiv.org/html/2606.24953#bib.bib245),[40](https://arxiv.org/html/2606.24953#bib.bib100)\]\.

These mutual dependencies generate feedback loops, where the parts involved continuously influence one another\. Importantly, feedback tends to blur the causal order\. There is no simple, one\-way causal chain from one part of the loop to another\. This also means that the parts that constrain the behaviour of others are themselves influenced by the very behaviours they constrain\. In particular, feedback challenges the traditional bottom\-up conception of causation, in which causal influence flows only from lower levels to higher levels\. In complex systems, micro\-level behaviour generates emergent system\-wide properties, which in turn feed back to constrain the micro\-level dynamics\[[40](https://arxiv.org/html/2606.24953#bib.bib100), pp\. 34\]\.

A classic example of a system with feedback is the predator\-prey interaction described by the Lotka–Volterra model\[[51](https://arxiv.org/html/2606.24953#bib.bib43),[37](https://arxiv.org/html/2606.24953#bib.bib44)\]\. This model formalizes the mutual dependency between predator and prey populations and shows that these two variables fluctuate in a periodic manner\. The Lotka\-Volterra model illustrates well the interaction between micro\-level dynamics and constraining system\-wide properties characteristic of feedback systems\. The micro\-level dynamics of successful hunting attempts of an individual predator or the successful reproduction of an individual prey influence both the sizes of the predator and prey populations, two system\-wide properties\. Yet the predator and prey population sizes, in turn, constrain the micro\-level interactions of successful hunting and successful reproduction\.

### 5\.1\. Feedback in NN Learning

In supervised NN training there is a feedback loop between the micro\-level dynamics of individual weight updates and the NN loss gradient, which is a system\-wide property\.

In supervised learning, we train a NN using a training algorithm that finds, based on labelled data, a NN configuration that can serve as a good prediction function on future data points\. Currently, the most widely used training algorithms for supervised learning are based on gradient optimization\.191919Here, we illustrate gradient\-based optimization in the case where only a single data point is used to update the network at each step\. For a general introduction to gradient\-based optimization, see\[[15](https://arxiv.org/html/2606.24953#bib.bib193), ch\. 8\]\.

A derivative of a functionf​\(x\)f\(x\)is the slope off​\(x\)f\(x\)at pointxx\. If a functionf​\(x,y\)f\(x,y\)depends on two variablesxxandyywe call the derivative with respect to a single variable thepartial derivative\. The concept of a derivative generalizes to a gradient in the multi\-dimensional setting\. Agradientis a vector containing all the partial derivatives of the respective function\. The negative gradient of a functionffat pointxxpoints to the steepest descent offfatxx\. The basic idea of gradient\-based optimization is to use the gradient to iteratively update the weights and biases of the NN, thereby finding a configuration that yields a low loss on the training data\.

The loss functionℒ\\mathcal\{L\}maps the current NN,fwf\_\{w\}, and a data point,\(x\(i\),y\(i\)\)\(x^\{\(i\)\},y^\{\(i\)\}\), to the corresponding loss valueℒ​\(fw,x\(i\),y\(i\)\)\\mathcal\{L\}\\big\(f\_\{w\},x^\{\(i\)\},y^\{\(i\)\}\\big\)and thus depends on the weight variableswiw\_\{i\}inww\. In the learning process, the researcher’s goal is to find a weight configuration such that the corresponding NN minimizes the loss on the training data\. For any given data point, the negative gradient of the loss function with respect to the network weights points in the direction in which the weights must be updated in order to minimize the loss most effectively\. By successively updating the weights of the network in this direction, we gradually approach a configuration with lower loss\.

Schematically, one update step in gradient\-based optimization for an individual weight valuewi\(n\)w^\{\(n\)\}\_\{i\}at stepnnis given by:

wi\(n\+1\)=wi\(n\)−ϵ​\(gw​ℒ​\(w\(n\),x\(i\),y\(i\)\)\)i\.w\_\{i\}^\{\(n\+1\)\}=w\_\{i\}^\{\(n\)\}\-\\epsilon\\Big\(g\_\{w\}\\mathcal\{L\}\\big\(w^\{\(n\)\},x^\{\(i\)\},y^\{\(i\)\}\\big\)\\Big\)\_\{i\}\.\(1\)We subtract from the weight valuewi\(n\)w\_\{i\}^\{\(n\)\}the value of theit​hi^\{th\}entry of the loss function’s gradientgw​ℒ​\(⋅,⋅,⋅\)g\_\{w\}\\mathcal\{L\}\\big\(\\cdot,\\cdot,\\cdot\\big\),202020The gradientgw​ℒ​\(⋅,⋅,⋅\)g\_\{w\}\\mathcal\{L\}\(\\cdot,\\cdot,\\cdot\)is the vector in which theit​hi^\{th\}entry,\(gw​ℒ​\(⋅,⋅,⋅\)\)i\\big\(g\_\{w\}\\mathcal\{L\}\(\\cdot,\\cdot,\\cdot\)\\big\)\_\{i\}, is the derivative ofℒ​\(⋅,⋅,⋅\)\\mathcal\{L\}\(\\cdot,\\cdot,\\cdot\)with respect towiw\_\{i\}\.evaluated on the current NN weight configurationw\(n\)w^\{\(n\)\}and the current data point\(x\(i\),y\(i\)\)\(x^\{\(i\)\},y^\{\(i\)\}\)\. This scalar value is scaled by the learning rateϵ\\epsilon\.

Equation[1](https://arxiv.org/html/2606.24953#S5.E1)shows that the size and direction of an updating step of an individual weight is constrained by theit​hi^\{th\}entry of the evaluated loss gradientgw​ℒ​\(w\(n\),x\(i\),y\(i\)\)g\_\{w\}\\mathcal\{L\}\\big\(w^\{\(n\)\},x^\{\(i\)\},y^\{\(i\)\}\)\. This scalar in turn depends on the general form of the loss gradientgw​ℒ​\(⋅,⋅,⋅\)g\_\{w\}\\mathcal\{L\}\\big\(\\cdot,\\cdot,\\cdot\\big\), which is calculated for the specific NN architecture\. Thus, the updating step of an individual weight depends on the distribution of weights across the whole NN\. Moreover, to update an individual weight the gradient is evaluated for the current weight value configuration, that means thatallcurrent weight values potentially affect an individual weight’s updating step\. The dependence of the gradient value on the architecture and the whole weight value configuration makes it a system\-wide property\.

Conversely, because the gradient is evaluated with respect to the overall weight configuration, a change in an individual weight can affect the value of the gradient\. This implies that micro\-level dynamics, individual weight updates, can influence the macro\-level dynamics of the gradient\.

In sum, the micro\-level dynamics of individual weight updates influence the overall NN weight configuration and thus the system\-wide property of the loss gradient value, which in turn constrains the further weight updates\.

### 5\.2\. Feedback and Computational Modelling

In natural complex systems, F prevents the system’s micro\-state time evolution from being representable in a tractable and theoretically useful way\. As a consequence, F limits researchers’ understanding of the target system’s behaviour, as they instead study its dynamics through an incomprehensible model\.

As mentioned above, the most accurate and fine\-grained way to represent a system’s state evolution is by modelling the dynamics of its constituents\. Classically, one would represent the system’s dynamics by deriving ananalytictime\-evolution function for these micro\-states\[\[, cf\.\]\]PARISI\. Deriving such a function requires solving the corresponding differential equation under fixed constraints on the system, expressed as constant boundary conditions\.

In complex systems, however, the constraints on the systems dynamics are not fixed\. This is because feedback occurs not only between the system’s components but also between the components and the very conditions that constrain their behaviour\. Recall the example of the Lotka\-Volterra model\. The overall prey population constrains the number of successful hunts of an individual predator\. Yet, the number of the predator’s successful hunts influences the overall prey population, which in turn influences the number of successful hunts of the individual predator\. The feedback between the system’s components and their constraints has thus the consequence that there are no constant boundary conditions and so it is often impossible to derive an analytic time\-evolution function that represents the dynamics of the complex system’s micro\-states in a theoretically useful way\[[21](https://arxiv.org/html/2606.24953#bib.bib174),[50](https://arxiv.org/html/2606.24953#bib.bib203),[24](https://arxiv.org/html/2606.24953#bib.bib166)\]\.212121Note that this does not imply that it is impossible to describe analytically the dynamics of morecoarse\-grainedstates of the system\. In the the Lotka\-Volterra model for example researchers are able to describe analytically the time\-evolution of the predator and preypopulations\.

Nevertheless, researchers studying target systems exhibiting feedback might have access to dynamical equations that are intended to represent the evolution of the system’s microstates\. Yet, as argued, these equations often lack analytic solutions\. In such cases, the standard approach for investigating the system’s behaviour is to rely on computer simulations\. Researchers may either approximate the solution of analytically intractable differential equations computationally, or directly simulate the target system using discrete difference equations from the outset\. The resulting simulated behaviour can then be treated as a model of the target system for the purpose of studying its dynamics\[\[, cf\.\]\]Weisberg\_Computer\_Simulation\_as\_model\.

Computer simulations are a source of uncertainty, since computers are finite, discrete machines and therefore subject to rounding errors\. Moreover, there are multiple ways to implement a given simulation, and the resulting behaviour may depend on the specific implementation details\[[23](https://arxiv.org/html/2606.24953#bib.bib85), p\. 37\]\. At a more philosophical level, it remains unclear how computer simulations contribute to our understanding of a target system’s behaviour\. One of the most widely discussed issues concerning the epistemic status of computer simulations is their epistemic opacity\[\[, cf\.\]\]Humphreys2009ThePN, Duran\_Computer\_Simulations, Humphreys2004\. Epistemic opacity arises from the fact that human agents, as finite reasoners, cannot have knowledge of all the epistemically relevant steps involved in producing a simulated state description of the target system\[[22](https://arxiv.org/html/2606.24953#bib.bib149)\]\. What counts as an ‘epistemically relevant’ step in the use of a computer simulation, however, is often vaguely defined\[\[, cf\.\]\]Beisbart2021\-Opacity\_CS\. On a broad conception, epistemically relevant steps may span the entire modelling process, including modelling assumptions such as parameter choices\[\[, cf\.\]\]WhyTrustSimulation\. On a narrower conception, they may be restricted to the computational steps internal to the simulation itself\[\[, cf\.\]\]Humphreys2009ThePN\. Whatever one takes to be the epistemically relevant steps, the epistemic opacity of computer simulations challenges researcher’s ability to explain how exactly simulated phenomena, such as emergent macro\-level feature, are generated by the computer simulation\. Understanding the behaviour of the real\-world target system becomes thus even more difficult, since the model that is used as the target system’s representation is not comprehensible\[\[, cf\.\]\]Humphreys2009ThePN\.

### 5\.3\. Feedback and Opacity

We now argue that the challenge posed to ML research by the absence of a tractable and theoretically useful representation does not lie in the use of incomprehensible models, but rather in the challenges of statistical inference\.

Although ML research is considered to lie at the interface of a formal and empirical science\[\[, cf\.\]\]Rethink\_empirical\_ML, NN researchers struggle to find a tractable and theoretical useful representation of the fine\-grained weight dynamics of realistic NN during learning\[\[, cf\.\]p\. 1\] Cohen\_understanding\_central\_flow\_ICML\. Given the affect of feedback in natural systems on the possibility of deriving a theoretical representation, feedback loops in gradient\-based optimization might be a source for this\. The difficulty in theoretically analysing the learning behaviour of realistic NNs significantly complicates the first type of inquiry and has led to an increased reliance on empirical approaches in ML\.

As pointed out in section[4\.3](https://arxiv.org/html/2606.24953#S4.SS3), ML researchers have full access to the target system\. An ML experiment consists in executing the learning process on a computer and directly observing how it unfolds\. Since the learning process is implemented on computers, uncertainties due to interaction effects between the hardware and the program, like rounding errors or the form of implementation, do play a role\[[10](https://arxiv.org/html/2606.24953#bib.bib187)\]\. Yet, due to the full access to the true dynamical description of the target system, empirical ML researchers are not required to represent or approximate the behaviour via a computer simulation\.222222We are not claiming that the features that render computer simulations epistemically opaque cannot also arise in the execution of NN learning processes on a computer\. However, we think that these problems are not generated by the complex properties that are the focus of the present paper\.

The reliance on direct observation in empirical ML has imported an inductive perspective such that researchers conduct \(repeated\) experiments and try to generalize their insight via statistical reasoning\[[20](https://arxiv.org/html/2606.24953#bib.bib122)\]\. Thus, the challenges that arise due to the lack of a tractable and theoretical useful representation of the fine\-grained weight dynamics of realistic NN are the epistemological challenges of empirical testing and statistical inference mentioned in section[4\.2](https://arxiv.org/html/2606.24953#S4.SS2)\.\[[20](https://arxiv.org/html/2606.24953#bib.bib122)\]identify four methodological challenges that the inductive perspective in ML research encounters\. First, the experiments in empirical ML are often biased in favour of the results the researchers want to demonstrate\. Second, there is an unilateral focus on performance experiments\. Third, the important abstract concepts in ML are often not well defined and it is not clear how they should be operationalized in experiments\. In addition to the three problems related to experimental design, empirical ML encounters the aforementioned problem that the theoretical assumptions underlying statistical tests are rarely satisfiable in practice, and the reliability of statistical test results is frequently overstated\. The authors argue that if researchers adopt the inductive perspective uncritically, these problems tend to go unnoticed, and the results end up being non\-reproducible, as has already happened in practice\[[20](https://arxiv.org/html/2606.24953#bib.bib122), sec\. 1\]\. The non\-replicability of certain results regarding the NN behaviour drawn from empirical studies suggests that the above epistemological challenges are hindering the understanding of the NN learning process, thus contributing to learning opacity\.

## 6\. Context Sensitivity

Context\-sensitivity\(CS\) refers to the dependence of the system’s behavior on conditions in its environment\[\[, cf\.\]\]Sandra\_Mitchell\. In other words, a context\-sensitive system is influenced not only by its internal parts and structure but also by external conditions\.

In physics, context\-sensitivity is understood as a system being open\. An open system is a system with a net influx of energy or matter to the system\. This influx usually originates from outside the system, which makes open systems driven by something external\. Open systems are generally not in thermodynamic equilibrium, though, some open systems with a compensating outflux can reach a dynamic equilibrium in which certain system properties remain approximately constant over time\. How the physical notions of open system and thermodynamic equilibrium are generally applicable to non physical systems, like financial markets, remains an open question\[[31](https://arxiv.org/html/2606.24953#bib.bib202)\]\.

An example of a natural system that is context sensitive is the foraging behaviour of ants\[[25](https://arxiv.org/html/2606.24953#bib.bib83)\]\. Researchers have found that, in some ant species, foraging behaviour is influenced by environmental temperature\. Specifically, the ants shift from a unimodal foraging pattern, that is, continuous foraging throughout the day, to a bimodal pattern, with two distinct foraging periods separated by a break, in order to avoid the high midday temperatures\. Furthermore, when a system is subjected to multiple external influences, the order in which these influences arrive matters\. For example, a plant’s successful growth depends on more than just eventually receiving water and sunlight\. Receiving excessive sunlight without sufficient water will dry the plant out, whereas too much water without adequate sunlight will drown it\.

### 6\.1\. Context Sensitivity in NN Learning

The learning dynamics of a NN is context sensitive, because its dynamics are sensitive to the training data distribution and order\.

For NNs, the training data are external conditions and thus its context: the data is generated in the world and processed by the researchers and not part of the network’s formal definition\. Importantly, the dynamics of learning are shaped both by the training distribution itself and by the order in which training examples are presented\.

That the training data distribution influences learning dynamics is a desired property\. Otherwise, learning different distributions would be impossible\. A clear illustration of this effect is provided by\[[57](https://arxiv.org/html/2606.24953#bib.bib58)\], who trained a NN for image classification\. They showed that altering the correlation between images and labels, i\.e\. modifying the training distribution, directly impacts convergence speed and generalization performance\.

The sensitivity of learning dynamics to the order of training data presentation is apparent incurriculum learning\[\[, cf\.\]\]ELMAN199371, Soviany2022\_Curriculum, Wang\_Curriculum\_Learning\. In curriculum learning, researchers reweight the training distribution over time, making certain examples more likely to appear at specific stages of training\. This imposes an order on the training data\. The process typically relies on two components: a difficulty measure that assigns a difficulty score to each training example, and a training scheduler that controls how examples of varying difficulty are presented over the course of training\.

\[[3](https://arxiv.org/html/2606.24953#bib.bib63)\]provide an early demonstration of curriculum learning in an image classification task with a NN trained via online stochastic gradient descent\.232323In online stochastic gradient descent, gradient optimization is performed on one randomly chosen training example at a time\[[15](https://arxiv.org/html/2606.24953#bib.bib193), ch\. 8\]\.The task was to classify shapes into three categories: rectangles, ellipses, and triangles\. The authors defined difficulty using shape variability\. The datasetGeomShapescontained rectangles, ellipses, and triangles with high variability, whileBasicShapescontained images of squares, circles, and equilateral triangles with less variability\. Thus, the task to classify images inBasicShapesis a supposedly simpler task than clssifying the images inGeomShapes\. Accordingly, the curriculum strategy was to first train onBasicShapesbefore switching toGeomShapes\. Bengio and co\-authors found that generalization performance was sensitive both to the time split between the two datasets and to the use of curriculum learning at all\.

The sensitivity of the learning dynamics to the order of data is not specific for online stochastic gradient descent nor limited to generalization performance\.\[[17](https://arxiv.org/html/2606.24953#bib.bib62)\]investigate curriculum learning in the context of mini\-batch stochastic gradient descent\. This is a technique, where stochastic gradient descent updates are performed on small, randomly chosen subsets of training data, the ‘mini\-batches’\[[15](https://arxiv.org/html/2606.24953#bib.bib193), ch\. 8\]\. To impose an order on the training data, the idea is to sample the mini\-batches non\-uniformly from the training set and regulate the type of batches that are presented to the NN in the mini\-batch stochastic gradient descent\.

The authors show that different curriculum strategies, defined by different difficulty measure and different training scheduler, affect not only generalization performance but also convergence speed\. In their experiments with convolutional NN for classification, the convergence behaviour varied substantially depending on the curriculum strategies, that is on the imposed training data order\.

### 6\.2\. Context Sensitivity and Reduced Unification

In natural complex systems, context\-Sensitivity \(CS\) poses a challenge for understanding because it reduces the degree of unification in explanations\.

Unifying explanations of a phenomenon contribute to understanding by allowing us to view the phenomenon as an instance of a greater scheme rather than as an isolated event\[\[, cf\.\]\]understanding\_scientific\_understanding\. There are two basic strategies by which unification can be achieved\. Michael\[[13](https://arxiv.org/html/2606.24953#bib.bib75)\]maintains that explanations consist of derivations of the phenomenon presented in the form of arguments\. He argues that the more comprehensive the generalization used in the argument is, i\.e\. the more situation it subsumes, the more unifying the explanation becomes\. Like Friedman,\[[26](https://arxiv.org/html/2606.24953#bib.bib74),[27](https://arxiv.org/html/2606.24953#bib.bib71)\]also considers explanations to be arguments\. Yet, he argues that the unifying power of an explanation lies not in the use of a comprehensive generalization, but in the application of an argument pattern that is as general as possible\. Two phenomena are unified in case their explanations instantiate similar argument patterns\.

In both accounts, CS limits our capacity to provide unifying explanations\. In general, the truth of a generalization is contingent on background conditions, and the degree of this contingency depends on the stability of the conditions under which the generalization holds\[[39](https://arxiv.org/html/2606.24953#bib.bib73),[55](https://arxiv.org/html/2606.24953#bib.bib6),[56](https://arxiv.org/html/2606.24953#bib.bib7)\]\.242424\[[55](https://arxiv.org/html/2606.24953#bib.bib6)\]calls a generalization ‘invariant’ if it would remain stable as various other conditions change\. The range of changes over which the generalization is invariant is the ‘domain of invariance’ \(p\. 205\)\.Whether a generalization can be applied to explain a target system’s behaviour depends on whether the necessary conditions under which the generalization holds true are satisfied in that specific case\. In scientific disciplines that investigate systems that are influenced by many external conditions to a high extent, such as biology, the stability of the relevant background conditions is presumably lower than in disciplines that study less context\-sensitive systems, such as physics\[\[, cf\.\]\] Mitchell2002\_lessonsfrombiology, woodward2010causation\. Consequently, in disciplines that study highly context\-sensitive target systems, the generalizations employed in explanations are less general, as their conditions are less often satisfied in concrete cases\.

Moreover, it is plausible to assume that more argument patterns are required when addressing context\-sensitive phenomena\. Argument patterns are abstracted from the particular arguments we use to derive, that is, to explain, the phenomenon in question\[[26](https://arxiv.org/html/2606.24953#bib.bib74), p\. 520\]\. However, different external conditions might give rise to different phenomena, which might require distinct arguments for their explanations\. The greater the number of arguments required to account for the observed phenomena, the more likely it is that these arguments instantiate different argument patterns\. Hence, the unifying power of a concrete explanation that instantiates a single general argument pattern is diminished\.

### 6\.3\. Training Data Sensitivity and Opacity

As with sensitivity to weight initialization, sensitivity to training data affects the researcher’s pragmatic understanding of the learning process\. The extent to which sensitivity to training data affects the possibility of providing unifying explanations of NN learning dynamics remains an open question\.

To assess the influence of training data sensitivity on the degree of unification of explanations of NN learning dynamics, studies are needed that investigate whether a given explanatory scheme, an argument pattern or explanations that contain a certain generalization, of a dynamical phenomenon is robust across distinct training data distributions and orders\. A robustness target is considered robust with respect to a given robustness modifier if the interventions within the robustness domain do not cause changes in the target that exceed the specified tolerance\[[12](https://arxiv.org/html/2606.24953#bib.bib192), p\. 8\]\. In studies examining the degree of unification of an explanatory scheme with respect to variations in training data distribution or order, the explanatory scheme would be treated as the robustness target, while the distributional or ordering variations of the training data would serve as robustness modifiers\.

Robustness analyses are ubiquitous in NN research\. However, the most common type focuses on the robustness of model performance, not on explanations\. To our knowledge, no studies have analysed the robustness ofexplanationsof dynamical phenomena in NN learning with respect to training data distribution or order\. Consequently, whether and to what extent the sensitivity of NN learning to training data affects our ability to provide unifying explanations of its dynamical behaviour remains an open question\.

Nevertheless, like sensitivity on the weight initialization strategy, sensitivity to the training data can lead to a lack of pragmatic understanding, since even expert practitioners are often unable to anticipate qualitatively characteristic effects of the training distribution or the order of presentation on the learning dynamics\. Consider, for example, adversarial training\[[28](https://arxiv.org/html/2606.24953#bib.bib98)\]\. The authors show that a visually imperceptible perturbation of a single training image can cause the prediction function to flip its predictions for 16 out of 30 test images compared to training on the unperturbed dataset\. This indicates that the resulting prediction functions are distinct, implying that the two learning processes, one with and one without the perturbation, followed different weight dynamics\. Yet because the perturbation is visually imperceptible, researchers cannot anticipate visually whether the two learning dynamics will differ qualitativelybeforerunning the training process\. Adversarial training thus illustrates that in some cases it is impossible to discern qualitatively significant consequences of the training data merely by visually inspecting it\.

Complex systemsNN learningFeedbackReliance on incomprehensible modelsChallenges of statistical inferenceSensitive dependence on initial conditionsChallenges of statistical inferenceLimited pragmatic understandingContext sensitivityReduced degree of unificationReduced unification?, limited pragmatic understandingTable 2:Challenges arising from complexity in natural complex systems and neural network learning\.

## 7\. Conclusion

This article shed light on how complexity contributes to learning opacity\. Table[2](https://arxiv.org/html/2606.24953#S6.T2)highlights our main findings\.

We argued that the learning process for a neural network \(NN\) is essentially a complex dynamical system\. In particular, as a complex system, it displays similar characteristics as natural complex systems, such as sensitivity to initial conditions \(weights\), feedback between micro\-states \(weights\) and system properties \(loss gradient\) and context \(data\) sensitivity\. We then showed how each of these properties contributes to the difficulty to understand the learning process\.

Feedback impedes a theoretical analysis of the learning process and leads to a shift towards empirical machine learning \(ML\) research\. This introduces the epistemological limitations inherent to statistical inference\. Sensitivity to the weight initialization strategy affects the researcher’s pragmatic understanding, as it becomes hard to qualitatively anticipate consequences of particular strategies\. Pragmatic understanding is also affected by sensitivity to the training data, but it remains an open question whether the latter also impedes unifying explanations of learning phenomena\.

Our main conclusion is that learning opacity, understood as difficulty to understanding crucial dynamical phenomena of NN learning, is due to the structure of complexity and the epistemological challenges that arise from it\. Thus learning opacity is not a matter of lack of access or cognitive limitation\. Rather, the intrinsic structure of ML is at the heart of opacity\. And we suspect that a similar analysis can be conducted for other ML learning algorithms\. Sensitivity to weight initialization, feedback in gradient based optimization and sensitivity to the training data are fundamental to the learning process\. Damping such complexity\-related properties or eliminating them would fundamentally alter how ML systems learn\. Consequently, some sources of opacity in ML may be irreducible\. And hence the research in finding methods to reduce opacity, other than by building interpretable ML algorithms from the outset, might be a pipe dream\. The conceptual bridge drawn between ML systems and complex natural systems, however, suggest that frameworks developed to study the latter may provide valuable approaches for addressing challenges of understanding the former, for example determining the appropriate level of coarse\-graining for analysing and simplifying ML systems might be a promising route\. The full potential of seeing ML learning as complex system, however, remains an open field for future research\.

## References

- \[1\]\(2016\)What is understanding? an overview of recent debates in epistemology and philosophy of science\.InExplaining Understanding: New Perspectives from Epistemology and Philosophy of Science,S\. R\. Grimm, C\. Baumberger, and S\. Ammon \(Eds\.\),pp\. 1–34\.External Links:ISBN 978\-1\-138\-92193\-1,[Document](https://dx.doi.org/10.4324/9781315686110)Cited by:[footnote 7](https://arxiv.org/html/2606.24953#footnote7)\.
- \[2\]M\. Belkin\(2021\)Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation\.External Links:2105\.14368,[Link](https://arxiv.org/abs/2105.14368)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1)\.
- \[3\]Y\. Bengio, J\. Louradour, R\. Collobert, and J\. Weston\(2009\)Curriculum learning\.InProceedings of the 26th Annual International Conference on Machine Learning,ICML ’09,New York, NY, USA,pp\. 41–48\.External Links:ISBN 9781605585161,[Link](https://doi.org/10.1145/1553374.1553380),[Document](https://dx.doi.org/10.1145/1553374.1553380)Cited by:[§6\.1](https://arxiv.org/html/2606.24953#S6.SS1.p5.1)\.
- \[4\]R\. C\. Bishop\(2011\)Metaphysical and epistemological issues in complex systems\.InPhilosophy of Complex Systems,C\. Hooker \(Ed\.\),Handbook of the Philosophy of Science, Vol\.10,pp\. 105–136\.External Links:ISSN 18789846,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/B978-0-444-52076-0.50003-1),[Link](https://www.sciencedirect.com/science/article/pii/B9780444520760500031)Cited by:[§3](https://arxiv.org/html/2606.24953#S3.p5.1),[§4\.2](https://arxiv.org/html/2606.24953#S4.SS2.p5.1),[§4](https://arxiv.org/html/2606.24953#S4.p2.1),[§4](https://arxiv.org/html/2606.24953#S4.p3.1)\.
- \[5\]F\. J\. Boge\(2022\-03\)Two dimensions of opacity and the deep learning predicament\.Minds and Machines32\(1\),pp\. 43–75\.External Links:ISSN 0924\-6495,[Link](https://doi.org/10.1007/s11023-021-09569-4),[Document](https://dx.doi.org/10.1007/s11023-021-09569-4)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p1.1),[§1](https://arxiv.org/html/2606.24953#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.24953#S2.SS1.p1.1)\.
- \[6\]S\. Bordt, M\. Finck, E\. Raidl, and U\. von Luxburg\(2022\)Post\-hoc explanations fail to achieve their purpose in adversarial contexts\.InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’22,New York, NY, USA,pp\. 891–905\.External Links:ISBN 9781450393522,[Link](https://doi.org/10.1145/3531146.3533153),[Document](https://dx.doi.org/10.1145/3531146.3533153)Cited by:[footnote 16](https://arxiv.org/html/2606.24953#footnote16)\.
- \[7\]O\. Buchholz and E\. Raidl\(2022\)A falsificationist account of artificial neural networks\.The British Journal for the Philosophy of Science\.External Links:[Document](https://dx.doi.org/10.1086/721797),[Link](https://doi.org/10.1086/721797),https://doi\.org/10\.1086/721797Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1)\.
- \[8\]J\. Burrell\(2016\)How the machine ‘thinks’: understanding opacity in machine learning algorithms\.Big Data & Society3\(1\),pp\. 2053951715622512\.External Links:[Document](https://dx.doi.org/10.1177/2053951715622512)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p1.1)\.
- \[9\]J\. Cohen, S\. Kaur, Y\. Li, J\. Z\. Kolter, and A\. Talwalkar\(2021\)Gradient descent on neural networks typically occurs at the edge of stability\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jh-rTtvkGeM)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1)\.
- \[10\]K\. A\. Creel\(2020\)Transparency in complex computational systems\.Philosophy of Science87\(4\),pp\. 568–589\.External Links:[Document](https://dx.doi.org/10.1086/709729)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p1.1),[§1](https://arxiv.org/html/2606.24953#S1.p2.1),[§5\.3](https://arxiv.org/html/2606.24953#S5.SS3.p3.1)\.
- \[11\]E\. Estrada\(2023\-05\)What is a complex system, after all?\.Foundations of Science,pp\. 1–28\.External Links:[Document](https://dx.doi.org/10.1007/s10699-023-09917-w)Cited by:[§5](https://arxiv.org/html/2606.24953#S5.p1.1)\.
- \[12\]T\. Freiesleben and T\. Grote\(2023\)Beyond generalization: a theory of robustness in machine learning\.Synthese202\(109\)\.Cited by:[§6\.3](https://arxiv.org/html/2606.24953#S6.SS3.p2.1)\.
- \[13\]M\. Friedman\(1974\)Explanation and scientific understanding\.The Journal of Philosophy71\(1\),pp\. 5–19\.Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p2.1)\.
- \[14\]X\. Glorot and Y\. Bengio\(2010\-01\)Understanding the difficulty of training deep feedforward neural networks\.Journal of Machine Learning Research \- Proceedings Track9,pp\. 249–256\.Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1),[§4\.1](https://arxiv.org/html/2606.24953#S4.SS1.p3.1),[§4\.1](https://arxiv.org/html/2606.24953#S4.SS1.p5.1)\.
- \[15\]I\. Goodfellow, Y\. Bengio, and A\. Courville\(2016\)Deep learning\.MIT Press\.Cited by:[§6\.1](https://arxiv.org/html/2606.24953#S6.SS1.p6.1),[footnote 19](https://arxiv.org/html/2606.24953#footnote19),[footnote 2](https://arxiv.org/html/2606.24953#footnote2),[footnote 23](https://arxiv.org/html/2606.24953#footnote23)\.
- \[16\]S\. Grimm\(2024\)Understanding\.InThe Stanford Encyclopedia of Philosophy,E\. N\. Zalta and U\. Nodelman \(Eds\.\),Note:[https://plato\.stanford\.edu/archives/win2024/entries/understanding/](https://plato.stanford.edu/archives/win2024/entries/understanding/)Cited by:[footnote 7](https://arxiv.org/html/2606.24953#footnote7)\.
- \[17\]G\. Hacohen and D\. Weinshall\(2019\-09–15 Jun\)On the power of curriculum learning in training deep networks\.InProceedings of the 36th International Conference on Machine Learning,K\. Chaudhuri and R\. Salakhutdinov \(Eds\.\),Proceedings of Machine Learning Research, Vol\.97,pp\. 2535–2544\.External Links:[Link](https://proceedings.mlr.press/v97/hacohen19a.html)Cited by:[§6\.1](https://arxiv.org/html/2606.24953#S6.SS1.p6.1)\.
- \[18\]J\. Hartmanis and R\. E\. Stearns\(1965\)On the computational complexity of algorithms\.Transactions of the American Mathematical Society117,pp\. 285–306\.External Links:[Document](https://dx.doi.org/10.2307/1994208)Cited by:[footnote 10](https://arxiv.org/html/2606.24953#footnote10)\.
- \[19\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2015\)Delving deep into rectifiers: surpassing human\-level performance on imagenet classification\.In2015 IEEE International Conference on Computer Vision \(ICCV\),Vol\.,pp\. 1026–1034\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2015.123)Cited by:[§4\.1](https://arxiv.org/html/2606.24953#S4.SS1.p3.1),[§4\.1](https://arxiv.org/html/2606.24953#S4.SS1.p5.1)\.
- \[20\]M\. Herrmann, F\. J\. D\. Lange, K\. Eggensperger, G\. Casalicchio, M\. Wever, M\. Feurer, D\. Rügamer, E\. Hüllermeier, A\. Boulesteix, and B\. Bischl\(2024\)Position: why we must rethink empirical research in machine learning\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§5\.3](https://arxiv.org/html/2606.24953#S5.SS3.p4.1)\.
- \[21\]Y\. Holovatch, R\. Kenna, and S\. Thurner\(2017\-02\)Complex systems: physics beyond physics\.European Journal of Physics38\(2\),pp\. 023002\.External Links:ISSN 1361\-6404,[Link](http://dx.doi.org/10.1088/1361-6404/aa5a87),[Document](https://dx.doi.org/10.1088/1361-6404/aa5a87)Cited by:[§5\.2](https://arxiv.org/html/2606.24953#S5.SS2.p3.1),[§5](https://arxiv.org/html/2606.24953#S5.p1.1)\.
- \[22\]P\. Humphreys\(2009\)The philosophical novelty of computer simulation methods\.Synthese169,pp\. 615–626\.External Links:[Link](https://api.semanticscholar.org/CorpusID:18181167)Cited by:[§5\.2](https://arxiv.org/html/2606.24953#S5.SS2.p5.1)\.
- \[23\]\(2011\)Introduction to philosophy of complex systems: a: part a: towards a framework for complex systems\.InPhilosophy of Complex Systems,C\. Hooker \(Ed\.\),Handbook of the Philosophy of Science, Vol\.10,pp\. 3–90\.External Links:ISSN 18789846Cited by:[§3](https://arxiv.org/html/2606.24953#S3.p5.1),[§5\.2](https://arxiv.org/html/2606.24953#S5.SS2.p5.1)\.
- \[24\]\(2011\)Introduction to philosophy of complex systems\.InPhilosophy of Complex Systems,C\. Hooker \(Ed\.\),Handbook of the Philosophy of Science, Vol\.10,pp\. 841–909\.External Links:ISSN 18789846Cited by:[§5\.2](https://arxiv.org/html/2606.24953#S5.SS2.p3.1)\.
- \[25\]P\. Jayatilaka, A\. Narendra, S\. F\. Reid, P\. Cooper, and J\. Zeil\(2011\-08\)Different effects of temperature on foraging activity schedules in sympatric myrmecia ants\.Journal of Experimental Biology214\(16\),pp\. 2730–2738\.External Links:ISSN 0022\-0949,[Document](https://dx.doi.org/10.1242/jeb.053710)Cited by:[§6](https://arxiv.org/html/2606.24953#S6.p3.1)\.
- \[26\]P\. Kitcher\(1981\)Explanatory unification\.Philosophy of Science48\(4\),pp\. 507–531\.Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p2.1),[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p4.1)\.
- \[27\]P\. Kitcher\(1989\)Explanatory unification and the causal strcuture of the world\.InScientific Explanation,P\. Kitcher and W\. C\. Salmon \(Eds\.\),Ideas in Context,pp\. 410–505\.Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p2.1)\.
- \[28\]P\. W\. Koh and P\. Liang\(2017\)Understanding black\-box predictions via influence functions\.InProceedings of the 34th International Conference on Machine Learning \- Volume 70,ICML’17,pp\. 1885–1894\.Cited by:[§6\.3](https://arxiv.org/html/2606.24953#S6.SS3.p4.1)\.
- \[29\]A\. N\. Kolmogorov\(1965\)Three approaches to the quantitative definition of information\.Problems of Information Transmission1\(1\),pp\. 1–7\.Cited by:[footnote 10](https://arxiv.org/html/2606.24953#footnote10)\.
- \[30\]J\. Koperski\(1998\)Models, confirmation, and chaos\.Philosophy of Science65\(4\),pp\. 624–648\.Cited by:[§4\.2](https://arxiv.org/html/2606.24953#S4.SS2.p5.1)\.
- \[31\]J\. Ladyman and K\. Wiesner\(2020\)What is a complex system?\.Yale University Press\.Cited by:[§5](https://arxiv.org/html/2606.24953#S5.p1.1),[§6](https://arxiv.org/html/2606.24953#S6.p2.1)\.
- \[32\]R\. Laymon\(1989\)Cartwright and the lying laws of physics\.The Journal of Philosophy86\(7\),pp\. 353–372\.External Links:ISSN 0022362XCited by:[§4\.2](https://arxiv.org/html/2606.24953#S4.SS2.p3.1)\.
- \[33\]J\. Lenhard\(2009\)The great deluge: simulation modeling and scientific understanding\.InScientific Understanding: Philosophical Perspectives,H\. W\. de Regt, S\. Leonelli, and K\. Eigner \(Eds\.\),pp\. 169–186\.Cited by:[§4\.3](https://arxiv.org/html/2606.24953#S4.SS3.p3.1)\.
- \[34\]J\. Lenhard\(2019\-03\)Calculated surprises: a philosophy of computer simulation\.Oxford University Press\.External Links:ISBN 9780190873288,[Document](https://dx.doi.org/10.1093/oso/9780190873288.001.0001),[Link](https://doi.org/10.1093/oso/9780190873288.001.0001)Cited by:[§4\.3](https://arxiv.org/html/2606.24953#S4.SS3.p3.1)\.
- \[35\]Z\. C\. Lipton\(2018\-09\)The mythos of model interpretability\.Commun\. ACM61\(10\),pp\. 36–43\.External Links:ISSN 0001\-0782,[Link](https://doi.org/10.1145/3233231),[Document](https://dx.doi.org/10.1145/3233231)Cited by:[§2\.1](https://arxiv.org/html/2606.24953#S2.SS1.p1.1)\.
- \[36\]E\. N\. Lorenz\(1963\)Deterministic nonperiodic flow\.Journal of Atmospheric Sciences20\(2\),pp\. 130 – 141\.External Links:[Document](https://dx.doi.org/10.1175/1520-0469%281963%29020%3C0130%3ADNF%3E2.0.CO%3B2),[Link](https://journals.ametsoc.org/view/journals/atsc/20/2/1520-0469_1963_020_0130_dnf_2_0_co_2.xml)Cited by:[§4](https://arxiv.org/html/2606.24953#S4.p4.11)\.
- \[37\]A\. J\. Lotka\(1925\)Elements of physical biology\.Williams & Wilkins Company,Baltimore\.Cited by:[§5](https://arxiv.org/html/2606.24953#S5.p3.1)\.
- \[38\]R\. May\(1976\)Simple mathematical models with very complicated dynamics\.Nature261,pp\. 459 – 467\.Cited by:[§4](https://arxiv.org/html/2606.24953#S4.p4.11)\.
- \[39\]S\. D\. Mitchell\(2000\)Dimensions of scientific law\.Philosophy of Science67\(2\),pp\. 242–265\.External Links:[Document](https://dx.doi.org/10.1086/392774)Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p3.1)\.
- \[40\]S\. D\. Mitchell\(2009\)Unsimple truths: science, complexity, and policy\.The University of Chicago Press\.Cited by:[§3](https://arxiv.org/html/2606.24953#S3.p5.1),[§5](https://arxiv.org/html/2606.24953#S5.p1.1),[§5](https://arxiv.org/html/2606.24953#S5.p2.1)\.
- \[41\]B\. Neyshabur, R\. Tomioka, and N\. Srebro\(2015\)In search of the real inductive bias: on the role of implicit regularization in deep learning\.External Links:1412\.6614,[Link](https://arxiv.org/abs/1412.6614)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1)\.
- \[42\]A\. Power, Y\. Burda, H\. Edwards, I\. Babuschkin, and V\. Misra\(2022\)Grokking: generalization beyond overfitting on small algorithmic datasets\.External Links:2201\.02177,[Link](https://arxiv.org/abs/2201.02177)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p3.1)\.
- \[43\]O\. Russakovsky, J\. Deng, H\. Su, J\. Krause, S\. Satheesh, S\. Ma, Z\. Huang, A\. Karpathy, A\. Khosla, M\. Bernstein, A\. C\. Berg, and L\. Fei\-Fei\(2015\)ImageNet large scale visual recognition challenge\.International Journal of Computer Vision115\(3\),pp\. 211–252\.External Links:[Document](https://dx.doi.org/10.1007/s11263-015-0816-y)Cited by:[footnote 14](https://arxiv.org/html/2606.24953#footnote14)\.
- \[44\]S\. Shalev\-Shwartz and S\. Ben\-David\(2014\)Understanding machine learning: from theory to algorithms\.Cambridge University Press\.Cited by:[footnote 2](https://arxiv.org/html/2606.24953#footnote2)\.
- \[45\]A\. Søgaard\(2023\)On the opacity of deep neural networks\.Canadian Journal of Philosophy53\(3\),pp\. 224–239\.External Links:[Document](https://dx.doi.org/10.1017/can.2024.1)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.24953#S2.SS1.p1.1),[footnote 8](https://arxiv.org/html/2606.24953#footnote8)\.
- \[46\]T\. F\. Sterkenburg and P\. D\. Grünwald\(2021\)The no\-free\-lunch theorems of supervised learning\.Synthese199\(3/4\),pp\. pp\. 9979–10015\.External Links:ISSN 00397857, 15730964,[Link](https://www.jstor.org/stable/48692436)Cited by:[footnote 2](https://arxiv.org/html/2606.24953#footnote2)\.
- \[47\]T\. F\. Sterkenburg\(2025\)Statistical learning theory and occam’s razor: the core argument\.Minds and Machines35\(3\)\.External Links:[Document](https://dx.doi.org/10.1007/s11023-024-09703-y),[Link](https://doi.org/10.1007/s11023-024-09703-y)Cited by:[footnote 2](https://arxiv.org/html/2606.24953#footnote2)\.
- \[48\]E\. Sullivan\(2022\)Inductive risk, understanding, and opaque machine learning models\.Philosophy of Science89\(5\),pp\. 1065–1074\.External Links:[Document](https://dx.doi.org/10.1017/psa.2022.62)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p1.1)\.
- \[49\]E\. Sullivan\(2022\)Understanding from machine learning models\.The British Journal for the Philosophy of Science73\(1\),pp\. 109–133\.External Links:[Document](https://dx.doi.org/10.1093/bjps/axz035)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p2.1)\.
- \[50\]S\. Thurner, P\. Klimek, and R\. Hanel\(2018\-09\)Introduction to the Theory of Complex Systems\.Oxford University Press\.External Links:ISBN 9780198821939Cited by:[§5\.2](https://arxiv.org/html/2606.24953#S5.SS2.p3.1),[§5](https://arxiv.org/html/2606.24953#S5.p1.1)\.
- \[51\]V\. Volterra\(1926\)Fluctuations in the abundance of a species considered mathematically\.Nature118,pp\. 558–560\.External Links:[Document](https://dx.doi.org/10.1038/118558a0)Cited by:[§5](https://arxiv.org/html/2606.24953#S5.p3.1)\.
- \[52\]K\. Vredenburgh\(2024\-11\)Transparency and explainability for public policy\.LSE Public Policy Review\.External Links:[Document](https://dx.doi.org/10.31389/lseppr.111)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p2.1)\.
- \[53\]S\. J\. Wetzel, S\. Ha, R\. Iten, M\. Klopotek, and Z\. Liu\(2025\)Interpretable machine learning in physics: a review\.External Links:2503\.23616,[Link](https://arxiv.org/abs/2503.23616)Cited by:[§1](https://arxiv.org/html/2606.24953#S1.p2.1)\.
- \[54\]W\. C\. Wimsatt\(2007\)False models as means to truer theories\.InRe\-Engineering Philosophy for Limited Beings: Piecewise Approximations to Reality,pp\. 94–132\.External Links:ISBN 9780674015456Cited by:[§4\.2](https://arxiv.org/html/2606.24953#S4.SS2.p3.1)\.
- \[55\]J\. Woodward\(2000\)Explanation and invariance in the special sciences\.British Journal for the Philosophy of Science51\(2\),pp\. 197–254\.Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p3.1),[footnote 24](https://arxiv.org/html/2606.24953#footnote24)\.
- \[56\]J\. Woodward\(2010\)Causation in biology: stability, specificity, and the choice of levels of explanation\.Biology & Philosophy25\(3\),pp\. 287–318\.External Links:[Document](https://dx.doi.org/10.1007/s10539-010-9200-z),[Link](https://doi.org/10.1007/s10539-010-9200-z)Cited by:[§6\.2](https://arxiv.org/html/2606.24953#S6.SS2.p3.1)\.
- \[57\]C\. Zhang, S\. Bengio, M\. Hardt, B\. Recht, and O\. Vinyals\(2017\)Understanding deep learning requires rethinking generalization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Sy8gdB9xx)Cited by:[§6\.1](https://arxiv.org/html/2606.24953#S6.SS1.p3.1)\.

Similar Articles

How are linear representations learned? Exact solutions to the dynamics of abstraction

arXiv cs.LG

This paper develops a framework to study how linear concept representations emerge during neural network training, providing exact solutions in linear networks and analyzing abstraction dynamics in nonlinear networks. The results reveal key principles governing abstraction and offer implications for interpretability and control.