The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era
Summary
This paper reviews 159 studies to assess if specialized machine learning architectures can be replaced by foundation models, finding that structural representation remains indispensable and language models often require supplementary non-linguistic components.
View Cached Full Text
Cached at: 09/01/26, 12:14 PM
# The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era
Source: [https://arxiv.org/html/2608.28980](https://arxiv.org/html/2608.28980)
###### Abstract
Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language\-based models? This question is examined through a review of 159 papers \(2016–2026\) across nine modalities, with predictive accuracy considered alongside structural representation and computation\. A distinction is made between*performing a task*and*preserving and computing the structure that makes the task tractable*, and existing approaches are organized into eight representational regimes, ranging from language\-only systems to fully specialized architectures\. Language\-mediated models are found to be highly competitive in specific settings, including extreme few\-shot prediction, discretized symbolic tasks, textually annotated knowledge graphs, and large\-scale single\-modality pretraining\. However, whenever structural representation or computation is directly evaluated rather than accuracy alone, no evidence of general architectural replacement is found\. Instead, a recurring pattern is observed across independent research communities: when language alone is insufficient, the missing structure is reintroduced through a graph module, structural tokens, specialized attention, or another non\-linguistic component\. In this sense, specialization more often*relocates*than disappears\. Moreover, although performance of language\-based models is improved by scaling, whether the gap to a structure\-aware architecture can eventually be eliminated remains untested\.
## 1Introduction
Three research communities with no shared authorship, benchmarks, or citation practice have recently made similar engineering decisions\. Protein\-language modeling reinjected Foldseek111Foldseek is a fast structural\-alignment tool that converts a protein’s 3D backbone conformation into a short discrete alphabet, letting structural similarity be encoded and searched as if it were a sequence\.\-derived structure tokens into sequence\-only models vocabularies once sequence\-only representations proved insufficient\[[77](https://arxiv.org/html/2608.28980#bib.bib87),[91](https://arxiv.org/html/2608.28980#bib.bib88)\]; tabular learning paired language\-model string embeddings with an explicit graph\-attention scaffold rather than relying on the language model alone\[[45](https://arxiv.org/html/2608.28980#bib.bib41)\]; and chemistry reintroduced graph\-based tools to SMILES\-based models after they struggled to reliably track ring systems and branching\[[11](https://arxiv.org/html/2608.28980#bib.bib81)\]\. In each case, a language\-based representation was tried first, found insufficient under careful evaluation, and then supplemented with a non\-linguistic structural channel\. Three unrelated communities converging on the same corrective strategy is evidence of something more general about what language\-based representations can and cannot carry\.
The reason specialized representations existed in the first place is central to this pattern\. Convolutional networks encode translation equivariance because natural images exhibit local spatial regularities\[[15](https://arxiv.org/html/2608.28980#bib.bib4)\]; permutation\-invariant architectures were developed for sets because their elements have no canonical ordering\[[47](https://arxiv.org/html/2608.28980#bib.bib5)\]; and message\-passing networks were framed within a broader theory of relational inductive biases\[[8](https://arxiv.org/html/2608.28980#bib.bib3)\]\. These choices encode assumptions about the hypothesis space a model should explore, and the no\-free\-lunch theorem explains why such assumptions are unavoidable: without assumptions about the data\-generating process, no learning algorithm can outperform all others across all possible tasks\[[99](https://arxiv.org/html/2608.28980#bib.bib1)\]\. For decades, one of the field’s principal strategies was therefore to place useful structural assumptions directly into the architecture\.
Large language models challenge this design paradigm\. At sufficient scale, a single pretrained model can process tables, graphs, time series, and other structured inputs after they have been rendered into a language\-mediated representation, often without a task\-specific architecture\. The important question is not whether such a model can produce a plausible prediction but whether the structural assumptions that a specialized architecture would encode explicitly are actually acquired and exploited when they are instead described, demonstrated, or implied through language\. We therefore ask:To what extent can language\-based models acquire and exploit structural inductive biases that have traditionally been encoded explicitly in specialized machine learning architectures?
This question cannot be answered by predictive accuracy alone\. It decomposes into several questions that can have different answers: Is language a general representation for structured data, or primarily an interface to systems that perform the underlying computation? Can an inductive bias that a specialized architecture enforces by construction instead be recovered by stating or demonstrating it in a prompt? If a language\-based model matches a specialized model’s accuracy, what exactly has been replaced? And can existing evidence distinguish genuine structural generalization from benchmark familiarity, encoding artifacts, or external computation? These questions motivate the framework developed throughout this review and are revisited directly in[Section7](https://arxiv.org/html/2608.28980#S7)\.
Read modality by modality, the empirical record appears difficult to reconcile\. Positive results are real: TabLLM narrows the gap to gradient\-boosted trees in the extreme few\-shot regime\[[34](https://arxiv.org/html/2608.28980#bib.bib31)\]; text\-based knowledge\-graph completion can outperform structural embedding baselines\[[94](https://arxiv.org/html/2608.28980#bib.bib53)\]; and text\-symbol spatial grids reach 84–91% accuracy where the corresponding pixel\-based formulation reaches only 60–73%\[[3](https://arxiv.org/html/2608.28980#bib.bib75)\]\. Yet controlled evaluations reveal a different picture\.[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)remove the language\-model component from three widely cited time\-series models and find that the resulting models, despite having roughly1000×1000\\timesfewer parameters, match or outperform the originals\.[Fatemi et al\. \[21\]](https://arxiv.org/html/2608.28980#bib.bib43)hold the graph, task, and model fixed while varying only the textual encoding, producing a 61\.8% change in accuracy on a property that a genuinely structure\-aware representation should preserve\.[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)show that causal positional encoding makes decoder\-only transformers non\-permutation\-invariant, meaning that a model can describe a structural property while its own computation violates it\. Finally,[Bordt et al\. \[10\]](https://arxiv.org/html/2608.28980#bib.bib35)show that some apparently strong few\-shot tabular results cannot be separated from benchmark memorization\. These findings point to different failure points in the same chain: a model may*describe*a structural regularity without*preserving*it, preserve information without*computing*the corresponding invariant, or achieve benchmark performance without efficiently*learning*the intended relationship\.
This distinction exposes the tension of*generality*and*structural fidelity*in the literature\. A language\-based system can be genuinely general\-purpose which the same model can process tables, graphs, and forecasts, while failing to implement the specific invariance or compositional operation required by a task\. Conversely, it can match a specialized model’s accuracy without matching its computational efficiency or robustness\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\]\. The term “replacement” therefore conceals several non\-equivalent claims, which we formalize in[Table1](https://arxiv.org/html/2608.28980#S3.T1)\. A model may functionally reproduce a specialized system’s outputs without reproducing its structural computation; it may match accuracy without matching robustness; or it may orchestrate a specialized model effectively without replacing the computation that model performs\. This last distinction is increasingly important because recent work suggests that language models are often most effective when used to*orchestrate*specialized tools and pipelines rather than replace them\[[1](https://arxiv.org/html/2608.28980#bib.bib90),[61](https://arxiv.org/html/2608.28980#bib.bib91)\]\. Meanwhile, specialized architectures continue to scale independently\[[29](https://arxiv.org/html/2608.28980#bib.bib29),[68](https://arxiv.org/html/2608.28980#bib.bib30),[5](https://arxiv.org/html/2608.28980#bib.bib59)\]\. Orchestration is therefore a genuine capability, but it should not be conflated with architectural substitution\.
Existing survey literature makes this distinction difficult to see because it is predominantly organized by modality or engineering design\. Surveys of language models for tabular data\[[20](https://arxiv.org/html/2608.28980#bib.bib103),[100](https://arxiv.org/html/2608.28980#bib.bib104)\], graphs\[[42](https://arxiv.org/html/2608.28980#bib.bib105),[72](https://arxiv.org/html/2608.28980#bib.bib106)\], time series\[[112](https://arxiv.org/html/2608.28980#bib.bib107)\], chemistry\[[33](https://arxiv.org/html/2608.28980#bib.bib82)\], proteins\[[103](https://arxiv.org/html/2608.28980#bib.bib108)\], and point clouds\[[85](https://arxiv.org/html/2608.28980#bib.bib109)\]provide valuable modality\-specific accounts, but offer little basis for asking whether the same phenomenon recurs across domains\.[Sun et al\. \[78\]](https://arxiv.org/html/2608.28980#bib.bib101)moves toward a cross\-modal perspective by organizing methods across three modalities according to tokenization, architecture, pretraining, and adaptation\. However, its taxonomy primarily describes how systems are constructed rather than what structural computations they can be shown to perform\. Related surveys of in\-context learning\[[17](https://arxiv.org/html/2608.28980#bib.bib111),[115](https://arxiv.org/html/2608.28980#bib.bib112),[57](https://arxiv.org/html/2608.28980#bib.bib113)\]focus on prompting and training strategies, and geometric deep learning\[[12](https://arxiv.org/html/2608.28980#bib.bib114)\]provides the theoretical foundations of specialized architectures without asking whether language can substitute for the inductive biases they encode\. To our knowledge, no existing survey systematically evaluates language\-mediated methods across modalities using a common distinction between describing, preserving, computing, and learning structural properties\.
This review addresses this gap by examining language\-mediated computation across nine modalities including tabular data, graphs, time series, vision, chemistry, code, knowledge graphs, point clouds, and protein structure, through a systematic corpus of 159 papers identified using the PRISMA\-based\[[65](https://arxiv.org/html/2608.28980#bib.bib2)\]search described in[Section2\.1](https://arxiv.org/html/2608.28980#S2.SS1)\. Our focus is not simply whether language models can achieve competitive performance, but where the structural information resides and what the model can actually do with it\. We therefore separate two questions that are frequently conflated\. The first concerns the*representational regime*: whether structure is encoded in language alone, supplied through demonstrations or generated programs, or provided by an external specialized component; this is captured by the eight\-level taxonomy introduced in[Section3](https://arxiv.org/html/2608.28980#S3)\. The second concerns the*computational role*of that information: whether a model can*describe*a structural regularity, whether its representation*preserves*it, whether its computation*respects*it, and whether it can*learn*it efficiently from realistic data\. Systems that contain a language model but retain jointly trained or specialized components, such as Chronicle\[[69](https://arxiv.org/html/2608.28980#bib.bib68)\], TabPFN\[[37](https://arxiv.org/html/2608.28980#bib.bib27)\], or PatchTST\[[62](https://arxiv.org/html/2608.28980#bib.bib64)\], provide important counterfactuals to claims of language\-mediated replacement because their performance cannot be attributed to language alone\.
In this work, results from the literature are connected to compare and distinguish genuine substitution from changes in interface, computation, or experimental conditions\. It also motivates a more precise interpretation of the recurring pattern that when structural information or computation is repeatedly reintroduced through a non\-linguistic channel after a language\-mediated approach reaches its limits, the relevant question is which stage of the representation\-to\-computation pipeline has failed, and whether the missing capability can in principle be recovered within the language\-mediated regime\.
The remainder of the paper is organized as follows\.[Section2](https://arxiv.org/html/2608.28980#S2)formalizes the distinction between describing, preserving, computing\.[Section3](https://arxiv.org/html/2608.28980#S3)introduces the representational\-regime taxonomy and the eight senses of “replacement” used throughout the review\.[Section4](https://arxiv.org/html/2608.28980#S4)applies these frameworks across modalities, while[Sections5](https://arxiv.org/html/2608.28980#S5)and[6](https://arxiv.org/html/2608.28980#S6)examine whether current benchmarks can distinguish genuine structural generalization from contamination, format familiarity, and information leakage\.[Section7](https://arxiv.org/html/2608.28980#S7)synthesizes the cross\-modal evidence and evaluates whether the recurring convergence toward hybrid systems is best understood as a mechanistic limitation rather than a collection of domain\-specific failures\. Finally,[Section8](https://arxiv.org/html/2608.28980#S8)identifies the central experiment that current evidence cannot yet resolve, holding information content and sample budget fixed while varying only the channel through which structural bias is supplied, and[Section9](https://arxiv.org/html/2608.28980#S9)concludes the work\.
## 2Scope, Methodology, and Problem Formulation
### 2\.1Survey methodology
We conducted this systematic literature review following the PRISMA\[[65](https://arxiv.org/html/2608.28980#bib.bib2)\]framework to identify, screen, and synthesize research on the use of large language models for structured and non\-linguistic data\. The literature search covered works published from 2016 through 2026 and was conducted across multiple complementary search streams covering tabular data, graphs, time series, vision and multimodal data, in\-context learning and learning theory, universal representations, and reported limitations of language\-based approaches\. The search queries combined terms referring to large language models and foundation models with modality\- and problem\-specific terms, including\("large language model" OR "LLM" OR "foundation model"\) AND \("tabular" OR "table"\),\("large l anguage model" OR "LLM"\) AND \("graph" OR "graph learning"\),\("large language model" OR "LLM"\) AND \("time series" OR "forecasting"\),\("large language model" OR "LLM"\) AND \("v ision" OR "multi modal"\), and corresponding queries targeting in\-context learning, structural generalization, inductive bias, universal representation, benchmark contamination, chemistry, code, knowledge graphs, point clouds, proteins, and related domains\. Searches were complemented by backward and forward citation tracing and targeted searches for theoretical and recent works published during 2025–2026\. Retrieved studies were screened for relevance based on their relationship to the review’s central question and were retained only after bibliographic and full\-text information was independently verified through resolvable scholarly sources, including arXiv, ACL Anthology, OpenReview, and publisher platforms\. The final corpus comprises 159 papers, which form the basis of the taxonomy, cross\-modal synthesis, and subsequent analysis\.
### 2\.2Positioning relative to existing surveys
The literature contains substantial surveys of LLMs for individual modalities, including tabular data\[[20](https://arxiv.org/html/2608.28980#bib.bib103),[100](https://arxiv.org/html/2608.28980#bib.bib104)\], graphs\[[42](https://arxiv.org/html/2608.28980#bib.bib105),[72](https://arxiv.org/html/2608.28980#bib.bib106)\], time series\[[112](https://arxiv.org/html/2608.28980#bib.bib107)\], chemistry\[[33](https://arxiv.org/html/2608.28980#bib.bib82)\], protein sequences\[[103](https://arxiv.org/html/2608.28980#bib.bib108)\], and point clouds\[[85](https://arxiv.org/html/2608.28980#bib.bib109)\]; these studies primarily organize the literature by modality\-specific architectures, representations, adaptation strategies, or downstream tasks\. Most notably,[Sun et al\. \[78\]](https://arxiv.org/html/2608.28980#bib.bib101)reviews LLM\-based methods for tabular data, time series, and graphs, providing a broader cross\-modal engineering taxonomy based on data representation, model architecture, pretraining, and adaptation\. They mentioned that reducing a task to a language\-mediated representation often loses explicit structural or relational inductive bias, without testing that claim directly\.
The graph\-specific surveys,[Liu et al\. \[50\]](https://arxiv.org/html/2608.28980#bib.bib102), and,[Wang et al\. \[96\]](https://arxiv.org/html/2608.28980#bib.bib51), treat graph foundation models on their own architectural terms without a language\-substitution comparison\.[Mahowald et al\. \[55\]](https://arxiv.org/html/2608.28980#bib.bib115)’s formal\-versus\-functional\-competence distinction is the closest interdisciplinary analogue to this review’s central framework, a comparable “knowing versus doing” structure applied to linguistic competence rather than to structured non\-linguistic data\.
Related surveys of in\-context learning\[[17](https://arxiv.org/html/2608.28980#bib.bib111),[115](https://arxiv.org/html/2608.28980#bib.bib112),[57](https://arxiv.org/html/2608.28980#bib.bib113)\], geometric deep learning\[[12](https://arxiv.org/html/2608.28980#bib.bib114)\]addresses complementary aspects of the problem but do not connect architectural inductive bias, representation, computation, and learning within a common framework\. No survey located in this search spans more than two or three of this review’s nine modalities under one consistent framework, and none proposes an analogue of the four\-way describe/preserve/compute/learn distinction\. The contribution of this review is therefore the use of that common framework to synthesize evidence across modalities, connect empirical findings with learning theory, and identify the controlled experimental gap that remains unresolved \([Sections3\.3](https://arxiv.org/html/2608.28980#S3.SS3),[6](https://arxiv.org/html/2608.28980#S6)and[8](https://arxiv.org/html/2608.28980#S8)\)\.
### 2\.3What the field actually disagrees about
Some disagreements in the literature are genuinely empirical, whereas others arise because different studies answer different questions while using similar language to describe their conclusions\. The question of whether language can replace specialized machine learning spans, at minimum, four stages of the learning process\. First, can a model*describe or recognize*a structural regularity, such as permutation invariance or relational structure? Second, does its language\-based representation*preserve*the information required to exploit that regularity? Third, does the model’s computation actually*respect*the relevant structural constraint, or can it describe a property correctly while violating it in its own predictions? Fourth, can the model*learn*and generalize that structure efficiently from a realistic amount of data?[Vafa et al\. \[88\]](https://arxiv.org/html/2608.28980#bib.bib21)’s inductive\-bias probe illustrates why these distinctions matter: a transformer achieves near\-perfect accuracy on orbital\-trajectory prediction while failing to internalize the underlying physical structure, breaking when evaluated on computationally different but physically equivalent targets\. High task accuracy can therefore coexist with failure to recover the structure that would make that accuracy robust\.[Section3](https://arxiv.org/html/2608.28980#S3)formalizes these four stages as the organizing framework for[Section4](https://arxiv.org/html/2608.28980#S4); here, the central point is that success at one stage should not be treated as evidence for success at another\.
Four recurring disagreements can then be distinguished by their source:*empirical*, when competing claims are supported by measured results that have not yet been reconciled under shared conditions;*mechanistic*, when researchers agree on what a system achieves but disagree about why;*definitional*, when the same term refers to different substantive claims; and*evaluative*, when the dispute concerns whether a reported result would survive a level of scrutiny that has not yet been systematically applied\. We separate these categories explicitly because they require different forms of resolution\.
The first disagreement is empirical and concerns the scope of the substitution claim\. One line of work advances the claim in a strong form:[Song and others \[76\]](https://arxiv.org/html/2608.28980#bib.bib96)present language models as universal regressors, while[Wang et al\. \[95\]](https://arxiv.org/html/2608.28980#bib.bib33)frame them as universal tabular classifiers\. A second line of work tests this claim under controlled conditions and finds important qualifications\. Removing the language\-model component from three widely cited time\-series forecasters does not reduce performance\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\]; changing only the textual encoding of an otherwise fixed graph task produces a 61\.8% change in accuracy\[[21](https://arxiv.org/html/2608.28980#bib.bib43)\]; and causal positional encoding makes standard decoder\-only transformers provably non\-permutation\-invariant regardless of what those models can describe\[[19](https://arxiv.org/html/2608.28980#bib.bib46)\]\. This is therefore an empirical and representational disagreement: both sides can point to measured results\. As discussed in[Sections4](https://arxiv.org/html/2608.28980#S4)and[7](https://arxiv.org/html/2608.28980#S7), the evidence favors a narrower interpretation of the strongest universal claims once computational cost, information availability, and contamination are controlled\. The important point here is that statements such as “language models are universal predictors” and “language models do not preserve structure” can currently both appear as conclusions in the literature because they are often established under different experimental conditions\.
The second disagreement is mechanistic\. Researchers broadly agree that specialized architectures can encode useful structural constraints, but disagree about whether those constraints must be embedded in the architecture itself\. Our working hypothesis is that structural assumptions generally need to be reflected in the computation rather than merely supplied at inference time\. This position, however, is not uncontested\.[Mittal et al\. \[60\]](https://arxiv.org/html/2608.28980#bib.bib20)show that a standard transformer, when supplied with an appropriate*inferential*bias through training and querying rather than an architectural bias, can match custom permutation\-invariant architectures on exchangeable sequence modeling\. This constitutes the most direct counter\-evidence to the stronger form of our working hypothesis identified in the search\. It suggests that, for at least some structural properties, the distinction between architectural and inferential bias may matter more than the distinction between architectural and language\-mediated bias\. Whether this generalizes across the modalities and structural properties considered here remains an open question addressed in[Section6](https://arxiv.org/html/2608.28980#S6)\.
The third disagreement is definitional\. Work on language\-model\-driven feature engineering, model search, and pipeline construction\[[1](https://arxiv.org/html/2608.28980#bib.bib90),[61](https://arxiv.org/html/2608.28980#bib.bib91)\], as well as position papers arguing that graph foundation models are “already here”\[[56](https://arxiv.org/html/2608.28980#bib.bib50)\], may use “replace” to mean that a language model now performs a role previously carried out by a human or specialized system\. In this review, by contrast, “replace” refers to whether the language model’s own computation implements the structural function, rather than delegating that function to an external tool\. We do not attempt to privilege one usage over the other\. Instead,[Section3](https://arxiv.org/html/2608.28980#S3)makes the distinction explicit, and[Table1](https://arxiv.org/html/2608.28980#S3.T1)separates eight senses of “replacement” precisely because many apparent disagreements are disagreements about which claim is being made\.
The fourth disagreement is evaluative and, in at least one prominent case, has already been addressed through direct re\-examination\. Tabula\-8B was originally reported to outperform XGBoost and TabPFN by 5–15 percentage points in few\-shot tabular prediction\[[23](https://arxiv.org/html/2608.28980#bib.bib34)\]\. A subsequent re\-analysis found that 92\.2% of the reported gain could be recovered by instruction\-tuning the same base model without tabular exposure, alongside evidence of train–test contamination that standard deduplication had failed to detect\[[28](https://arxiv.org/html/2608.28980#bib.bib36)\]\. Similar concerns have been raised through subsequent re\-analyses\[[75](https://arxiv.org/html/2608.28980#bib.bib37),[10](https://arxiv.org/html/2608.28980#bib.bib35)\]\. Yet comparable audits have not been performed across most tabular and time\-series benchmarks in[Table3](https://arxiv.org/html/2608.28980#S5.T3), or across the graph, vision, and adjacent\-modality benchmarks considered in this review\. For these settings, the question of how much positive evidence would survive equivalent scrutiny remains open rather than being settled by the cases that have already been audited\.
These disagreements are therefore not constructed to organize the discussion\. The empirical disagreement requires controlled experiments such as the one specified in[Section8](https://arxiv.org/html/2608.28980#S8); the mechanistic disagreement requires further theoretical analysis; the definitional disagreement requires explicit terminology; and the evaluative disagreement requires audits that have not yet been conducted systematically\. Understanding why these particular questions emerged requires examining the representational history from which the current tension arose\.
### 2\.4Historical evolution
The four disagreements above are rooted in a representational history that predates LLMs\.[Figure1](https://arxiv.org/html/2608.28980#S2.F1)summarize this history as a sequence of eras defined by what researchers believed a given representational strategy could substitute for\. The eras overlap, and the transitions shown in both should therefore be interpreted as approximate\.
Figure 1:Historical evolution of research on language\-mediated learning of structured data from 2016 to 2026\. The timeline traces the transition from explicitly encoded structural inductive biases, through language\-based representations and prompting, to controlled ablations, contamination analyses, theoretical investigations, and renewed interest in specialized and jointly pretrained foundation models\.Between 2016 and 2021, the dominant strategy was to encode known structure directly into the computation graph, because doing so remained one of the most reliable ways to generalize from limited domain\-specific data\. Group\-equivariant convolutions incorporated symmetry into visual representations\[[15](https://arxiv.org/html/2608.28980#bib.bib4)\]; permutation\-invariant architectures were developed for sets\[[47](https://arxiv.org/html/2608.28980#bib.bib5)\]; and message\-passing networks were unified under a broader framework of relational inductive biases\[[8](https://arxiv.org/html/2608.28980#bib.bib3)\]\. In parallel, another line of work pursued generality through architecture rather than language\. Perceiver IO\[[40](https://arxiv.org/html/2608.28980#bib.bib95)\]and Gato\[[71](https://arxiv.org/html/2608.28980#bib.bib94)\]demonstrated that a common architecture could process substantially different input and output types, showing that the aspiration toward general\-purpose models predates, and does not require, a language\-based interface\. What these approaches could not readily provide was transfer of a learned solution across modalities: each new domain could still require substantial architectural specialization\. That cost created the conditions for the next era’s central hypothesis\.
From roughly 2022 onward, researchers increasingly investigated whether this specialization could be reduced by representing structured objects as text and prompting or lightly fine\-tuning pretrained language models\. Examples include TabLLM\[[34](https://arxiv.org/html/2608.28980#bib.bib31)\]for tabular data, LLMTime\[[31](https://arxiv.org/html/2608.28980#bib.bib60)\]for time series, and early graph\-verbalization approaches\[[92](https://arxiv.org/html/2608.28980#bib.bib45),[21](https://arxiv.org/html/2608.28980#bib.bib43)\]for relational data\. The appeal was straightforward: if a sufficiently capable pretrained model could acquire useful structural behavior from pretraining or contextual information, some of the architectural specialization of the preceding era might no longer be necessary\. The optimism of this period is reflected in position work arguing that graph foundation models were already emerging as a distinct paradigm\[[56](https://arxiv.org/html/2608.28980#bib.bib50)\]\. What many early results did not establish, however, was whether the language model itself was responsible for the observed gains, rather than the information supplied to it or the particular format in which that information was encoded\.
This question motivated the next phase of the literature\. During 2024–2025, controlled ablations increasingly tested whether the language\-model component was actually load\-bearing\. Most prominently, removing or reinitializing the language\-model component of three widely cited time\-series forecasters did not reduce performance\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\]\. In parallel, contamination audits examined whether apparent few\-shot advantages instead reflected information already present in pretraining\[[27](https://arxiv.org/html/2608.28980#bib.bib97),[10](https://arxiv.org/html/2608.28980#bib.bib35)\]\. Theoretical work supplied a complementary account of why language\-model components might fail to provide the expected structural advantage:[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)established a width\- and precision\-independent ceiling on the compositional operations achievable by a single attention layer\. Together, these results changed the evidentiary standard\. High benchmark accuracy remained important, but it was no longer sufficient evidence that a language model had acquired the structural properties required by the task\.
The field’s subsequent development reflects this more cautious standard\. By 2025–2026, specialized non\-linguistic architectures continued to scale independently, including TabPFN\-2\.5, TabICLv2, GraphBFF, and TiRex\[[29](https://arxiv.org/html/2608.28980#bib.bib29),[68](https://arxiv.org/html/2608.28980#bib.bib30),[9](https://arxiv.org/html/2608.28980#bib.bib26),[5](https://arxiv.org/html/2608.28980#bib.bib59)\]\. A different strategy also emerged: instead of adapting a pretrained language model to structured data after the fact, Chronicle jointly pretrains language and time\-series representations through shared attention blocks\[[69](https://arxiv.org/html/2608.28980#bib.bib68)\]\. This represents a distinct hypothesis from both explicit specialization and post hoc language\-mediated adaptation\. At the same time, cases in which language\-mediated systems provide clear practical value increasingly involve orchestration, using language models to direct specialized models or tools, rather than replacing the underlying computation \([Section4\.6](https://arxiv.org/html/2608.28980#S4.SS6)\)\.
The resulting trajectory is therefore better understood as a single question becoming progressively more precise\. The field first asked whether structured data could be expressed through language; it then asked whether language models could perform competitively on such representations; subsequently, it began asking why they succeeded or failed; and it now faces the more precise question of what, exactly, is transferred, preserved, computed, and learned through a language interface\. The disagreements identified above are the unresolved residue of this progression\.[Section3](https://arxiv.org/html/2608.28980#S3)introduces the framework used throughout the remainder of the review to analyze them\.
## 3A Unifying Taxonomy: Representation Regimes and the Four\-Way Distinction
A central challenge in interpreting this literature is that “language\-based” does not describe a single modeling regime\. A system that uses an LLM to generate Python code and invoke XGBoost, a model that projects visual tokens into an LLM’s embedding space, and a model that receives a structured object only as plain text may all be described as LLM\-based, despite relying on fundamentally different mechanisms\. Consequently, evidence that supports one regime is often used to make claims about another\. Resolving this requires two taxonomies that answer two different questions\. The first, developed in[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1), is structural: it asks*where*the computation that must exploit a task’s structure actually happens\. The second, developed in[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3), is functional: given that a result occupies a particular structural regime, it asks*what*the system does with the information available there: describes it, preserves it, computes over it, or learns from it\.[Section3\.2](https://arxiv.org/html/2608.28980#S3.SS2)sits between the two, because it is impossible to say what a result has “replaced” without knowing both where its computation is located and what it does there\.
### 3\.1Eight representational regimes
We organize existing approaches into eight representational regimes, summarized in[Figure2](https://arxiv.org/html/2608.28980#S3.F2)\. A regime is defined by a single question: given a task that requires exploiting some structural property of the input, where does the computation that exploits it actually take place? A system can serialize a graph as plain text while using an ordinary decoder\-only transformer trained by next\-token prediction, and the regime question is still open until we ask whether that transformer’s own forward pass, an artifact it generates, an external tool it calls, or a separate encoder entirely is responsible for the structural computation\.
Because this question has three qualitatively different answers, the eight regimes group into three tiers rather than tracking one quantity end to end\. In Levels 1–4, the language model’s own forward computation over natural\-language tokens is asked to carry out the structural computation directly, and what changes across these four levels is how much linguistic scaffolding is supplied to that same computation: nothing beyond the task description \(Level 1\); a structured serialization of the object itself \(Level 2\); an explicit natural\-language statement of the relevant structural property \(Level 3\); and worked demonstrations \(Level 4\)\. This is the one genuinely monotonic sub\-ordering in the taxonomy, and it is monotonic in a narrow, specific sense that each level strictly increases the task\-relevant information available to the same computational mechanism, not the quality of that mechanism’s reasoning\. In Levels 5–6, responsibility for computation moves outside the language model’s own forward pass: at Level 5 the model generates an executable artifact, typically code, that performs the transformation; at Level 6 the model instead searches over, selects, or calls an external tool or specialized model\[[1](https://arxiv.org/html/2608.28980#bib.bib90),[61](https://arxiv.org/html/2608.28980#bib.bib91),[109](https://arxiv.org/html/2608.28980#bib.bib92)\]\. This boundary overlaps with, but is not identical to, the reasoning\-versus\-tool\-augmentation axis that[Mialon et al\. \[59\]](https://arxiv.org/html/2608.28980#bib.bib110)use to organize a broader survey of augmented language models\. Their axis asks whether a capability comes from a reasoning strategy or an external tool call, largely independent of any particular non\-linguistic modality, whereas the regime ladder asks specifically where the computation carrying a task’s*structural*bias is located, which is why Level 7, a case their reasoning/tool axis does not distinguish from ordinary multimodal input, receives a dedicated tier here rather than being folded into either of theirs\. Level 7 is not a further increment along the Level 5–6 axis in any case\. Here the structural computation is performed by a separate, typically pretrained encoder \(a visual tokenizer\[[87](https://arxiv.org/html/2608.28980#bib.bib72)\], a graph\-encoder projector\[[82](https://arxiv.org/html/2608.28980#bib.bib48),[13](https://arxiv.org/html/2608.28980#bib.bib49)\], or a point\-cloud encoder\[[104](https://arxiv.org/html/2608.28980#bib.bib84)\]\), and language never had causal access to the structural transformation in the first place; it operates only on the encoder’s output\. Level 8 removes the language model entirely\. The three\-tier grouping in[Figure2](https://arxiv.org/html/2608.28980#S3.F2)reflects this directly: it is a labeled ordinal scale tracking where computational responsibility has moved, not a claim that every adjacent pair differs by the same underlying quantity\. Level 7 in particular should not be read as “more” of whatever Level 5 or 6 are; it is a different mechanism, not a further degree of the same one\.
Figure 2:The eight representational regimes considered in this review\. The regimes are grouped into three broad tiers according to the role of language: language\-only approaches \(Levels 1–3\), language\-based approaches with additional scaffolding or specialized components \(Levels 4–6\), and approaches in which the task\-relevant structure is represented outside natural language \(Levels 7–8\)\. The boundary between Levels 6 and 7 marks the transition from language as the computational medium to non\-linguistic representations as the primary carrier of structural information\.Two boundaries matter most for interpreting the evidence in[Section4](https://arxiv.org/html/2608.28980#S4)\. The first separates language\-based reasoning \(Levels 1–4\) from language\-based orchestration \(Level 6\): an LLM that generates feature transformations, searches over model configurations, or constructs and refines machine\-learning pipelines demonstrates that language can effectively*direct*specialized machine learning, but the structural computation remains with the specialized models or tools it invokes, and this is evidence for orchestration, not for replacement\. The second separates language\-mediated representations \(Levels 1–6\) from non\-linguistic ones \(Levels 7–8\): a system that contains an LLM is not thereby a language\-mediated system in the sense this review tests, if the structural information it exploits is carried by a separate encoder the LLM never processes as text\.
The regimes are not mutually exclusive at the level of a paper\.[Table2](https://arxiv.org/html/2608.28980#S4.T2)tags several methods at two levels simultaneously, and it does so because the underlying results genuinely occupy two regimes rather than one: TabLLM\[[34](https://arxiv.org/html/2608.28980#bib.bib31)\]reports both a plain\-serialization baseline and a few\-shot demonstration condition \(Levels 2/4\); SimKGC’s textual scoring model sits at the same boundary\[[94](https://arxiv.org/html/2608.28980#bib.bib53)\]; LLM\-FE\[[1](https://arxiv.org/html/2608.28980#bib.bib90)\]both generates code and searches with it \(Levels 5/6\)\. This is resolved the same way[Section2\.1](https://arxiv.org/html/2608.28980#S2.SS1)resolves it for the review as a whole that the unit classified is the reported result, not the paper, so a single method can legitimately occupy different cells for different components or conditions\. The more informative boundary case is Time\-LLM\[[43](https://arxiv.org/html/2608.28980#bib.bib61)\], marketed and named around “reprogramming” time series through language\.[Table2](https://arxiv.org/html/2608.28980#S4.T2)classifies its actual mechanism at Level 7, not Levels 1–4, because the structural computation it performs on the input sequence does not run through interpretable language tokens at all, and the ablation evidence reviewed in[Section4\.3](https://arxiv.org/html/2608.28980#S4.SS3), that removing its language\-model component does not hurt performance\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\], is exactly the pattern this classification predicts\. if the mechanism were genuinely Level 1–4, removing the language component should remove the computation itself, not leave it intact\. This is the sense in which the taxonomy is explanatory and it generates a testable expectation before the ablation is run\.
The same boundary recurs across every modality examined in[Section4](https://arxiv.org/html/2608.28980#S4), not only time series: tabular prediction separates plain serialization \(Level 2\) from numeric in\-context learning with no language at all \(Level 8, the TabPFN family\); graph learning separates textual encodings of the same graph \(Level 2\) from learned graph\-encoder projections \(Level 7, GraphGPT, LLaGA\); and vision separates frontier vision\-language models operating on learned visual tokens \(Level 7\) from text\-symbol grid representations in which structure has already been discretized into language before the model ever sees it \(Levels 2–3\)\. That the same boundary does the same explanatory work in tabular, graph, time\-series, and vision literatures sharing no common authorship or terminology is the basis for treating it as a property of the representational regime rather than of any single modality\.
The same check, extended to the five adjacent modalities in[Section4\.5](https://arxiv.org/html/2608.28980#S4.SS5)\. Point clouds and knowledge graphs reproduce the boundary directly: PointLLM pairs a specialized point\-cloud encoder with a language interface, placing the structural representation at Level 7 while language handles only captioning and interaction\[[104](https://arxiv.org/html/2608.28980#bib.bib84)\], and knowledge\-graph agents with extensive tool access remain at Level 6 regardless of that access, continuing to underperform specialized graph models at Level 8\[[80](https://arxiv.org/html/2608.28980#bib.bib56)\]\. Protein structure produces a genuine boundary case that[Su et al\. \[77\]](https://arxiv.org/html/2608.28980#bib.bib87)and[Wang and others \[91\]](https://arxiv.org/html/2608.28980#bib.bib88)inject Foldseek\-derived structure tokens \(the output of a separate, non\-linguistic structural\-alignment tool\) directly into a sequence model’s vocabulary, so the structural computation originates at Level 7 but is consumed as ordinary sequence tokens at Level 2\. The eight discrete levels do not separate this cleanly, and they should not be forced to: it is a compositional case, a Level 7 computation feeding a Level 2 representation, read the same way[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1)already reads TabLLM and SimKGC above, and it sharpens rather than weakens the taxonomy by showing what a compositional reading looks like when the composition happens inside the vocabulary rather than across separate components\. Code is the weakest fit of the nine modalities: the reviewed evidence\[[32](https://arxiv.org/html/2608.28980#bib.bib83)\]concerns whether models correctly simulate program execution, a describe\-versus\-compute question in the sense of[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3)rather than a regime\-boundary one, and no non\-linguistic Level 8 baseline is reported for comparison\. For code, the regime boundary is therefore untested rather than confirmed, and it is reported here as such\.
### 3\.2What “replace” means
Knowing which regime a result occupies is not sufficient to say what, if anything, it has replaced\. “Replacement” is not one claim:[Table1](https://arxiv.org/html/2608.28980#S3.T1)distinguishes eight senses, and they do not reduce to points on a single scale\.
Table 1:Eight distinct senses in which a language\-mediated model could be said to “replace” a specialized architecture \([Section3](https://arxiv.org/html/2608.28980#S3)\)\. Treating these as one claim is the second most common overclaim identified in this review, after conflating representational levels \([Figure2](https://arxiv.org/html/2608.28980#S3.F2)\)\.Four of the eight track distinct properties a specialized model has: functional \(predictive performance\), statistical \(sample efficiency\), computational \(FLOPs and latency for equal accuracy\), and practical \(deployment cost, latency, calibration\)\. A language\-mediated system can match a specialized model on any one of these without matching it on the others\.[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)’s ablated forecasters match the functional performance of their language\-mediated counterparts at roughly1000×1000\\timesfewer parameters, which is simultaneously evidence*against*computational replacement for the original models \(they needed far more compute for the same result\) and silent on statistical replacement \(neither variant was tested for sample efficiency against a matched\-bias architecture\)\. Robustness and universality are broader claims again, concerning behavior outside the training distribution and generality across task classes respectively, and are correspondingly rarer to find satisfied even where the narrower properties hold\.
*Structural*replacement concerns whether a language\-mediated system reproduces the inductive bias that a specialized architecture would explicitly encode\. In other words, similar outputs alone are not sufficient: the question is whether the underlying structural principle is also preserved\. This notion has an asymmetric relationship with the other seven forms of replacement\. That asymmetry is not simply an empirical pattern found in the reviewed literature; the two directions of the relationship can each be justified by an independent argument\. The forward direction follows the same generalization\-bound logic that motivates encoding inductive bias architecturally in the first place \([Section2\.4](https://arxiv.org/html/2608.28980#S2.SS4)\): a computation that genuinely respects a structural constraint operates over an effectively smaller, constrained hypothesis class, and standard results relate that constraint directly to sample efficiency and to correct behavior across the full orbit of structure\-preserving transformations, not merely the specific instances a benchmark happens to test\[[79](https://arxiv.org/html/2608.28980#bib.bib15),[105](https://arxiv.org/html/2608.28980#bib.bib12)\]\. Structural replacement is therefore a principled basis for expecting the other properties to follow, grounded in the same complexity argument that justifies architectural bias generally specific to this review\. The reverse direction is not inductive at all but a basic identifiability point: an output value alone underdetermines the mechanism that produced it\. Matching a specialized system’s accuracy on a fixed evaluation set is equally consistent with genuine structural computation, with pretrained world knowledge the comparison did not control for\[[10](https://arxiv.org/html/2608.28980#bib.bib35)\], with a representation that preserves the needed information without the model computing anything invariant over it\[[19](https://arxiv.org/html/2608.28980#bib.bib46)\], or with an information advantage unavailable to the specialized baseline\[[94](https://arxiv.org/html/2608.28980#bib.bib53)\]; no quantity of matched output values on a finite test set can distinguish among these, and only a direct test of the invariance itself, of the kind discussed in[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1), can\.[Section4](https://arxiv.org/html/2608.28980#S4)shows this is not a hypothetical concern: nearly every reviewed case of apparent functional or statistical replacement, when a direct structural test has actually been run, either fails it or has simply never been subjected to one\.
The eighth sense, orchestration, is not a weaker form of replacement but a different relationship, and treating it as a point on the same scale is the source of much of the field’s loosest rhetoric\. When a language model searches over feature transformations\[[1](https://arxiv.org/html/2608.28980#bib.bib90)\]or refines a modeling pipeline\[[61](https://arxiv.org/html/2608.28980#bib.bib91)\], it satisfies none of the other seven senses, because the specialized computation it directs is not performed by the language model at all; this is Level 6 of[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1), and it corresponds to the sense of “replace” that position pieces describing graph or tabular foundation models as already here typically have in mind\[[56](https://arxiv.org/html/2608.28980#bib.bib50)\]\. That usage is locally coherent \(something has changed about who designs the pipeline\), but it answers a different question than the one this review asks, which is whether the language model’s own computation implements the structural function, not whether a language model is present somewhere in the loop that produces it\.
A related and equally common confusion is treating the disappearance of a*named*specialized architecture as evidence that specialization itself has disappeared\. It has more often relocated\.[Su et al\. \[77\]](https://arxiv.org/html/2608.28980#bib.bib87)and[Wang and others \[91\]](https://arxiv.org/html/2608.28980#bib.bib88)remove nothing from a protein language model’s sequence\-only design without also reinjecting Foldseek\-derived structure tokens into its vocabulary;[Kim et al\. \[45\]](https://arxiv.org/html/2608.28980#bib.bib41)pairs language\-model string embeddings with an explicit graph\-attention scaffold rather than a bare LLM; and chemistry language models trained directly on SMILES strings needed graph\-based tools reattached once sequence\-only representations proved unable to track ring systems and branching\[[11](https://arxiv.org/html/2608.28980#bib.bib81)\]\. In each case, a specialized, non\-linguistic component was removed from the headline architecture, and a specialized, non\-linguistic component was added back somewhere else in the system: in the vocabulary, in an attention scaffold, in a reattached tool\. None of these is evidence of representational, architectural, or computational replacement in the senses defined above;each is evidence that specialization remained exactly as necessary as before, relocated to a place the system’s name no longer advertises\.
### 3\.3The central distinction: describe, preserve, compute, learn
Where a result sits in[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1)says nothing about what the system does with the information once it has it\. A result can occupy Level 4 \(language, with demonstrations, as the computational medium\) and still fail a task for any of several unrelated reasons, and distinguishing those reasons requires a second set of questions\. Claims that a model “understands” a structure conflate at least four of them:
1. \(i\)Describe\.Can the model correctly state or explain a structural property, such as permutation invariance or graph isomorphism, independent of any particular input? This is a claim about declarative knowledge, evaluable by asking the model to explain the property in the abstract, and it is the weakest of the four: a model can pass it while being given, on a specific instance, a representation that does not contain the information the property requires\.
2. \(ii\)Preserve\.After a structured object is serialized or otherwise transformed for a specific instance, does the resulting representation retain the information the target computation needs, not all information about the object, but the information relevant to the property in question?
3. \(iii\)Compute\.Given a representation that does preserve the relevant information, does the model’s forward computation implement a function that actually respects the structural property, or can it violate the property in its output even when the input was sufficient to respect it?
4. \(iv\)Learn\.Given a realistic sample and compute budget, can the model acquire and generalize the relevant structural function, as opposed to needing substantially more data than an architecture that encodes the corresponding bias directly? This is a claim about sample and compute efficiency, not about whether learning happens at all\.
These four properties are best read as dimensions a system can satisfy in any combination, not as a strict sequence in which each stage is a prerequisite for the next, and the literature already documents failures in every direction\. Describe does not imply preserve:[Fatemi et al\. \[21\]](https://arxiv.org/html/2608.28980#bib.bib43)hold the graph, task, and model fixed and vary only the textual encoding, producing a 61\.8\-point swing in accuracy that is not explicable by any change in what the model can describe \(the description is unaffected\) but only by what the specific encoding did or did not preserve\. Preserve does not imply compute:[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)show that causal positional encoding makes standard decoder\-only transformers provably non\-permutation\-invariant regardless of whether the input representation preserves the relevant relational information, so a model can receive everything it needs and still compute the wrong function\. Compute does not require describe: mechanistic evidence shows transformers implement specific computational procedures, induction mechanisms for copying, task\-level representations for simple mappings, gradient\-descent\-like computation in linear attention\[[64](https://arxiv.org/html/2608.28980#bib.bib9),[89](https://arxiv.org/html/2608.28980#bib.bib7),[2](https://arxiv.org/html/2608.28980#bib.bib8)\], with no accompanying natural\-language account of what they are doing\. And learn does not require any of the other three to be separately established, as when transformers acquire simple function classes purely from in\-context examples\[[24](https://arxiv.org/html/2608.28980#bib.bib6)\]\. Reporting only one of these four properties and treating it as evidence for the others is, together with conflating regimes in[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1), the most common source of overclaiming identified in this review\. Theoretical results reviewed in[Section6](https://arxiv.org/html/2608.28980#S6)correspond to the computation and learning questions specifically: architectural expressivity and computational limitations address \(iii\), while sample\-complexity and phase\-transition results address \(iv\)\[[46](https://arxiv.org/html/2608.28980#bib.bib14),[79](https://arxiv.org/html/2608.28980#bib.bib15),[105](https://arxiv.org/html/2608.28980#bib.bib12),[26](https://arxiv.org/html/2608.28980#bib.bib18)\]\.
The relationship between this distinction and the eight regimes is complementary\. The regime ladder is a structural taxonomy and it asks where the computation is located\. The describe/preserve/compute/learn distinction is a functional taxonomy: given a location, it asks what the system actually does there\. The two are logically independent: a Level 2 serialization can fail on preservation alone \([Fatemi et al\.](https://arxiv.org/html/2608.28980#bib.bib43)’s result\), while a Level 7 learned\-encoder system can succeed on preservation and computation simultaneously because a separate, specialized network is responsible for both, with the language component contributing nothing to either\. This is precisely why the two taxonomies are necessary together: regime alone would misclassify Time\-LLM as a language\-computation success because language is nominally present, and the four\-way distinction alone has no way to explain*why*removing that presence does not remove the computation, since it never states where a given describe, preserve, compute, or learn result is actually coming from\.
[Figure3](https://arxiv.org/html/2608.28980#S3.F3)reports the resulting synthesis across the modalities in[Section4](https://arxiv.org/html/2608.28980#S4): descriptive competence is close to uniform, while the fraction of cases in which preservation, computation, and learning are also established falls sharply and unevenly\. The figure should be read as a synthesis of the evidence reviewed later no meta\-analysis\.
Figure 3:The four\-stage distinction applied across the modalities reviewed in[Section4](https://arxiv.org/html/2608.28980#S4)\. Cells represent the qualitative evidence synthesized in this review rather than meta\-analysis\. The framework separates descriptive competence from representation preservation, structural computation, and efficient learning, allowing evidence at each stage to be evaluated independently\.This framework does not imply that language\-mediated models are incapable of structural computation in general;[Section6](https://arxiv.org/html/2608.28980#S6)reviews theoretical results establishing exactly when they can\. The narrower claim is that whether a model computes or merely describes a given structure is task\-, representation\-, and architecture\-dependent, and is not reliably predicted by model scale or language fluency alone, which is why[Section4](https://arxiv.org/html/2608.28980#S4)evaluates each of the four properties separately for every modality, rather than inferring three of them from evidence about the fourth\.
## 4Major Methodological Families: A Modality\-by\-Modality Synthesis
We apply the same analytical template across modalities: first, we identify the specialized architecture and inductive bias traditionally used for the task; second, we examine the language\-mediated alternative and determine which representational regime it actually occupies \([Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1)\); third, we assess the strongest positive evidence for that alternative alongside controlled, adversarial, or contamination\-audited evidence that qualifies it; fourth, we identify what specialized computation remains once the language component is isolated; and finally, following the distinction in[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3), we ask whether the evidence supports the claims that the system can describe, preserve, compute, or learn the relevant structure\.[Table2](https://arxiv.org/html/2608.28980#S4.T2)summarizes the resulting classification across modalities, and[Section4\.8](https://arxiv.org/html/2608.28980#S4.SS8)synthesizes the mechanisms that recur across them\.
Table 2:Representative methods across the four core modalities, focusing on their representational regime and principal limitation\. The representation levels follow the taxonomy in[Figure2](https://arxiv.org/html/2608.28980#S3.F2)\.ModalityMethodRepresentationMain limitationTabularTabPFN\-2\.5\[[29](https://arxiv.org/html/2608.28980#bib.bib29)\]Numeric ICL \(L8\)Not language; limited to small–medium tablesTabICLv2\[[68](https://arxiv.org/html/2608.28980#bib.bib30)\]Numeric ICL \(L8\)Not language; primarily classification\-focusedTabLLM\[[34](https://arxiv.org/html/2608.28980#bib.bib31)\]Text serialization \(L2/4\)Overtaken by tree\-based models beyond∼\\sim8 shotsTabula\-8B\[[23](https://arxiv.org/html/2608.28980#bib.bib34)\]Text serialization \(L4\)Most reported advantage traced to contamination\[[28](https://arxiv.org/html/2608.28980#bib.bib36)\]LLM\-FE\[[1](https://arxiv.org/html/2608.28980#bib.bib90)\]Language \+ code \(L5/6\)Not a language\-native predictor; delegates to specialized searchGraphGraphBFF\[[9](https://arxiv.org/html/2608.28980#bib.bib26)\]Numeric transformer \(L8\)Not language; relies on industrial/proprietary dataTalk like a Graph\[[21](https://arxiv.org/html/2608.28980#bib.bib43)\]Text serialization \(L2\)Performance varies substantially with encoding choiceGraphGPT\[[82](https://arxiv.org/html/2608.28980#bib.bib48)\]Learned projector \(L7\)Structural information is carried by a non\-linguistic encoderSimKGC\[[94](https://arxiv.org/html/2608.28980#bib.bib53)\]Text serialization \(L2/4\)Gains rely on textual side information rather than graph topologyTime seriesChronos\[[4](https://arxiv.org/html/2608.28980#bib.bib58)\]Numeric tokenization \(L8\)Not language; the “language” framing is metaphoricalTiRex\[[5](https://arxiv.org/html/2608.28980#bib.bib59)\]Numeric, recurrent \(L8\)Demonstrates the value of specialized architecture rather than LLM lineageTime\-LLM\[[43](https://arxiv.org/html/2608.28980#bib.bib61)\]Reprogrammed embeddings \(L7\)Ablated version outperforms the full model\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\]Chronicle\[[69](https://arxiv.org/html/2608.28980#bib.bib68)\]Joint pretraining \(L8\)Small scaleVisionCNN/ViT specialistsPixel/patch input \(L8\)Domain\-specific; lacks a general language interfaceGPT\-4V / frontier VLMsLearned visual tokens \(L7\)Weak performance on several low\-level visual reasoning tasks\[[22](https://arxiv.org/html/2608.28980#bib.bib70),[70](https://arxiv.org/html/2608.28980#bib.bib69)\]SpatialVLMLearned tokens \+ specialist\-CV labels \(L7\)Depends on specialized vision models for supervisionText\-symbol grid encoding\[[3](https://arxiv.org/html/2608.28980#bib.bib75)\]Structured serialization \(L2/3\)Limited to layouts that can be discretized into symbolic representations### 4\.1Tabular Data
Tabular learning has traditionally relied on tree\-based ensembles and, more recently, transformers pretrained on synthetic priors\. Both approaches encode inductive biases suited to sparse and nonlinear feature interactions rather than the smooth sequential structure on which language models are pretrained\. TabPFN\[[37](https://arxiv.org/html/2608.28980#bib.bib27)\]provides a strong non\-linguistic baseline: it performs in\-context learning directly over feature–value pairs, without language, and can match tuned gradient\-boosting ensembles on small datasets at low inference cost\. TabPFN\-v2, TabPFN\-2\.5, and TabICLv2 extend this paradigm to datasets with up to 10,000 rows\[[38](https://arxiv.org/html/2608.28980#bib.bib28),[29](https://arxiv.org/html/2608.28980#bib.bib29),[68](https://arxiv.org/html/2608.28980#bib.bib30)\]\. Thus, the baseline against which language\-mediated approaches should be compared is itself a strong general\-purpose learner rather than a narrow classical model\.
Language\-mediated approaches show their clearest advantage in the opposite regime: when labeled data are extremely scarce\. TabLLM\[[34](https://arxiv.org/html/2608.28980#bib.bib31)\]and LIFT\[[16](https://arxiv.org/html/2608.28980#bib.bib32)\], which serialize rows and use either prompting or few\-shot demonstrations \(Levels 2 and 4\), can exploit pretrained knowledge when only a handful of labeled examples are available\. The advantage, however, is narrow\. As more labeled data become available, gradient\-boosted trees overtake LLM\-based prediction while requiring substantially less computation\[[39](https://arxiv.org/html/2608.28980#bib.bib42)\]\. More importantly, stronger claims of broad superiority have not survived independent scrutiny\. Tabula\-8B initially reported a 5–15 percentage\-point advantage over XGBoost and TabPFN in few\-shot settings\[[23](https://arxiv.org/html/2608.28980#bib.bib34)\];[Gorla and Puduppully \[28\]](https://arxiv.org/html/2608.28980#bib.bib36)subsequently found that 92\.2% of the reported improvement could be recovered by instruction\-tuning the same base model without tabular exposure, while also identifying train–test contamination that standard deduplication had failed to detect\. Independently,[Silvestri et al\. \[75\]](https://arxiv.org/html/2608.28980#bib.bib37)identified contamination in four of eight widely used tabular benchmarks, and[Bordt et al\. \[10\]](https://arxiv.org/html/2608.28980#bib.bib35)showed that LLMs can memorize popular benchmarks while performing poorly on genuinely in\-context statistical tasks as dimensionality increases\.
Broader evaluations further narrow the positive claim\. Using 142 datasets spanning IID, temporal, and grouped distribution shifts,[Purucker et al\. \[67\]](https://arxiv.org/html/2608.28980#bib.bib38)found tabular foundation models to be strongest in the small\-to\-medium IID regime that dominates existing benchmarks, while tree\-based and deep\-learning methods remain competitive or superior under other conditions\. A separate, as yet unreplicated study reports that LLM performance degrades systematically with dimensionality while classical baselines remain comparatively stable\[[25](https://arxiv.org/html/2608.28980#bib.bib39)\]\.
Tabular data consequently illustrates a narrow form of functional replacement\. TabLLM and LIFT demonstrate comparable predictive performance in extreme few\-shot settings, but only within a regime in which strong non\-linguistic baselines do not necessarily have the same advantage\. Tabula\-8B’s stronger replacement claim did not survive contamination analysis, and no study in this modality directly demonstrates structural replacement\. Thus, the evidence supports descriptive competence and, under restricted conditions, functional performance; it does not establish that language\-mediated systems preserve or compute the same inductive biases as specialized tabular learners, nor that they learn those biases more efficiently\.
### 4\.2Graph\-Structured Data
Graph neural networks encode permutation invariance by construction that graph properties should remain unchanged under permutations of node labels or input ordering, and message\-passing architectures are designed to respect this constraint\.[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)show why decoder\-only language models do not inherit this guarantee automatically\. Causal positional encoding and input ordering introduce dependencies that permutation\-invariant computation must eliminate\. The failure is empirical as well as architectural\. Holding the graph, task, and model fixed while changing only the textual encoding,[Fatemi et al\. \[21\]](https://arxiv.org/html/2608.28980#bib.bib43)report a 61\.8\-point swing in accuracy\.[Thushalika et al\. \[86\]](https://arxiv.org/html/2608.28980#bib.bib47)similarly find reliable graph\-isomorphism detection under one node labeling but substantial degradation after relabeling the same graph\.[Herbst et al\. \[35\]](https://arxiv.org/html/2608.28980#bib.bib44)further show that fine\-tuning does not simply eliminate this sensitivity that larger non\-fine\-tuned models can be more robust to serialization changes, whereas fine\-tuning may reduce sensitivity to node relabeling while increasing sensitivity to formatting and structural changes\. The issue is therefore not merely the choice of prompt; the representation itself may fail to preserve the invariances required by the task\.
Strong results from Level 7 systems do not contradict this conclusion\. GraphGPT and LLaGA route structural information through learned graph encoders that project graph representations into the LLM’s embedding space\[[82](https://arxiv.org/html/2608.28980#bib.bib48),[13](https://arxiv.org/html/2608.28980#bib.bib49)\]\. Their success therefore provides evidence for specialized structural encoders coupled to language models, not for natural language as the computational substrate of graph reasoning\. Recent comparative evaluations reinforce this interpretation: GraphInfer\-Bench finds LLM\-based approaches weaker than conventional GNNs on graph\-comparison tasks\[[66](https://arxiv.org/html/2608.28980#bib.bib55)\], while a large agentic benchmark shows that LLM agents with extensive tool access \(Level 6\) continue to struggle with sophisticated graph analysis\[[80](https://arxiv.org/html/2608.28980#bib.bib56)\]\.
Knowledge\-graph completion provides the clearest positive result and, simultaneously, one of the clearest confounds\. SimKGC\[[94](https://arxiv.org/html/2608.28980#bib.bib53)\]substantially outperforms TransE, ComplEx, and RotatE using a text\-based scoring model\. However, its inputs include textual descriptions of entities and relations that are unavailable to the structural baselines\. The comparison therefore introduces an information advantage and does not establish that language recovers graph topology from triples alone\. Similarly,[Yao et al\. \[106\]](https://arxiv.org/html/2608.28980#bib.bib54)find that fine\-tuned smaller models can outperform zero\-shot frontier models on the same task, suggesting that task\-specific adaptation rather than general linguistic reasoning accounts for much of the observed advantage\.
Graph data therefore provides the clearest illustration of the distinction between describing and preserving structure\. Models can describe permutation invariance fluently while violating it in their own predictions when the representation and computation are tested directly\. In the terminology of[Table1](https://arxiv.org/html/2608.28980#S3.T1), this is one of the few modalities in which structural replacement has been subjected to a direct test, and the test fails\. SimKGC’s apparent functional replacement is also weakened once its textual information advantage is controlled, while the strongest graph–LLM systems generally retain an explicit non\-linguistic structural pathway at Levels 7–8\.
### 4\.3Time\-Series Forecasting
Time\-series forecasting relies on inductive biases concerning temporal order, locality, and dependence across time\. Specialized architectures encode these properties directly, whereas a pretrained language model has no inherent reason to respect them\. This modality therefore provides some of the strongest controlled evidence against the claim that a pretrained language model automatically supplies the required temporal inductive bias\.
[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)evaluate three widely used LLM\-based forecasters \(Time\-LLM\[[43](https://arxiv.org/html/2608.28980#bib.bib61)\], GPT4TS/OneFitsAll\[[114](https://arxiv.org/html/2608.28980#bib.bib62)\], and CALF\[[52](https://arxiv.org/html/2608.28980#bib.bib63)\]\) and systematically remove or replace their language\-model components\. Across thirteen datasets and two metrics, the resulting language\-free variants outperform the original models in most comparisons while using roughly1000×1000\\timesfewer parameters and up to three orders of magnitude less training time\. Randomly reinitializing the pretrained weights also matches or exceeds the performance of the pretrained models\. Most strikingly, shuffling the input sequence has little effect on predictions despite temporal order being central to the task\. The result is decisive for the three architectures tested, although whether the finding generalizes to the broader family of LLM\-adapted forecasters remains open\.
Time\-series foundation models reinforce this conclusion from the non\-linguistic side\. Chronos is described using a “language of time series” framing but operates on quantized numerical values through a T5 architecture without natural\-language input\[[4](https://arxiv.org/html/2608.28980#bib.bib58)\]\. TiRex uses an xLSTM architecture without an LLM component and achieves state\-of\-the\-art performance on GIFT\-Eval\[[5](https://arxiv.org/html/2608.28980#bib.bib59)\]\. Independent evaluations likewise report that LLM\-based approaches can underperform numerical methods in epidemic forecasting\[[41](https://arxiv.org/html/2608.28980#bib.bib67)\], while train–test leakage and temporal correlation have been identified as sources of inflated performance in time\-series foundation\-model benchmarks\[[58](https://arxiv.org/html/2608.28980#bib.bib66)\]\.
A genuinely different direction has also emerged\. Rather than adapting a pretrained language model to time series, Chronicle jointly pretrains language and time\-series representations using shared attention blocks\[[69](https://arxiv.org/html/2608.28980#bib.bib68)\]\. This approach is not contradicted by the ablation results of[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57); it implements a different hypothesis in which the two modalities are learned jointly rather than one being attached to a pretrained language model\. The result remains preliminary, however, given the model’s relatively small scale and the absence of independent replication\.
Time series therefore provides the clearest separation between describing a temporal structure and computing over it\. Representing a sequence in a language\-model\-compatible format is straightforward, but[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)directly test whether the language component contributes to the computation and find little evidence that it does for the architectures examined\. The result is correspondingly negative for computational replacement: removing the language component yields equal or better accuracy at substantially lower cost\. It is also negative for structural replacement, since a basic perturbation that should matter to a genuinely temporal computation \(input shuffling\) has little effect\.
### 4\.4Vision and Spatial Reasoning
Vision requires distinguishing at least three regimes that are often conflated: systems that convert visual information into language before reasoning \(Levels 2–3\); vision\-language models that process images through learned visual tokens connected to a language model \(Level 7\); and specialized vision architectures \(Level 8\)\. Most current vision\-language systems occupy the second regime\. Their performance therefore cannot be interpreted as evidence that language itself preserves or computes visual structure\.
BLINK reports only 51\.26% accuracy for the best evaluated vision\-language model across fourteen computer\-vision tasks reformulated as visual question answering, compared with 95\.7% for humans\[[22](https://arxiv.org/html/2608.28980#bib.bib70)\]\. A complementary study reports roughly 58% average accuracy across four frontier models on simple pixel\-geometry tasks, again substantially below human performance\[[70](https://arxiv.org/html/2608.28980#bib.bib69)\]\. A larger and more recent benchmark reaches a similar conclusion: across spatial\-reasoning tasks involving depth, orientation, scale, and navigation, the best evaluated model achieves 54\.93% accuracy in multiple\-choice settings and 40\.93% in open\-ended settings, compared with 87\.57% and 64\.93% for humans, respectively\[[97](https://arxiv.org/html/2608.28980#bib.bib76)\]\. Mechanistic evidence provides a possible explanation that predictions change relatively little when visual patches are randomly permuted, suggesting that the downstream language model does not fully exploit the spatial organization encoded by the visual representation\[[48](https://arxiv.org/html/2608.28980#bib.bib78)\]\.
Grid2Matrix identifies a more specific bottleneck\. It describes a “Digital Agnosia” phenomenon in which the visual encoder retains substantially more information than is expressed through the model’s language output\[[113](https://arxiv.org/html/2608.28980#bib.bib74)\]\. Controlled experiments varying encoder architecture, positional encoding, and training objective across the LLaVA family report persistent spatial\-reasoning deficits, suggesting that scale or encoder selection alone does not resolve the problem\[[3](https://arxiv.org/html/2608.28980#bib.bib75)\]\. This is precisely the preservation\-versus\-computation distinction introduced in[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3): a visual encoder may preserve spatial information that the subsequent language\-mediated computation fails to exploit reliably\.
Positive results nevertheless demonstrate that language can work effectively once spatial structure has already been discretized\. Text\-symbol grid representations reach 84–91% accuracy on a spatial\-localization task for which the identical task posed in pixels achieves 60–73%\[[3](https://arxiv.org/html/2608.28980#bib.bib75)\]\. SpatialEval similarly finds text\-only reasoning outperforming vision\-augmented models on synthetic maze and map tasks\[[93](https://arxiv.org/html/2608.28980#bib.bib73)\]\. These approaches explicitly encode spatial structure into symbolic representations before language processing begins\. They therefore provide evidence for language as a substrate for already\-discretized structure, rather than for language as a general replacement for continuous visual computation\.
Vision consequently sharpens the distinction between representation and computation\. For raw visual inputs, structural replacement has been tested directly through perturbations such as patch permutation and has not been established\. For already\-discretized spatial structure, text\-symbol representations come substantially closer to functional replacement because the representation has been designed to preserve the information required by the task before language processing begins\. Describing and reasoning over spatial relations is therefore within language models’ capabilities once the relevant structure is made symbolic; preserving and exploiting continuous geometric structure through language\-mediated computation remains substantially less reliable\.
### 4\.5Adjacent Modalities: Does the Pattern Generalize?
Five additional modalities provide a broader test of whether the preceding pattern is specific to tabular, graph, time\-series, and vision data or recurs across other forms of structured information\. The evidence is considerably thinner\. So the conclusions below should be regarded as directional\.
#### Chemistry\.
Chemistry provides the strongest positive result in the review\. MoLFormer, a 1\.1\-billion\-molecule transformer trained on SMILES without explicit graph structure, outperforms supervised and self\-supervised graph neural networks across ten molecular\-property benchmarks\[[73](https://arxiv.org/html/2608.28980#bib.bib79)\]\. This constitutes a genuine case in which a sequence representation can functionally substitute for a graph\-specific architecture at sufficient scale\. The result, however, does not generalize uniformly within the modality\. A 2026 study finds that smaller SMILES\-only models exhibit “structural blindness” to rings and branching, with performance recovering only after graph\-based information is reintroduced\[[11](https://arxiv.org/html/2608.28980#bib.bib81)\]\. A 2025 survey similarly concludes that general\-purpose LLMs remain insufficient for many scientific chemistry tasks, with recent systems increasingly orchestrating specialized tools rather than replacing them\[[33](https://arxiv.org/html/2608.28980#bib.bib82)\]\. Chemistry therefore illustrates both the potential and the limitation of scale\. MoLFormer’s result establishes functional replacement at large scale, but structural replacement has not been directly tested\[[73](https://arxiv.org/html/2608.28980#bib.bib79)\]\.
#### Code\.
Code provides a particularly clean test of the describe\-versus\-compute distinction because code is already a symbolic language\. CRUXEval finds frontier models achieving only 67% and 63% accuracy when predicting the input/output behavior of short Python functions, despite strong performance on code generation\[[32](https://arxiv.org/html/2608.28980#bib.bib83)\]\. Fluent code generation therefore does not imply reliable simulation of program execution\. No non\-linguistic specialized baseline is reported for this task in the reviewed literature, so the regime boundary for code remains untested\.
#### Point Clouds and 3D Data\.
Point\-cloud results closely parallel the vision findings\. Point\-cloud\-augmented multimodal models outperform text\-only models on some spatial\-reasoning tasks but continue to fail on basic binary spatial relations\[[111](https://arxiv.org/html/2608.28980#bib.bib86)\]\. PointLLM achieves strong 3D captioning by combining a specialized point\-cloud encoder with a language model\[[104](https://arxiv.org/html/2608.28980#bib.bib84)\]\. This is a Level 7 configuration: the structural representation remains in the non\-linguistic encoder, while language primarily supplies the interaction interface\.
#### Protein Structure\.
SaProt\[[77](https://arxiv.org/html/2608.28980#bib.bib87)\]and S\-PLM\[[91](https://arxiv.org/html/2608.28980#bib.bib88)\]reintroduce structure\-derived information into sequence\-based protein models after sequence\-only representations proved insufficient for capturing relevant structural relationships\. The mechanism parallels the corrective pattern observed in chemistry and graph learning that once the structural information becomes important to the task, an explicit structural channel is restored\. These systems therefore illustrate relocated specialization rather than its replacement\.
The adjacent modalities support the same qualification observed in the four core domains\. Language\- or sequence\-based representations can work well when the relevant structure is shallow, redundant with the representation, or recoverable through sufficiently large domain\-specific pretraining\. When the task depends critically on topology, execution semantics, or geometric relationships, the strongest systems continue to retain or restore an explicit structural channel\.
### 4\.6Language as Orchestrator: The Genuine Growth Area
The strongest recent evidence for the practical value of language\-mediated systems concerns orchestration and not replacement\. At Level 6 of[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1), LLMs generate, select, evaluate, and refine specialized computational components rather than performing the underlying structural computation themselves\. LLM\-FE generates and evaluates feature\-transformation programs within an evolutionary search and consistently improves over previous automated feature\-engineering methods\[[1](https://arxiv.org/html/2608.28980#bib.bib90)\]\. MLE\-STAR retrieves baseline models and iteratively refines them through targeted, ablation\-guided experimentation, earning medals in 64% of evaluated Kaggle competitions\[[61](https://arxiv.org/html/2608.28980#bib.bib91)\]\. At industrial scale, a planner\-guided multi\-agent system reduced feature\-engineering time from three weeks to one day\[[84](https://arxiv.org/html/2608.28980#bib.bib93)\]\.
These results establish a meaningful and increasingly practical role for language models that searching over specialized methods, composing tools, evaluating alternatives, and refining machine\-learning pipelines\. The underlying structural computation, however, remains with the specialized models being selected or invoked\. The evidence therefore supports orchestration in the specific sense of[Table1](https://arxiv.org/html/2608.28980#S3.T1), not the other seven senses of replacement defined there\. An important open question follows: whether the component\-ablation logic that isolated the contribution of language models in time\-series forecasting\[[81](https://arxiv.org/html/2608.28980#bib.bib57)\]can similarly determine how much of the observed orchestration benefit is attributable to the language model itself rather than to the search and tool\-use process it enables\.
### 4\.7Why These Comparisons Are \(Not\) Scientifically Valid
Claims that one method “outperforms” another are meaningful only when the conditions of comparison are explicit\. Three factors recur across the modalities reviewed above and determine whether such comparisons are informative\. First is the*data regime*\. Language\-mediated methods can exhibit advantages in extreme few\-shot settings but lose those advantages as labeled data accumulate\[[39](https://arxiv.org/html/2608.28980#bib.bib42)\]\. Second is the*compute budget*\. The ablation of[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)shows that removing the language\-model component can reduce parameter count by roughly1000×1000\\timeswithout reducing predictive performance, making accuracy\-only comparisons potentially misleading\. Third is*information availability*\. SimKGC and MoLFormer both benefit from information \(textual descriptions in the former and large domain\-specific pretraining corpora in the latter\) that may not be available to their specialized baselines\. In such cases, a comparison confounds representational choice with information access\. Comparisons that fail to control or report data regime, compute budget, and information availability should therefore be interpreted cautiously\. A reported performance difference may reflect unequal data, computation, or information rather than a genuine advantage of one representational strategy over another\.
### 4\.8Cross\-Modal Synthesis
Considered individually, the nine modalities above appear to constitute nine distinct application areas\. Considered comparatively, however, they converge on a much smaller set of recurring mechanisms\. These mechanisms provide the strongest evidence for the review’s central argument\.
The clearest convergence is a three\-part pattern that recurs wherever the evidence is sufficiently developed to permit controlled comparison\. First, a strong non\-linguistic specialized baseline can outperform language\-mediated alternatives when evaluated under matched conditions: TabPFN in tabular learning, TiRex in time series, GNNs in graph learning, and specialized vision architectures in visual reasoning\. Second, language\-mediated systems can obtain narrow advantages under particular conditions, such as extreme few\-shot settings in tabular learning or access to additional textual information in knowledge\-graph completion\. Third, when direct controls are available, either the apparent advantage disappears under closer scrutiny, as in Tabula\-8B, or the language component turns out not to be responsible for the underlying computation, as in the ablation of[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)\.
A second convergence is even more striking that at least three modalities independently adopted the same corrective strategy\. CARTE introduces a graph\-attention scaffold around language\-derived representations; SaProt and S\-PLM reintroduce structure\-derived tokens; and chemistry systems reattach graph\-based tools after sequence\-only representations prove insufficient\. These are not identical architectures, but they implement the same methodological response that when a language\- or sequence\-based representation fails to capture a task\-relevant structural property, an explicit non\-linguistic structural channel is restored\.[Section7\.1](https://arxiv.org/html/2608.28980#S7.SS1)examines whether this recurring response is better explained as a mechanistic property of the representational regime than as a collection of domain\-specific engineering decisions\.
The modalities diverge, however, in a way that is consistent with the taxonomy\. When task\-relevant structure can be discretized into a symbolic representation before language processes it \(as in text\-symbol spatial grids or SMILES\-based molecular representations\), language models can perform well because the representation preserves much of the information required by the task\. When the relevant structure is continuous, relational, or defined by an invariance that serialization does not automatically preserve \(as in raw visual geometry, graph topology, or temporal ordering\), the same strategy is substantially less reliable, and strong systems tend to retain or restore a non\-linguistic structural pathway\.
This pattern suggests that*discretizability*may be an important moderator of when language\-mediated representations can substitute for specialized architectures\. The critical question is not simply whether an object can be written as a sequence, but whether the task\-relevant structural relations survive that transformation in a form that the subsequent computation can exploit\. Chemistry provides an important qualification\. MoLFormer’s success at very large scale, together with the structural failures observed in smaller SMILES\-only models, suggests that sufficiently large domain\-specific pretraining can partially compensate for representational limitations\. Whether such compensation remains possible under matched compute, data, and information budgets is not yet established\.
Finally, no modality in this review provides unambiguous evidence of*structural replacement*\(that is, of a language\-mediated computation implementing the same structural constraint that a specialized architecture would encode explicitly\) that survives a direct test\. The strength of this negative conclusion differs across modalities\. Time series, graphs, and raw\-pixel vision subject the claim to direct tests: component ablation, controlled serialization and permutation analysis, and patch\-permutation experiments, respectively\. In each case, structural replacement fails the test\. By contrast, tabular data, chemistry, code, and many of the benchmarks summarized in[Table3](https://arxiv.org/html/2608.28980#S5.T3)have not undergone an equivalent direct structural test\. Their apparent successes may instead be weakened by information advantages, narrow data regimes, or benchmark contamination, but these findings do not constitute direct tests of structural replacement\.
The evidence for the review’s central claim is consequently strongest where the relevant structural property has been tested directly and weaker where the literature has relied primarily on benchmark performance\.[Section7](https://arxiv.org/html/2608.28980#S7)returns to this distinction and asks the harder question left open by the modality\-level evidence that whether the observed convergence reflects limitations of current language models or a more fundamental limitation of language\-mediated computation itself\.
## 5Datasets and Benchmarks
A dataset, a task, and a benchmark are not the same thing, and the difference matters for what follows\. A*dataset*is a collection of examples; a*task*is the prediction problem defined over it; a*benchmark*is a standardized combination of dataset, task, split, metric, and baseline that licenses a specific comparison\. A benchmark score is therefore always a claim about performance under one particular combination of these choices, not a direct measurement of representation, computation, or replacement in the sense of[Section3](https://arxiv.org/html/2608.28980#S3)\. This distinction is not pedantry because[Section3\.2](https://arxiv.org/html/2608.28980#S3.SS2)shows that a single accuracy number cannot, on its own, distinguish functional replacement from statistical, computational, or structural replacement, and a benchmark’s design determines which of these a given score can possibly speak to before any model is ever run on it\.
Table 3:Benchmarks referenced across the review, selected to illustrate what each protocol can and cannot distinguish \([Section5](https://arxiv.org/html/2608.28980#S5)\)\.[Table3](https://arxiv.org/html/2608.28980#S5.T3)summarizes the benchmarks this review relies on most, chosen to illustrate what each protocol can and cannot distinguish\. Read against the regime ladder and the four\-way distinction, a pattern emerges that this section treats as a finding in its own right that almost every entry directly evaluates predictive performance, and almost none evaluates anything else\. TabArena, TALENT, and GIFT\-Eval measure in\-domain accuracy; BLINK and MMVP measure accuracy on reformatted or targeted subsets; WN18RR and FB15k\-237 measure link\-prediction accuracy under a comparison that, as discussed below, is not actually matched\. Only two exceptions in the entire reviewed corpus depart from pure task\-level evaluation toward representation\- or replacement\-level evidence:[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s component ablation, which is a one\-off experimental protocol rather than a standing benchmark, and the direct invariance\-violation metric used by[Yoon and others \[108\]](https://arxiv.org/html/2608.28980#bib.bib52)and[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)and discussed further in[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1)\. Everything else in this review’s evidence base infers describe, preserve, compute, and learn from a single number that was designed to measure none of them individually\.
#### Benchmark bias\.
Where benchmark design does shape the comparison, the evidence located in this review points in one direction more than the other\.*Regime selection*favors language\-mediated methods where it is least informative to do so\. TabArena and TALENT are dominated by small\-to\-medium IID tabular datasets, exactly the regime in which tabular foundation models perform most strongly, and[Purucker et al\. \[67\]](https://arxiv.org/html/2608.28980#bib.bib38)’s broader, purpose\-built alternative shows that advantage weakening substantially under temporal and grouped distribution shift\.*Contamination*inflates reported performance in the same direction that independent audits found train–test overlap in widely used tabular benchmarks\[[28](https://arxiv.org/html/2608.28980#bib.bib36),[75](https://arxiv.org/html/2608.28980#bib.bib37)\]and leakage in time\-series foundation\-model benchmarks\[[58](https://arxiv.org/html/2608.28980#bib.bib66)\], in each case a bias that specifically favors whichever pretrained model had prior exposure to the test distribution\.*Information asymmetry*does the same in knowledge\-graph completion\. WN18RR and FB15k\-237 supply textual entity and relation descriptions to text\-based methods that structural embedding baselines such as TransE, ComplEx, and RotatE never receive, so SimKGC’s reported advantage reflects an unequal comparison, not solely a representational one\.
The reverse direction is real but operates through a different mechanism\. A dedicated search of the wider evaluation\-methodology literature, beyond this review’s own 159\-paper corpus, finds systematic output\-format bias documented across state\-of\-the\-art LLMs on exact\-match and structured\-output tasks\[[53](https://arxiv.org/html/2608.28980#bib.bib71)\]\. A scoring artifact that penalizes a linguistically fluent but format\-variable answer regardless of whether the underlying computation was correct, and that a specialized model producing output in the required format by construction never incurs\. This is not the same claim as “benchmarks favor specialized architectures because the tasks are numerically or structurally demanding”; it is a claim about the scoring mechanism itself disadvantaging language\-mediated output independent of the reasoning it reflects, and it is the clearest evidence this review locates running against language\-mediated methods\. The two biases are not symmetric in the sense of canceling each other out: regime selection, contamination, and information asymmetry each shape which*tasks*get chosen or how*information*is distributed, while output\-format bias shapes how a correct answer gets*scored*, a narrower but real effect operating alongside, not instead of, the biases documented above\.
A separate class of benchmark simply cannot support a replacement claim in either direction, which is a different problem from bias\. NLGraph and GraphQA report no GNN baseline in most comparisons, so no result on either benchmark can establish whether language\-mediated reasoning is competitive with, better than, or worse than the specialized computation it would need to replace\. MLE\-Bench’s Lite subset is easier than the full benchmark it is drawn from, so medal\-rate results such as[Nam et al\.](https://arxiv.org/html/2608.28980#bib.bib91)’s 64% should not be read as transferring automatically to the harder setting\. BLINK’s multiple\-choice format may bottleneck scores independently of the visual competence it is meant to measure\. Each of these is a documented limitation in[Table3](https://arxiv.org/html/2608.28980#S5.T3)\.
#### What current benchmarks do not test\.
Two gaps follow directly from the pattern above and connect this section forward to[Section8](https://arxiv.org/html/2608.28980#S8)\. First, information preservation and structural computation, the second and third stages of[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3), are evaluated almost nowhere:[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1)shows only two papers in this review’s core matrix use a metric designed to detect a preserved\-but\-uncomputed structural property rather than an accuracy proxy that cannot distinguish the two\. Second, replacement in the structural or computational sense of[Table1](https://arxiv.org/html/2608.28980#S3.T1)is tested almost nowhere:[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s ablation is the only genuine component\-level test located in this review, and[Section4\.6](https://arxiv.org/html/2608.28980#S4.SS6)notes the same gap for Level\-6 orchestration systems, whose reported end\-task success cannot currently be decomposed into the language model’s contribution versus the specialized tool’s\. Coverage is also uneven by modality: chemistry, protein structure, and point clouds now each have exactly one benchmark entry in[Table3](https://arxiv.org/html/2608.28980#S5.T3)\(MoleculeNet, ProteinGym, and ModelNet40 respectively\) against several apiece for the four core modalities, which mirrors rather than resolves the thinner evidence base[Section4\.5](https://arxiv.org/html/2608.28980#S4.SS5)already discloses for those three; a single benchmark supports far less scrutiny of contamination, baseline fairness, or regime selection than the multiple entries available for tabular, graph, time\-series, and vision data\. Code remains the one modality with no benchmark entry at all, consistent with[Section4\.5](https://arxiv.org/html/2608.28980#S4.SS5)’s finding that no non\-linguistic baseline is reported for it in the reviewed literature\. No benchmark in this review’s corpus, in any modality, was purpose\-built to test serialization sensitivity or invariance directly\.[Section8\.1](https://arxiv.org/html/2608.28980#S8.SS1)specifies the controlled experiment this absence motivates\. A design that holds information content and sample budget fixed while measuring the direct invariance\-violation metric[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1)shows\.
## 6Evaluation Methodology and Theoretical Foundations
### 6\.1What current metrics can and cannot distinguish
A benchmark score is an observation about model behavior, not a direct measurement of the mechanism that produced it\.[Table4](https://arxiv.org/html/2608.28980#S6.T4)summarizes the evaluation metrics used across the literature reviewed in[Section4](https://arxiv.org/html/2608.28980#S4), and reading it against[Sections3\.1](https://arxiv.org/html/2608.28980#S3.SS1)and[3\.3](https://arxiv.org/html/2608.28980#S3.SS3)exposes a hierarchy the metrics themselves do not make explicit\. Accuracy, error, and task\-specific scores measure*output performance*means whether a prediction matches a label\. Zero\- and few\-shot deltas measure a form of*generalization behavior*: whether performance survives reduced adaptation\. Neither level says anything about*representation fidelity*\(whether the information a task requires survived the transformation into a language\-mediated form\) or about*computational behavior*\(whether the model’s forward pass respects the structural property that fidelity would make available\)\. The strongest level,*mechanistic replacement*, asks whether a specialized component can actually be removed without loss under matched conditions, and almost nothing in[Table4](https://arxiv.org/html/2608.28980#S6.T4)operates at this level\. Two systems can therefore be*predictively equivalent*\(similar accuracy\) while differing entirely in whether they are representationally equivalent \(retain the same task\-relevant information\), computationally equivalent \(implement the same function\), or architecturally substitutable in the stricter senses[Table1](https://arxiv.org/html/2608.28980#S3.T1)distinguishes\. Collapsing these into a single accuracy comparison is the same conflation[Section3\.2](https://arxiv.org/html/2608.28980#S3.SS2)identifies at the conceptual level, located here in the measurement instrument itself\.
Direct tests of structural behavior narrow this gap only rarely\. Measuring the maximum change in a model’s output under a structure\-preserving transformation,maxπ\|f\(x\)−f\(π\(x\)\)\|\\max\_\{\\pi\}\\left\|f\(x\)\-f\(\\pi\(x\)\)\\right\|, is a representation\- and computation\-level test\. It asks whether the model’s behavior is invariant\. Only two studies in the reviewed literature use such a measure\[[108](https://arxiv.org/html/2608.28980#bib.bib52),[19](https://arxiv.org/html/2608.28980#bib.bib46)\]\. The strongest available evidence at the mechanistic\-replacement level is not a metric at all but a controlled ablation \(removing or reinitializing a specific component while holding data, task, and evaluation fixed, as in[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)\) because only a counterfactual comparison of this kind can establish that a component was unnecessary\. Aggregate benchmark scores compound the problem that averaging across datasets of different structural difficulty, as TabArena and TALENT do for tabular data \([Section5](https://arxiv.org/html/2608.28980#S5)\), can report a healthy mean while concealing failure concentrated in the harder, less\-represented cases\[[67](https://arxiv.org/html/2608.28980#bib.bib38)\]\. The literature has therefore extensively measured whether models are accurate, and rarely measured whether they compute the structure that would explain that accuracy\.[Section8](https://arxiv.org/html/2608.28980#S8)returns to this gap and specifies the controlled design needed to close it\.
Table 4:Evaluation metrics used across the reviewed literature\. Accuracy\-based metrics dominate but cannot, by construction, distinguish “the model computed the right structure” from “the model reached the right answer by another route”; the direct invariance\-violation metric is the only one in this table designed specifically to make that distinction \([Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3)\)\.
### 6\.2Why the empirical pattern has a theoretical floor
“Theoretical floor” names a specific, bounded claim here, not a claim of impossibility in general\. For certain structural properties, formal results establish a computational or sample\-complexity cost that does not vanish merely by increasing model scale within the same representational regime, unless the model also acquires the relevant structure through its architecture, its representation, or its training data\. This is weaker than saying language\-mediated computation cannot in principle solve these tasks, and stronger than an empirical trend that might simply reflect insufficient scale to date; the distinction matters because the two readings license different predictions about what more compute would do, and this subsection is careful throughout about which of the two a given result actually supports\. The starting point is the No Free Lunch theorem: without assumptions about the target function, no learner has a universal advantage, so reliable generalization from finite data requires an inductive bias matched to the task’s structure\[[99](https://arxiv.org/html/2608.28980#bib.bib1)\]\. This result is formal but general; it does not say which architectures pay which costs for which structures, which is what the results below establish case by case\.
A clarification is necessary before comparing biases at all\. A language model has inductive biases of its own, arising from its attention architecture, its positional encoding, its tokenization, its pretraining objective, and the distribution of its pretraining data\. The relevant question is therefore not specialized bias versus no bias, but explicit, domain\-matched bias built into an architecture by construction versus general\-purpose bias whose match to a specific structural property is incidental and must be checked, not assumed\. Every result below should be read as characterizing whether the bias transformers already have happens to match a given structural property, not as showing that transformers compute unconstrained by any bias\.
For some function classes, formal results say the match is close\.[Kim et al\. \[44\]](https://arxiv.org/html/2608.28980#bib.bib10)prove that transformers can be minimax\-optimal for function classes such as smooth nonparametric regression and Bayesian posterior mixtures, provided the pretraining distribution is matched to the target tasks\.[Wakayama and Suzuki \[90\]](https://arxiv.org/html/2608.28980#bib.bib17)prove that the posterior\-variance component of in\-context risk decays exponentially with the number of demonstrations under stated assumptions\. Both are formal results for specific function classes, not empirical regularities and not claims about in\-context learning in general; together they provide a theoretical account for the positive results this review documents in low\-data settings such as tabular prediction \([Section4\.1](https://arxiv.org/html/2608.28980#S4.SS1)\), where a broad pretraining distribution can supply a useful statistical prior without a task\-specific architecture\.
This positive picture is itself contested, and a dedicated search for counterevidence located a direct challenge\. A body of theoretical work argues that in\-context regression succeeds because transformers implement a known algorithm \(ordinary least squares or gradient descent\) during the forward pass\.[Hill et al\. \[36\]](https://arxiv.org/html/2608.28980#bib.bib11)report empirical evidence against this specific mechanistic story: transformers trained for in\-context least\-squares regression fail to generalize once the prompt distribution shifts, a pattern inconsistent with genuinely implementing OLS, and their behavior instead correlates with spectral signatures of the training distribution, consistent with a form of memorization rather than algorithm execution\. This work does not contradict[Kim et al\.](https://arxiv.org/html/2608.28980#bib.bib10)’s minimax\-optimality result directly \(the two examine different function classes and different notions of what “implementing an algorithm” requires\), but it is a direct, located counterexample to the stronger informal claim, made elsewhere in this literature, that strong in\-context regression performance is evidence the model has learned a generalizable procedure rather than a distribution\-specific shortcut\. The theoretical literature on why in\-context learning succeeds is not settled even for the function classes where it is reported to succeed\.
The situation is different for stronger compositional or relational structure\.[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)formally prove that one\-layer softmax attention cannot solve three\-way matching, function composition, or binary\-relation composition, regardless of width or numerical precision, a limitation independent of scale by construction, since the proof holds for arbitrarily large width and precision\. The same paper constructs an alternative mechanism, Strassen attention, and proves it solves all three tasks, which is the detail that makes this a claim about architecture rather than about capacity: the barrier is removable, but only by changing what the attention mechanism computes, not by enlarging the existing one\. This is the clearest formal theorem in this review’s theoretical corpus and the closest direct theoretical counterpart to the empirical failures documented in graph\-structured data \([Section4\.2](https://arxiv.org/html/2608.28980#S4.SS2)\)\.
Permutation invariance illustrates the same architecture\-versus\-scale distinction with a matching efficiency result\.[Tabaghi and Wang \[79\]](https://arxiv.org/html/2608.28980#bib.bib15)prove that a permutation\-invariant architecture can universally represent identifiable multiset functions using a latent dimension of2DN2DN\(vector dimensionDD, multiset sizeNN\), a formal, constructive bound, and a substantial improvement over theO\(ND\)O\(N^\{D\}\)bound implied by earlier sum\-decomposable models\. A standard transformer carries no equivalent guarantee and must learn permutation\-invariant behavior from data if it acquires it at all\.[Chiu et al\. \[14\]](https://arxiv.org/html/2608.28980#bib.bib13)report the empirical consequence that explicitly encoding permutation invariance matches a transformer’s in\-context\-learning performance on a permutation\-invariant task using roughly an order of magnitude fewer parameters\. A formal bound and an empirical efficiency result, read together, do not show that transformers cannot learn structured functions; they show that an architecture encoding the relevant structure by construction reaches the same performance at a measurably lower parameter cost\.
Two further results including,[Webson and Pavlick \[98\]](https://arxiv.org/html/2608.28980#bib.bib100)show experimentally that language models can learn effectively from prompts whose stated content is actively misleading, indicating that a prompt’s semantic content is not reliably coupled to the computation the model performs\.[Yoon and others \[108\]](https://arxiv.org/html/2608.28980#bib.bib52)report a directly relevant engineering finding that prompting and formatting alone were insufficient, in their experiments, to reliably induce order\-invariant behavior, and achieving it required modifying the positional mechanism\. Neither is a theorem; both are consistent with, and offer empirical support for, the distinction in[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3)between describing a structural concept and computationally implementing it;[Yoon and others](https://arxiv.org/html/2608.28980#bib.bib52)’s result specifically locates where architecturally encoded bias would need to be restored \(the positional mechanism\) once description and prompting are shown insufficient on their own\.
The pretraining distribution is a further candidate locus for structural bias, and the evidence here is again a mix of formal and empirical results applying to a narrower setting than in\-context learning in general\.[Goddard et al\. \[26\]](https://arxiv.org/html/2608.28980#bib.bib18)empirically characterize, through a phase diagram built from simulated regression tasks, a phase transition in task diversity: below a critical level, in\-context learning remains tied to the pretraining task distribution; above it, the learned procedure generalizes to a broader task space\.[Azizian and Hasan \[6\]](https://arxiv.org/html/2608.28980#bib.bib19)prove that the choice of pretraining\-task distribution carries a formal trade\-off between robustness to distribution shift and sample efficiency, so widening the pretraining distribution is not a free way to acquire more structure\. It substitutes one cost, task\-specific data, for another, a less sample\-efficient prior\. Universal\-approximation results establish that transformer\-based in\-context learning is expressive enough, in principle, to represent a broad class of functions\[[49](https://arxiv.org/html/2608.28980#bib.bib16)\], but expressivity is a different property from the sample and compute efficiency the results above address, and this specific gap \(between what transformers can represent and what they can efficiently learn\) is exactly where the current theoretical literature remains incomplete\. No result surveyed here characterizes the sample complexity of learning a permutation\-invariant or graph\-structured function through pretraining\-distribution diversity alone, which is the missing connective theory[Section8](https://arxiv.org/html/2608.28980#S8)treats as this review’s central open question\.
### 6\.3Scaling is real, sub\-linear, and does not by itself close the gap
This subsection’s title makes three separable claims, and each should be checked against what the referenced studies measure\. Whether scaling improves performance at all; that any such improvement is sub\-linear, with respect to a specified quantity; and that scaling closes the gap to a matched\-bias specialized architecture rather than merely improving in isolation\. The literature supports the first claim directly, the second precisely once the scaled quantity is specified, and only a substantially weaker version of the third\.
[Ma et al\. \[54\]](https://arxiv.org/html/2608.28980#bib.bib23)fit a power\-law exponent of approximately 0\.4 for tabular in\-context learning with respect to both model parameter count and pretraining data size, considered separately, an exponent below 1 in the sense ofL\(N\)∝N−αL\(N\)\\propto N^\{\-\\alpha\}, so loss falls more slowly than scale grows, which is the only sense in which this review uses “sub\-linear\.” For time series,[Yao et al\. \[107\]](https://arxiv.org/html/2608.28980#bib.bib24)report that encoder\-only architectures scale more favorably than decoder\-only architectures, and that changes improving in\-distribution performance can reduce out\-of\-distribution scalability, a finding about which architecture scales best, not only whether scaling helps\. In graph learning,[Liu et al\. \[51\]](https://arxiv.org/html/2608.28980#bib.bib25)find that model depth, rather than parameter count, is the dominant determinant of scaling behavior, so the relevant scaled quantity is not even the same across studies\.[Tay et al\. \[83\]](https://arxiv.org/html/2608.28980#bib.bib22)show more broadly that the architecture with the best scaling behavior changes as scale increases, so an exponent measured at one scale range need not extrapolate to another\. None of these four studies measures the same quantity, over the same scale range, against the same baseline, which is why[Figure4](https://arxiv.org/html/2608.28980#S6.F4)presents them side by side rather than pooled into a single comparison; pooling them would not be a valid inference from what any one of them reports\.
Figure 4:Reported scaling exponents for structured\-data foundation models, with the Chinchilla language\-modeling exponent shown as an illustrative reference\[[83](https://arxiv.org/html/2608.28980#bib.bib22)\]\. The reported exponents are positive, indicating that scale improves performance, but remain sub\-linear\. Existing studies do not directly measure how the performance gap relative to a matched\-bias architecture changes with scale and structural complexity\.None of the four measures captures the quantity that this review’s central question actually requires\. The relevant quantity is notL\(N\)L\(N\), the loss of the language\-mediated model as scaleNNincreases, butΔ\(N\)=Lgeneral\(N\)−Lspecialized\(N\)\\Delta\(N\)=L\_\{\\text\{general\}\}\(N\)\-L\_\{\\text\{specialized\}\}\(N\), the performance gap to a matched\-bias architecture as a function of scale\. A positive, shrinkingL\(N\)L\(N\)is compatible withΔ\(N\)\\Delta\(N\)shrinking toward zero, staying constant, or even growing, depending on how the specialized baseline itself scales over the same range;[Tay et al\.](https://arxiv.org/html/2608.28980#bib.bib22)’s finding that the best\-scaling architecture changes with size is a direct reason to expectΔ\(N\)\\Delta\(N\)need not behave monotonically even where each individualL\(N\)L\(N\)does\.
The evidence located in this review supports a conservative reading that scaling reduces loss within a representational regime, at a rate that is real but sub\-linear in the specific senses reported above, and no study identified here measures whether that reduction eliminatesΔ\(N\)\\Delta\(N\)asymptotically, because no study measuresΔ\(N\)\\Delta\(N\)directly\. This claim rests on a dedicated search of the wider scaling\-law literature conducted specifically for this review, looking for any controlled comparison that scales a language\-mediated and a matched\-bias specialized architecture together and reports the gap between them as a function of scale\. No genuineΔ\(N\)\\Delta\(N\)\-measuring study was located in either direction \(neither one showing scaling closing a matched\-bias gap for structured data specifically, nor one showing scaling failing to close one under a controlled comparison\.
## 7Comparative Synthesis
[Section4](https://arxiv.org/html/2608.28980#S4)establishes the patterns that recur across nine modalities; this section asks what those recurrences imply, answers the motivating questions posed in[Section1](https://arxiv.org/html/2608.28980#S1), and states explicitly what the evidence does and does not establish\.
### 7\.1Why the cross\-modality pattern is mechanistic?
[Section4\.8](https://arxiv.org/html/2608.28980#S4.SS8)identifies the same corrective move in at least three research communities with no shared authorship, benchmarks, or citation practice\. The recurrence itself is initially only phenomenological\. Three communities arriving at similar remedies could reflect a shared underlying constraint, unrelated domain\-specific failures that happen to admit similar repairs, diffusion of ideas across neighboring fields, or simply the cases this review was able to document most clearly\.
The mechanistic interpretation is nevertheless the most economical explanation of the evidence\. First, the correction is not attributable to scale, pretraining data, or compute alone\. In each case, the decisive intervention is architectural or representational \(an attention scaffold, a vocabulary extension, or a reattached structural tool\) rather than simply a larger model or longer training\. Second, the effect is not explained by prompting or superficial engineering\. None of the three studies characterizes its corrective mechanism as a prompting intervention, while[Yoon and others](https://arxiv.org/html/2608.28980#bib.bib52)’s attempt to induce order\-invariant behavior through prompting and formatting alone, discussed in[Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2), failed until the positional mechanism itself was modified\. Third, the convergence is not an artifact of trivially weak baselines\. MoLFormer’s SMILES\-only representation already outperformed graph neural networks at 1\.1 billion molecules of pretraining scale\[[73](https://arxiv.org/html/2608.28980#bib.bib79)\]; the subsequent structural correction therefore addresses a specific failure \(structural blindness to rings and branching\) within an otherwise strong sequence model\. Finally, the fact that the successful systems are hybrid rather than purely linguistic does not weaken the argument\. Hybridization is precisely the recurring response that once a sequence or serialization representation is shown to omit information required by the task, the missing structural channel is restored\.
The mechanistic interpretation is further supported by[Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2), where two relevant constraints are established independently of the application domains\.[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)’s VC\-dimension ceiling concerns one\-layer softmax attention as a function class\. The limitation applies to tasks requiring the specific compositional operations considered in the proof, regardless of whether the underlying inputs represent molecules, proteins, tables, or other objects\. Likewise,[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)’s non\-permutation\-invariance result follows from positional encoding and causal ordering, properties that arise whenever a decoder\-only transformer is applied to set\- or graph\-structured inputs\. These results provide a domain\-agnostic mechanism that predicts why structurally different applications can encounter the same class of failure\. The independently observed convergence of corrective strategies is therefore consistent with a common architectural constraint and is more explained by it\.
The argument also does not exclude methodological diffusion between adjacent fields\. CARTE\[[45](https://arxiv.org/html/2608.28980#bib.bib41)\]and SaProt\[[77](https://arxiv.org/html/2608.28980#bib.bib87)\], for example, necessarily draw on ideas from broader graph\-neural\-network and structural\-learning literatures\. The narrower claim is that such diffusion is not sufficient to explain why the same type of correction is effective\. A domain\-agnostic architectural mechanism independently provides a reason for these communities to require an explicit structural channel\. Thus, the convergence documented in[Section4\.8](https://arxiv.org/html/2608.28980#S4.SS8)is best interpreted as evidence for a shared mechanism affecting at least two of the four failure modes in[Table6](https://arxiv.org/html/2608.28980#S7.T6), establishing a common cause for all four\.[Section7\.4](https://arxiv.org/html/2608.28980#S7.SS4)returns to this distinction\.
A dedicated search for a counterexample to this interpretation \(a pure, architecturally unmodified language model solving a permutation\-sensitive or compositional task at scale without an external structural channel or modification to its own attention or positional mechanism\) did not identify one\. Instead, it revealed a refinement already contained in the reviewed evidence\.[Egressy and Stühmer \[19\]](https://arxiv.org/html/2608.28980#bib.bib46)’s Set\-LLM, introduced above as evidence that standard decoder\-only transformers are not permutation\-invariant, also provides a direct architectural repair that it replaces the model’s positional encoding and causal attention mask with set\-specific alternatives, proves invariance from the resulting construction, and validates it empirically\. This is an internal modification of the language model rather than the addition of an external non\-linguistic encoder of the kind used by CARTE, SaProt, and the chemistry approach\. The mechanistic claim is therefore more precise than the statement that language\-mediated systems must always attach a separate structural component\. The evidence instead identifies at least two ways of restoring the missing inductive bias: modify the attention mechanism itself, as in[Kozachinskiy et al\.](https://arxiv.org/html/2608.28980#bib.bib14)’s Strassen attention and[Egressy and Stühmer](https://arxiv.org/html/2608.28980#bib.bib46)’s Set\-LLM, or reintroduce a separate structural pathway, as in CARTE, SaProt, and the chemistry correction\. What the reviewed evidence does not provide is a case in which scale, additional data, or prompting alone resolves the underlying structural limitation while the architecture and representation remain otherwise unchanged\.
### 7\.2The claim map
[Table5](https://arxiv.org/html/2608.28980#S7.T5)consolidates the evidence from[Sections4](https://arxiv.org/html/2608.28980#S4)and[6](https://arxiv.org/html/2608.28980#S6)into ten claims that recur, in different formulations, throughout the reviewed literature\. Each claim is paired with its strongest supporting and counter\-evidence\.The strongest claim is that specialized inductive biases remain necessary\. Its evidence is consistently supportive across the modalities reviewed, although the demonstrations differ in form: TabPFN achieves a 230×\\timesinference speedup over tuned boosting ensembles without using language\[[37](https://arxiv.org/html/2608.28980#bib.bib27)\]; PatchTST reaches state\-of\-the\-art long\-horizon forecasting using a language\-free patch\-based architecture\[[62](https://arxiv.org/html/2608.28980#bib.bib64)\]; Set Transformer explicitly encodes permutation\-invariant attention for set\-structured inputs\[[47](https://arxiv.org/html/2608.28980#bib.bib5)\]; and GraphBFF demonstrates predictable scaling in a billion\-parameter, non\-linguistic graph architecture\[[9](https://arxiv.org/html/2608.28980#bib.bib26)\]\. This is the claim most closely aligned with the theoretical distinction in[Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2)that specialized structure is not merely an implementation preference but can provide measurable computational or statistical advantages\.
A second tier consists of the claims that language provides a genuine interface and orchestration capability, and that predictive competitiveness is real in specific regimes\. These claims are well supported but conditional\. The evidence in[Section4](https://arxiv.org/html/2608.28980#S4)establishes concrete settings in which language\-mediated systems are competitive, including extreme few\-shot prediction, symbolic or discretized representations, and orchestration of specialized tools\. It does not establish that these advantages extend across data regimes, structural complexity, or matched computational budgets\.
A third tier contains claims about generalization to unseen synthetic structural tasks and statistical efficiency\. Here the evidence is genuinely contested\. Positive results exist, but so do controlled failures and evidence of information asymmetry or contamination\. Once those factors are accounted for \([Section4\.7](https://arxiv.org/html/2608.28980#S4.SS7)\), the evidence in this review weighs against broad interpretations of these results, but it does not justify treating every such claim as definitively disproven\.
The weakest claims are those asserting general replacement of specialized architectures\. No located result supporting such a broad claim survives direct scrutiny across the relevant dimensions of representation, computation, information availability, and structural invariance\. These claims are therefore better characterized as refuted within this review’s evidence base\.
Table 5:Claim map: the ten claims in the literature, with the strongest evidence and counter evidence located on each side \([Section7\.2](https://arxiv.org/html/2608.28980#S7.SS2)\)\. “Status” is this review’s synthesis judgment\.The minimum defensible thesis that explains this hierarchy is narrower than either extreme in the field’s public discourse\. Language\-mediated systems can reliably achieve behavioral competitiveness in identifiable regimes \(including extreme few\-shot prediction, discretizable spatial layouts, textually annotated knowledge graphs, and large\-scale single\-modality pretraining\), but that competitiveness does not generalize into representational, computational, or architectural replacement in the stricter senses distinguished in[Table1](https://arxiv.org/html/2608.28980#S3.T1)\. Notably, the direction of recent research already reflects this distinction more clearly that the strongest systems increasingly use language as an interface to, or orchestrator of, specialized structural computation rather than treating language itself as a universal substitute for that computation \([Section2\.4](https://arxiv.org/html/2608.28980#S2.SS4)\)\.
### 7\.3Answers to the central questions
The confidence labels below are qualitative judgments about this review’s evidence base\.*High*indicates that multiple independent studies, or a formal proof, directly address the specific claim\.*Moderate*indicates that the evidence directly supports a closely related question but does not fully establish the exact claim\.*Low*indicates that the conclusion depends primarily on a single unreplicated result or on the absence of counterevidence\. This convention is consistent with the qualitative interpretation used in this work \([Figures2](https://arxiv.org/html/2608.28980#S3.F2)and[3](https://arxiv.org/html/2608.28980#S3.F3)\)\.
Is language a general representation for structured data, or primarily an interface to systems that perform the underlying computation?Primarily an interface, with high confidence within the structural regimes directly tested\. Every domain in[Section4](https://arxiv.org/html/2608.28980#S4)can be described, queried, or orchestrated through language \([Section4\.6](https://arxiv.org/html/2608.28980#S4.SS6)\), but interface universality is not equivalent to representational sufficiency\. The failures documented in[Section4](https://arxiv.org/html/2608.28980#S4)concentrate precisely at this boundary that language\-mediated representations do not reliably preserve permutation invariance \([Section4\.2](https://arxiv.org/html/2608.28980#S4.SS2)\), temporal structure \([Section4\.3](https://arxiv.org/html/2608.28980#S4.SS3)\), or continuous spatial geometry \([Section4\.4](https://arxiv.org/html/2608.28980#S4.SS4)\) where these properties have been directly tested\.
Can an inductive bias that a specialized architecture enforces by construction instead be recovered by stating or demonstrating it in a prompt?Rarely, and not at the level of exact computational invariance directly measured in the reviewed cases\. Confidence is high for the cases actually tested, without extending the conclusion beyond them\.[Yoon and others](https://arxiv.org/html/2608.28980#bib.bib52)’s attempt to induce order\-invariant behavior through prompting and formatting alone failed on real listwise tasks and required modification of the positional\-encoding mechanism \([Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2)\)\.[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)provide a theoretical reason for expecting such failures in broader settings that their structural VC\-dimension ceiling for one\-layer attention cannot be removed by additional instructions or demonstrations because it follows from the computational form of the attention mechanism itself\.
If a language\-based model matches a specialized model’s accuracy, what exactly has been replaced?At most functional replacement in the narrow sense of[Table1](https://arxiv.org/html/2608.28980#S3.T1), with moderate\-to\-high confidence\. This conclusion is partly definitional but has important empirical consequences\. A predictive match establishes equivalence only under the data regime, compute budget, and information conditions of the comparison \([Section4\.7](https://arxiv.org/html/2608.28980#S4.SS7)\)\. Structural or computational replacement requires additional evidence, such as a direct invariance test or a component\-level ablation\. Where this review locates such tests \(including[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s ablation and[Egressy and Stühmer](https://arxiv.org/html/2608.28980#bib.bib46)’s architectural analysis\), the evidence does not establish structural replacement\.
Can existing evidence distinguish genuine structural generalization from benchmark familiarity, encoding artifacts, or external computation?Sometimes, with confidence that is high where a direct test exists and low where it does not\. The distinction is not uniformly available across the corpus\. Controlled ablation, independent contamination audits, and information\-matched comparisons can separate these explanations when they are performed:[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s ablation isolates the contribution of the language\-model component;[Gorla and Puduppully](https://arxiv.org/html/2608.28980#bib.bib36)and[Silvestri et al\.](https://arxiv.org/html/2608.28980#bib.bib37)provide evidence distinguishing genuine generalization from benchmark familiarity; and[Fatemi et al\.](https://arxiv.org/html/2608.28980#bib.bib43)directly tests the sensitivity of performance to the textual encoding of an otherwise fixed graph\. But[Section5](https://arxiv.org/html/2608.28980#S5)shows that such protocols remain exceptional\. For most reported results, the correct conclusion is therefore not that structural generalization has been demonstrated or ruled out, but that the evaluation was never designed to distinguish it from these alternatives\.
### 7\.4Limitations and failure modes
The conclusions of this review are bounded by two distinct classes of limitation\. The first concerns the evidence base itself\. Coverage is uneven across modalities: the four core modalities contain approximately fifteen to twenty\-five papers each, whereas the five adjacent modalities contain only five to seven each \([Section4\.5](https://arxiv.org/html/2608.28980#S4.SS5)\)\. More importantly,[Section5](https://arxiv.org/html/2608.28980#S5)shows that most benchmarks in this literature were not designed to test representation preservation or replacement directly\. Consequently, an absence of positive evidence has two possible interpretations: in some cases, a claim has been directly tested and failed; in others, the relevant test was simply never performed\.[Section7\.3](https://arxiv.org/html/2608.28980#S7.SS3)makes this distinction explicit\. This is a limitation of what the field has measured\.
The second limitation concerns the failure modes of language\-mediated learning themselves\.[Table6](https://arxiv.org/html/2608.28980#S7.T6)organizes these failures by mechanism rather than symptom because a mechanism\-first organization allows a more informative question: can the failure plausibly be resolved by scale, by changing the representation, or only by restoring specialized computation? The evidence does not support the same answer for all four mechanisms\.
For two mechanisms, the claim that scale alone does not resolve the failure is supported by formal results rather than by an empirical failure to find a sufficiently large model\. Non\-permutation\-invariance \([Section4\.2](https://arxiv.org/html/2608.28980#S4.SS2)\) follows from positional encoding and causal ordering, and[Kozachinskiy et al\.](https://arxiv.org/html/2608.28980#bib.bib14)’s ceiling holds for arbitrarily large width by construction\. The compositional\-reasoning ceiling has the same status: the limitation follows from the computational form of the attention mechanism rather than from insufficient parameter count\. These are therefore high\-confidence claims under the confidence convention of[Section7\.3](https://arxiv.org/html/2608.28980#S7.SS3)\.
The remaining two mechanisms require more cautious interpretation\. Benchmark contamination \([Sections4\.1](https://arxiv.org/html/2608.28980#S4.SS1)and[4\.3](https://arxiv.org/html/2608.28980#S4.SS3)\) is an evaluation\-mismatch problem: additional data, scale, or architectural changes cannot retrospectively make a contaminated benchmark measure clean generalization\. But this conclusion follows from the definition and consequences of contamination rather than from a theorem about attention or representation learning\. The vision\-to\-language information bottleneck \([Section4\.4](https://arxiv.org/html/2608.28980#S4.SS4)\) is weaker still\. It appears representation\-solvable in the specific sense that discretized symbolic spatial encodings substantially improve performance on the tested tasks, while[Alam et al\. \[3\]](https://arxiv.org/html/2608.28980#bib.bib75)find no resolution through the encoder and training\-objective variations they examine\. That negative result is informative, but it is a single controlled study rather than a proof that scaling or other architectural changes could never close the gap\.
[Section7\.1](https://arxiv.org/html/2608.28980#S7.SS1)further argues that two of these four mechanisms \(non\-permutation\-invariance and the compositional ceiling\) have a domain\-agnostic mechanistic explanation rooted in properties of attention\. This should not be generalized to all four mechanisms\. The available proofs establish why two failure modes recur; they do not constitute a formal cross\-domain meta\-analysis demonstrating that every failure in[Table6](https://arxiv.org/html/2608.28980#S7.T6)has the same cause\. The recurrence is therefore strong evidence for a shared mechanism in these two cases, while the remaining mechanisms should be treated with the weaker confidence warranted by their empirical evidence\.
The conceptual framework introduced by this review carries a corresponding limitation\. The four\-way distinction in[Section3\.3](https://arxiv.org/html/2608.28980#S3.SS3)is a synthesis proposed here\. Its components nevertheless correspond to distinctions already present in the theoretical literature reviewed in[Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2): describing a structure and computationally implementing it are different properties, and the contribution of this review is to apply that distinction systematically across structured non\-linguistic data modalities\. What the framework does not establish is a quantitative relationship between structural complexity and the resulting performance or efficiency gap\.[Section8](https://arxiv.org/html/2608.28980#S8)specifies the controlled experiment required to measure that relationship\.
Table 6:Recurring limitations of language\-mediated learning, organized by underlying mechanism rather than surface symptom \([Section7\.4](https://arxiv.org/html/2608.28980#S7.SS4)\)\. Across the four core modalities, the same mechanisms recur, suggesting that these limitations arise from the representational regime rather than from any single domain\.
## 8Open Problems and Future Work
### 8\.1The precise gap
Existing empirical work evaluates language\-based representations within individual structural modalities and individual inductive biases, almost always on real\-world benchmarks confounded by semantic priors, memorization, and contamination \([Section5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1)\)\. Learning theory has independently supplied several of the relevant ingredients \(sample\-complexity bounds for transformers\[[105](https://arxiv.org/html/2608.28980#bib.bib12)\], universal\-approximation costs of permutation\-invariant architectures\[[79](https://arxiv.org/html/2608.28980#bib.bib15)\], a VC\-dimension ceiling on compositional tasks\[[46](https://arxiv.org/html/2608.28980#bib.bib14)\], and a phase transition governing when in\-context learning generalizes structurally\[[26](https://arxiv.org/html/2608.28980#bib.bib18)\]\), but no study has connected these ingredients to a controlled, cross\-regime empirical measurement\. A dedicated search for such a study, conducted specifically for this work and did not locate one\. The resulting gap is therefore an untested empirical relationship\.
The missing experiment can be stated compactly, and is best understood as one cross\-section of a larger surface that this review has already partially identified\. LetLgeneralL\_\{\\text\{general\}\}denote the loss of a language\-mediated model on a task andLspecializedL\_\{\\text\{specialized\}\}the loss of an architecture designed to encode that task’s structure explicitly, with training and evaluation otherwise matched\.[Section6\.3](https://arxiv.org/html/2608.28980#S6.SS3)definesΔ\(N\)=Lgeneral\(N\)−Lspecialized\(N\)\\Delta\(N\)=L\_\{\\text\{general\}\}\(N\)\-L\_\{\\text\{specialized\}\}\(N\)as the performance gap as a function of model or data scaleNN, and shows that no existing study measures it\. The complementary quantity is the same gap as a function of structural complexityccat fixed scale:Δ\(c\)=Lgeneral\(c\)−Lspecialized\(c\)\\Delta\(c\)=L\_\{\\text\{general\}\}\(c\)\-L\_\{\\text\{specialized\}\}\(c\), with information content and sample budget held fixed while the source of structural bias is varied \(an architectural constraint, an explicit natural\-language instruction, or an in\-context demonstration\) across multiple structural regimes\. The evaluation should use synthetic tasks with known ground truth and the direct invariance\-violation metric of[Table4](https://arxiv.org/html/2608.28980#S6.T4)\. Together,Δ\(N\)\\Delta\(N\)andΔ\(c\)\\Delta\(c\)define a single unmeasured surface\. The contribution proposed here is therefore to specify theΔ\(c\)\\Delta\(c\)cross\-section precisely enough to make it experimentally testable\.
The gap is falsifiable\. A flat or non\-monotonicΔ\(c\)\\Delta\(c\)curve under the controls above would refute the central empirical hypothesis developed in[Section6](https://arxiv.org/html/2608.28980#S6)\. It is also experimentally tractable using the metric already defined in[Table4](https://arxiv.org/html/2608.28980#S6.T4)\. More importantly, it provides a common axis for several apparently conflicting results in[Section4](https://arxiv.org/html/2608.28980#S4): SimKGC and MoLFormer appear to support substitution, whereas[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s and[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)’s results appear to challenge it\. These results occupy different, previously unconnected points on the structural\-complexity axis\. No existing paper, however, measures that axis directly\.
### 8\.2Research questions and hypotheses
Each question below sets competing hypotheses against one another that make distinct, observable predictions\.
RQ1\.For a fixed synthetic function class, how does the sample efficiency of in\-context learning compare with that of a matched specialized architecture as structural complexity \(permutation\-group size, locality radius, interaction sparsity, or compositional depth\) increases?[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)prove the one\-layer case for standard softmax attention: none of three compositional tasks \(three\-way matching, function composition, binary\-relation composition\) is solvable, regardless of width or precision, whereas their proposed Strassen attention solves all three at one layer\. Their proof does not establish what standard, unmodified multi\-layer attention can do on the same tasks; their own discussion identifies this as future work\.H1 \(architectural ceiling\):the impossibility persists at any constant depth for standard attention, so the barrier reflects what softmax attention computes rather than the width of a single layer\.H2 \(depth suffices\):additional layers, without changing the attention mechanism itself, recover the missing compositional operations, making the one\-layer result a depth limitation rather than a limitation of the mechanism\. These hypotheses make different predictions as depth increases at fixed parameter count: H1 predicts no improvement on any of the three tasks regardless of depth; H2 predicts improvement attributable to depth before parameter count alone would explain it\.Falsification of H1:a many\-layer standard\-attention model closing the gap on any of the three tasks without a Strassen\-style architectural change\.Confound:an apparent depth effect caused by the additional parameters introduced by increasing depth controllable by matching parameter count across depths\.
RQ2\.When a structural invariance is stated explicitly in language, does the resulting output distribution become measurably closer to invariant under the direct metric of[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1)?H1 \(bias is inert\):the residual gap remains flat or increases with dimensionality, consistent with[Yoon and others](https://arxiv.org/html/2608.28980#bib.bib52)’s prompting\-only attempt before positional\-encoding surgery was required\.H2 \(bias is dilute but real\):instruction narrows the gap gradually, at a rate too small to have been detected in[Yoon and others](https://arxiv.org/html/2608.28980#bib.bib52)’s particular tasks but detectable in aggregate\.Critical experiment:measure the invariance metric directly across a dimensionality sweep rather than infer invariance from downstream task accuracy, which cannot distinguish H1 from H2 at the required resolution\.Falsification of H1:a measurable narrowing that remains stable across dimensionality\.
RQ3\.Do demonstrations close more of the gap than explicit instruction alone, and can their effect be distinguished from memorization?H1 \(in\-support generalization\):[Wakayama and Suzuki](https://arxiv.org/html/2608.28980#bib.bib17)’s exponential\-decay result predicts that demonstrations provide rapidly diminishing risk when the target structure lies within the support of the pretraining task mixture\.H2 \(out\-of\-support failure\):[Goddard et al\.](https://arxiv.org/html/2608.28980#bib.bib18)’s phase transition instead predicts a qualitative break once the target structure leaves that support, instead of a smooth continuation of H1’s decay\. In their linear\-function setting, a transformer pretrained on tasks spanning less than approximately120∘120^\{\\circ\}of task\-angle diversity fails on unseen tasks outside that span, whereas diversity beyond roughly120∘120^\{\\circ\}\(shifting to approximately135∘135^\{\\circ\}under added label noise\) yields a solution that generalizes across the full test range\. This is a measured threshold, not merely a qualitative claim of “sharpness\.”Critical experiment:vary the target task’s distance from the pretraining mixture directly in a setting with a matched structural\-complexity axis rather than[Goddard et al\.](https://arxiv.org/html/2608.28980#bib.bib18)’s task\-angle axis, and test whether the empirical curve follows H1’s smooth decay or instead exhibits a comparable transition\.Confound:apparent in\-support generalization that is actually retrieval of a memorized near\-duplicate, controllable through the contamination canary in[Section8\.3](https://arxiv.org/html/2608.28980#S8.SS3)\.
RQ4\.Which measurable task properties \(symmetry\-group order, minimum description length, or discretizability\) predict in advance whether the gap will be small or catastrophic? The discrete\-versus\-continuous split reported in[Section4\.4](https://arxiv.org/html/2608.28980#S4.SS4), where text\-symbol grids outperform pixel grids on identical tasks while both fail on continuous geometry, provides the strongest existing empirical indication that discretizability may be an important moderator\.[Tabaghi and Wang](https://arxiv.org/html/2608.28980#bib.bib15)’s2DN2DNparameter\-cost formula provides a natural, citable complexity variable against which this hypothesis can be tested\.H1is that discretizability alone predicts the gap’s magnitude\.H2is that discretizability is one of several independent moderators, including symmetry\-group order and minimum description length, that cannot be reduced to a single variable\.Critical experiment:hold discretizability fixed while varying symmetry\-group order and MDL independently, and test whether the gap changes\.
[Table7](https://arxiv.org/html/2608.28980#S8.T7)restates these questions as five independently addressable sub\-problems, each paired with the reason it remains open and a concrete experimental direction\.
Table 7:Open problems this review identifies for future directions \([Section8](https://arxiv.org/html/2608.28980#S8)\)\. The first row is the paper’s central gap statement; the remaining rows are narrower, independently addressable sub\-problems\.
### 8\.3Minimal experimental design
The design follows directly from the gap defined in[Section8\.1](https://arxiv.org/html/2608.28980#S8.SS1)and is intended to resolve RQ1–RQ4 through a common protocol\. For each of a small set of task families with known generating functions \(permutation invariance; a direct replication of all three of[Kozachinskiy et al\.](https://arxiv.org/html/2608.28980#bib.bib14)’s compositional impossibility tasks, three\-way matching, function composition, and binary\-relation composition, tested separately because their one\-layer proof treats them as distinct and their multi\-layer behavior remains unknown for all three; locality; sparse tabular interaction; and graph neighborhood aggregation\), five representation arms are compared:
1. 1\.a specialized architecture matched to the target bias;
2. 2\.naïve text serialization with no structural hint;
3. 3\.serialization plus an explicit statement of the relevant invariance;
4. 4\.serialization plus demonstrations only \(no explicit statement\), isolating whether demonstrations alone can convey the bias;
5. 5\.a lossless canonical serialization, providing the best\-case upper bound attainable if representation loss were the sole source of any observed gap\.
Within the taxonomy of[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1), arms 2–4 occupy Levels 1–4 and arm 1 occupies Level 8\. The design deliberately excludes Levels 5–7 because the question is whether the language\-mediated computation itself acquires the relevant bias, rather than whether an orchestrated tool or separate encoder can supply it, a question[Section4\.6](https://arxiv.org/html/2608.28980#S4.SS6)and[Table1](https://arxiv.org/html/2608.28980#S3.T1)’s orchestration row already answer affirmatively for a different reason\. Contamination is controlled by sampling fresh generating functions on every run, contrasting randomized\-symbol feature names with real\-sounding names to isolate world\-knowledge leakage, and holding out one structural regime never referenced during development as a contamination canary\. The primary dependent variable is the direct invariance\-violation metric defined in[Section6\.1](https://arxiv.org/html/2608.28980#S6.SS1), not accuracy\. Compute cost per prediction is reported alongside it so that any apparent advantage remains interpretable against[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s demonstration that accuracy parity can coexist with a1000×1000\\timescompute disadvantage\. Together, these measurements populate the dimensions of[Table1](https://arxiv.org/html/2608.28980#S3.T1)that end\-task accuracy alone cannot\. Functional replacement \(arm comparison at fixed accuracy\), computational replacement \(cost per prediction\), and structural replacement \(the direct invariance metric\) are each measured in this case\.
### 8\.4Concrete future directions
Beyond the primary experiment in[Section8\.3](https://arxiv.org/html/2608.28980#S8.SS3), four narrower and independently valuable directions follow directly from gaps identified in[Sections4](https://arxiv.org/html/2608.28980#S4)and[6](https://arxiv.org/html/2608.28980#S6)\.
1. 1\.Report structured\-data scaling laws asΔ\(N\)\\Delta\(N\)\-versus\-scale curves relative to a matched\-bias baseline instead of isolated loss\-versus\-scale curves\.[Section6\.3](https://arxiv.org/html/2608.28980#S6.SS3)shows that this connective analysis is absent from every scaling\-law paper located in the review, even though it requires no new training run and it can be obtained by re\-analyzing quantities that most of these papers already collected and reported\.
2. 2\.Test empirically whether depth alone, while keeping the attention mechanism otherwise standard, closes the gap on[Kozachinskiy et al\.](https://arxiv.org/html/2608.28980#bib.bib14)’s three compositional tasks \(three\-way matching, function composition, and binary\-relation composition\), treating them separately because their one\-layer proof treats them as distinct and their paper identifies the multi\-layer case as open for all three\. This is feasible as an empirical study on the small synthetic tasks already specified by their proof and requires modest compute, well before the harder theoretical question of proving a corresponding multi\-layer impossibility result, if H1 in RQ1 is correct, is resolved\. The experimental design in[Section8\.3](https://arxiv.org/html/2608.28980#S8.SS3)is constructed to run precisely this test\.
3. 3\.Perform an ablation on SimKGC analogous to[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)’s time\-series ablation by removing textual entity and relation descriptions while keeping the underlying triples fixed\. This would determine how much of SimKGC’s advantage over embedding\-based methods arises from information availability rather than genuine exploitation of structure, a distinction currently inferred instead of directly tested in[Section4\.2](https://arxiv.org/html/2608.28980#S4.SS2)\. This review searched specifically for a verified prior instance of this exact test and did not locate one meeting the evidentiary standard applied elsewhere in the review; the direction is therefore treated as fully open\. It is highly feasible because SimKGC’s codebase and the WN18RR/FB15k\-237 benchmarks are public and small, making the ablation a relatively inexpensive\.
4. 4\.Apply a[Tan et al\.](https://arxiv.org/html/2608.28980#bib.bib57)\-style component ablation to Level\-6 orchestration agents \([Section4\.6](https://arxiv.org/html/2608.28980#S4.SS6)\), decomposing end\-task success into the contribution attributable to the language model’s search and planning versus that of the specialized tool being orchestrated\. Current evaluations report primarily end\-task metrics and therefore cannot distinguish genuine orchestration capability from cases in which the tool performs the substantive computation largely independently of the orchestrator’s quality\.
## 9Conclusion
Across nine modalities and four methodological lenses, this review finds that behavioral competitiveness should not be conflated with replacement\. Language\-mediated systems can match specialized architectures on predictive performance, but where representation, computation, or architecture has been tested directly, such parity rarely extends beyond the narrower senses of replacement distinguished in[Table1](https://arxiv.org/html/2608.28980#S3.T1)\. Language can substitute for explicit inductive bias in identifiable regimes, including extreme few\-shot prediction, discretizable spatial layouts, textually annotated knowledge graphs, and settings where sufficiently large pretraining compensates for limited task\-specific structure\. Beyond these regimes, apparent success is often associated with contamination, format familiarity, or information supplied through channels that are not purely linguistic\. The central methodological lesson is therefore simple: accuracy, statistical efficiency, computational cost, and structural equivalence are different claims, and benchmark performance alone cannot establish the latter\.
The evidence also points to a recurring cross\-modal pattern\. When language\-only representations prove insufficient, successful systems frequently reintroduce non\-linguistic structure through tools, specialized tokens, or architectural scaffolds\. What appears to be the disappearance of specialization is therefore often its relocation\. This pattern recurs across independent research communities, although the available evidence does not establish it as a universal law\. Similarly, the literature provides clear evidence that language can serve as an effective interface for directing and composing specialized components, but this should be understood as a distinct form of integration rather than as architectural replacement\. Scaling is also real, but the relevant unresolved quantity is not whetherL\(N\)L\(N\)decreases with scale; it is whether the gap to a matched\-bias architecture,Δ\(N\)\\Delta\(N\), decreases and eventually closes \([Section6\.3](https://arxiv.org/html/2608.28980#S6.SS3)\)\.
The central open question is consequently more precise than whether language can replace specialized machine learning\. Existing theoretical results, including the compositional\-task VC\-dimension ceiling of[Kozachinskiy et al\. \[46\]](https://arxiv.org/html/2608.28980#bib.bib14)and the pretraining\-diversity phase transition of[Goddard et al\. \[26\]](https://arxiv.org/html/2608.28980#bib.bib18), provide important constraints in their respective settings, but neither directly predicts the performance gap as a function of scale or structural complexity \([Sections8\.1](https://arxiv.org/html/2608.28980#S8.SS1)and[6\.3](https://arxiv.org/html/2608.28980#S6.SS3)\)\. The next step is therefore to identify what predicts, quantitatively and in advance, where language\-mediated systems cease to substitute for specialized structure\.[Section8](https://arxiv.org/html/2608.28980#S8)outlines a falsifiable framework for answering this question and for moving the field from demonstrations of benchmark parity toward a theory of when, why, and at what structural cost such parity is achievable\.
## References
- \[1\]N\. Abhyankar, P\. Shojaee, and C\. K\. Reddy\(2025\)LLM\-FE: automated feature engineering for tabular data with LLMs as evolutionary optimizers\.Transactions on Machine Learning Research\.External Links:2503\.14434Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p5.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p4.1),[§4\.6](https://arxiv.org/html/2608.28980#S4.SS6.p1.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.6.2.1.1)\.
- \[2\]E\. Akyürek, D\. Schuurmans, J\. Andreas, T\. Ma, and D\. Zhou\(2023\)What learning algorithm is in\-context learning? investigations with linear models\.InInternational Conference on Learning Representations \(ICLR\),External Links:2211\.15661Cited by:[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1)\.
- \[3\]N\. Alam, L\. K\. Murali, S\. Bharadwaj, P\. Liu, T\. Chung, D\. Sharma, A\. A\., K\. Kiran, W\. Tam, and B\. K\. S\. Vegesna\(2026\)Spatial reasoning is not a free lunch: a controlled study on LLaVA\.External Links:2603\.12545Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p3.1),[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p4.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.18.2.1.1),[§7\.4](https://arxiv.org/html/2608.28980#S7.SS4.p4.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.5.4.1.1)\.
- \[4\]A\. F\. Ansari, L\. Stella, C\. Turkmen, X\. Zhang,et al\.\(2024\)Chronos: learning the language of time series\.External Links:2403\.07815Cited by:[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p3.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.11.2.1.1)\.
- \[5\]A\. Auer, P\. Podest, D\. Klotz, S\. Böck, G\. Klambauer, and S\. Hochreiter\(2025\)TiRex: zero\-shot forecasting across long and short horizons with enhanced in\-context learning\.Note:NeurIPS 2025External Links:2505\.23719Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p5.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p3.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.12.2.1.1)\.
- \[6\]W\. Azizian and A\. Hasan\(2025\)How does the pretraining distribution shape in\-context learning? a fundamental trade\-off\.External Links:2510\.01163Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p8.1)\.
- \[7\]M\. Balesni, T\. Korbak, and O\. Evans\(2024\)The two\-hop curse: LLMs trained on A→\\toB, B→\\toC fail to learn A→\\toC\.External Links:2411\.16353Cited by:[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.3.4.1.1)\.
- \[8\]P\. W\. Battaglia, J\. B\. Hamrick, V\. Bapst,et al\.\(2018\)Relational inductive biases, deep learning, and graph networks\.External Links:1806\.01261Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p2.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p2.1)\.
- \[9\]M\. Bechler\-Speicher, Y\. Gottlieb, A\. Isakov, D\. Abensur, A\. Tavory, D\. Haimovich, I\. Guy, and U\. Weinsberg\(2026\)Billion\-scale graph foundation models\.External Links:2602\.04768Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p5.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.7.2.1.1),[§7\.2](https://arxiv.org/html/2608.28980#S7.SS2.p1.1)\.
- \[10\]S\. Bordt, H\. Nori, V\. Rodrigues, B\. Nushi, and R\. Caruana\(2024\)Elephants never forget: memorization and learning of tabular data in large language models\.InConference on Language Modeling \(COLM\),External Links:2404\.06209Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p6.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p4.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p3.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1)\.
- \[11\]K\. Bougiatiotis, D\. Kelesis, and G\. Paliouras\(2026\)Improving molecular property prediction in small language models using graph\-based tools\.External Links:2607\.13115Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p1.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p5.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px1.p1.1)\.
- \[12\]M\. M\. Bronstein, J\. Bruna, T\. Cohen, and P\. Veličković\(2021\)Geometric deep learning: grids, groups, graphs, geodesics, and gauges\.External Links:2104\.13478Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p3.1)\.
- \[13\]R\. Chen, T\. Zhao, A\. Jaiswal, N\. Shah, and Z\. Wang\(2024\)LLaGA: large language and graph assistant\.InInternational Conference on Machine Learning \(ICML\),External Links:2402\.08170Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p2.1)\.
- \[14\]S\. Chiu, J\. Hong, and U\. Braga\-Neto\(2024\)DeepOSets: non\-autoregressive in\-context learning with permutation\-invariance inductive bias\.External Links:2410\.09298Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p6.1)\.
- \[15\]T\. Cohen and M\. Welling\(2016\)Group equivariant convolutional networks\.InInternational Conference on Machine Learning \(ICML\),External Links:1602\.07576Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p2.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p2.1)\.
- \[16\]T\. Dinh, Y\. Zeng, R\. Zhang, Z\. Lin, M\. Gira, S\. Rajput, J\. Sohn, D\. Papailiopoulos, and K\. Lee\(2022\)LIFT: language\-interfaced fine\-tuning for non\-language machine learning tasks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2206\.06565Cited by:[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1)\.
- \[17\]Q\. Dong, L\. Li, D\. Dai, C\. Zheng, J\. Ma, R\. Li, H\. Xia, J\. Xu, Z\. Wu, T\. Liu, B\. Chang, X\. Sun, L\. Li, and Z\. Sui\(2024\)A survey on in\-context learning\.InEmpirical Methods in Natural Language Processing \(EMNLP\),External Links:2301\.00234Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p3.1)\.
- \[18\]N\. Dziri, X\. Lu, M\. Sclar,et al\.\(2023\)Faith and fate: limits of transformers on compositionality\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Spotlight,External Links:2305\.18654Cited by:[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.9.3.1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.3.4.1.1)\.
- \[19\]B\. Egressy and J\. Stühmer\(2025\)Set\-LLM: a permutation\-invariant LLM\.External Links:2505\.15433Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p3.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.6.3.1.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p1.1),[§5](https://arxiv.org/html/2608.28980#S5.p2.1),[§6\.1](https://arxiv.org/html/2608.28980#S6.SS1.p2.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p3.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p5.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p4.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.3.3.1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.2.4.1.1)\.
- \[20\]X\. Fang, W\. Xu, F\. A\. Tan, J\. Zhang, Z\. Hu, Y\. Qi, S\. Nickleach, D\. Socolinsky, S\. Sengamedu, and C\. Faloutsos\(2024\)Large language models \(llms\) on tabular data: prediction, generation, and understanding – a survey\.Note:Transactions on Machine Learning ResearchExternal Links:2402\.17944Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[21\]B\. Fatemi, J\. Halcrow, and B\. Perozzi\(2024\)Talk like a graph: encoding graphs for large language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:2310\.04560Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p3.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p4.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p1.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.8.2.1.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p5.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.4.3.1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.2.4.1.1),[35](https://arxiv.org/html/2608.28980#bib.bib44)\.
- \[22\]X\. Fu, Y\. Hu, B\. Li, Y\. Feng, H\. Wang, X\. Lin, D\. Roth, N\. A\. Smith, W\. Ma, and R\. Krishna\(2024\)BLINK: multimodal large language models can see but not perceive\.InEuropean Conference on Computer Vision \(ECCV\),External Links:2404\.12390Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.16.4.1.1)\.
- \[23\]J\. Gardner, J\. C\. Perdomo, and L\. Schmidt\(2024\)Large scale transfer learning for tabular data via language modeling\.Note:NeurIPS 2024; central claim independently overturned by[Gorla and Puduppully \[28\]](https://arxiv.org/html/2608.28980#bib.bib36)External Links:2406\.12031Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p6.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.5.2.1.1)\.
- \[24\]S\. Garg, D\. Tsipras, P\. Liang, and G\. Valiant\(2022\)What can transformers learn in\-context? a case study of simple function classes\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.9.2.1.1)\.
- \[25\]M\. Garnelo and W\. M\. Czarnecki\(2026\)Why large language models fail at tabular prediction\.Note:Unverified: no independent replication found at time of reviewExternal Links:2608\.02412Cited by:[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p3.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.10.3.1.1)\.
- \[26\]B\. Goddard, K\. Smith, V\. Ngampruetikorn, and D\. J\. Schwab\(2025\)When can in\-context learning generalize out of task distribution?\.InInternational Conference on Machine Learning \(ICML\),External Links:2506\.05574Cited by:[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p8.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p1.1),[§8\.2](https://arxiv.org/html/2608.28980#S8.SS2.p4.1),[§9](https://arxiv.org/html/2608.28980#S9.p3.1)\.
- \[27\]S\. Golchin and M\. Surdeanu\(2024\)Time travel in LLMs: tracing data contamination in large language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:2308\.08493Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p4.1),[Table 4](https://arxiv.org/html/2608.28980#S6.T4.5.7.3.1.1)\.
- \[28\]A\. Gorla and R\. Puduppully\(2026\)The illusion of generalization in tabular language models\.InInternational Conference on Machine Learning \(ICML\),External Links:2602\.04031Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p6.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.5.4.1.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p1.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p5.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.2.3.1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.4.4.1.1),[23](https://arxiv.org/html/2608.28980#bib.bib34)\.
- \[29\]L\. Grinsztajn, K\. Flöge, O\. Key,et al\.\(2025\)TabPFN\-2\.5: advancing the state of the art in tabular foundation models\.Note:Prior LabsExternal Links:2511\.08667Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p5.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.2.2.1.1)\.
- \[30\]L\. Grinsztajn, E\. Oyallon, and G\. Varoquaux\(2022\)Why do tree\-based models still outperform deep learning on typical tabular data?\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,External Links:2207\.08815Cited by:[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.10.3.1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.2.3.1.1)\.
- \[31\]N\. Gruver, M\. Finzi, S\. Qiu, and A\. G\. Wilson\(2023\)Large language models are zero\-shot time series forecasters\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:LLMTimeExternal Links:2310\.07820Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p3.1)\.
- \[32\]A\. Gu, B\. Rozière, H\. Leather, A\. Solar\-Lezama, G\. Synnaeve, and S\. I\. Wang\(2024\)CRUXEval: a benchmark for code reasoning, understanding and execution\.InInternational Conference on Machine Learning \(ICML\),External Links:2401\.03065Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p6.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px2.p1.1)\.
- \[33\]Y\. Han, Z\. Wan, L\. Chen, K\. Yu, and X\. Chen\(2025\)From generalist to specialist: a survey of large language models for chemistry\.InInternational Conference on Computational Linguistics \(COLING\),External Links:2412\.19994Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px1.p1.1)\.
- \[34\]S\. Hegselmann, A\. Buendia, H\. Lang, M\. Agrawal, X\. Jiang, and D\. Sontag\(2023\)TabLLM: few\-shot classification of tabular data with large language models\.InInternational Conference on Artificial Intelligence and Statistics \(AISTATS\),External Links:2210\.10723Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p3.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p4.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.4.2.1.1)\.
- \[35\]D\. Herbst, L\. Karbevska, D\. Kumar, A\. Ahuja, F\. G\. Nasrabadi, and F\. Frasca\(2025\)Lost in serialization: invariance and generalization of LLM graph reasoners\.Note:Verified via arXiv abstract page \(title, authors, and finding cross\-checked directly; not a full\-text read of the PDF\)\. Finds larger non\-fine\-tuned models more robust to serialization changes than fine\-tuned variants, and that fine\-tuning trades reduced node\-relabeling sensitivity for increased sensitivity to structure/format changes – extends[Fatemi et al\. \[21\]](https://arxiv.org/html/2608.28980#bib.bib43)rather than merely repeating itExternal Links:2511\.10234Cited by:[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p1.1)\.
- \[36\]J\. Hill, B\. Eyre, and E\. Creager\(2025\)Transformers don’t in\-context learn least squares regression\.Note:ICML 2025 Workshop on Reliable and Responsible Foundation Models – workshop tier, not full peer review\. Verified via direct fetch of the arXiv abstract\. Empirical counterevidence to the claim that transformers implement OLS/gradient\-descent\-like algorithms in\-context: shows ICL transformers fail to generalize under prompt\-distribution shift and instead appear to memorize training\-distribution spectral signatures, located during a dedicated adversarial search for counterevidence to the positive in\-context\-learning results cited in[Section6\.2](https://arxiv.org/html/2608.28980#S6.SS2)External Links:2507\.09440Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p4.1)\.
- \[37\]N\. Hollmann, S\. Müller, K\. Eggensperger, and F\. Hutter\(2023\)TabPFN: a transformer that solves small tabular classification problems in a second\.InInternational Conference on Learning Representations \(ICLR\),External Links:2207\.01848Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p8.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p1.1),[§7\.2](https://arxiv.org/html/2608.28980#S7.SS2.p1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.8.2.1.1)\.
- \[38\]N\. Hollmann, S\. Müller, L\. Purucker, A\. Krishnakumar, M\. Körfer, S\. B\. Hoo, R\. T\. Schirrmeister, and F\. Hutter\(2025\)Accurate predictions on small data with a tabular foundation model\.Nature637,pp\. 319–326\.Cited by:[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p1.1)\.
- \[39\]C\. Huertas\(2024\)Gradient boosting trees and large language models for tabular data few\-shot learning\.Note:FedCSIS 2024 Data Mining CompetitionExternal Links:2411\.04324Cited by:[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1),[§4\.7](https://arxiv.org/html/2608.28980#S4.SS7.p1.1)\.
- \[40\]A\. Jaegle, S\. Borgeaud, J\. Alayrac,et al\.\(2021\)Perceiver IO: a general architecture for structured inputs & outputs\.External Links:2107\.14795Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p2.1)\.
- \[41\]A\. Jafari, J\. Fox, G\. C\. Fox, M\. Marathe, and A\. Adiga\(2026\)Understanding key features of time series foundation models from epidemic forecasting\.External Links:2606\.19560Cited by:[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p3.1)\.
- \[42\]B\. Jin, G\. Liu, C\. Han, M\. Jiang, H\. Ji, and J\. Han\(2023\)Large language models on graphs: a comprehensive survey\.External Links:2312\.02783Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[43\]M\. Jin, S\. Wang, L\. Ma,et al\.\(2024\)Time\-LLM: time series forecasting by reprogramming large language models\.InInternational Conference on Learning Representations \(ICLR\),Note:Falsified by ablation in[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)External Links:2310\.01728Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p4.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.13.2.1.1)\.
- \[44\]J\. Kim, T\. Nakamaki, and T\. Suzuki\(2024\)Transformers are minimax optimal nonparametric in\-context learners\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2408\.12186Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p3.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p4.1)\.
- \[45\]M\. J\. Kim, L\. Grinsztajn, and G\. Varoquaux\(2024\)CARTE: pretraining and transfer for tabular learning\.InInternational Conference on Machine Learning \(ICML\),External Links:2402\.16785Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p1.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p5.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p4.1)\.
- \[46\]A\. Kozachinskiy, F\. Urrutia, H\. Jimenez, T\. Steifer, G\. Pizarro, I\. Fuentes, D\. Meza, F\. Calderon, and C\. Rojas\(2025\)Strassen attention, split VC dimension and compositionality in transformers\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2501\.19215Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p4.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.6.3.1.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p5.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p3.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p5.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p3.1),[§7\.4](https://arxiv.org/html/2608.28980#S7.SS4.p3.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.3.4.1.1),[item 2](https://arxiv.org/html/2608.28980#S8.I2.i2.p1.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p1.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p3.1),[§8\.2](https://arxiv.org/html/2608.28980#S8.SS2.p2.1),[§8\.3](https://arxiv.org/html/2608.28980#S8.SS3.p1.1),[§9](https://arxiv.org/html/2608.28980#S9.p3.1)\.
- \[47\]J\. Lee, Y\. Lee, J\. Kim, A\. Kosiorek, S\. Choi, and Y\. W\. Teh\(2019\)Set transformer: a framework for attention\-based permutation\-invariant neural networks\.InInternational Conference on Machine Learning \(ICML\),External Links:1810\.00825Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p2.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p2.1),[§7\.2](https://arxiv.org/html/2608.28980#S7.SS2.p1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.8.2.1.1)\.
- \[48\]Y\. Lee, B\. Ko, H\. Kim, Y\. Hwang, and H\. Choi\(2024\)Intriguing properties of large language and vision models\.InarXiv preprint,External Links:2410\.04751Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p2.1)\.
- \[49\]X\. Li, Y\. Jiao, J\. Huang, Y\. Wei, and X\. Chen\(2025\)Transformers meet in\-context learning: a universal approximation theory\.External Links:2506\.05200Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p8.1)\.
- \[50\]J\. Liu, C\. Yang, Z\. Lu, J\. Chen, Y\. Li, M\. Zhang, T\. Bai, Y\. Fang, L\. Sun, P\. S\. Yu, and C\. Shi\(2025\)Graph foundation models: concepts, opportunities and challenges\.IEEE Transactions on Pattern Analysis and Machine Intelligence\.Note:Cited as reference \[14\] within[Sun et al\. \[78\]](https://arxiv.org/html/2608.28980#bib.bib101); located via citation\-chaining from that survey’s primary text rather than an independent search passCited by:[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p2.1)\.
- \[51\]J\. Liu, H\. Mao, Z\. Chen, T\. Zhao, N\. Shah, and J\. Tang\(2024\)Towards neural scaling laws on graphs\.External Links:2402\.02054Cited by:[§6\.3](https://arxiv.org/html/2608.28980#S6.SS3.p2.1)\.
- \[52\]P\. Liu, H\. Guo, T\. Dai,et al\.\(2024\)CALF: aligning LLMs for time series forecasting via cross\-modal fine\-tuning\.Note:Falsified by ablation in[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)External Links:2403\.07300Cited by:[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p2.1)\.
- \[53\]D\. X\. Long, H\. N\. Ngoc, T\. Sim, H\. Dao, S\. Joty, K\. Kawaguchi, N\. F\. Chen, and M\. Kan\(2025\)LLMs are biased towards output formats\! systematically evaluating and mitigating output format bias of LLMs\.InConference of the North American Chapter of the Association for Computational Linguistics \(NAACL\),pp\. 299–330\.Note:Verified via ACL Anthology \(2025\.naacl\-long\.15\); located during a dedicated search of the wider evaluation\-methodology literature, outside this review’s core 195\-paper corpus, for evidence of benchmark bias running against rather than toward language\-mediated methods – see[Section5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1)External Links:2408\.08656Cited by:[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p2.1)\.
- \[54\]J\. Ma, V\. Thomas, R\. Hosseinzadeh, H\. Kamkari, A\. Labach, J\. C\. Cresswell, K\. Golestan, G\. Yu, A\. L\. Caterini, and M\. Volkovs\(2025\)TabDPT: scaling tabular foundation models on real data\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2410\.18164Cited by:[§6\.3](https://arxiv.org/html/2608.28980#S6.SS3.p2.1)\.
- \[55\]K\. Mahowald, A\. A\. Ivanova, I\. A\. Blank, N\. Kanwisher, J\. B\. Tenenbaum, and E\. Fedorenko\(2024\)Dissociating language and thought in large language models\.Trends in Cognitive Sciences\.Note:Feature ReviewExternal Links:2301\.06627Cited by:[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p2.1)\.
- \[56\]H\. Mao, Z\. Chen, W\. Tang, J\. Zhao, Y\. Ma, T\. Zhao, N\. Shah, M\. Galkin, and J\. Tang\(2024\)Position: graph foundation models are already here\.InInternational Conference on Machine Learning \(ICML\), Position Paper Track,External Links:2402\.02216Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p5.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p3.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p4.1)\.
- \[57\]H\. Mao, G\. Liu, Y\. Ma, R\. Wang, K\. Johnson, and J\. Tang\(2025\)A survey to recent progress towards understanding in\-context learning\.InFindings of the Association for Computational Linguistics: NAACL,External Links:2402\.02212Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p3.1)\.
- \[58\]M\. Meyer, S\. Kaltenpoth, K\. Zalipski, and O\. Müller\(2025\)Rethinking evaluation in the era of time series foundation models: \(un\)known information leakage challenges\.External Links:2510\.13654Cited by:[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p3.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.4.4.1.1)\.
- \[59\]G\. Mialon, R\. Dessì, M\. Lomeli, C\. Nalmpantis, R\. Pasunuru, R\. Raileanu, B\. Rozière, T\. Schick, J\. Dwivedi\-Yu, A\. Celikyilmaz, É\. Grave, Y\. LeCun, and T\. Scialom\(2023\)Augmented language models: a survey\.Transactions on Machine Learning Research\.Note:Verified via dblp \(journals/tmlr/MialonDLNPRRSDC23\); distinguishes reasoning\-augmentation from tool/action\-augmentation along a different organizing axis than[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1)’s location\-of\-computation axis – see[Section3\.1](https://arxiv.org/html/2608.28980#S3.SS1)External Links:2302\.07842Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1)\.
- \[60\]S\. Mittal, Y\. Li, I\. Yen, D\. Guetta, and H\. Namkoong\(2025\)Architectural and inferential inductive biases for exchangeable sequence modeling\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2503\.01215Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p4.1)\.
- \[61\]J\. Nam, J\. Yoon, J\. Chen, J\. Shin, S\. Ö\. Arık, and T\. Pfister\(2025\)MLE\-STAR: machine learning engineering agent via search and targeted refinement\.Note:GoogleExternal Links:2506\.15692Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p5.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p4.1),[§4\.6](https://arxiv.org/html/2608.28980#S4.SS6.p1.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p3.1)\.
- \[62\]Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. Kalagnanam\(2023\)A time series is worth 64 words: long\-term forecasting with transformers\.InInternational Conference on Learning Representations \(ICLR\),Note:PatchTSTExternal Links:2211\.14730Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p8.1),[§7\.2](https://arxiv.org/html/2608.28980#S7.SS2.p1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.8.2.1.1)\.
- \[63\]P\. Notin, A\. W\. Kollasch, D\. Ritter, L\. van Niekerk, S\. Paul, H\. Spinner, N\. Rollins, A\. Shaw, R\. Orenbuch, R\. Weitzman, J\. Frazer, M\. Dias, D\. Franceschi, Y\. Gal, and D\. S\. Marks\(2023\)ProteinGym: large\-scale benchmarks for protein fitness prediction and design\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,Note:Verified via the NeurIPS proceedings record;[Su et al\. \[77\]](https://arxiv.org/html/2608.28980#bib.bib87)adopts this benchmark, alongside a ten\-task suite including Thermostability, Metal Ion Binding, DeepLoc, EC/GO annotation, and HumanPPI, for downstream evaluationCited by:[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.13.1.1)\.
- \[64\]C\. Olsson, N\. Elhage, N\. Nanda,et al\.\(2022\)In\-context learning and induction heads\.Note:Transformer Circuits ThreadExternal Links:2209\.11895Cited by:[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1)\.
- \[65\]M\. J\. Page, J\. E\. McKenzie, P\. M\. Bossuyt, I\. Boutron, T\. C\. Hoffmann, C\. D\. Mulrow, L\. Shamseer, J\. M\. Tetzlaff, E\. A\. Akl, S\. E\. Brennan, R\. Chou, J\. Glanville, J\. M\. Grimshaw, A\. Hróbjartsson, M\. M\. Lalu, T\. Li, E\. W\. Loder, E\. Mayo\-Wilson, S\. McDonald, L\. A\. McGuinness, L\. A\. Stewart, J\. Thomas, A\. C\. Tricco, V\. A\. Welch, P\. Whiting, and D\. Moher\(2021\)The PRISMA 2020 statement: an updated guideline for reporting systematic reviews\.BMJ372,pp\. n71\.External Links:[Document](https://dx.doi.org/10.1136/bmj.n71)Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p8.1),[§2\.1](https://arxiv.org/html/2608.28980#S2.SS1.p1.1)\.
- \[66\]Z\. Peng, J\. Jiang, H\. Gu, L\. Fan, and Y\. Yang\(2026\)GraphInfer\-Bench: benchmarking LLM’s inference capability on graphs\.External Links:2606\.11562Cited by:[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p2.1)\.
- \[67\]L\. Purucker, A\. Tschalzev, N\. Erickson,et al\.\(2026\)Beyond IID: how general are tabular foundation models, really?\.External Links:2606\.30410Cited by:[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p3.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.2.4.1.1),[§6\.1](https://arxiv.org/html/2608.28980#S6.SS1.p2.1)\.
- \[68\]J\. Qu, D\. Holzmüller, G\. Varoquaux, and M\. Le Morvan\(2026\)TabICLv2: a better, faster, scalable, and open tabular foundation model\.External Links:2602\.11139Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p5.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.3.2.1.1)\.
- \[69\]P\. Quinlan, J\. Levasseur, Q\. Li, and X\. Zhu\(2026\)Chronicle: a multimodal foundation model for joint language and time series understanding\.External Links:2605\.20268Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p8.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p5.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p4.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.14.2.1.1)\.
- \[70\]P\. Rahmanzadehgervi, L\. Bolton, M\. R\. Taesiri, and A\. T\. Nguyen\(2024\)Vision language models are blind: failing to translate detailed visual features into words\.InAsian Conference on Computer Vision \(ACCV\),External Links:2407\.06581Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.16.4.1.1)\.
- \[71\]S\. Reed, K\. Zolna, E\. Parisotto,et al\.\(2022\)A generalist agent\.Transactions on Machine Learning Research\.Note:GatoExternal Links:2205\.06175Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p2.1)\.
- \[72\]X\. Ren, J\. Tang, D\. Yin, N\. Chawla, and C\. Huang\(2024\)A survey of large language models for graphs\.External Links:2405\.08011Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[73\]J\. Ross, B\. Belgodere, V\. Chenthamarakshan, I\. Padhi, Y\. Mroueh, and P\. Das\(2022\)Large\-scale chemical language representations capture molecular structure and properties\.Nature Machine Intelligence4\(12\),pp\. 1256–1264\.Note:MoLFormerExternal Links:2106\.09553Cited by:[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.12.4.1.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p2.1),[101](https://arxiv.org/html/2608.28980#bib.bib80)\.
- \[74\]R\. Saxena, A\. Suglia, and P\. Minervini\(2026\)VLM\-RobustBench: a comprehensive benchmark for robustness of vision\-language models\.External Links:2603\.06148Cited by:[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.5.3.1.1)\.
- \[75\]M\. Silvestri, F\. Veglianti, F\. Giorgi, F\. Silvestri, and G\. Tolomei\(2026\)When large language models know the table: a framework for assessing data contamination in tabular datasets\.External Links:2510\.20351Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p6.1),[§4\.1](https://arxiv.org/html/2608.28980#S4.SS1.p2.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px1.p1.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p5.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.4.4.1.1)\.
- \[76\]X\. Songet al\.\(2024\)OmniPred: language models as universal regressors\.Transactions on Machine Learning Research\.Note:Google DeepMindExternal Links:2402\.14547Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p3.1)\.
- \[77\]J\. Su C\. Hanet al\.\(2024\)SaProt: protein language modeling with structure\-aware vocabulary\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p6.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p5.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px4.p1.1),[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.13.4.1.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p4.1),[63](https://arxiv.org/html/2608.28980#bib.bib89)\.
- \[78\]Q\. Sun, H\. Yuan, Y\. Huang, Z\. Zhang, X\. Fu, R\. Wang, H\. Zhou, J\. Wu, J\. Li, and P\. S\. Yu\(2026\)A survey on foundation models for structured data: tabular, time series, and graphs\.Note:Preprints\.orgdoi:10\.20944/preprints202605\.1532\.v2\. Posted 28 May 2026; explicitly labeled “Not peer\-reviewed version” by the platform \(Preprints\.org is an MDPI\-affiliated preprint server, not IEEE TKDE – an earlier draft of this manuscript misattributed the venue based on an ambiguous secondary source and has been corrected after obtaining and reading the primary PDF directly; see[Section2\.2](https://arxiv.org/html/2608.28980#S2.SS2)\)\. Full text read in full for this review\.Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1),[50](https://arxiv.org/html/2608.28980#bib.bib102)\.
- \[79\]P\. Tabaghi and Y\. Wang\(2024\)Universal representation of permutation\-invariant functions on vectors and tensors\.InInternational Conference on Machine Learning \(ICML\),External Links:2310\.13829Cited by:[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p6.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p1.1),[§8\.2](https://arxiv.org/html/2608.28980#S8.SS2.p5.1)\.
- \[80\]J\. Tan, Z\. Zhang, Y\. Guo, J\. Liu, Y\. Xing, M\. Zhang, C\. Yang, and C\. Shi\(2026\)GABench: a comprehensive benchmark for evaluating LLM agents on graph analysis tasks\.External Links:2608\.01684Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p6.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p2.1)\.
- \[81\]M\. Tan, M\. A\. Merrill, V\. Gupta, T\. Althoff, and T\. Hartvigsen\(2024\)Are language models actually useful for time series forecasting?\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Spotlight,External Links:2406\.16964Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§1](https://arxiv.org/html/2608.28980#S1.p6.1),[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p4.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.4.3.1.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p2.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p4.1),[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p5.1),[§4\.6](https://arxiv.org/html/2608.28980#S4.SS6.p2.1),[§4\.7](https://arxiv.org/html/2608.28980#S4.SS7.p1.1),[§4\.8](https://arxiv.org/html/2608.28980#S4.SS8.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.13.4.1.1),[§5](https://arxiv.org/html/2608.28980#S5.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2608.28980#S5.p2.1),[§6\.1](https://arxiv.org/html/2608.28980#S6.SS1.p2.1),[Table 4](https://arxiv.org/html/2608.28980#S6.T4.5.4.3.1.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p4.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p5.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.11.3.1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.5.3.1.1),[item 3](https://arxiv.org/html/2608.28980#S8.I2.i3.p1.1),[item 4](https://arxiv.org/html/2608.28980#S8.I2.i4.p1.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p3.1),[§8\.3](https://arxiv.org/html/2608.28980#S8.SS3.p3.1),[Table 7](https://arxiv.org/html/2608.28980#S8.T7.5.6.3.1.1),[43](https://arxiv.org/html/2608.28980#bib.bib61),[52](https://arxiv.org/html/2608.28980#bib.bib63),[114](https://arxiv.org/html/2608.28980#bib.bib62)\.
- \[82\]J\. Tang, Y\. Yang, W\. Wei, L\. Shi, L\. Su, S\. Cheng, D\. Yin, and C\. Huang\(2024\)GraphGPT: graph instruction tuning for large language models\.InInternational ACM SIGIR Conference on Research and Development in Information Retrieval,External Links:2310\.13023Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p2.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.9.2.1.1)\.
- \[83\]Y\. Tay M\. Dehghaniet al\.\(2023\)Scaling laws vs model architectures: how does inductive bias influence scaling?\.InFindings of the Association for Computational Linguistics: EMNLP,External Links:2207\.10551Cited by:[Figure 4](https://arxiv.org/html/2608.28980#S6.F4),[§6\.3](https://arxiv.org/html/2608.28980#S6.SS3.p2.1),[§6\.3](https://arxiv.org/html/2608.28980#S6.SS3.p3.1)\.
- \[84\]H\. Thakur, A\. Kamath, A\. Muthyala, D\. Sanmukhani, S\. Mukund, and J\. Katukuri\(2026\)Towards reliable ML feature engineering via planning in constrained\-topology of LLM agents\.External Links:2601\.10820Cited by:[§4\.6](https://arxiv.org/html/2608.28980#S4.SS6.p1.1)\.
- \[85\]V\. Thengane, X\. Zhu, S\. Bouzerdoum, S\. L\. Phung, and Y\. Li\(2025\)Foundational models for 3d point clouds: a survey and outlook\.External Links:2501\.18594Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[86\]Y\. Thushalika, S\. Kishanthan, and A\. Hevapathige\(2026\)Detecting differences is not understanding structure: large language models fail at graph isomorphism\.Note:Unverified: very recent preprint, not yet independently replicatedExternal Links:2606\.09484Cited by:[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p1.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.3.3.1.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.2.4.1.1)\.
- \[87\]S\. Tong, Z\. Liu, Y\. Zhai, Y\. Ma, Y\. LeCun, and S\. Xie\(2024\)Eyes wide shut? exploring the visual shortcomings of multimodal LLMs\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Note:MMVPExternal Links:2401\.06209Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1)\.
- \[88\]K\. Vafa, J\. Y\. Chang, A\. Rambachan, and S\. Mullainathan\(2025\)What has a foundation model found? using inductive bias to probe for world models\.InInternational Conference on Machine Learning \(ICML\),External Links:2507\.06952Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p1.1)\.
- \[89\]J\. von Oswald, E\. Niklasson, E\. Randazzo, J\. Sacramento, A\. Mordvintsev, A\. Zhmoginov, and M\. Vladymyrov\(2023\)Transformers learn in\-context by gradient descent\.InInternational Conference on Machine Learning \(ICML\),External Links:2212\.07677Cited by:[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1)\.
- \[90\]T\. Wakayama and T\. Suzuki\(2025\)In\-context learning is provably bayesian inference: a generalization theory for meta\-learning\.External Links:2510\.10981Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p3.1),[§8\.2](https://arxiv.org/html/2608.28980#S8.SS2.p4.1)\.
- \[91\]D\. Wanget al\.\(2025\)S\-PLM: structure\-aware protein language model via contrastive learning between sequence and structure\.Advanced Science\.Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p6.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p5.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px4.p1.1)\.
- \[92\]H\. Wang, S\. Feng, T\. He, Z\. Tan, X\. Han, and Y\. Tsvetkov\(2023\)Can language models solve graph problems in natural language?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2305\.10037Cited by:[§2\.4](https://arxiv.org/html/2608.28980#S2.SS4.p3.1)\.
- \[93\]J\. Wang, Y\. Ming, Z\. Shi, V\. Vineet, X\. Wang, Y\. Li, and N\. Joshi\(2024\)Is a picture worth a thousand words? delving into spatial reasoning for vision language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:SpatialEvalExternal Links:2406\.14852Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p4.1)\.
- \[94\]L\. Wang, W\. Zhao, Z\. Wei, and J\. Liu\(2022\)SimKGC: simple contrastive knowledge graph completion with pre\-trained language models\.External Links:2203\.02167Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p5.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p3.1),[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p3.1),[Table 2](https://arxiv.org/html/2608.28980#S4.T2.5.10.2.1.1)\.
- \[95\]R\. Wang, Z\. Wang, and J\. Sun\(2023\)UniPredict: large language models are universal tabular classifiers\.External Links:2310\.03266Cited by:[§2\.3](https://arxiv.org/html/2608.28980#S2.SS3.p3.1)\.
- \[96\]Z\. Wang, Z\. Liu, T\. Ma,et al\.\(2025\)Graph foundation models: a comprehensive survey\.Note:Author list corrected from an earlier draft’s Wang, Xin and Liu, Sike and Ma, Yushun, which did not match the arXiv record; verified against arxiv\.org/abs/2505\.15116External Links:2505\.15116Cited by:[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p2.1)\.
- \[97\]A\. T\. Wasi, W\. Faisal, A\. Rahman, M\. A\. Anik, M\. Shahriar, M\. M\. Topu, S\. T\. Meem, R\. N\. Priti, S\. A\. Mitu, Md\. I\. Hoque, S\. Z\. Ridoy, M\. E\. Ali, M\. Hawasly, M\. Raza, and M\. R\. Parvez\(2026\)SpatiaLab: can vision\-language models perform spatial reasoning in the wild?\.InInternational Conference on Learning Representations \(ICLR\),Note:Verified via arXiv abstract page; ICLR 2026 venue confirmed\. Note: an earlier internal draft of this review’s notes mistakenly conflated this paper with a differently\-titled benchmark \("Spatial\-DISE," arXiv:2510\.13394\) found in the same search pass – that conflation has been corrected; this entry cites only the paper actually verified at arXiv:2602\.03916External Links:2602\.03916Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p2.1)\.
- \[98\]A\. Webson and E\. Pavlick\(2022\)Do prompt\-based models really understand the meaning of their prompts?\.External Links:2109\.01247Cited by:[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p7.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.6.3.1.1)\.
- \[99\]D\. H\. Wolpert and W\. G\. Macready\(1997\)No free lunch theorems for optimization\.IEEE Transactions on Evolutionary Computation1\(1\),pp\. 67–82\.Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p2.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p1.1)\.
- \[100\]X\. Wu, A\. Ritter, and W\. Xu\(2025\)Tabular data understanding with large language models: a survey of recent advances and challenges\.External Links:2508\.00217Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[101\]Z\. Wu, B\. Ramsundar, E\. N\. Feinberg, J\. Gomes, C\. Geniesse, A\. S\. Pappu, K\. Leswing, and V\. Pande\(2018\)MoleculeNet: a benchmark for molecular machine learning\.Chemical Science9\(2\),pp\. 513–530\.Note:Verified via the publisher record \(Royal Society of Chemistry\); this is the benchmark suite[Ross et al\. \[73\]](https://arxiv.org/html/2608.28980#bib.bib79)’s “ten molecular\-property benchmarks” draws on \(BBBP, Tox21, ClinTox, HIV, BACE, SIDER, QM8, QM9, ESOL, FreeSolv, lipophilicity\)External Links:[Document](https://dx.doi.org/10.1039/c7sc02664a)Cited by:[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.12.1.1)\.
- \[102\]Z\. Wu, S\. Song, A\. Khosla, F\. Yu, L\. Zhang, X\. Tang, and J\. Xiao\(2015\)3D ShapeNets: a deep representation for volumetric shapes\.InConference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 1912–1920\.Note:Verified via the CVPR open\-access proceedings; introduces the ModelNet benchmark, whose ModelNet40 test split[Xu et al\. \[104\]](https://arxiv.org/html/2608.28980#bib.bib84)uses for generative 3D object classificationCited by:[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.14.1.1)\.
- \[103\]Y\. Xiao W\. Zhaoet al\.\(2025\)Protein large language models: a comprehensive survey\.Note:EMNLP 2025 FindingsExternal Links:2502\.17504Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[104\]R\. Xu, X\. Wang, T\. Wang, Y\. Chen, J\. Pang, and D\. Lin\(2024\)PointLLM: empowering large language models to understand point clouds\.InEuropean Conference on Computer Vision \(ECCV\),External Links:2308\.16911Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1),[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p6.1),[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px3.p1.1),[102](https://arxiv.org/html/2608.28980#bib.bib85)\.
- \[105\]Y\. Yang, N\. Srebro, and Y\. Li\(2026\)Tight sample complexity of transformers\.InConference on Learning Theory \(COLT\),External Links:2606\.09731Cited by:[§3\.2](https://arxiv.org/html/2608.28980#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.28980#S3.SS3.p3.1),[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.3.3.1.1),[§8\.1](https://arxiv.org/html/2608.28980#S8.SS1.p1.1)\.
- \[106\]L\. Yao, J\. Peng, C\. Mao, and Y\. Luo\(2025\)Exploring large language models for knowledge graph completion\.InIEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),External Links:2308\.13916Cited by:[§4\.2](https://arxiv.org/html/2608.28980#S4.SS2.p3.1)\.
- \[107\]Q\. Yao, C\. H\. Yang, R\. Jiang, Y\. Liang, M\. Jin, and S\. Pan\(2025\)Towards neural scaling laws for time series foundation models\.InInternational Conference on Learning Representations \(ICLR\),External Links:2410\.12360Cited by:[§6\.3](https://arxiv.org/html/2608.28980#S6.SS3.p2.1)\.
- \[108\]S\. Yoonet al\.\(2025\)Towards more reliable responses for order\-invariant inputs\.InAssociation for Computational Linguistics \(ACL\),Note:RoToRExternal Links:2502\.08662Cited by:[Table 1](https://arxiv.org/html/2608.28980#S3.T1.6.6.3.1.1),[§5](https://arxiv.org/html/2608.28980#S5.p2.1),[§6\.1](https://arxiv.org/html/2608.28980#S6.SS1.p2.1),[§6\.2](https://arxiv.org/html/2608.28980#S6.SS2.p7.1),[§7\.1](https://arxiv.org/html/2608.28980#S7.SS1.p2.1),[§7\.3](https://arxiv.org/html/2608.28980#S7.SS3.p3.1),[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.6.3.1.1),[§8\.2](https://arxiv.org/html/2608.28980#S8.SS2.p3.1)\.
- \[109\]Z\. Yuan, T\. Liu, Y\. Yang, Y\. Wang, F\. Qi, K\. Rangadurai, B\. Li, and S\. Yang\(2025\)ArchPilot: a proxy\-guided multi\-agent approach for machine learning engineering\.External Links:2511\.03985Cited by:[§3\.1](https://arxiv.org/html/2608.28980#S3.SS1.p2.1)\.
- \[110\]A\. Zeng, M\. Chen, L\. Zhang, and Q\. Xu\(2023\)Are transformers effective for time series forecasting?\.InAAAI Conference on Artificial Intelligence,Note:DLinearExternal Links:2205\.13504Cited by:[Table 5](https://arxiv.org/html/2608.28980#S7.T5.5.8.2.1.1)\.
- \[111\]W\. Zhanget al\.\(2025\)The point, the vision and the text: does point cloud boost spatial reasoning of large language models? a bias\-controlled study\.External Links:2504\.04540Cited by:[§4\.5](https://arxiv.org/html/2608.28980#S4.SS5.SSS0.Px3.p1.1),[Table 3](https://arxiv.org/html/2608.28980#S5.T3.5.14.4.1.1)\.
- \[112\]X\. Zhang, R\. R\. Chowdhury, R\. K\. Gupta, and J\. Shang\(2024\)Large language models for time series: a survey\.External Links:2402\.01801Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p1.1)\.
- \[113\]Y\. Zhang, L\. Li, Y\. Cui, X\. Ruan, Z\. Zheng, K\. Chen, Y\. Zhang, and D\. Yang\(2026\)Grid2Matrix: revealing digital agnosia in vision\-language models\.External Links:2604\.09687Cited by:[§4\.4](https://arxiv.org/html/2608.28980#S4.SS4.p3.1),[Table 6](https://arxiv.org/html/2608.28980#S7.T6.5.5.4.1.1)\.
- \[114\]T\. Zhou, P\. Niu, X\. Wang, L\. Sun, and R\. Jin\(2023\)One fits all: power general time series analysis by pretrained LM\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Spotlight,Note:GPT4TS; falsified by ablation in[Tan et al\. \[81\]](https://arxiv.org/html/2608.28980#bib.bib57)External Links:2302\.11939Cited by:[§4\.3](https://arxiv.org/html/2608.28980#S4.SS3.p2.1)\.
- \[115\]Y\. Zhou, J\. Li, Y\. Xiang, H\. Yan, L\. Gui, and Y\. He\(2024\)The mystery of in\-context learning: a comprehensive survey on interpretation and analysis\.InEmpirical Methods in Natural Language Processing \(EMNLP\),External Links:2311\.00237Cited by:[§1](https://arxiv.org/html/2608.28980#S1.p7.1),[§2\.2](https://arxiv.org/html/2608.28980#S2.SS2.p3.1)\.Similar Articles
Can Large Language Models Reinvent Foundational Algorithms?
Researchers introduce 'Unlearn-and-Reinvent', a pipeline that removes knowledge of foundational algorithms (e.g., Dijkstra's, Euclid's) from LLMs via unlearning, then tests whether models can independently reinvent them. Results show LLMs can reinvent algorithms with intuitive structures but struggle with those requiring non-obvious data structures or counterintuitive invariants.
On the use of foundation models in cognitive science
This perspective paper from arXiv articulates a four-stage inferential framework for evaluating foundation models as cognitive and developmental models, emphasizing that behavioral alignment alone is insufficient and must be embedded within theoretical commitments and contrastive evaluation.
AI for Games in the Foundation Model Era
This paper surveys the use of foundation models in game AI across roles like playing, modeling, design, and evaluation, highlighting transferability challenges and the need for game-specific validation.
Small Foundation Models of Human Cognition and Behaviour
This paper investigates whether small foundation models fine-tuned on human behavioral data can serve as cognitive proxies, finding that scale matters little in-distribution but larger models generalize better out-of-distribution.
Generalized Multimodal Foundation Model
This paper introduces a generalized multimodal foundation model capable of handling arbitrary modality combinations and prediction tasks, achieving competitive performance through training on large-scale synthetic datasets with diverse causal structures.