A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks

arXiv cs.AI Papers

Summary

This survey provides a comprehensive review of wireless foundation models, covering architectures, pre-training paradigms, applications, and challenges for AI-native 6G networks.

arXiv:2608.14694v1 Announce Type: new Abstract: Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning models that are trained for individual applications, wireless foundation models (WFMs) learn generalized representations from large-scale heterogeneous wireless data and can be efficiently adapted to communication, sensing, localization, and network optimization tasks with minimal task-specific supervision. Despite rapid progress, current research remains fragmented across architectures, training paradigms, and application domains, with no unified survey dedicated to the design, learning, and deployment of WFMs. This survey presents a comprehensive and unified review of wireless foundation models. We first establish the fundamental concepts of WFMs and introduce a taxonomy that organizes the field according to model architectures, pre-training paradigms, and applications. We then review representative architectures, self-supervised pre-training strategies, parameter-efficient adaptation methods, datasets, benchmarks, and evaluation methodologies, highlighting their roles in enabling transferable wireless intelligence. Furthermore, we examine emerging applications spanning physical-layer signal processing, network intelligence, and cross-layer optimization, and discuss the key challenges of data availability, generalization, interpretability, efficient edge deployment, and standardization. Finally, we outline future research directions toward scalable, trustworthy, and general-purpose wireless intelligence for AI-native 6G networks. This survey provides a comprehensive reference for researchers and practitioners developing next-generation intelligent wireless systems.
Original Article
View Cached Full Text

Cached at: 08/18/26, 10:01 AM

# A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Source: [https://arxiv.org/html/2608.14694](https://arxiv.org/html/2608.14694)
Naveed Khan, Besan Al Sbeihi, Maryam Alshehhi, , and Nasir SaeedThe authors are with the Department of Electrical and Communication Engineering, College of Engineering, United Arab Emirates University \(UAEU\), Al Ain, UAE \(e\-mail: mr\.nasir\.saeed@ieee\.org\)\.

###### Abstract

Foundation models are emerging as a transformative paradigm for AI\-native sixth\-generation \(6G\) wireless networks by enabling scalable, transferable, and data\-efficient intelligence across diverse communication tasks\. Unlike conventional deep learning models that are trained for individual applications, wireless foundation models \(WFMs\) learn generalized representations from large\-scale heterogeneous wireless data and can be efficiently adapted to communication, sensing, localization, and network optimization tasks with minimal task\-specific supervision\. Despite rapid progress, current research remains fragmented across architectures, training paradigms, and application domains, with no unified survey dedicated to the design, learning, and deployment of WFMs\. This survey presents a comprehensive and unified review of wireless foundation models\. We first establish the fundamental concepts of WFMs and introduce a taxonomy that organizes the field according to model architectures, pre\-training paradigms, and applications\. We then review representative architectures, self\-supervised pre\-training strategies, parameter\-efficient adaptation methods, datasets, benchmarks, and evaluation methodologies, highlighting their roles in enabling transferable wireless intelligence\. Furthermore, we examine emerging applications spanning physical\-layer signal processing, network intelligence, and cross\-layer optimization, and discuss the key challenges of data availability, generalization, interpretability, efficient edge deployment, and standardization\. Finally, we outline future research directions toward scalable, trustworthy, and general\-purpose wireless intelligence for AI\-native 6G networks\. This survey provides a comprehensive reference for researchers and practitioners developing next\-generation intelligent wireless systems\.

## IIntroduction

Wireless communication systems are evolving from model\-driven signal processing toward data\-driven intelligence to meet the stringent performance, scalability, and adaptability requirements of AI\-native sixth\-generation \(6G\) networks\[[31](https://arxiv.org/html/2608.14694#bib.bib152)\]\. For decades, analytical channel models, estimation theory, optimization techniques, and signal processing algorithms have successfully supported wireless tasks such as channel estimation, signal detection, beamforming, resource allocation, and network management\[[112](https://arxiv.org/html/2608.14694#bib.bib153)\]\. However, emerging technologies, including massive multiple\-input multiple\-output \(MIMO\), ultra\-dense heterogeneous networks, integrated sensing and communication \(ISAC\), semantic communications, and AI\-native 6G systems, have introduced levels of complexity that increasingly challenge conventional analytical models\[[32](https://arxiv.org/html/2608.14694#bib.bib48),[177](https://arxiv.org/html/2608.14694#bib.bib49),[76](https://arxiv.org/html/2608.14694#bib.bib11)\]\.

Deep learning has significantly advanced wireless communications by learning complex signal representations directly from data, with early successes in physical\-layer tasks such as channel estimation, signal detection, and beamforming\[[32](https://arxiv.org/html/2608.14694#bib.bib48)\]\. Its use has subsequently expanded to localization, modulation recognition, intelligent wireless receivers, and network\-level resource optimization\[[177](https://arxiv.org/html/2608.14694#bib.bib49),[36](https://arxiv.org/html/2608.14694#bib.bib50)\]\. Nevertheless, most existing learning\-based approaches remain task\-specific, requiring separate models and training procedures for individual communication problems\. Moreover, their dependence on large labeled datasets and environment\-specific training can make adaptation to new deployment conditions costly\[[125](https://arxiv.org/html/2608.14694#bib.bib52)\]\. Models trained under particular channel, mobility, or hardware conditions may also generalize poorly when the underlying data distribution changes\[[7](https://arxiv.org/html/2608.14694#bib.bib5),[105](https://arxiv.org/html/2608.14694#bib.bib51)\]\. As wireless systems evolve toward increasingly dynamic and heterogeneous environments, repeatedly collecting labeled data, retraining models, and maintaining independent learning pipelines for individual tasks becomes inefficient and difficult to scale\.

These limitations have motivated a transition from task\-specific learning toward foundation models\. Rather than optimizing separate models for individual tasks, foundation models are pre\-trained on large\-scale and heterogeneous datasets, often using self\-supervised or unsupervised objectives, to learn transferable representations that can be efficiently adapted to multiple downstream applications\[[184](https://arxiv.org/html/2608.14694#bib.bib154)\]\. This paradigm has already reshaped natural language processing and computer vision through improved data efficiency, generalization, and knowledge transfer, and is now attracting growing interest in wireless communications\[[1](https://arxiv.org/html/2608.14694#bib.bib155)\]\.

Wireless systems are particularly well suited to this paradigm because they generate diverse and complementary data modalities\. Channel\-centric measurements, such as channel state information \(CSI\) and channel impulse responses \(CIR\), provide rich representations of propagation characteristics and spatial structure\[[22](https://arxiv.org/html/2608.14694#bib.bib53)\]\. At the signal level, in\-phase and quadrature \(IQ\) samples, spectrograms, and other radio\-frequency \(RF\) observations enable foundation models to learn reusable representations directly from raw or transformed wireless signals\[[10](https://arxiv.org/html/2608.14694#bib.bib69),[172](https://arxiv.org/html/2608.14694#bib.bib8)\]\. More recent approaches further incorporate heterogeneous sensing and communication modalities to capture complementary information across wireless environments\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]\. The integration of channel, signal, sensing, and network\-level observations therefore provides a natural basis for multimodal representation learning and the development of general\-purpose wireless foundation models\[[181](https://arxiv.org/html/2608.14694#bib.bib21)\]\.

Wireless foundation models \(WFMs\) have consequently emerged as a promising paradigm for enabling scalable and general\-purpose wireless intelligence\. Rather than developing independent models for communication, sensing, localization, and network optimization, a single pre\-trained backbone can be efficiently adapted to multiple downstream tasks through fine\-tuning, parameter\-efficient adaptation, prompt tuning, or in\-context learning\[[84](https://arxiv.org/html/2608.14694#bib.bib156)\]\. Recent models, including WirelessGPT, Large Wireless Model \(LWM\), WavesFM, and emerging multimodal WFMs, demonstrate that large\-scale pre\-training substantially improves knowledge transfer, parameter efficiency, and cross\-domain generalization while reducing dependence on large labeled datasets\[[163](https://arxiv.org/html/2608.14694#bib.bib157)\]\. Figure\.[1](https://arxiv.org/html/2608.14694#S1.F1)summarizes the fundamental learning paradigm of WFMs\. Unlike conventional deep learning, where separate neural networks are independently trained for individual communication tasks, WFMs decouple large\-scale representation learning from task\-specific optimization\. During pre\-training, heterogeneous wireless data including CSI, IQ samples, CIRs, RF signals, sensing measurements, and network observations are used to learn a shared representation that captures common propagation characteristics across diverse wireless environments\. The resulting backbone is subsequently adapted to multiple downstream tasks through lightweight techniques such as fine\-tuning, parameter\-efficient adaptation, or prompt tuning\. This unified workflow enables communication, sensing, localization, and network optimization to share a common representation, substantially improving scalability, transferability, and data efficiency while reducing retraining costs\.

### I\-AMotivation and Contributions

Despite these rapid advances, research on wireless foundation models remains fragmented across different communities and application domains\. Existing surveys primarily focus on large AI models for communications, transfer learning, multimodal wireless intelligence, or specific communication applications, without treating WFMs as a unified learning paradigm\[[161](https://arxiv.org/html/2608.14694#bib.bib159),[34](https://arxiv.org/html/2608.14694#bib.bib158)\]\. Consequently, there remains a lack of a comprehensive survey that systematically integrates architectural design, self\-supervised pre\-training, adaptation strategies, datasets, benchmark frameworks, evaluation methodologies, deployment considerations, and emerging wireless applications within a single coherent framework\. Furthermore, the rapid emergence of representative WFMs since 2024 has fundamentally reshaped the research landscape, highlighting the need for a timely and comprehensive survey dedicated specifically to wireless foundation models\.

Motivated by these developments, this survey presents a unified treatment of wireless foundation models, covering their fundamental principles, architectural design, learning paradigms, adaptation strategies, deployment considerations, evaluation methodologies, and applications across AI\-native 6G networks\. Beyond summarizing the existing literature, this survey synthesizes the rapidly evolving research landscape, identifies common design principles, compares representative models and learning strategies, discusses emerging research trends, and outlines future directions toward scalable and trustworthy general\-purpose wireless intelligence\. The main contributions of this survey are summarized as follows:

- •We establish a unified conceptual framework for wireless foundation models by formalizing their definition, distinguishing them from conventional task\-specific deep learning, and describing the transition toward large\-scale transferable representation learning for AI\-native 6G networks\.
- •We develop a comprehensive taxonomy that organizes wireless foundation models from three complementary perspectives, namely model architectures, learning paradigms, and applications, providing a systematic framework for analyzing existing developments and future research directions\.
- •We present a comprehensive review and critical comparison of representative wireless foundation models, covering Transformer\-based, physics\-informed, graph\-based, and multimodal architectures together with self\-supervised pre\-training, parameter\-efficient adaptation, prompt\-based learning, and transfer learning strategies\.
- •We systematically review publicly available datasets, simulation platforms, benchmark methodologies, and evaluation metrics, and discuss their roles in enabling scalable, transferable, and reproducible wireless foundation models\.
- •We provide a holistic analysis of emerging applications spanning physical\-layer signal processing, integrated sensing and communication, localization, resource management, and network intelligence, highlighting how a shared pre\-trained backbone enables general\-purpose wireless intelligence across heterogeneous wireless tasks\.
- •We identify key research challenges and future opportunities, including data scarcity, domain generalization, multimodal learning, interpretability, efficient edge deployment, continual adaptation, and standardization, and present a research roadmap toward scalable, trustworthy, and AI\-native 6G wireless systems\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/figS1a.png)Figure 1:Conceptual workflow of a wireless foundation model\. Large\-scale self\-supervised pre\-training learns a shared wireless representation from heterogeneous wireless data, which is subsequently adapted to diverse downstream communication, sensing, localization, and network intelligence tasks through lightweight adaptation mechanisms\.
### I\-BOrganization

The remainder of this survey is organized as follows\. Section II reviews the background of model\-based and learning\-based wireless communications together with the fundamental concepts of wireless foundation models\. Section III introduces a unified taxonomy of wireless foundation models from the perspectives of architecture, training strategy, and protocol layer\. Section IV reviews representative model architectures, while Section V discusses large\-scale pre\-training strategies and adaptation techniques\. Section VI summarizes the major application domains of wireless foundation models\. Section VII presents publicly available datasets, benchmark platforms, and evaluation methodologies\. Section VIII discusses the major open research challenges and future research directions, followed by a roadmap toward AI\-native 6G systems in Section IX\. Finally, Section X concludes the survey\.

## IIEvolution Toward Wireless Foundation Models

WFMs represent the latest stage in the evolution of artificial intelligence for wireless communications\. Their emergence has been driven by the limitations of both conventional model\-based signal processing and task\-specific deep learning to address the increasing complexity of AI\-native 6G networks\[[138](https://arxiv.org/html/2608.14694#bib.bib160)\]\. Understanding this evolution provides the necessary context to appreciate the design philosophy of WFMs and the motivation behind the large\-scale transferable representation learning\.

This section first reviews the foundations of model\-based wireless communications and the transition toward deep learning\. It then discusses how the success of foundation models in natural language processing \(NLP\) and computer vision \(CV\) has inspired a new generation of wireless learning systems capable of supporting multiple communication tasks through a shared pre\-trained backbone\.

### II\-AModel\-Based Wireless Communications

Wireless communication systems have traditionally relied on analytical models, estimation theory, optimization, and statistical signal processing to solve fundamental problems such as channel estimation, signal detection, synchronization, beamforming, equalization, and resource allocation\[[51](https://arxiv.org/html/2608.14694#bib.bib190)\]\. Classical approaches, including least\-squares \(LS\), minimum mean square error \(MMSE\), maximum likelihood \(ML\), and convex optimization, provide mathematically interpretable solutions with well\-understood theoretical guarantees and have formed the cornerstone of successive wireless generations\[[158](https://arxiv.org/html/2608.14694#bib.bib56),[45](https://arxiv.org/html/2608.14694#bib.bib57)\]\.

The effectiveness of these techniques depends on accurate mathematical descriptions of wireless propagation and network behavior\. For example, channel estimation algorithms assume specific statistical channel models, while beamforming and resource allocation require reliable CSI and tractable optimization formulations\[[9](https://arxiv.org/html/2608.14694#bib.bib185)\]\. Under these assumptions, model\-driven approaches achieve predictable performance and can often be analyzed rigorously using estimation theory, information theory, and convex optimization\[[82](https://arxiv.org/html/2608.14694#bib.bib58),[159](https://arxiv.org/html/2608.14694#bib.bib59)\]\.

However, the transition toward AI\-native 6G networks has significantly increased the complexity of wireless systems\[[73](https://arxiv.org/html/2608.14694#bib.bib193)\]\. Emerging technologies, including massive MIMO, millimeter\-wave and terahertz communications, ISAC, ultra\-dense heterogeneous deployments, and highly dynamic propagation environments, introduce nonlinear interactions and uncertainties that are increasingly difficult to capture using analytical models alone\[[135](https://arxiv.org/html/2608.14694#bib.bib161)\]\. Consequently, deriving accurate mathematical models and designing optimal signal processing algorithms have become considerably more challenging, motivating a gradual transition toward data\-driven learning approaches capable of extracting knowledge directly from wireless measurements\.

### II\-BDeep Learning for Wireless Communications

Deep learning has transformed wireless communications by enabling data\-driven models to learn complex nonlinear relationships directly from wireless observations, complementing conventional approaches that rely on analytical channel and signal models\[[46](https://arxiv.org/html/2608.14694#bib.bib63)\]\. Its effectiveness has been demonstrated across a broad range of physical\-layer tasks, including channel estimation, signal detection, beamforming, localization, modulation recognition, and resource allocation\[[33](https://arxiv.org/html/2608.14694#bib.bib1)\]\. For example, deep neural networks have been successfully applied to joint channel estimation and signal detection in OFDM systems, illustrating their ability to learn communication functions directly from received signals\[[174](https://arxiv.org/html/2608.14694#bib.bib2)\]\. Such learning capabilities are particularly relevant as wireless systems evolve toward increasingly high\-dimensional configurations involving technologies such as massive MIMO and millimeter\-wave communications\[[19](https://arxiv.org/html/2608.14694#bib.bib3)\]\.

Despite these advances, current learning\-based wireless systems remain predominantly task\-specific\. Individual models are typically trained for a single communication task under particular channel conditions and deployment assumptions\. Consequently, their performance often degrades under changes in propagation environments, antenna configurations, user mobility, hardware impairments, or other distribution shifts\[[106](https://arxiv.org/html/2608.14694#bib.bib7)\]\. Recovering performance usually requires collecting additional labeled data and retraining or fine\-tuning the model, limiting scalability in practical deployments\.

Another major limitation is the dependence on large labeled datasets\. Accurate supervision often requires pilot\-assisted measurements, extensive simulations, or costly measurement campaigns, making dataset construction both expensive and time\-consuming\[[124](https://arxiv.org/html/2608.14694#bib.bib4)\]\. In addition, modern deep neural networks typically require substantial computational resources, memory, and training time, posing significant challenges for latency\-sensitive and resource\-constrained wireless devices\. Their black\-box nature also limits interpretability and complicates verification in safety\-critical wireless applications\[[37](https://arxiv.org/html/2608.14694#bib.bib6)\]\.

Collectively, these limitations highlight that task\-specific deep learning alone is insufficient to support the scalability, adaptability, and generalization required by AI\-native 6G networks\. Rather than repeatedly training independent models for every communication task, a more scalable paradigm is needed one that learns reusable wireless representations capable of transferring knowledge across heterogeneous tasks and deployment scenarios\. This requirement has motivated the emergence of wireless foundation models\.

## IIIDefinitions and Taxonomy of Wireless Foundation Models

WFMs represent a new generation of learning\-based wireless intelligence that extends beyond conventional task\-specific deep learning\. Instead of developing independent models for individual communication tasks, a wireless foundation model is pre\-trained on large\-scale heterogeneous wireless data to learn transferable representations that can be efficiently adapted to multiple downstream applications\[[130](https://arxiv.org/html/2608.14694#bib.bib162)\]\. Owing to the diversity of existing model architectures, pre\-training strategies, and application domains, a systematic taxonomy is essential for understanding the rapidly evolving research landscape\. As illustrated in Fig\.[2](https://arxiv.org/html/2608.14694#S3.F2), wireless foundation models can be classified from three complementary perspectives:*\(i\)*model architecture, which determines how wireless representations are learned;*\(ii\)*pre\-training strategy, which defines how transferable knowledge is acquired from large\-scale wireless datasets; and*\(iii\)*deployment layer, which categorizes the downstream communication and networking tasks supported by the pre\-trained model\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/Figures/figS1b.png)Figure 2:Taxonomy of wireless foundation models based on three complementary dimensions: \(i\) model architecture; \(ii\) pre\-training strategy; and \(iii\) deployment, including physical\-layer, MAC\-layer, network\-layer, and cross\-layer optimization tasks\.The proposed taxonomy is organized around these three dimensions because they collectively describe the complete lifecycle of a wireless foundation model\. The architectural dimension specifies the underlying neural backbone responsible for representation learning, the pre\-training dimension characterizes how transferable knowledge is acquired from large\-scale wireless data, and the application dimension reflects how the learned representations are adapted to downstream wireless communication and networking tasks\. Other possible classifications, such as data modality or deployment platform, typically represent specific aspects of these three fundamental dimensions rather than independent design categories\. Therefore, this taxonomy provides a unified and systematic framework for analyzing, comparing, and organizing existing wireless foundation models while accommodating future developments in AI\-native 6G systems\.

### III\-ADefinition of Wireless Foundation Models

WFMs extend conventional task\-specific learning by introducing a unified pre\-trained backbone that can be adapted to diverse wireless applications\. Unlike traditional machine learning and deep learning models, which are typically optimized for a single communication task or deployment scenario, WFMs learn transferable representations from large\-scale heterogeneous wireless data and reuse the acquired knowledge across multiple downstream tasks with limited task\-specific supervision\[[21](https://arxiv.org/html/2608.14694#bib.bib14)\]\. This paradigm enables a single model to support communication, sensing, localization, and network intelligence within a unified learning framework, thereby improving scalability, data efficiency, and cross\-task generalization\.

The operation of a WFM follows a two\-stage learning paradigm consisting of large\-scale pre\-training and downstream adaptation\. During the pre\-training stage, heterogeneous wireless measurements, including CSI, IQ samples, CIRs, RF signals, spectrograms, sensing measurements, and network observations, are used to learn a shared wireless representation through self\-supervised or unsupervised objectives such as masked reconstruction, contrastive learning, or generative modeling\[[172](https://arxiv.org/html/2608.14694#bib.bib8),[6](https://arxiv.org/html/2608.14694#bib.bib9),[10](https://arxiv.org/html/2608.14694#bib.bib69)\]\. Rather than learning task\-specific features, the objective is to capture common spatial, temporal, spectral, and propagation characteristics that are transferable across different wireless environments\.

The learned representation is subsequently adapted to downstream applications using lightweight techniques, including full fine\-tuning, parameter\-efficient adaptation, prompt tuning, or task\-specific prediction heads\. Representative applications include channel estimation, channel prediction, beamforming, MIMO detection, CSI feedback, localization, integrated sensing, wireless signal classification, and resource allocation\[[183](https://arxiv.org/html/2608.14694#bib.bib97),[29](https://arxiv.org/html/2608.14694#bib.bib12)\]\. Compared with independently trained deep learning models, this shared\-backbone paradigm substantially reduces labeled data requirements, retraining costs, and deployment complexity while improving robustness across heterogeneous wireless scenarios\.

Fig\.[1](https://arxiv.org/html/2608.14694#S1.F1)summarizes the fundamental workflow of a wireless foundation model\. Large\-scale self\-supervised pre\-training first learns a shared representation from heterogeneous wireless data, after which the pre\-trained backbone is efficiently adapted to diverse downstream communication tasks through lightweight adaptation mechanisms\. This separation of representation learning from task\-specific optimization constitutes the defining characteristic of WFMs and provides the conceptual basis for the taxonomy presented in the remainder of this section\.

### III\-BTaxonomy by Architecture

The architectural backbone of a wireless foundation model largely determines its representation learning capability, scalability, computational efficiency, and adaptation performance\. As illustrated in Fig\.[2](https://arxiv.org/html/2608.14694#S3.F2), current WFMs can be broadly categorized into four architectural families: Transformer\-based, convolutional neural network \(CNN\)\-based, physics\-informed, and graph\-based models\. Rather than representing competing solutions, these architectures offer complementary design philosophies for learning generalized wireless representations across different communication scenarios\. Transformer\-based architectures have emerged as the dominant backbone for wireless foundation models owing to their ability to model long\-range spatial, temporal, and frequency\-domain dependencies through self\-attention mechanisms\[[116](https://arxiv.org/html/2608.14694#bib.bib163)\]\. Their scalability and strong representation learning capability make them particularly well suited for large\-scale self\-supervised pre\-training using heterogeneous wireless measurements, CSI, IQ samples, CIRs, spectrograms, and RF signals\. Consequently, most recent WFMs, including WirelessGPT, LWM, and WavesFM, adopt Transformer\-based architectures as their primary learning backbone\[[160](https://arxiv.org/html/2608.14694#bib.bib60),[39](https://arxiv.org/html/2608.14694#bib.bib61)\]\.

CNN\-based architectures remain attractive for applications where computational efficiency and low\-latency inference are critical\. By exploiting local spatial and temporal correlations, CNNs provide lightweight yet effective feature extraction for tasks such as channel estimation, modulation recognition, spectrum sensing, and wireless signal classification\. Although they generally exhibit lower representation capacity than Transformers, their efficiency makes them well suited for edge deployment and resource\-constrained wireless devices\[[94](https://arxiv.org/html/2608.14694#bib.bib62),[46](https://arxiv.org/html/2608.14694#bib.bib63)\]\. Physics\-informed architectures combine data\-driven learning with communication\-domain knowledge by embedding analytical models, optimization algorithms, or signal processing principles into trainable neural networks\. Representative approaches, such as deep unfolding, preserve the interpretability of conventional communication algorithms while improving robustness, sample efficiency, and generalization\. These models are particularly attractive for wireless tasks where reliable physical models are available but require adaptive learning capabilities\[[55](https://arxiv.org/html/2608.14694#bib.bib64),[58](https://arxiv.org/html/2608.14694#bib.bib65)\]\.

Graph\-based architectures naturally represent wireless networks as graphs, where nodes correspond to users, base stations, access points, or network entities, and edges describe communication, interference, or connectivity relationships\. By explicitly modeling network topology, graph neural networks effectively capture spatial interactions that are difficult to represent using conventional neural architectures, making them particularly suitable for resource allocation, routing, user association, interference management, and network optimization in large\-scale AI\-native 6G systems\[[165](https://arxiv.org/html/2608.14694#bib.bib66),[113](https://arxiv.org/html/2608.14694#bib.bib67)\]\. Fig\.[2](https://arxiv.org/html/2608.14694#S3.F2)highlights that these architectural families differ primarily in how they encode wireless information rather than in their overall learning objective\. Transformer models emphasize global contextual modeling, CNNs focus on local feature extraction, physics\-informed models integrate communication\-domain knowledge, whereas graph\-based models explicitly exploit network topology\. In practice, emerging WFMs increasingly combine multiple architectural paradigms to balance representation quality, computational efficiency, interpretability, and scalability, suggesting that future wireless foundation models are likely to evolve toward hybrid architectures rather than relying on a single neural backbone\.

### III\-CTaxonomy by Learning Paradigm

Besides architectural design, wireless foundation models can also be categorized according to their learning paradigm, which determines how transferable wireless representations are acquired from large\-scale heterogeneous data\. Unlike conventional supervised learning, where models are optimized for a single task using manually annotated datasets, WFMs primarily rely on self\-supervised pre\-training to exploit the abundance of unlabeled wireless measurements\. As illustrated in Fig\.[2](https://arxiv.org/html/2608.14694#S3.F2), current learning paradigms can be broadly grouped into self\-supervised learning, contrastive representation learning, masked modeling, and multimodal pre\-training\[[42](https://arxiv.org/html/2608.14694#bib.bib164),[181](https://arxiv.org/html/2608.14694#bib.bib21)\]\.

Self\-supervised learning \(SSL\) forms the foundation of most contemporary WFMs\. Instead of relying on manually generated labels, SSL derives supervisory signals directly from the input data through carefully designed pretext tasks\. This learning paradigm is particularly attractive for wireless communications because large volumes of CSI, IQ samples, RF signals, spectrograms, and channel impulse responses are continuously generated during network operation, whereas obtaining accurate labels often requires costly measurements, simulations, or annotation\. Recent WFMs illustrate the effectiveness of this approach across different wireless modalities\. WirelessGPT\[[172](https://arxiv.org/html/2608.14694#bib.bib8)\]and LWM\[[10](https://arxiv.org/html/2608.14694#bib.bib69)\]exploit self\-supervised pre\-training to learn transferable representations from heterogeneous wireless observations, while WavesFM\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]extends representation learning across diverse wireless signals and tasks\. Similarly, IQFM\[[117](https://arxiv.org/html/2608.14694#bib.bib16)\]focuses on reusable representations from IQ data, whereas CSI2Vec\[[128](https://arxiv.org/html/2608.14694#bib.bib15)\]learns transferable representations from CSI\. Collectively, these models demonstrate that self\-supervised pre\-training can reduce dependence on task\-specific labels while enabling learned representations to be reused across multiple downstream communication tasks\.

Within the SSL framework, different pre\-training objectives encourage models to capture complementary characteristics of wireless signals\. Contrastive learning, in particular, learns discriminative representations by increasing the similarity between related wireless observations while separating unrelated samples in the latent space\. Positive pairs are typically constructed from different augmentations of the same signal or measurements obtained under similar propagation conditions, whereas negative pairs correspond to unrelated channel or signal realizations\. Such objectives can improve representation robustness to signal and channel variations and facilitate transfer across downstream tasks\. ContraWiMAE\[[49](https://arxiv.org/html/2608.14694#bib.bib13)\], for example, combines contrastive principles with masked representation learning, while CSI2Vec\[[128](https://arxiv.org/html/2608.14694#bib.bib15)\]applies representation learning specifically to CSI\. Related approaches extend this principle to other wireless modalities, with IQFM targeting IQ signals\[[117](https://arxiv.org/html/2608.14694#bib.bib16)\]and CSI\-CLIP exploiting contrastive alignment for CSI representations\[[78](https://arxiv.org/html/2608.14694#bib.bib18)\]\. These developments illustrate how contrastive objectives can be tailored to different wireless data modalities while retaining a common goal of learning transferable and discriminative representations\.

Masked modeling has recently emerged as an effective self\-supervised objective for WFMs\. Inspired by masked language modeling and masked image modeling, portions of the wireless input are intentionally hidden and the model is trained to reconstruct the missing information\. This approach is well suited to wireless data because signals often exhibit strong spatial, temporal, and frequency\-domain correlations, allowing models to learn propagation and signal structure by reconstructing masked CSI matrices, IQ sequences, OFDM resource grids, or spectrogram patches without requiring explicit labels\. WavesFM, for example, employs reconstruction\-oriented pre\-training to learn generalizable representations from wireless signals\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]\. Scalable masked channel modeling further exploits the inherent structure of channel observations to support transferable channel representations\[[50](https://arxiv.org/html/2608.14694#bib.bib17)\], while WiFo applies masked pre\-training to learn reusable wireless features for channel\-related downstream tasks\[[103](https://arxiv.org/html/2608.14694#bib.bib19)\]\. Collectively, these approaches demonstrate the potential of masked modeling for applications such as channel estimation, channel prediction, and CSI feedback\.

Beyond single\-modality learning, multimodal pre\-training extends representation learning by jointly modeling heterogeneous wireless and contextual information\. This capability is particularly relevant to future AI\-native 6G networks, where communication, sensing, localization, and network management are expected to become increasingly integrated\[[181](https://arxiv.org/html/2608.14694#bib.bib21)\]\. Such systems may need to jointly exploit wireless modalities such as CSI, IQ samples, and CIRs together with radar measurements, LiDAR, RGB images, GPS information, traffic statistics, and environmental context\. Multimodal WFM explores the integration of heterogeneous wireless information within a shared representation space\[[5](https://arxiv.org/html/2608.14694#bib.bib10)\], while MuSE\-FM extends this principle toward multimodal sensing and environmental representations\[[183](https://arxiv.org/html/2608.14694#bib.bib97)\]\. CSI\-CLIP further demonstrates contrastive alignment between CSI and complementary modalities, enabling semantically aligned representations across heterogeneous observations\[[78](https://arxiv.org/html/2608.14694#bib.bib18)\]\. These developments suggest that multimodal pre\-training can provide richer representations by exploiting complementary information that is unavailable to models trained on individual modalities alone\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/Figures/application_layer_taxonomy.png)Figure 3:Protocol\-layer taxonomy of wireless foundation models\. A shared pre\-trained backbone is adapted across the wireless protocol stack, enabling communication, resource management, network intelligence, and cross\-layer optimization through transferable wireless representations\.The learning paradigms shown in Fig\.[2](https://arxiv.org/html/2608.14694#S3.F2)should therefore be viewed as complementary rather than mutually exclusive\. Contemporary WFMs frequently combine multiple objectives during pre\-training, e\.g, integrating masked reconstruction with contrastive representation learning or multimodal alignment to improve robustness, transferability, and generalization\. Consequently, the evolution of WFMs is moving toward unified pre\-training frameworks that simultaneously exploit multiple learning objectives instead of relying on a single self\-supervised strategy\.

### III\-DTaxonomy by Deployment

Besides architectural design and learning paradigm, wireless foundation models can also be categorized according to the layer of the wireless protocol stack in which they are deployed\. This perspective is particularly important because each protocol layer operates on different data sources, optimization objectives, latency requirements, and decision timescales\. As illustrated in Fig\.[3](https://arxiv.org/html/2608.14694#S3.F3), a shared pre\-trained backbone can support applications spanning the entire wireless stack, from physical\-layer signal processing to network\-wide intelligence and cross\-layer optimization\[[100](https://arxiv.org/html/2608.14694#bib.bib23)\]\. Fig\.[3](https://arxiv.org/html/2608.14694#S3.F3)highlights that WFMs enable a common representation to be reused across protocol layers rather than developing independent AI models for individual networking functions\. This shared representation allows knowledge acquired from one wireless task to improve performance in related applications, thereby reducing retraining effort while improving scalability and transferability throughout AI\-native 6G systems\.

At the physical layer, WFMs primarily support signal processing tasks that operate directly on wireless measurements such as CSI, IQ samples, CIRs, and RF signals\. Representative applications include channel estimation, channel prediction, CSI feedback, MIMO detection, beam prediction, RF signal classification, localization, and integrated sensing\. Since these tasks exhibit strong spatial, temporal, and frequency\-domain correlations, they benefit significantly from large\-scale pre\-training and transferable representations\. Consequently, the physical layer has become the most mature application domain for WFMs, with representative systems including WirelessGPT, LWM, WavesFM, and WiFo demonstrating substantial improvements in transferability and data efficiency\[[43](https://arxiv.org/html/2608.14694#bib.bib165),[3](https://arxiv.org/html/2608.14694#bib.bib166)\]\.

Moving beyond signal processing, the MAC layer focuses on intelligent radio resource management\. WFMs can learn network\-level representations that support adaptive spectrum access, scheduling, beam management, interference coordination, power allocation, and link adaptation\. Recent LLM\-based approaches, for example, have explored intelligent resource allocation and radio\-access optimization using contextual network information\[[95](https://arxiv.org/html/2608.14694#bib.bib25),[126](https://arxiv.org/html/2608.14694#bib.bib26)\]\. Other studies investigate task prediction and adaptation through shared wireless representations\[[150](https://arxiv.org/html/2608.14694#bib.bib27)\], while LLM\-assisted MAC frameworks extend this capability to MAC\-layer decision making and coordination\[[154](https://arxiv.org/html/2608.14694#bib.bib30)\]\. These developments indicate a shift from independently optimized functions toward shared models that can incorporate network state, user behavior, and service requirements to support more adaptive resource management in dynamic wireless environments\.

At the network layer, WFMs operate on broader contextual information, including traffic statistics, mobility patterns, topology information, network telemetry, service requirements, and quality\-of\-service indicators\. Such information can support functions ranging from traffic prediction and mobility management to routing, anomaly detection, network slicing, and intent\-driven networking\. Digital twin\-assisted task offloading, for instance, illustrates how learned network representations and contextual information can facilitate adaptive decisions across cloud, edge, and hybrid deployments\[[8](https://arxiv.org/html/2608.14694#bib.bib186)\]\. More broadly, large AI models are increasingly being investigated as a foundation for intelligent wireless network operation\[[66](https://arxiv.org/html/2608.14694#bib.bib45)\], including network management and orchestration\[[164](https://arxiv.org/html/2608.14694#bib.bib28)\]\. Emerging frameworks further extend this direction toward greater wireless autonomy\[[101](https://arxiv.org/html/2608.14694#bib.bib46)\]and world\-model\-based representations of telecommunication environments\[[187](https://arxiv.org/html/2608.14694#bib.bib31)\]\. Together with task\-oriented WFMs\[[150](https://arxiv.org/html/2608.14694#bib.bib27)\], these developments point toward increasingly self\-managing and self\-optimizing AI\-native 6G networks\.

Although protocol layers are conventionally designed as distinct functional entities, their decisions are inherently coupled\. Physical\-layer channel conditions influence MAC\-layer scheduling and resource allocation, while traffic dynamics and service requirements at higher layers affect radio resource allocation, energy management, and communication reliability\. This coupling motivates cross\-layer WFMs that jointly represent wireless signals, network state, environmental context, and service objectives\. Such cross\-layer intelligence is particularly relevant to ISAC, where sensing, communication, and resource\-management decisions are closely interconnected\[[137](https://arxiv.org/html/2608.14694#bib.bib188),[167](https://arxiv.org/html/2608.14694#bib.bib32)\]\. Multimodal WFMs provide another pathway by integrating heterogeneous wireless and contextual information within shared representations\[[5](https://arxiv.org/html/2608.14694#bib.bib10),[183](https://arxiv.org/html/2608.14694#bib.bib97)\], while emerging models are extending this concept toward integrated communication and sensing\[[111](https://arxiv.org/html/2608.14694#bib.bib33)\]\. Similar cross\-layer principles are also relevant to semantic communications, digital twins, edge intelligence, and autonomous network management, where decisions increasingly span multiple functional layers and timescales\[[187](https://arxiv.org/html/2608.14694#bib.bib31)\]\.

## IVArchitectures for Wireless Foundation Models

The architectural backbone of a wireless foundation model largely determines its representation learning capability, computational efficiency, scalability, and adaptation performance\. Unlike conventional deep learning models that are designed for specific communication tasks, WFM architectures aim to learn transferable wireless representations that can be efficiently reused across diverse downstream applications\[[151](https://arxiv.org/html/2608.14694#bib.bib167)\]\. Although several neural architectures have been explored, Transformer\-based models have emerged as the dominant backbone owing to their superior ability to capture long\-range dependencies and scale to large self\-supervised pre\-training datasets\. Meanwhile, physics\-informed, graph\-based, and lightweight architectures complement Transformers by incorporating communication\-domain knowledge, network topology, or computational efficiency\[[108](https://arxiv.org/html/2608.14694#bib.bib168)\]\. This section reviews these architectural paradigms and discusses their suitability for AI\-native 6G systems\.

### IV\-ATransformer\-Based Architectures

Transformer\-based architectures have become the primary backbone of wireless foundation models because they combine scalable representation learning with excellent transferability across heterogeneous wireless tasks\. Their self\-attention mechanism enables the joint modeling of spatial, temporal, and frequency\-domain correlations that naturally exist in wireless measurements, making them particularly suitable for learning generalized representations from large\-scale heterogeneous datasets\[[79](https://arxiv.org/html/2608.14694#bib.bib169)\]\.

Unlike conventional task\-specific neural networks, Transformer\-based WFMs separate large\-scale representation learning from downstream task adaptation\. During self\-supervised pre\-training, the model learns reusable wireless representations from heterogeneous measurements, including CSI, IQ samples, CIRs, spectrograms, and RF signals\[[28](https://arxiv.org/html/2608.14694#bib.bib170)\]\. These representations are subsequently adapted to diverse communication, sensing, and networking applications through lightweight techniques such as full fine\-tuning, Low\-Rank Adaptation \(LoRA\), adapters, or prompt tuning, substantially reducing labeled data requirements and retraining costs\[[2](https://arxiv.org/html/2608.14694#bib.bib171)\]\. The flexibility of the Transformer architecture allows it to accommodate multiple wireless data modalities\. Sequential representations, such as IQ samples and CSI time series, are naturally processed using sequence Transformers, whereas image\-like representations, including CSI matrices, spectrograms, beamspace maps, and orthogonal frequency\-division multiplexing \(OFDM\) resource grids, are effectively modeled using Vision Transformers \(ViTs\)\[[39](https://arxiv.org/html/2608.14694#bib.bib61)\]\. This unified representation framework enables the direct application of self\-supervised objectives such as masked modeling and contrastive learning, allowing WFMs to exploit large volumes of unlabeled wireless data while learning representations that generalize across multiple propagation environments and communication tasks\.

Several representative models illustrate the potential of this paradigm\. WirelessGPT employs large\-scale self\-supervised pre\-training to learn general\-purpose wireless representations that can be adapted to different downstream tasks\[[172](https://arxiv.org/html/2608.14694#bib.bib8)\]\. LWM similarly explores transferable representation learning for channel modeling and broader communication intelligence\[[10](https://arxiv.org/html/2608.14694#bib.bib69)\]\. WavesFM extends this direction by combining a Vision Transformer backbone with masked pre\-training and LoRA\-based adaptation to support communication, sensing, and localization tasks within a shared framework\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]\. Collectively, these studies highlight the potential of Transformer\-based WFMs to improve representation reuse, parameter\-efficient adaptation, and transferability across wireless tasks compared with independently trained task\-specific models\.

Despite these advantages, several challenges remain\. The quadratic computational and memory complexity of global self\-attention can limit scalability when processing long CSI sequences, high\-dimensional channel observations, or large wireless resource grids\. Moreover, standard Transformer architectures do not inherently encode communication\-domain characteristics such as sparse multipath propagation, delay\-Doppler structure, channel reciprocity, and antenna geometry\[[14](https://arxiv.org/html/2608.14694#bib.bib172)\]\. Efficient attention mechanisms have therefore been investigated to reduce the computational burden associated with conventional self\-attention\[[155](https://arxiv.org/html/2608.14694#bib.bib68)\]\. Within wireless systems, emerging architectures increasingly incorporate scalable and domain\-aware designs to better capture the structure of wireless observations, as exemplified by AirFM\[[15](https://arxiv.org/html/2608.14694#bib.bib70)\]\. Lightweight Transformer variants provide another direction for reducing inference and deployment overhead in resource\-constrained wireless environments\[[30](https://arxiv.org/html/2608.14694#bib.bib71)\]\. These developments suggest that future Transformer\-based WFMs will increasingly combine scalable attention with wireless\-domain inductive biases to support efficient representation learning for AI\-native 6G systems\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/fig4.png)Figure 4:Conceptual architecture of a Transformer\-based wireless foundation model\.Fig\.[4](https://arxiv.org/html/2608.14694#S4.F4)summarizes the learning paradigm of a Transformer\-based wireless foundation model\. Instead of functioning as a task\-specific predictor, the Transformer serves as a reusable representation learner that acquires generalized wireless knowledge during large\-scale self\-supervised pre\-training\. The resulting backbone is subsequently specialized to downstream communication, sensing, and localization tasks through lightweight adaptation mechanisms, enabling a single pre\-trained model to support diverse wireless applications while preserving the knowledge acquired during pre\-training\. This decoupling of representation learning from task\-specific optimization constitutes the defining characteristic of Transformer\-based wireless foundation models\.

### IV\-BPhysics\-Informed and Model\-Driven Architectures

Although Transformer\-based architectures provide powerful representation learning capabilities, purely data\-driven models often require large training datasets and may struggle to generalize under unseen propagation environments\. Wireless communication systems, however, are governed by well\-established physical principles, including channel propagation, estimation theory, signal detection, antenna array processing, and optimization\. Physics\-informed and model\-driven architectures seek to bridge these complementary strengths by integrating communication\-domain knowledge into modern foundation models, thereby improving data efficiency, interpretability, robustness, and generalization while preserving the scalability of large\-scale representation learning\[[55](https://arxiv.org/html/2608.14694#bib.bib64),[152](https://arxiv.org/html/2608.14694#bib.bib72)\]\.

The central idea is to embed analytical models and physical constraints directly into the learning process rather than relying solely on statistical correlations\. Early model\-driven approaches achieved this through deep unfolding, where iterative communication algorithms were transformed into trainable neural network layers\. Representative examples include unfolded approximate message passing \(AMP\), projected gradient descent \(PGD\), and weighted minimum mean\-square error \(WMMSE\), which preserve the mathematical structure of conventional optimization algorithms while allowing trainable parameters to be learned from data\[[58](https://arxiv.org/html/2608.14694#bib.bib65),[55](https://arxiv.org/html/2608.14694#bib.bib64)\]\. These hybrid architectures demonstrated that incorporating communication\-domain knowledge significantly improves convergence, robustness, and sample efficiency compared with purely data\-driven models\.

Recent wireless foundation models extend this philosophy beyond individual communication tasks by incorporating physical priors directly into large\-scale self\-supervised pre\-training\. Rather than learning generic statistical representations, these models exploit communication\-specific characteristics such as channel reciprocity, sparse multipath propagation, delay\-Doppler structure, antenna geometry, and optimization constraints to guide representation learning toward physically meaningful solutions that generalize across heterogeneous deployment scenarios\[[15](https://arxiv.org/html/2608.14694#bib.bib70),[179](https://arxiv.org/html/2608.14694#bib.bib41)\]\. Representative systems illustrate this evolution\. AirFM\-DDA learns wireless representations in the delay\-Doppler\-angle domain to better capture propagation characteristics, while Adaptive 3D\-RoPE introduces physics\-aware positional encoding that preserves spatial and temporal relationships within wireless measurements\. These studies demonstrate that integrating communication\-domain knowledge into Transformer\-based foundation models substantially improves representation quality and robustness without sacrificing scalability\[[15](https://arxiv.org/html/2608.14694#bib.bib70),[179](https://arxiv.org/html/2608.14694#bib.bib41)\]\.

### IV\-CMultimodal Architectures

Future AI\-native 6G networks are expected to simultaneously support communication, sensing, localization, and network intelligence, requiring wireless foundation models to process information from multiple heterogeneous sources rather than a single wireless modality\. Consequently, multimodal architectures have emerged as an important extension of Transformer\-based foundation models by jointly learning representations from communication signals and contextual information\. Compared with unimodal architectures, multimodal models provide a richer understanding of the wireless environment, enabling improved generalization, robustness, and task transfer across diverse wireless applications\[[169](https://arxiv.org/html/2608.14694#bib.bib148),[88](https://arxiv.org/html/2608.14694#bib.bib149)\]\. Modern multimodal wireless foundation models integrate various input modalities, including CSI, IQ samples, CIR, spectrograms, RF signals, localization information, and environmental observations\. These complementary modalities capture different aspects of the wireless environment, allowing the model to jointly exploit spatial, temporal, spectral, and contextual information that cannot be fully represented by a single data source\[[129](https://arxiv.org/html/2608.14694#bib.bib150)\]\. Representative architectures demonstrate different approaches to multimodal representation learning\. WavesFM adopts a shared Vision Transformer \(ViT\) backbone capable of processing image\-like wireless representations, such as CSI matrices, spectrograms, and OFDM resource grids, while employing LoRA for parameter\-efficient task adaptation\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]\. WirelessGPT, in contrast, focuses on learning generalized channel representations through large\-scale self\-supervised pre\-training, enabling efficient transfer across multiple communication and sensing tasks\[[171](https://arxiv.org/html/2608.14694#bib.bib151)\]\. More recent multimodal wireless foundation models further extend this paradigm by incorporating complementary information from radar, LiDAR, cameras, GPS, network telemetry, and environmental measurements, thereby improving situational awareness and enabling joint communication, sensing, and localization\[[75](https://arxiv.org/html/2608.14694#bib.bib20),[181](https://arxiv.org/html/2608.14694#bib.bib21)\]\.

Architecturally, multimodal foundation models differ from conventional wireless models because they require dedicated modality encoders together with fusion mechanisms that align heterogeneous feature spaces before learning a shared representation\. Early fusion, late fusion, cross\-attention, and multimodal Transformer encoders are among the most commonly adopted strategies for integrating complementary information while preserving modality\-specific characteristics\. The resulting shared representation enables efficient adaptation to multiple downstream tasks without requiring separate models for each sensing modality\.

Despite their considerable potential, multimodal architectures introduce additional challenges, including modality alignment, synchronization, missing or incomplete observations, increased computational complexity, and limited availability of large\-scale multimodal datasets\. Addressing these issues will require scalable multimodal pre\-training strategies, efficient cross\-modal representation learning, and standardized benchmark datasets\. These developments are expected to play a key role in enabling robust and context\-aware wireless foundation models for future AI\-native 6G systems\.

### IV\-DScalability Considerations

Scalability is a fundamental requirement for practical wireless foundation models because future AI\-native 6G networks must support diverse communication tasks, heterogeneous wireless environments, and resource\-constrained deployment platforms\. Unlike conventional deep learning models designed for a single communication scenario, wireless foundation models are expected to operate across different channel configurations, frequency bands, antenna arrays, sampling rates, and sensing modalities while maintaining efficient training and inference\. Consequently, scalability should be evaluated not only in terms of model accuracy, but also with respect to parameter count, memory consumption, computational complexity, adaptation cost, and inference latency, all of which directly influence practical deployment at cloud servers, base stations, edge nodes, and user equipment\[[26](https://arxiv.org/html/2608.14694#bib.bib34),[66](https://arxiv.org/html/2608.14694#bib.bib45),[40](https://arxiv.org/html/2608.14694#bib.bib47)\]\. Model size and memory consumption remain among the most important scalability considerations\. Larger backbone models generally provide stronger representation learning and improved transferability across wireless tasks; however, they also increase storage requirements, communication overhead, and fine\-tuning costs\.

To address these challenges, recent architectures increasingly combine shared pre\-trained backbones with parameter\-efficient adaptation techniques, allowing multiple downstream tasks to reuse a common model without full retraining\. LoRA, for example, updates only low\-rank parameter components and substantially reduces the number of trainable parameters during adaptation\[[64](https://arxiv.org/html/2608.14694#bib.bib42)\]\. This strategy has been incorporated into WFMs such as WavesFM to enable efficient adaptation across different wireless tasks\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\], while related multimodal frameworks, including MuSE\-FM, further explore efficient adaptation of shared representations across heterogeneous inputs\[[183](https://arxiv.org/html/2608.14694#bib.bib97)\]\. Beyond adaptation, model compression provides another pathway toward practical deployment\. Pruning, quantization, and other lightweight model\-design techniques can reduce memory, computation, and inference overhead, as demonstrated by compact wireless foundation models such as TinyWiFo\[[182](https://arxiv.org/html/2608.14694#bib.bib37)\]and TinyFWFM\[[53](https://arxiv.org/html/2608.14694#bib.bib38)\]\. These techniques are particularly important for deploying WFMs on resource\-constrained edge and IoT devices, where computational and energy budgets are limited\[[181](https://arxiv.org/html/2608.14694#bib.bib21)\]\.

Training scalability presents another important challenge because WFMs require large and heterogeneous datasets together with substantial computational resources during pre\-training\. Variations across propagation environments, hardware platforms, frequency bands, and communication standards further complicate the learning of representations that remain transferable across deployment conditions\. One emerging approach is to incorporate communication\-domain structure directly into the learning process\. Masked channel modeling exploits correlations within wireless channel observations to improve representation learning\[[50](https://arxiv.org/html/2608.14694#bib.bib17)\], while delay\-Doppler\-aware architectures incorporate the underlying structure of time\-varying wireless channels\[[16](https://arxiv.org/html/2608.14694#bib.bib40)\]\. Physics\-aware positional encoding, including adaptive three\-dimensional positional representations, provides another mechanism for embedding spatial and propagation structure into the model architecture\[[179](https://arxiv.org/html/2608.14694#bib.bib41)\]\. Such domain\-aware designs can reduce the burden on purely data\-driven learning by introducing inductive biases that better reflect the physical structure of wireless environments\.

Inference efficiency is equally critical because many wireless decisions must be completed within channel coherence times or stringent scheduling intervals\. Although standard Transformer architectures provide powerful representation\-learning capabilities, the quadratic computational and memory complexity of global self\-attention can become prohibitive for long wireless sequences and high\-dimensional observations\[[156](https://arxiv.org/html/2608.14694#bib.bib43)\]\. Consequently, recent WFMs increasingly explore architectures that reduce inference overhead while preserving the ability to capture long\-range dependencies\. WiMamba, for example, employs state\-space modeling as an alternative to conventional attention for efficient processing of wireless observations\[[140](https://arxiv.org/html/2608.14694#bib.bib36)\]\. ComHymba adopts a hybrid architecture that combines complementary sequence\-modeling mechanisms to balance representation capability and computational efficiency\[[170](https://arxiv.org/html/2608.14694#bib.bib39)\]\. AirFM further incorporates structured and efficient processing tailored to wireless channel representations, including delay\-Doppler\-aware modeling\[[16](https://arxiv.org/html/2608.14694#bib.bib40)\]\. In parallel, lightweight encoders, windowed attention, and patch\-based processing provide additional mechanisms for limiting sequence length and attention cost\. Collectively, these architectural directions indicate a shift toward computation\-aware WFMs that balance representation quality, latency, and memory requirements for practical deployment in AI\-native 6G systems\. Table[I](https://arxiv.org/html/2608.14694#S4.T1)summarizes the principal architectural trade\-offs among representative wireless foundation model designs, highlighting the balance between representation capability, computational efficiency, and deployment suitability across different wireless scenarios\.

TABLE I:Scalability Comparison of Representative Wireless Foundation Model ArchitecturesArchitecture CategoryRepresentative ModelsScalability AdvantagesMain LimitationsTypical DeploymentTransformer / Vision TransformerWirelessGPT, LWM, WavesFMExcellent representation learning; strong multi\-task transfer; captures long\-range spatial and temporal dependenciesLarge parameter count; quadratic self\-attention complexity; high memory consumptionCloud servers, base stations, GPU\-enabled edge nodesMasked AutoencoderWavesFM, WiFo, ContraWiMAELabel\-efficient pre\-training; learns hidden channel structures; improves data efficiencySensitive to masking strategy; requires diverse pre\-training dataChannel estimation, CSI feedback, channel predictionPrompt\-guided Encoder–DecoderMuSE\-FMSupports heterogeneous task formats; enables efficient task adaptationAdditional architectural complexity; prompt design remains challengingMulti\-task wireless intelligenceParameter\-Efficient AdaptationLoRA, Adapter\-based WFMsVery small number of trainable parameters; low adaptation cost; efficient model updatesPerformance may degrade under highly mismatched domainsEdge adaptation, continual learning, model updatesLightweight EncodersTiny\-WiFo, Lightweight Time\-Series FMLow latency; reduced memory footprint; suitable for embedded devicesLimited global context modelingIoT devices, mobile terminals, real\-time inferenceState\-Space / MambaWiMamba, ComHymbaNear\-linear computational complexity; efficient long\-sequence modelingImmature ecosystem; limited wireless benchmarksLong CSI sequences, real\-time wireless inferenceDomain\-informed AttentionAirFM\-DDA, ComHymbaExploits wireless\-domain priors; lower attention complexityRequires domain\-specific preprocessingLarge CSI tensors, physical\-layer tasksCompression TechniquesTiny Federated WFM, Quantized WFMsReduces model size, memory, and energy consumptionPotential accuracy degradation after compressionResource\-constrained edge devices and federated learningMultimodal Foundation ModelsMultimodal WFM, WavesFM, MuSE\-FMJointly learns communication, sensing, and localization representationsHigher computational and memory requirements due to modality fusionISAC, AI\-native 6G, semantic communicationsWhile each architectural family offers unique advantages, no single architecture is universally optimal for all wireless applications\. Transformer\-based models provide the strongest representation learning capability and multi\-task transferability by capturing long\-range spatial and temporal dependencies, but their quadratic self\-attention complexity limits efficient deployment on resource\-constrained devices\. CNN\-based architectures remain attractive for applications requiring low computational complexity and real\-time inference, although they are less effective in modeling global wireless dependencies\. Physics\-informed architectures improve data efficiency, interpretability, and robustness by incorporating communication\-domain knowledge into the learning process, but they may sacrifice flexibility when operating under highly diverse propagation environments\. Multimodal architectures further enhance representation quality by jointly processing heterogeneous wireless observations; however, this comes at the cost of increased computational complexity and larger data requirements\. Consequently, the choice of architecture should be guided by the target application, available computational resources, and deployment constraints rather than by representation accuracy alone\.

TABLE II:Comparison of major wireless foundation model architectures\.Table[II](https://arxiv.org/html/2608.14694#S4.T2)highlights that the architectural design of a wireless foundation model involves balancing representation capability, computational efficiency, data availability, and deployment requirements\. Transformer\-based models are well suited for large\-scale pre\-training and multi\-task learning, whereas CNN\-based and physics\-informed architectures remain attractive for latency\-sensitive and resource\-constrained applications\. Multimodal architectures provide the richest representations for AI\-native 6G systems but require significantly larger datasets and computational resources\. Future wireless foundation models are therefore expected to combine the strengths of multiple architectural paradigms rather than relying on a single backbone\.

## VPre\-Training Strategies

### V\-ASelf\-Supervised Learning

SSL has become the dominant pre\-training strategy for wireless foundation models because it enables representation learning from large volumes of unlabeled wireless data\. Unlike supervised learning, which depends on manually annotated datasets, SSL generates supervisory signals directly from the input data through carefully designed pretext tasks\. This paradigm is particularly attractive for wireless communications, where raw measurements such as CSI, IQ samples, CIR, and spectrograms are readily available, whereas obtaining accurate labels is expensive and time\-consuming\[[80](https://arxiv.org/html/2608.14694#bib.bib73),[110](https://arxiv.org/html/2608.14694#bib.bib74)\]\. The primary objective of SSL is to learn generalized wireless representations that capture the spatial, temporal, and frequency\-domain characteristics of radio signals\. After large\-scale pre\-training, these representations can be efficiently transferred to downstream tasks, including channel estimation, channel prediction, signal detection, beam prediction, localization, and wireless sensing, using only limited labeled data\. Consequently, SSL significantly reduces annotation costs while improving robustness and generalization across diverse wireless environments\.

Among the various SSL objectives, masked signal modeling has emerged as one of the most effective approaches for wireless foundation models\. Inspired by masked language modeling and masked image modeling, portions of the wireless input are intentionally hidden, and the model is trained to reconstruct the missing information\. Depending on the application, the masked regions may correspond to CSI elements, IQ samples, OFDM resource grids, or time\-frequency patches\. By reconstructing these missing measurements, the model learns the intrinsic spatial, temporal, and spectral structure of wireless signals without requiring explicit supervision\. Representative examples include WavesFM and scalable masked channel models, which employ reconstruction\-based pre\-training for channel estimation, CSI feedback, and channel prediction\[[56](https://arxiv.org/html/2608.14694#bib.bib75),[50](https://arxiv.org/html/2608.14694#bib.bib17)\]\. Another widely adopted SSL objective is contrastive representation learning, which learns discriminative feature embeddings by maximizing the similarity between different augmented views of the same wireless sample while separating unrelated samples in the latent space\. Positive pairs are typically generated through signal augmentations or multiple observations of the same propagation environment, whereas negative pairs correspond to unrelated wireless measurements\.

This objective encourages invariant representation learning and improves robustness to noise, mobility, and distribution shifts\. Recent studies have demonstrated its effectiveness for CSI representation learning, localization, and RF signal classification\[[24](https://arxiv.org/html/2608.14694#bib.bib76),[48](https://arxiv.org/html/2608.14694#bib.bib77)\]\. Rather than relying on a single objective, contemporary wireless foundation models frequently combine masked reconstruction and contrastive learning during pre\-training\. Reconstruction objectives encourage the model to capture the structural characteristics of wireless signals, whereas contrastive objectives improve representation discrimination and transferability\. This combination has become the prevailing pre\-training strategy for modern wireless foundation models\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/fig5S1.png)Figure 5:Self\-supervised pre\-training workflow for wireless foundation models\.As illustrated in Fig\.[5](https://arxiv.org/html/2608.14694#S5.F5), heterogeneous wireless measurements are first collected and transformed through signal\-domain augmentation to generate multiple views of the same wireless observation\. The augmented data are then used to optimize complementary self\-supervised objectives, including masked signal reconstruction and contrastive representation learning, enabling the model to capture both structural and discriminative characteristics of wireless signals\. The resulting shared wireless representation forms the pre\-trained backbone of the foundation model and can subsequently be adapted to downstream tasks through lightweight techniques such as LoRA, adapters, prompt tuning, or task\-specific prediction heads\. This unified pre\-training paradigm substantially reduces the dependence on labeled datasets while improving scalability and transferability across a broad range of wireless communication applications\.

### V\-BData Generation and Simulation\-Based Pre\-Training

The effectiveness of wireless foundation models depends critically on the availability of large\-scale and diverse pre\-training datasets\. However, collecting real\-world wireless measurements is expensive, time\-consuming, and often constrained by deployment cost, hardware availability, and privacy considerations\. Furthermore, constructing labeled datasets for applications such as channel estimation, beam prediction, and localization requires extensive measurement campaigns under different propagation environments, antenna configurations, and mobility conditions\. Consequently, synthetic data generation has become an indispensable component of the pre\-training pipeline for wireless foundation models\[[121](https://arxiv.org/html/2608.14694#bib.bib140),[139](https://arxiv.org/html/2608.14694#bib.bib78)\]\.

Simulation\-based pre\-training relies on realistic wireless channel simulators to generate large numbers of channel realizations under diverse propagation conditions\. Classical stochastic channel models, including Rayleigh, Rician, Nakagami\-mm, WINNER II, and the 3GPP spatial channel model, are widely used to emulate fading, path loss, shadowing, Doppler effects, and multipath propagation\. These simulators produce CSI, CIR, and received signal samples over a broad range of SNRs, carrier frequencies, antenna configurations, and mobility scenarios, enabling foundation models to learn from considerably more diverse wireless environments than would be feasible through field measurements alone\[[120](https://arxiv.org/html/2608.14694#bib.bib79),[93](https://arxiv.org/html/2608.14694#bib.bib80)\]\. Deterministic ray\-tracing has further enhanced the realism of synthetic wireless datasets by explicitly modeling electromagnetic wave propagation within site\-specific environments\. Unlike stochastic channel models, ray\-tracing captures reflections, diffraction, scattering, blockage, and other propagation phenomena arising from the physical geometry of buildings and surrounding objects\. Modern ray\-tracing platforms therefore provide highly realistic datasets for millimeter\-wave, terahertz, localization, beam management, and ISAC applications\[[141](https://arxiv.org/html/2608.14694#bib.bib81),[63](https://arxiv.org/html/2608.14694#bib.bib82)\]\. Rather than relying exclusively on either simulated or measured data, many recent wireless foundation models adopt hybrid pre\-training strategies that combine both sources\. Synthetic datasets provide broad coverage across propagation conditions and communication scenarios, whereas real\-world measurements reduce the simulation\-to\-reality \(Sim2Real\) gap and improve deployment robustness\. By systematically varying channel models, antenna arrays, carrier frequencies, mobility patterns, hardware impairments, and interference conditions, simulation\-based pre\-training substantially improves representation diversity while reducing overfitting to a particular wireless environment\[[145](https://arxiv.org/html/2608.14694#bib.bib173)\]\.

Despite these advantages, simulation\-based pre\-training also presents important challenges\. Synthetic datasets inevitably simplify real propagation environments and cannot fully capture hardware non\-idealities, environmental dynamics, or unexpected interference sources\. Bridging the resulting Sim2Real gap remains an active research topic, motivating the development of higher\-fidelity simulators, domain adaptation techniques, and hybrid datasets that combine simulation with real\-world measurements\[[11](https://arxiv.org/html/2608.14694#bib.bib174)\]\. Figure[6](https://arxiv.org/html/2608.14694#S5.F6)illustrates the general workflow of simulation\-based pre\-training for wireless foundation models\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/figS5b.png)Figure 6:Simulation\-based data generation and pre\-training pipeline for wireless foundation models\.As illustrated in Fig\.[6](https://arxiv.org/html/2608.14694#S5.F6), the workflow begins with stochastic channel modeling and deterministic ray\-tracing, which generate large\-scale synthetic wireless measurements under configurable propagation environments\. The generated CSI, IQ samples, CIR, and spectrograms are subsequently augmented and integrated with real\-world measurements to form a comprehensive pre\-training corpus\. This hybrid dataset enables foundation models to learn generalized wireless representations that transfer effectively across different channel conditions, antenna configurations, frequency bands, and mobility scenarios\. Following large\-scale self\-supervised pre\-training, the learned representations can be efficiently adapted to downstream tasks, including channel estimation, channel prediction, beam prediction, localization, and integrated sensing, thereby providing a scalable and cost\-effective alternative to relying solely on extensive real\-world measurement campaigns\.

### V\-CTransfer Learning and Fine\-Tuning

Transfer learning enables a pre\-trained wireless foundation model to adapt its learned representations to new communication tasks using significantly less labeled data than training from scratch\. Instead of independently developing models for channel estimation, signal detection, beam prediction, or localization, a single pre\-trained backbone can be efficiently transferred to multiple downstream applications through task\-specific adaptation\. This paradigm substantially reduces training cost, improves data efficiency, and accelerates deployment across diverse wireless environments\[[131](https://arxiv.org/html/2608.14694#bib.bib83)\]\.

The most direct adaptation strategy is full fine\-tuning, in which all model parameters are updated using task\-specific data\. Although this approach generally provides the highest task\-specific performance, it also requires considerable computational resources, memory, and storage because a separate model must be maintained for each downstream application\[[61](https://arxiv.org/html/2608.14694#bib.bib85)\]\. Such requirements become increasingly impractical for AI\-native 6G systems that are expected to support numerous communication and sensing tasks simultaneously\. To improve adaptation efficiency, recent research has focused on parameter\-efficient fine\-tuning \(PEFT\), where only a small subset of parameters is optimized while the majority of the pre\-trained backbone remains fixed\. Representative techniques include LoRA, adapters, prefix tuning, and bias\-only optimization\. These methods significantly reduce computational complexity, communication overhead, and memory consumption while preserving most of the performance achieved by full fine\-tuning\[[60](https://arxiv.org/html/2608.14694#bib.bib86),[97](https://arxiv.org/html/2608.14694#bib.bib87)\]\. For example, WavesFM adopts LoRA modules to efficiently adapt a shared backbone across communication, sensing, and localization tasks without maintaining multiple independently trained models\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]\.

### V\-DPrompting and In\-Context Learning

Prompt\-based learning has recently emerged as an alternative adaptation strategy that further reduces the computational cost of transferring foundation models to downstream tasks\. Instead of modifying the model parameters, prompt tuning introduces a small set of learnable prompt embeddings or task instructions that guide the pre\-trained model toward a target application while preserving the knowledge acquired during pre\-training\[[107](https://arxiv.org/html/2608.14694#bib.bib88)\]\. Representative approaches include soft prompt tuning, which learns continuous prompt embeddings, and prefix tuning, which injects trainable vectors into intermediate Transformer layers without updating the backbone parameters\[[98](https://arxiv.org/html/2608.14694#bib.bib89),[97](https://arxiv.org/html/2608.14694#bib.bib87)\]\. Another emerging paradigm is in\-context learning, where a pre\-trained model performs a new task by conditioning on a small number of demonstration examples rather than updating its parameters\. Although originally proposed for large language models, this capability has attracted growing interest in wireless communications because it enables rapid adaptation to unseen channel conditions, network configurations, or communication objectives with little or no retraining\[[20](https://arxiv.org/html/2608.14694#bib.bib90),[38](https://arxiv.org/html/2608.14694#bib.bib91)\]\. Such flexibility is particularly attractive for AI\-native wireless systems operating in highly dynamic environments\.

These adaptation strategies represent a progression from computationally intensive retraining toward increasingly lightweight and flexible learning mechanisms\. As wireless foundation models continue to evolve, parameter\-efficient fine\-tuning, prompt\-based adaptation, and in\-context learning are expected to become key technologies for enabling scalable multi\-task wireless intelligence across heterogeneous AI\-native 6G systems\.

## VIApplications of Wireless Foundation Models

### VI\-AChannel Estimation and Equalization

Channel estimation and equalization are among the most mature application domains of wireless foundation models because they directly determine the reliability, throughput, and spectral efficiency of wireless communication systems\. Conventional deep learning approaches generally train separate models for specific channel models, antenna configurations, or SNR ranges\. Consequently, their performance often degrades under unseen propagation environments, mobility conditions, or hardware configurations, requiring extensive retraining and limiting their scalability in practical wireless networks\[[44](https://arxiv.org/html/2608.14694#bib.bib139)\]\.

Wireless foundation models overcome these limitations by learning generalized channel representations through large\-scale self\-supervised pre\-training\. Instead of optimizing for a single propagation scenario, the pre\-trained backbone captures common spatial, temporal, and frequency\-domain characteristics shared across diverse wireless environments\. These learned representations can subsequently be adapted to channel estimation, channel prediction, CSI feedback, and equalization using only limited task\-specific supervision, thereby improving generalization while significantly reducing retraining cost\[[77](https://arxiv.org/html/2608.14694#bib.bib143),[10](https://arxiv.org/html/2608.14694#bib.bib69)\]\.

Recent studies demonstrate the effectiveness of this paradigm\. WirelessGPT employs a Transformer\-based backbone to learn reusable channel representations that support multiple physical\-layer tasks, while the LWM exploits large\-scale channel datasets to improve robustness across heterogeneous propagation environments\. Similarly, WavesFM combines masked signal modeling with a Vision Transformer architecture to jointly support channel estimation, CSI feedback, localization, and sensing using a shared pre\-trained model\. More recently, channel foundation models have incorporated masked channel reconstruction and contrastive learning objectives to further improve robustness under channel variations and limited labeled data\[[50](https://arxiv.org/html/2608.14694#bib.bib17),[134](https://arxiv.org/html/2608.14694#bib.bib175)\]\. Channel equalization has also benefited from transferable wireless representations\. Rather than designing separate equalizers for individual channel conditions, wireless foundation models learn generalized propagation characteristics that enable more robust compensation for channel distortion before signal demodulation\. Such shared representations improve equalization performance across varying propagation environments, mobility conditions, and hardware impairments while reducing the need for repeated model retraining\[[148](https://arxiv.org/html/2608.14694#bib.bib176)\]\. As wireless communication systems continue to evolve toward highly dynamic AI\-native 6G networks, wireless foundation models are expected to provide unified physical\-layer intelligence capable of jointly supporting channel estimation, equalization, prediction, CSI compression, and related signal processing tasks within a single scalable framework\.

### VI\-BSignal Detection and Modulation Recognition

Signal detection and modulation recognition represent another important application domain of wireless foundation models\. Accurate signal detection is essential for recovering transmitted symbols in MIMO and OFDM systems\[[90](https://arxiv.org/html/2608.14694#bib.bib184)\], whereas automatic modulation recognition supports adaptive communication, spectrum monitoring, cognitive radio, and intelligent wireless network management\. Conventional deep learning solutions typically train separate models for individual modulation formats, antenna configurations, or channel conditions, limiting their ability to generalize across heterogeneous wireless environments and communication standards\[[35](https://arxiv.org/html/2608.14694#bib.bib93)\]\.

Wireless foundation models address these limitations by learning generalized signal representations from large collections of unlabeled IQ samples, spectrograms, and channel measurements\. Through large\-scale self\-supervised pre\-training, the model captures invariant signal characteristics that remain robust under noise, fading, interference, and hardware impairments\. These representations can subsequently be adapted to signal detection, modulation classification, wireless technology recognition, and RF fingerprinting using only limited task\-specific supervision, substantially reducing the dependence on large labeled datasets\[[117](https://arxiv.org/html/2608.14694#bib.bib16),[27](https://arxiv.org/html/2608.14694#bib.bib94)\]\. Representative studies demonstrate the effectiveness of this approach\. IQFM learns universal feature representations directly from raw IQ streams, enabling efficient transfer across multiple wireless signal processing tasks\. Similarly, unified wireless foundation models for wireless technology recognition and localization employ a shared pre\-trained backbone that simultaneously supports signal classification and positioning without requiring independent task\-specific networks\. Compared with conventional deep learning models\[[143](https://arxiv.org/html/2608.14694#bib.bib183)\], these approaches exhibit improved robustness under unseen propagation environments because they learn reusable wireless representations rather than features tailored to a single dataset or communication scenario\[[122](https://arxiv.org/html/2608.14694#bib.bib177)\]\.

Beyond traditional communication systems, wireless foundation models are expected to play an increasingly important role in cognitive radio, spectrum sensing, RF environment awareness, and intelligent network monitoring\. Their ability to rapidly adapt to new modulation formats, communication protocols, and spectrum conditions makes them well suited for AI\-native 6G systems, where communication environments evolve continuously and require scalable, transferable, and context\-aware wireless intelligence\.

### VI\-CBeamforming and Resource Allocation

Beamforming and resource allocation are fundamental components of modern wireless communication systems because they directly influence network capacity, spectral efficiency, energy efficiency, and QoS\[[96](https://arxiv.org/html/2608.14694#bib.bib187)\]\. Conventional approaches formulate these problems as non\-convex optimization tasks that require accurate CSI together with iterative optimization algorithms\. As wireless networks evolve toward massive MIMO, millimeter\-wave \(mmWave\) communications, ultra\-dense deployments, and ISAC, solving these optimization problems in real time becomes increasingly challenging\[[18](https://arxiv.org/html/2608.14694#bib.bib95),[87](https://arxiv.org/html/2608.14694#bib.bib145),[149](https://arxiv.org/html/2608.14694#bib.bib96)\]\. Wireless foundation models offer a scalable alternative by learning generalized representations of wireless channels, user mobility, traffic patterns, and network dynamics through large\-scale pre\-training\. Instead of optimizing beamforming vectors or resource allocation policies independently for each deployment scenario, a shared pre\-trained backbone can be efficiently adapted to different propagation environments, antenna configurations, and service requirements using lightweight adaptation techniques\[[185](https://arxiv.org/html/2608.14694#bib.bib178)\]\. This paradigm reduces computational complexity while improving robustness and transferability across heterogeneous wireless networks\.

Recent studies demonstrate the effectiveness of this approach\. MuSE\-FM jointly exploits wireless measurements and environmental context to support multi\-task communication optimization, while wireless foundation models for multi\-task prediction learn shared representations for traffic forecasting, mobility prediction, and resource allocation\. More recently, large AI models have explored intent\-driven wireless optimization by incorporating natural\-language reasoning into resource management and autonomous network control, providing a new direction for intelligent radio access networks\[[114](https://arxiv.org/html/2608.14694#bib.bib138),[150](https://arxiv.org/html/2608.14694#bib.bib27),[65](https://arxiv.org/html/2608.14694#bib.bib98)\]\. Compared with conventional optimization methods that solve individual problems independently, wireless foundation models enable unified decision\-making across multiple network functions\. By jointly optimizing beamforming, power allocation, interference mitigation, spectrum management, and user scheduling within a single pre\-trained framework, they provide a scalable foundation for adaptive and autonomous wireless resource management\. This capability is expected to become increasingly important as AI\-native 6G systems evolve toward fully self\-optimizing communication networks\.

### VI\-DLocalization and Integrated Sensing

Wireless localization and integrated sensing have emerged as key application domains for wireless foundation models because future 6G networks are expected to simultaneously provide high\-precision positioning, environmental perception, and reliable communication through ISAC\[[144](https://arxiv.org/html/2608.14694#bib.bib180)\]\. Achieving these capabilities requires the joint processing of heterogeneous wireless measurements, including CSI, CIR, radar echoes, and IQ signals, under dynamic propagation environments characterized by mobility, blockage, and NLOS conditions\[[166](https://arxiv.org/html/2608.14694#bib.bib99),[104](https://arxiv.org/html/2608.14694#bib.bib100)\]\. Wireless foundation models address these challenges by learning unified representations from multimodal wireless observations through self\-supervised pre\-training\. Rather than developing separate models for localization, sensing, and communication, a shared backbone captures common spatial, temporal, and propagation characteristics that can be efficiently adapted to multiple downstream tasks\. This unified learning paradigm improves robustness across diverse environments while reducing the dependence on large annotated positioning datasets\.

Recent studies demonstrate the effectiveness of this approach\. WavesFM employs a Vision Transformer backbone with parameter\-efficient adaptation to jointly support communication, localization, and sensing\. Similarly, multimodal wireless foundation models combine CSI, IQ signals, radar measurements, and environmental information to enhance positioning accuracy and sensing performance\. More recent frameworks further integrate multimodal sensing and communication within a unified architecture, enabling joint localization, environmental perception, and channel prediction using a single pre\-trained model\[[6](https://arxiv.org/html/2608.14694#bib.bib9),[176](https://arxiv.org/html/2608.14694#bib.bib144)\]\. The integration of localization and sensing within a common foundation model represents an important step toward context\-aware wireless intelligence\. By jointly exploiting heterogeneous sensing modalities, wireless foundation models can support applications such as autonomous driving, digital twins, smart manufacturing, and intelligent transportation systems, where communication, positioning, and environmental awareness must operate seamlessly\. This unified capability is expected to become a core building block of AI\-native 6G networks, enabling scalable and adaptive ISAC services across diverse deployment scenarios\.

### VI\-ESemantic and Goal\-Oriented Communications

Semantic and goal\-oriented communications have emerged as key paradigms for next\-generation wireless networks by shifting the communication objective from faithfully transmitting every bit to delivering information that is most relevant for a target task\[[52](https://arxiv.org/html/2608.14694#bib.bib189)\]\. Unlike conventional communication systems that optimize metrics such as bit error rate \(BER\), throughput, or spectral efficiency, semantic communications focus on preserving the meaning of transmitted information, whereas goal\-oriented communications further optimize transmission according to the receiver’s task or decision objective\[[136](https://arxiv.org/html/2608.14694#bib.bib103),[86](https://arxiv.org/html/2608.14694#bib.bib146)\]\. Wireless foundation models provide a natural framework for this paradigm because they learn high\-level representations that capture relationships among wireless signals, environmental context, user behavior, and network states\. Rather than processing communication signals independently, a shared pre\-trained model can jointly encode semantic information across heterogeneous modalities, enabling more efficient compression, robust semantic encoding, and adaptive decision\-making while reducing unnecessary wireless transmissions\.

Recent studies have demonstrated the potential of large foundation models for semantic\-aware wireless systems\. large language models \(LLMs\) have been explored for user\-intent understanding, semantic reasoning, and autonomous network management, while multimodal wireless foundation models combine wireless measurements with vision, radar, and environmental information to support semantic communication and integrated sensing\[[89](https://arxiv.org/html/2608.14694#bib.bib192)\]\. These capabilities improve communication efficiency in applications such as autonomous driving, intelligent transportation, industrial automation, extended reality \(XR\), and digital twins, where successful task completion is often more important than accurate symbol recovery\[[99](https://arxiv.org/html/2608.14694#bib.bib104),[74](https://arxiv.org/html/2608.14694#bib.bib105),[85](https://arxiv.org/html/2608.14694#bib.bib147)\]\. Although semantic wireless foundation models remain at an early stage of development, they represent an important evolution beyond conventional physical\-layer optimization\. Future research will require unified semantic representations for wireless data, evaluation metrics that quantify task effectiveness rather than bit\-level accuracy, and closer integration between semantic reasoning and physical\-layer signal processing\[[69](https://arxiv.org/html/2608.14694#bib.bib195)\]\. These advances are expected to establish semantic and goal\-oriented communications as one of the defining applications of wireless foundation models in AI\-native 6G systems\.

The representative applications discussed above demonstrate that wireless foundation models are evolving from specialized physical\-layer solutions into unified learning frameworks capable of supporting communication, sensing, localization, and network optimization\. Despite sharing the common objective of learning reusable wireless representations, existing models differ considerably in their architectures, pre\-training strategies, adaptation mechanisms, supported protocol layers, and target applications\. Table[III](https://arxiv.org/html/2608.14694#S6.T3)summarizes representative wireless foundation models and highlights the key design trends driving this evolution\.

TABLE III:Comparison of Representative Wireless Foundation ModelsFoundation ModelArchitecturePre\-training StrategyDatasetWireless LayerApplicationTransfer LearningEvaluation MetricMain ContributionWirelessGPT\[[172](https://arxiv.org/html/2608.14694#bib.bib8)\]TransformerSelf\-supervisedCSI, IQPHYChannel estimation, detectionFine\-tuningNMSE, BERGeneral\-purpose Transformer backbone for multi\-task wireless communications\.Large Wireless Model \(LWM\)\[[10](https://arxiv.org/html/2608.14694#bib.bib69)\]TransformerSSLLarge wireless channel datasetPHYChannel predictionFine\-tuningNMSELearns transferable channel representations across propagation environments\.WavesFM\[[6](https://arxiv.org/html/2608.14694#bib.bib9)\]Vision TransformerMasked ModelingCSI, IQ, SpectrogramsPHYChannel estimation, localization, sensingLoRANMSE, Localization ErrorUnified foundation model supporting communication, sensing, and localization\.MuSE\-FM\[[150](https://arxiv.org/html/2608.14694#bib.bib27)\]TransformerMulti\-modal SSLCSI \+ Environmental DataCross\-layerBeamforming, Resource AllocationLoRASpectral EfficiencyEnvironment\-aware multi\-task optimization framework\.WiFo\[[103](https://arxiv.org/html/2608.14694#bib.bib19)\]TransformerMasked Channel ModelingCSIPHYChannel predictionFine\-tuningNMSELong\-term channel prediction using pre\-trained representations\.CSI2Vec\[[128](https://arxiv.org/html/2608.14694#bib.bib15)\]Transformer EncoderContrastive LearningCSIPHYLocalizationFine\-tuningPosition ErrorUniversal CSI embedding for localization and channel charting\.IQFM\[[117](https://arxiv.org/html/2608.14694#bib.bib16)\]TransformerSelf\-supervisedIQ StreamsPHYSignal DetectionFine\-tuningBERLearns transferable representations directly from raw IQ samples\.Unified WFM\[[27](https://arxiv.org/html/2608.14694#bib.bib94)\]TransformerSSLWireless Technology DatasetPHYTechnology RecognitionFine\-tuningClassification AccuracyJoint wireless technology recognition and localization\.Multimodal WFM\[[4](https://arxiv.org/html/2608.14694#bib.bib101)\]Multi\-modal TransformerMulti\-modal SSLCSI, IQ, RadarCross\-layerISAC, LocalizationLoRADetection AccuracyJoint communication and sensing representation learning\.WiFo\-MISAC\[[111](https://arxiv.org/html/2608.14694#bib.bib33)\]TransformerMulti\-modal SSLWireless Sensing DatasetCross\-layerIntegrated Sensing and CommunicationLoRADetection AccuracyUnified multimodal sensing and communication foundation model\.AirFM\-DDA\[[15](https://arxiv.org/html/2608.14694#bib.bib70)\]Physics\-informed TransformerSelf\-supervisedDelay–Doppler Channel DatasetPHYChannel EstimationPrompt / Fine\-tuningNMSEIntroduces physics\-aware positional encoding for wireless foundation models\.

Table[III](https://arxiv.org/html/2608.14694#S6.T3)reveals several important trends in the current development of wireless foundation models\. First, Transformer\-based architectures have become the dominant backbone owing to their ability to model long\-range spatial and temporal dependencies while supporting scalable multi\-task learning\. Although convolutional and physics\-informed architectures remain valuable for specific applications, most recent foundation models favor Transformer variants because of their flexibility across heterogeneous wireless modalities\.

Second, self\-supervised learning has emerged as the predominant pre\-training paradigm\. Techniques such as masked modeling, contrastive learning, and multimodal representation learning effectively exploit the abundance of unlabeled wireless measurements, reducing the dependence on costly annotated datasets while improving model generalization across diverse deployment scenarios\.

Third, adaptation strategies are shifting from conventional full\-parameter fine\-tuning toward parameter\-efficient methods, particularly LoRA\. These lightweight techniques significantly reduce computational and memory overhead while enabling a single pre\-trained backbone to support multiple downstream communication and sensing tasks\.

Finally, the scope of wireless foundation models is expanding beyond traditional physical\-layer applications\. While early studies primarily focused on channel estimation, prediction, and signal detection, more recent models increasingly support beamforming, resource allocation, localization, ISAC, semantic communications, and cross\-layer network intelligence\. This evolution reflects a broader transition from task\-specific deep learning models toward unified wireless intelligence capable of serving diverse communication, sensing, and network management functions within a common foundation\-model framework\.

## VIIDatasets, Benchmarks, and Evaluation Metrics

### VII\-APublic Wireless Datasets

Large\-scale datasets are the cornerstone of wireless foundation models because they provide the diverse observations required for large\-scale pre\-training, benchmarking, and downstream adaptation\. Unlike conventional supervised learning, which relies on task\-specific labeled datasets, wireless foundation models benefit from heterogeneous collections of CSI, IQ samples, CIR, spectrograms, localization measurements, sensing observations, and network telemetry collected under diverse propagation environments\[[175](https://arxiv.org/html/2608.14694#bib.bib196)\]\. The diversity of these datasets enables the learning of transferable wireless representations that generalize across communication tasks, deployment scenarios, and network configurations\.

Several public datasets have become standard benchmarks for wireless AI research\. DeepMIMO is one of the most widely adopted ray\-tracing\-based datasets for millimeter\-wave and massive MIMO communications\. Generated using Remcom Wireless InSite, it provides configurable channel measurements suitable for beam prediction, channel estimation, localization, and deep learning\-based wireless communications\[[13](https://arxiv.org/html/2608.14694#bib.bib106)\]\. Raymobtime further extends this concept by incorporating user mobility, making it particularly valuable for evaluating beam tracking and channel prediction algorithms under dynamic propagation conditions\[[92](https://arxiv.org/html/2608.14694#bib.bib107)\]\.

Public datasets have also been developed for signal recognition and multimodal wireless intelligence\. RadioML provides labeled IQ samples covering multiple modulation formats across different SNRs, making it one of the most widely used benchmarks for automatic modulation recognition and RF signal classification\[[127](https://arxiv.org/html/2608.14694#bib.bib108)\]\. More recently, DeepSense 6G has introduced a multimodal benchmark that combines wireless communication, sensing, localization, and environmental information, reflecting the growing emphasis on AI\-native 6G applications\[[180](https://arxiv.org/html/2608.14694#bib.bib109)\]\. Similarly, ViWi integrates wireless channel measurements with visual information to support vision\-aided beamforming and localization\[[12](https://arxiv.org/html/2608.14694#bib.bib110)\], while Sionna RT generates physics\-based synthetic datasets through differentiable ray tracing for wireless channel modeling and sensing applications\[[62](https://arxiv.org/html/2608.14694#bib.bib111)\]\.

Despite their importance, current wireless datasets remain significantly smaller and less diverse than the datasets used to train foundation models in natural language processing and computer vision\. Most existing benchmarks focus on specific communication tasks, frequency bands, or propagation environments, limiting their ability to support truly general\-purpose wireless intelligence\. Future wireless foundation models will therefore require large\-scale multimodal datasets that combine synthetic channel simulations, real\-world measurements, sensing observations, environmental context, and network management information under standardized benchmark protocols\. Such datasets will play a critical role in improving model generalization, enabling fair comparison among wireless foundation models, and accelerating the development of scalable AI\-native 6G systems\. Table[IV](https://arxiv.org/html/2608.14694#S7.T4)summarizes representative public datasets currently used for the development and evaluation of wireless foundation models\.

TABLE IV:Representative Public Datasets for Wireless Foundation ModelsDatasetModalitySourceRepresentative ApplicationsTypical MetricsKey FeaturesDeepMIMO\[[13](https://arxiv.org/html/2608.14694#bib.bib106)\]CSIRay tracingBeam prediction, channel estimation, localizationNMSE, Spectral EfficiencyConfigurable mmWave channel dataset with multiple scenarios\.Raymobtime\[[92](https://arxiv.org/html/2608.14694#bib.bib107)\]CSIRay tracing \+ MobilityBeam tracking, channel predictionNMSECaptures temporal channel evolution under user mobility\.RadioML\[[127](https://arxiv.org/html/2608.14694#bib.bib108)\]IQ SamplesSimulationModulation recognition, RF classificationAccuracyWidely adopted benchmark for wireless signal classification\.DeepSense 6G\[[180](https://arxiv.org/html/2608.14694#bib.bib109)\]CSI, Radar, Sensor DataReal \+ SimulationLocalization, ISACLocalization Error, Detection AccuracyLarge\-scale multimodal benchmark for AI\-native 6G\.ViWi\[[12](https://arxiv.org/html/2608.14694#bib.bib110)\]CSI \+ ImagesRay tracingVision\-aided beamforming, localizationBeam Prediction AccuracyCombines wireless channels with visual information\.Sionna RT\[[62](https://arxiv.org/html/2608.14694#bib.bib111)\]CSI, CIRDifferentiable Ray TracingChannel estimation, localization, sensingNMSEPhysics\-aware synthetic dataset generation framework\.
### VII\-BPerformance Metrics

Evaluating wireless foundation models requires a broader set of metrics than those traditionally used in wireless communications\. Unlike conventional communication algorithms that are optimized for a single task, wireless foundation models are expected to support multiple downstream applications while maintaining strong generalization, computational efficiency, and adaptability across heterogeneous wireless environments\[[173](https://arxiv.org/html/2608.14694#bib.bib197)\]\. Consequently, their evaluation should simultaneously consider communication performance, representation quality, transferability, and deployment efficiency\. For physical\-layer applications, the most commonly adopted metrics remain the BER and the normalized mean square error \(NMSE\)\. BER quantifies the reliability of signal detection by measuring the proportion of incorrectly detected bits, whereas NMSE evaluates the accuracy of estimated channel coefficients relative to the ground truth\. Lower BER and NMSE values generally indicate improved communication reliability and channel estimation performance\[[45](https://arxiv.org/html/2608.14694#bib.bib57)\]\.

Higher\-layer applications are typically evaluated using network\-oriented metrics, including spectral efficiency, energy efficiency, throughput, latency, and localization error\. Spectral efficiency, usually expressed in bits/s/Hz, measures the effectiveness of spectrum utilization, whereas localization accuracy is commonly quantified using the RMSE or the mean positioning error\. For ISAC, additional sensing metrics such as detection probability, false alarm rate, and sensing accuracy are widely adopted\[[104](https://arxiv.org/html/2608.14694#bib.bib100)\]\. Beyond application\-specific performance, wireless foundation models must also be evaluated according to their ability to generalize across diverse deployment scenarios\.

Unlike conventional deep learning models, foundation models are expected to transfer knowledge across different propagation environments, carrier frequencies, antenna configurations, mobility conditions, and communication tasks\. Consequently, cross\-domain evaluation, few\-shot adaptation, transfer learning performance, and out\-of\-distribution \(OOD\) robustness have become essential benchmark criteria for measuring the effectiveness of transferable wireless representations\[[168](https://arxiv.org/html/2608.14694#bib.bib198)\]\. Computational efficiency represents another critical evaluation dimension because wireless foundation models are expected to operate across cloud servers, edge platforms, and resource\-constrained user devices\. Common efficiency metrics include the number of trainable parameters, floating\-point operations \(FLOPs\), inference latency, memory consumption, and energy usage\[[67](https://arxiv.org/html/2608.14694#bib.bib199)\]\. In addition, parameter\-efficient adaptation methods, such as LoRA and adapter\-based tuning, are often assessed by comparing the percentage of trainable parameters and computational overhead relative to conventional full fine\-tuning\[[186](https://arxiv.org/html/2608.14694#bib.bib200)\]\.

Table[V](https://arxiv.org/html/2608.14694#S7.T5)summarizes the most commonly adopted evaluation metrics for wireless foundation models, together with their primary objectives and representative application domains\.

TABLE V:Common Evaluation Metrics for Wireless Foundation Models
### VII\-CGeneralization and Robustness Evaluation

Unlike conventional deep learning models that are typically evaluated under fixed training and testing conditions, wireless foundation models are expected to operate across diverse propagation environments, network configurations, hardware platforms, and communication standards\. Consequently, evaluating only application\-specific metrics such as BER or NMSE is insufficient\[[132](https://arxiv.org/html/2608.14694#bib.bib201)\]\. A comprehensive evaluation must also assess the model’s ability to generalize across unseen wireless scenarios while maintaining robustness against distribution shifts, environmental variations, and adversarial perturbations\.

Generalization evaluation measures how effectively a pre\-trained wireless foundation model transfers to deployment conditions that differ from those encountered during pre\-training\. Representative evaluation scenarios include changes in SNR, carrier frequency, antenna configuration, user mobility, propagation environment, and hardware impairments\. Since wireless channels are inherently dynamic, cross\-domain evaluation has become an essential benchmark for measuring transferability across heterogeneous communication environments\. Consequently, transfer learning performance, few\-shot adaptation, and cross\-scenario generalization are increasingly adopted as key evaluation criteria for wireless foundation models\[[7](https://arxiv.org/html/2608.14694#bib.bib5),[123](https://arxiv.org/html/2608.14694#bib.bib84)\]\. OOD testing provides a complementary assessment by intentionally separating the training and testing distributions\. Typical OOD scenarios include previously unseen propagation environments, different antenna arrays, new modulation formats, changing mobility patterns, and hardware variations\. Models that maintain stable performance under these conditions are considered more suitable for practical AI\-native wireless systems, where operating conditions continuously evolve\[[105](https://arxiv.org/html/2608.14694#bib.bib51),[57](https://arxiv.org/html/2608.14694#bib.bib112)\]\.

Robustness evaluation further examines the resilience of wireless foundation models against noise, interference, adversarial attacks, and hardware non\-idealities\. Adversarial robustness is commonly evaluated using attack methods such as the Fast Gradient Sign Method \(FGSM\) and PGD, together with measurements of performance degradation under different perturbation strengths\. In practical wireless deployments, robustness is also assessed under channel estimation errors, synchronization offsets, phase noise, nonlinear hardware distortion, quantization effects, and imperfect channel state information, providing a more realistic measure of deployment reliability\[[47](https://arxiv.org/html/2608.14694#bib.bib113),[115](https://arxiv.org/html/2608.14694#bib.bib114)\]\. Although considerable progress has been achieved, standardized evaluation protocols for wireless foundation models remain limited\. Existing studies often employ different datasets, channel models, and experimental settings, making direct comparison difficult\. Future benchmark suites should therefore integrate multi\-domain datasets, standardized OOD evaluation, adversarial robustness testing, continual learning scenarios, and computational efficiency metrics to provide a comprehensive and reproducible assessment of wireless foundation models\.

## VIIIOpen Challenges and Research Directions

### VIII\-AData Scarcity and Domain Shift

The development of wireless foundation models is fundamentally constrained by the availability of large\-scale, diverse, and representative wireless datasets\. Unlike natural language processing and computer vision, where foundation models are trained using billions of publicly available text documents and images, wireless communications lack standardized datasets of comparable scale and diversity\[[68](https://arxiv.org/html/2608.14694#bib.bib202)\]\. Existing public datasets are typically collected for specific communication tasks, propagation environments, carrier frequencies, or hardware platforms, limiting their suitability for learning truly general\-purpose wireless representations\. Consequently, many current wireless foundation models rely heavily on synthetic datasets generated using stochastic channel models or ray\-tracing simulators, which cannot fully capture the complexity of real\-world wireless environments\[[162](https://arxiv.org/html/2608.14694#bib.bib204)\]\.

An even more fundamental challenge arises from the non\-stationary nature of wireless environments\. Wireless channels continuously evolve due to user mobility, environmental dynamics, network reconfiguration, hardware impairments, and spectrum utilization\. As a result, the statistical distribution of wireless data changes over time, violating the independent and identically distributed \(IID\) assumption underlying most current pre\-training strategies\[[178](https://arxiv.org/html/2608.14694#bib.bib203)\]\. Although large\-scale pre\-training improves representation learning, it cannot completely eliminate the performance degradation caused by distribution shifts between the pre\-training and deployment environments\. For example, a model trained on urban sub\-6 GHz channels may perform poorly when deployed in millimeter\-wave, terahertz, satellite, or industrial communication scenarios without additional adaptation\.

This challenge distinguishes wireless foundation models from their counterparts in natural language processing and computer vision\. While linguistic structures and visual features remain relatively stable across domains, wireless signals are governed by physical propagation mechanisms, antenna configurations, operating frequencies, and hardware characteristics that vary substantially across deployment scenarios\. Consequently, simply increasing dataset size is insufficient for achieving robust generalization\. Future wireless foundation models must instead learn representations that remain invariant to environmental changes while preserving information relevant to downstream communication tasks\.

Current research primarily addresses domain shift through transfer learning, domain adaptation, and parameter\-efficient fine\-tuning\. Although these approaches improve adaptation efficiency, they remain reactive because they require additional data from the target domain after deployment\. A more scalable research direction is proactive domain generalization, where the pre\-training objective explicitly encourages invariant representation learning across diverse propagation environments before deployment\. Integrating self\-supervised learning with physics\-informed constraints, causal representation learning, and meta\-learning may further improve the ability of wireless foundation models to generalize across previously unseen scenarios\[[83](https://arxiv.org/html/2608.14694#bib.bib205),[123](https://arxiv.org/html/2608.14694#bib.bib84)\]\. Another critical challenge is the lack of standardized benchmark datasets and evaluation protocols\. Existing studies employ different channel models, simulation platforms, antenna configurations, mobility patterns, and experimental settings, making direct comparison among wireless foundation models difficult\. Similar to the role of ImageNet in computer vision, the wireless community requires open, large\-scale benchmark datasets that integrate synthetic simulations, real\-world measurements, multimodal sensing observations, and network management information under standardized evaluation methodologies\. Such benchmarks would not only enable fair comparison among competing models but also accelerate the development of reproducible and scalable wireless foundation models\.

Looking ahead, future dataset development should extend beyond communication signals alone\. AI\-native wireless networks increasingly generate heterogeneous information, including CSI, IQ samples, radar observations, LiDAR data, environmental maps, mobility traces, traffic statistics, and network telemetry\. Constructing multimodal datasets that jointly capture communication, sensing, localization, and networking information will therefore be a key prerequisite for developing truly general\-purpose wireless foundation models capable of supporting diverse wireless intelligence tasks within a unified learning framework\.

### VIII\-BInterpretability and Physical Consistency

The remarkable performance of wireless foundation models has been driven by increasingly expressive neural architectures, particularly large Transformer\-based models\. However, this improved representation capability is accompanied by reduced interpretability\. Unlike conventional communication algorithms, whose behavior can be explained through estimation theory, optimization, or information theory, the internal decision\-making process of large pre\-trained models remains largely opaque\. Consequently, it is often difficult to determine whether a model has learned physically meaningful propagation characteristics or merely exploited statistical correlations present in the training data\[[142](https://arxiv.org/html/2608.14694#bib.bib115),[147](https://arxiv.org/html/2608.14694#bib.bib116)\]\. This limitation is particularly significant in wireless communications because wireless systems are governed by well\-established physical principles\. Channel reciprocity, electromagnetic propagation, antenna geometry, and environmental interactions impose constraints that should be respected by any learning\-based solution\. Models that violate these physical properties may achieve strong performance on benchmark datasets while exhibiting poor reliability when deployed under realistic operating conditions\. Therefore, unlike foundation models developed for natural language processing or computer vision, wireless foundation models must satisfy both statistical learning objectives and communication\-theoretic constraints\.

Current wireless foundation models are primarily optimized using data\-driven objectives, such as masked reconstruction and contrastive learning, without explicitly enforcing physical consistency\. Although these objectives improve downstream task performance, they do not guarantee that the learned latent representations preserve meaningful channel characteristics or propagation behavior\. As model size and complexity continue to increase, this discrepancy between statistical optimization and physical consistency may become an important obstacle to reliable deployment across heterogeneous wireless environments\. A promising research direction is the integration of physics\-informed learning into large\-scale pre\-training\. Rather than treating communication theory and deep learning as independent paradigms, future wireless foundation models should incorporate analytical knowledge directly into their architectures, training objectives, and adaptation strategies\. Examples include embedding channel reciprocity, propagation constraints, delay\-Doppler sparsity, antenna array geometry, and optimization objectives into the representation learning process\. Such physics\-aware inductive biases can improve both generalization and data efficiency while encouraging physically consistent model behavior\[[81](https://arxiv.org/html/2608.14694#bib.bib117),[152](https://arxiv.org/html/2608.14694#bib.bib72)\]\.

Interpretability itself remains another open research challenge\. Existing explainable artificial intelligence \(XAI\) techniques, including saliency maps, feature attribution, and attention visualization, provide only limited insight into the physical meaning of learned wireless representations\. Future research should therefore develop wireless\-specific interpretability methods capable of linking latent features to measurable physical phenomena, such as dominant propagation paths, multipath components, beam directions, or interference patterns\. Such capabilities would not only increase confidence in model predictions but also facilitate debugging, system verification, and communication algorithm design\.

Looking ahead, interpretability and physical consistency should become fundamental evaluation criteria rather than optional properties\. Beyond conventional performance metrics such as BER and NMSE, future benchmark frameworks should quantify physical plausibility, constraint satisfaction, uncertainty calibration, and decision reliability under realistic deployment conditions\. Establishing standardized interpretability benchmarks will be essential for developing trustworthy wireless foundation models capable of supporting safety\-critical AI\-native 6G applications, including autonomous transportation, industrial automation, and integrated sensing and communication\.

### VIII\-CRobustness and Security

The deployment of wireless foundation models in future AI\-native wireless networks introduces security and robustness challenges that extend beyond those encountered in conventional communication systems\. Unlike task\-specific deep learning models, a single foundation model may simultaneously support communication, sensing, localization, and network management\[[146](https://arxiv.org/html/2608.14694#bib.bib181)\]\. Consequently, vulnerabilities affecting the shared backbone can propagate across multiple downstream applications, making robustness and security fundamental design requirements rather than optional deployment considerations\[[133](https://arxiv.org/html/2608.14694#bib.bib118),[23](https://arxiv.org/html/2608.14694#bib.bib119)\]\. One of the most important challenges is adversarial robustness\. Deep neural networks are known to be vulnerable to carefully crafted perturbations that can significantly alter model predictions while remaining difficult to detect\. In wireless systems, such perturbations may arise from malicious signal injections, spoofing attacks, jamming, or manipulated CSI\. These attacks can degrade channel estimation, mislead beam prediction, disrupt signal detection, and compromise localization accuracy\[[109](https://arxiv.org/html/2608.14694#bib.bib182)\]\. Since foundation models learn shared representations for multiple tasks, adversarial perturbations introduced during inference may affect a broader range of applications than in conventional task\-specific models\[[91](https://arxiv.org/html/2608.14694#bib.bib120)\]\.

Robustness must also be evaluated under realistic wireless operating conditions rather than only adversarial attacks\. Practical communication systems are affected by channel estimation errors, synchronization offsets, carrier frequency offset \(CFO\), phase noise, nonlinear power amplifier distortion, quantization effects, and hardware mismatches that are often absent from simulation\-based training datasets\. Models trained under limited propagation conditions may therefore experience substantial performance degradation when deployed in real networks\. Incorporating such non\-idealities into pre\-training and evaluation pipelines represents an important step toward improving deployment reliability\. Another critical challenge concerns the integrity of the pre\-training pipeline\. Wireless foundation models rely on extremely large datasets collected from heterogeneous devices, sensors, and communication infrastructures, making them vulnerable to data poisoning, malicious dataset contamination, and label manipulation\. Unlike conventional supervised learning, where corrupted samples primarily affect a single downstream model, compromised pre\-training data may influence the shared representations learned by the entire foundation model, potentially affecting all subsequent applications\. Developing trusted data collection mechanisms and effective poisoning detection methods therefore remains an important research direction\[[17](https://arxiv.org/html/2608.14694#bib.bib121),[70](https://arxiv.org/html/2608.14694#bib.bib122)\]\.

Privacy protection represents an additional challenge because wireless foundation models may be trained using sensitive communication measurements, user mobility traces, localization information, and network telemetry\. Without appropriate safeguards, pre\-trained models may unintentionally reveal information about the training data through model inversion or membership inference attacks\. Federated learning, secure aggregation, differential privacy, and trusted execution environments have therefore emerged as promising techniques for privacy\-preserving foundation model training and deployment\[[119](https://arxiv.org/html/2608.14694#bib.bib123),[41](https://arxiv.org/html/2608.14694#bib.bib124)\]\. Looking ahead, robustness and security should become intrinsic objectives throughout the entire lifecycle of wireless foundation models, from data collection and pre\-training to adaptation and deployment\. Future research should jointly optimize communication performance, robustness, interpretability, privacy, and security rather than treating them as independent design objectives\. Integrating adversarial training, uncertainty estimation, physics\-informed learning, secure federated optimization, and continual adaptation offers a promising path toward trustworthy wireless foundation models capable of supporting safety\-critical AI\-native 6G applications\.

### VIII\-DEnergy Efficiency and Edge Deployment

Although wireless foundation models have demonstrated remarkable performance across a wide range of communication tasks, their large computational and memory requirements remain a major obstacle to practical deployment\. Most existing models contain millions or even billions of parameters, requiring substantial computational resources for both pre\-training and inference\. Although cloud infrastructures can support models of this scale, many future wireless applications, including autonomous vehicles, industrial automation, UAVs, and IoT networks, require real\-time intelligence on edge devices operating under stringent latency, energy, and hardware constraints\[[71](https://arxiv.org/html/2608.14694#bib.bib194)\]\. Consequently, improving deployment efficiency without sacrificing representation quality has become a critical research challenge\[[54](https://arxiv.org/html/2608.14694#bib.bib125),[25](https://arxiv.org/html/2608.14694#bib.bib126)\]\. A fundamental limitation arises from the mismatch between model complexity and device capability\. Large Transformer\-based architectures are designed to maximize representational capacity, whereas edge devices are constrained by limited processing power, memory, storage, and battery capacity\. Simply deploying cloud\-scale foundation models at the network edge is therefore impractical because communication latency, transmission overhead, and energy consumption may outweigh the benefits of local intelligence\. This observation suggests that deployment efficiency should be treated as a primary design objective rather than an optimization applied after model development\.

To address this challenge, recent research has explored lightweight model adaptation and compression techniques, including model pruning, quantization, knowledge distillation, low\-rank approximation, and PEFT\. Methods such as LoRA and adapter modules allow multiple downstream tasks to share a common pre\-trained backbone while updating only a small subset of parameters, substantially reducing computational cost and memory requirements\[[59](https://arxiv.org/html/2608.14694#bib.bib127)\]\. Nevertheless, their effectiveness under highly dynamic wireless environments, where communication tasks and network conditions continuously evolve, remains insufficiently understood\. Distributed and federated learning provide another promising direction for scalable deployment\. Instead of relying exclusively on centralized cloud infrastructures, future wireless networks may collaboratively train and adapt foundation models across edge servers, base stations, and user devices\. Such decentralized learning can reduce communication overhead, improve privacy preservation, and enable localized adaptation\. However, heterogeneous device capabilities, intermittent connectivity, limited wireless bandwidth, and synchronization overhead introduce new optimization challenges that are largely absent from conventional centralized training\[[119](https://arxiv.org/html/2608.14694#bib.bib123),[25](https://arxiv.org/html/2608.14694#bib.bib126)\]\.

Energy efficiency should likewise become an explicit optimization objective throughout the model lifecycle\. Existing wireless foundation models are primarily optimized for communication performance and transferability, whereas computational energy is often considered only during deployment\. Future research should jointly optimize model architecture, pre\-training strategy, adaptation mechanism, and inference scheduling to balance communication accuracy with computational efficiency\. Hardware\-aware neural architecture search, dynamic model scaling, early\-exit inference, and energy\-aware scheduling represent promising directions for developing adaptive foundation models capable of adjusting their computational complexity according to available hardware resources and application requirements\[[153](https://arxiv.org/html/2608.14694#bib.bib128),[102](https://arxiv.org/html/2608.14694#bib.bib129)\]\. Looking ahead, scalable deployment will require a holistic co\-design of algorithms, communication systems, and hardware platforms\. Rather than treating edge deployment as a model compression problem alone, future research should develop integrated software\-hardware frameworks that jointly optimize latency, energy consumption, memory usage, communication overhead, and model accuracy\. Such co\-design principles will be essential for enabling practical, sustainable, and energy\-efficient wireless foundation models capable of supporting large\-scale AI\-native 6G networks\.

### VIII\-EStandardization and Practical Deployment

Despite the rapid progress of wireless foundation models, their transition from research prototypes to operational wireless systems remains limited\. Most existing studies emphasize algorithm development and simulation\-based validation, whereas practical deployment requires standardized model interfaces, interoperable architectures, unified evaluation methodologies, and compatibility with existing wireless communication standards\. Without these supporting frameworks, comparing different wireless foundation models, reproducing experimental results, and integrating foundation models into commercial communication systems remain challenging\.

One of the most significant barriers is the absence of standardized datasets and benchmark suites specifically designed for wireless foundation models\. Existing studies employ different channel models, simulation platforms, antenna configurations, carrier frequencies, and evaluation protocols, making direct performance comparison difficult\. Unlike computer vision, where ImageNet established a common benchmark for foundation model development, wireless communications still lack a universally accepted large\-scale benchmark capable of simultaneously supporting communication, sensing, localization, and network management tasks\. Establishing open benchmark datasets together with standardized evaluation methodologies will therefore be essential for enabling reproducible research and accelerating the development of scalable wireless foundation models\[[157](https://arxiv.org/html/2608.14694#bib.bib141),[118](https://arxiv.org/html/2608.14694#bib.bib142)\]\.

Interoperability represents another important deployment challenge\. Future AI\-native wireless networks will consist of heterogeneous devices, multiple radio access technologies, cloud\-edge collaboration, and distributed intelligence across the communication infrastructure\. Wireless foundation models must therefore operate seamlessly across different hardware platforms and communication standards while maintaining consistent performance\. Achieving this objective requires standardized model interfaces, common representation formats, and efficient mechanisms for model exchange, compression, and adaptation across heterogeneous network entities\.

Another practical challenge concerns model lifecycle management\. Unlike conventional communication algorithms, foundation models require continual updates to accommodate evolving propagation environments, new spectrum bands, emerging communication services, and changing network topologies\[[72](https://arxiv.org/html/2608.14694#bib.bib191)\]\. Consequently, future deployment frameworks should support secure model distribution, version control, continual adaptation, validation, rollback mechanisms, and compatibility with legacy communication systems throughout the model lifecycle\. Practical deployment also requires trustworthy AI frameworks that address reliability, transparency, robustness, privacy, and security\. Since wireless foundation models may support safety\-critical applications such as autonomous transportation, industrial automation, healthcare, and public safety, future standardization efforts should incorporate explainability, uncertainty estimation, robustness evaluation, and privacy\-preserving learning as integral components rather than optional enhancements\.

Looking ahead, standardization should evolve alongside technological innovation rather than follow it\. Close collaboration among academia, industry, and standardization organizations, including the 3rd Generation Partnership Project \(3GPP\), the International Telecommunication Union \(ITU\), the European Telecommunications Standards Institute \(ETSI\), and the O\-RAN Alliance will be essential for defining common datasets, benchmark methodologies, deployment architectures, AI\-native interfaces, and interoperability requirements\. Such coordinated efforts will provide the technological foundation for scalable, interoperable, and trustworthy wireless foundation models capable of supporting future AI\-native 6G systems\.

## IXFuture Roadmap Toward AI\-Native 6G Systems

Wireless foundation models are expected to become one of the key enabling technologies for AI\-native 6G wireless networks\. Unlike current wireless systems, where artificial intelligence primarily assists individual communication tasks, future wireless infrastructures are expected to employ foundation models as shared intelligence layers capable of jointly supporting communication, sensing, localization, network optimization, and autonomous management\. Realizing this vision requires coordinated advances in data infrastructure, model architectures, deployment frameworks, and standardization\.

![Refer to caption](https://arxiv.org/html/2608.14694v1/figSaN.png)Figure 7:Future roadmap toward AI\-native 6G systems\. The roadmap illustrates the expected evolution of wireless foundation models from large\-scale data collection and self\-supervised pre\-training to multimodal wireless intelligence, cloud\-edge collaboration, autonomous network management, and fully AI\-native 6G communication systems\.As illustrated in Fig\.[7](https://arxiv.org/html/2608.14694#S9.F7), the near\-term phase \(2026\-2028\) is expected to focus on establishing the foundations of wireless intelligence through the construction of large\-scale multimodal datasets, standardized benchmark suites, and scalable self\-supervised pre\-training strategies\. During this stage, advances in simulation\-assisted data generation, parameter\-efficient adaptation, and benchmark standardization are expected to improve representation learning while reducing the dependence on task\-specific labeled datasets\.

The medium\-term phase \(2028\-2031\) is anticipated to witness the emergence of general\-purpose wireless foundation models capable of supporting multiple communication, sensing, and networking tasks using a unified backbone\. Progress in physics\-informed learning, continual adaptation, distributed training, and cloud\-edge collaboration will enable more efficient deployment across heterogeneous wireless environments\. At the same time, multimodal representation learning will become increasingly important as communication signals are jointly processed with sensing observations, environmental context, network telemetry, and user behavior\.

In the long term \(2031\-2035\), wireless foundation models are expected to evolve into autonomous wireless intelligence engines capable of continuously learning from network observations and adapting to changing operating conditions with minimal human intervention\. Instead of optimizing communication modules independently, future systems will jointly perform communication, sensing, localization, mobility management, resource allocation, and network orchestration using shared representations\. Such capabilities will enable intent\-driven networking, semantic communications, digital twins, ISAC, and autonomous radio access networks that continuously optimize their operation without extensive human supervision\.

Achieving this vision requires advances beyond model scaling alone\. Future research must simultaneously address trustworthy and explainable AI, continual learning, privacy\-preserving collaborative training, energy\-efficient edge intelligence, interoperable deployment frameworks, and standardized evaluation methodologies\. Progress in these complementary areas will determine whether wireless foundation models can successfully transition from promising research prototypes to practical communication infrastructures\.

Ultimately, the evolution toward AI\-native 6G represents a paradigm shift from task\-specific optimization to general\-purpose wireless intelligence\. Rather than serving as independent learning models for individual communication functions, wireless foundation models are expected to become shared knowledge engines that continuously learn, adapt, and collaborate across heterogeneous wireless environments\. Such an evolution has the potential to transform future communication systems into scalable, autonomous, and trustworthy intelligent networks capable of supporting the diverse services envisioned for beyond\-5G and 6G ecosystems\.

## XConclusion

WFMs represent a significant evolution in the application of artificial intelligence to wireless communications\. By replacing isolated task\-specific models with large\-scale pre\-trained models capable of learning transferable wireless representations, WFMs offer a unified framework for communication, sensing, localization, and network intelligence\. This paradigm enables improved generalization across heterogeneous environments, reduces dependence on large labeled datasets, and supports efficient adaptation to diverse downstream wireless tasks, making it a promising foundation for AI\-native 6G systems\.

This survey presented a comprehensive review of the emerging WFM landscape from the perspectives of model architectures, learning paradigms, deployment across the wireless protocol stack, datasets, pre\-training strategies, adaptation techniques, applications, and evaluation methodologies\. Rather than viewing these components independently, the survey highlighted how scalable pre\-training, self\-supervised learning, parameter\-efficient adaptation, multimodal representation learning, and physics\-informed modeling collectively form the technological foundation of next\-generation wireless intelligence\. The discussion also demonstrated the ongoing transition from narrowly optimized communication algorithms toward unified, reusable models capable of supporting multiple wireless functions within a single learning framework\.

Despite the rapid progress in this field, several fundamental challenges remain before WFMs can be deployed at scale\. These include constructing large, diverse, and standardized wireless datasets, improving robustness under distribution shifts, integrating communication\-domain knowledge into foundation models, enabling efficient cloud\-edge deployment, and establishing reliable evaluation protocols and benchmarks\. Addressing these challenges will require closer collaboration between the wireless communications, machine learning, networking, and standardization communities to develop interoperable datasets, reproducible benchmarks, trustworthy AI frameworks, and scalable deployment strategies\.

Looking forward, future research is expected to move beyond single\-modality and task\-specific learning toward multimodal, physics\-aware, and continually adaptive wireless foundation models capable of reasoning across communication, sensing, localization, and network management\. As these capabilities mature, WFMs are expected to become the intelligence backbone of AI\-native 6G networks, enabling autonomous, context\-aware, and self\-optimizing wireless systems\. Their successful development will fundamentally reshape wireless network design, shifting from collections of independently optimized algorithms to unified, scalable intelligence platforms that continuously learn, adapt, and evolve with dynamic communication environments\.

## References

- \[1\]\(2026\)Paradigm shift toward distributed learning in iot intelligence: a comprehensive survey of opportunities and challenges\.IEEE Internet of Things Journal\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p3.1)\.
- \[2\]A\. A\. Abdullah, A\. Zubiaga, S\. Mirjalili, A\. H\. Gandomi, F\. Daneshfar, M\. Amini, A\. S\. Mohammed, and H\. Veisi\(2025\)Evolution of meta’s llama models and parameter\-efficient fine\-tuning of large language models: a survey\.arXiv preprint arXiv:2510\.12178\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p2.1)\.
- \[3\]M\. Abo Zahhad, E\. Ismail, A\. Ali,et al\.\(2026\)Enhancing wireless physical layer performance with deep\-learning techniques: a comprehensive review\.JES\. Journal of Engineering Sciences54\(3\),pp\. 110–138\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p2.1)\.
- \[4\]A\. Aboulfotouh and H\. Abou\-Zeid\(2025\)Multimodal wireless foundation models\.arXiv preprint\.Cited by:[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.10.10.1.1.1)\.
- \[5\]A\. Aboulfotouh and H\. Abou\-Zeid\(2026\)Multimodal wireless foundation models\.InICC 2026\-IEEE International Conference on Communications,pp\. 1–6\.Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p5.1),[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1)\.
- \[6\]A\. Aboulfotouh, E\. Mohammed, and H\. Abou\-Zeid\(2025\)6G wavesfm: a foundation model for sensing, communication, and localization\.IEEE Open Journal of the Communications Society\(\),pp\.\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2025.3600616)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p4.1),[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p4.1),[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p3.1),[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1),[§V\-C](https://arxiv.org/html/2608.14694#S5.SS3.p2.1),[§VI\-D](https://arxiv.org/html/2608.14694#S6.SS4.p2.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.4.4.1.1.1)\.
- \[7\]M\. Akrout, A\. Feriani, F\. Bellili, A\. Mezghani, and E\. Hossain\(2023\)Domain generalization in machine learning models for wireless communications: concepts, state\-of\-the\-art, and open issues\.\(\),pp\. 1–22\.External Links:[Document](https://dx.doi.org/)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p2.1),[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p2.1)\.
- \[8\]S\. Al\-Shareeda, N\. Saeed, K\. Redmill, B\. A\. Jabr, Y\. B\. Salamah, A\. Al\-Dubai, and F\. Ozguner\(2026\)Novel digital twin\-assisted vehicular task offloading for cloud, edge, and hybrid deployments\.IEEE Transactions on Vehicular Technology75\(4\),pp\. 6565–6581\.External Links:[Document](https://dx.doi.org/10.1109/TVT.2025.3620810)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1)\.
- \[9\]Z\. Albataineh, M\. Al Bataineh, M\. F\. Tamimi, and N\. Saeed\(2026\)Robust adaptive beam tracking for terahertz beamspace MIMO: an uncertainty\-aware kalman filter\.IEEE Open Journal of Vehicular Technology7,pp\. 1168–1182\.External Links:[Document](https://dx.doi.org/10.1109/OJVT.2026.3678133)Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p2.1)\.
- \[10\]S\. Alikhani, G\. Charan, and A\. Alkhateeb\(2024\)Large wireless model \(lwm\): a foundation model for wireless channels\.arXiv preprint arXiv:2411\.08872\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p4.1),[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p2.1),[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p3.1),[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p2.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.3.3.1.1.1)\.
- \[11\]T\. Alkhalifah, H\. Wang, and O\. Ovcharenko\(2022\)MLReal: bridging the gap between training on synthetic data and real data applications in machine learning\.Artificial Intelligence in Geosciences3,pp\. 101–114\.Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p3.1)\.
- \[12\]A\. Alkhateebet al\.\(2020\)ViWi: a deep learning dataset framework for vision\-aided wireless communications\.IEEE Vehicular Technology Conference\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p3.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.6.6.1.1.1)\.
- \[13\]A\. Alkhateeb\(2019\)DeepMIMO: a generic deep learning dataset for millimeter wave and massive mimo applications\.Information10\(7\),pp\. 228\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p2.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.2.2.1.1.1)\.
- \[14\]L\. Balaji, A\. Dhanalakshmi, B\. Chempavathy, S\. Aswini, S\. Z\. Parvez, and D\. Sridhar\(2025\)Spatio\-temporal transformer framework for next\-generation wireless channel estimation\.Telecommunications and Radio Engineering\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p4.1)\.
- \[15\]K\. Bian, M\. Tao, J\. Mo, Z\. Chen, and L\. Chen\(2026\)AirFM\-dda: air\-interface foundation model in the delay\-doppler\-angle domain for ai\-native 6g\.arXiv preprint arXiv:2605\.00020\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p4.1),[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p3.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.12.12.1.1.1)\.
- \[16\]K\. Bian, M\. Tao, J\. Mo, Z\. Chen, and L\. Chen\(2026\)AirFM\-dda: air\-interface foundation model in the delay\-doppler\-angle domain for ai\-native 6g\.arXiv preprint arXiv:2605\.00020\.Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p3.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p4.1)\.
- \[17\]B\. Biggio and F\. Roli\(2018\)Wild patterns: ten years after the rise of adversarial machine learning\.Pattern Recognition84,pp\. 317–331\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p2.1)\.
- \[18\]E\. Björnson, J\. Hoydis, and L\. Sanguinetti\(2017\)Massive mimo networks: spectral, energy, and hardware efficiency\.Now Publishers\.Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p1.1)\.
- \[19\]F\. Boccardi, R\. W\. Heath, A\. Lozano, T\. L\. Marzetta, and P\. Popovski\(2014\)Five disruptive technology directions for 5g\.IEEE Communications Magazine52\(2\),pp\. 74–80\.External Links:[Document](https://dx.doi.org/)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p1.1)\.
- \[20\]T\. B\. Brownet al\.\(2020\)Language models are few\-shot learners\.Advances in Neural Information Processing Systems33\.Cited by:[§V\-D](https://arxiv.org/html/2608.14694#S5.SS4.p1.1)\.
- \[21\]D\. Buffelli, S\. Das, Y\. Lin, S\. Vakili, C\. Wang, M\. Attarifar, P\. Nath, and D\. Shiu\(2025\)Towards a foundation model for communication systems\.arXiv preprint arXiv:2505\.14603\.Cited by:[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p1.1)\.
- \[22\]D\. Buffelli, S\. Das, Y\. Lin, S\. Vakili, C\. Wang, M\. Attarifar, P\. Nath, and D\. Shiu\(2025\)Towards a foundation model for communication systems\.arXiv preprint arXiv:2505\.14603\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p4.1)\.
- \[23\]N\. Carlini and D\. Wagner\(2019\)Adversarial machine learning at scale\.IEEE Symposium on Security and Privacy\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p1.1)\.
- \[24\]T\. Chen, S\. Kornblith, M\. Norouzi, and G\. Hinton\(2020\)A simple framework for contrastive learning of visual representations\.InProceedings of the International Conference on Machine Learning \(ICML\),pp\. 1597–1607\.Cited by:[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p3.1)\.
- \[25\]X\. Cheng, B\. Liu, X\. Liu, and X\. Cai\(2024\)Distributed foundation models for multi\-modal learning in 6g wireless networks\.IEEE Wireless Communications31\(3\),pp\. 20–30\.Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p1.1),[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p2.1)\.
- \[26\]X\. Cheng, B\. Liu, X\. Liu, and X\. Cai\(2026\)Large wireless foundation models: stronger over bigger\.arXiv preprint arXiv:2601\.10963\.Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p1.1)\.
- \[27\]M\. Cheraghinia, E\. D\. Poorter, J\. Fontaine,et al\.\(2025\)A unified foundation model for wireless technology recognition and localization\.arXiv preprint\.Cited by:[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p2.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.9.9.1.1.1)\.
- \[28\]M\. Cheraghinia, E\. De Poorter, J\. Fontaine, M\. Debbah, and A\. Shahid\(2025\)A foundation model for wireless technology recognition and localization tasks\.IEEE Open Journal of the Communications Society6,pp\. 9879–9896\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p2.1)\.
- \[29\]M\. Cheraghinia, E\. De Poorter, J\. Fontaine, M\. Debbah, and A\. Shahid\(2025\)A unified foundation model for wireless technology recognition and localization\.arXiv preprint arXiv:2505\.19390\.Cited by:[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p3.1)\.
- \[30\]M\. Cheraghinia, E\. De Poorter, J\. Fontaine, K\. S\. Kim, M\. Debbah, and A\. Shahid\(2025\)Lightweight foundation model for wireless time series downstream tasks on edge devices\.In2025 IEEE Globecom Workshops \(GC Wkshps\),pp\. 104–109\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p4.1)\.
- \[31\]Q\. Cui, X\. You, N\. Wei, G\. Nan, X\. Zhang, J\. Zhang, X\. Lyu, M\. Ai, X\. Tao, Z\. Feng,et al\.\(2025\)Overview of ai and communication for 6g network: fundamentals, challenges, and future research opportunities\.Science China Information Sciences68\(7\),pp\. 171301\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p1.1)\.
- \[32\]L\. Dai, R\. Jiao, F\. Adachi, H\. V\. Poor, and L\. Hanzo\(2020\)Deep learning for wireless communications: an emerging interdisciplinary paradigm\.IEEE Wireless Communications\.Note:Early version available as arXiv:2007\.05952External Links:2007\.05952Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p1.1),[§I](https://arxiv.org/html/2608.14694#S1.p2.1)\.
- \[33\]L\. Dai, R\. Jiao, F\. Adachi, H\. V\. Poor, and L\. Hanzo\(2020\)Deep learning for wireless communications: An emerging interdisciplinary paradigm\.IEEE Communications Magazine\(\),pp\. 1–7\.External Links:[Document](https://dx.doi.org/)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p1.1)\.
- \[34\]Z\. Decan, Y\. Junteng, D\. Zhao, Z\. Lechi, Z\. Xu, C\. Anjie, W\. Lin, C\. Wenchi, D\. Qinghe, and L\. Li\(2026\)AI for wireless waveform recognition: a survey from a component perspective\.Electronics15\(10\),pp\. 2112\.Cited by:[§I\-A](https://arxiv.org/html/2608.14694#S1.SS1.p1.1)\.
- \[35\]S\. R\. Doha and A\. Abdelhadi\(2025\)Deep learning in wireless communication receivers: a survey\.IEEE Access13,pp\. 113586–113606\.Cited by:[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p1.1)\.
- \[36\]S\. R\. Doha and A\. Abdelhadi\(2025\)Deep learning in wireless communication receivers: a survey\.IEEE Access13,pp\. 113586–113605\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p2.1)\.
- \[37\]S\. R\. Doha and A\. Abdelhadi\(2025\)Deep learning in wireless communication receivers: a survey\.IEEE Access13\(\),pp\. 113586–113600\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3584000)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p3.1)\.
- \[38\]Q\. Donget al\.\(2024\)A survey on in\-context learning\.ACM Computing Surveys\.Cited by:[§V\-D](https://arxiv.org/html/2608.14694#S5.SS4.p1.1)\.
- \[39\]A\. Dosovitskiyet al\.\(2021\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p1.1),[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p2.1)\.
- \[40\]J\. Du, T\. Lin, C\. Jiang, Q\. Yang, C\. F\. Bader, and Z\. Han\(2024\)Distributed foundation models for multi\-modal learning in 6g wireless networks\.IEEE Wireless Communications31\(3\),pp\. 20–30\.External Links:[Document](https://dx.doi.org/10.1109/MWC.009.2300501)Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p1.1)\.
- \[41\]C\. Dwork and A\. Roth\(2014\)The algorithmic foundations of differential privacy\.Now Publishers\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p3.1)\.
- \[42\]L\. Ericsson, H\. Gouk, C\. C\. Loy, and T\. M\. Hospedales\(2022\)Self\-supervised representation learning: introduction, advances, and challenges\.IEEE Signal Processing Magazine39\(3\),pp\. 42–62\.Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p1.1)\.
- \[43\]J\. Fontaine, A\. Shahid, and E\. De Poorter\(2024\)Towards a wireless physical\-layer foundation model: challenges and strategies\.In2024 IEEE International Conference on Communications Workshops \(ICC Workshops\),pp\. 1–7\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p2.1)\.
- \[44\]R\. Gadhafi, Y\. Himeur, and J\. Purushothama\(2026\)Machine learning and deep learning for antenna and electromagnetic systems: a structured review of methods, applications, and benchmarking\.IEEE Access\.Cited by:[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p1.1)\.
- \[45\]A\. Goldsmith\(2005\)Overview of wireless communications\.Wireless communications,pp\. 1–26\.Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p1.1),[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p1.1)\.
- \[46\]I\. Goodfellow, Y\. Bengio, A\. Courville, and Y\. Bengio\(2016\)Deep learning\.Vol\.1,MIT press Cambridge\.Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p1.1),[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p2.1)\.
- \[47\]I\. J\. Goodfellow, J\. Shlens, and C\. Szegedy\(2015\)Explaining and harnessing adversarial examples\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p3.1)\.
- \[48\]J\. Grill, F\. Strub, F\. Altché, C\. Tallec, P\. H\. Richemond, E\. Buchatskaya, C\. Doersch, B\. Á\. Pires, Z\. Guo, M\. G\. Azar, B\. Piot, K\. Kavukcuoglu, R\. Munos, and M\. Valko\(2020\)Bootstrap your own latent: a new approach to self\-supervised learning\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\. F\. Balcan, and H\. T\. Lin \(Eds\.\),Vol\.33,pp\. 21271–21284\.Cited by:[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p3.1)\.
- \[49\]B\. Guler, G\. Geraci, and H\. Jafarkhani\(2025\)A multi\-task foundation model for wireless channel representation using contrastive and masked autoencoder learning\.External Links:2505\.09160Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p3.1)\.
- \[50\]J\. Guo, Z\. Deng, Z\. Qiao, J\. Zhang, J\. Xue, D\. Niyato, and Z\. Xu\(2026\)Scalable pre\-trained masked channel model of wireless communications\.IEEE Transactions on Communications74\(\),pp\. 6197–6212\.External Links:[Document](https://dx.doi.org/10.1109/TCOMM.2026.3675420)Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p4.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p3.1),[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p2.1),[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p3.1)\.
- \[51\]S\. Guo, B\. Lu, M\. Wen, S\. Dang, and N\. Saeed\(2022\)Customized 5G and beyond private networks with integrated URLLC, eMBB, mMTC, and positioning for industrial verticals\.IEEE Communications Standards Magazine6\(1\),pp\. 52–58\.External Links:[Document](https://dx.doi.org/10.1109/MCOMSTD.0001.2100041)Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p1.1)\.
- \[52\]S\. Guo, Y\. Wang, S\. Li, and N\. Saeed\(2023\)Semantic importance\-aware communications using pre\-trained language models\.IEEE Communications Letters27\(9\),pp\. 2328–2332\.External Links:[Document](https://dx.doi.org/10.1109/LCOMM.2023.3293805)Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p1.1)\.
- \[53\]M\. Hallaq, F\. M\. A\. Khan, A\. AboElfotouh, S\. A\. Hassan, K\. Dev, M\. T\. Quasim, and H\. Abou\-Zeid\(2025\)Tiny federated wireless foundation models for resource\-constrained devices\.IEEE Internet of Things Journal12\(19\),pp\. 39197–39210\.External Links:[Document](https://dx.doi.org/10.1109/JIOT.2025.3591169)Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1)\.
- \[54\]K\. Han, S\. Han,et al\.\(2024\)Edge ai: on\-demand accelerating deep neural network inference via edge computing\.IEEE Communications Surveys & Tutorials26\(2\),pp\. 1023–1051\.Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p1.1)\.
- \[55\]H\. He, C\. Wen, S\. Jin, and G\. Y\. Li\(2018\)A model\-driven deep learning network for mimo detection\.IEEE Global Communications Conference \(GLOBECOM\)\.Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p2.1),[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p1.1),[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p2.1)\.
- \[56\]K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollár, and R\. Girshick\(2022\)Masked autoencoders are scalable vision learners\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 16000–16009\.Cited by:[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p2.1)\.
- \[57\]D\. Hendrycks and K\. Gimpel\(2021\)A baseline for detecting misclassified and out\-of\-distribution examples in neural networks\.International Conference on Learning Representations \(ICLR\)\.Cited by:[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p2.1)\.
- \[58\]J\. R\. Hershey, J\. L\. Roux, and F\. Weninger\(2014\)Deep unfolding: model\-based inspiration of novel deep architectures\.InarXiv preprint arXiv:1409\.2574,Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p2.1),[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p2.1)\.
- \[59\]G\. Hinton, O\. Vinyals, and J\. Dean\(2015\)Distilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p2.1)\.
- \[60\]N\. Houlsbyet al\.\(2019\)Parameter\-efficient transfer learning for nlp\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§V\-C](https://arxiv.org/html/2608.14694#S5.SS3.p2.1)\.
- \[61\]J\. Howard and S\. Ruder\(2018\-07\)Universal language model fine\-tuning for text classification\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 328–339\.Cited by:[§V\-C](https://arxiv.org/html/2608.14694#S5.SS3.p2.1)\.
- \[62\]J\. Hoydiset al\.\(2023\)Sionna rt: differentiable ray tracing for radio propagation modeling\.arXiv preprint arXiv:2311\.03606\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p3.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.7.7.1.1.1)\.
- \[63\]J\. Hoydiset al\.\(2023\)Sionna rt: differentiable ray tracing for radio propagation modeling\.arXiv preprint arXiv:2311\.03606\.Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p2.1)\.
- \[64\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.Iclr1\(2\),pp\. 3\.Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1)\.
- \[65\]C\. Huang, G\. Chen, P\. Xiao,et al\.\(2026\)Large artificial intelligence models for future wireless communications\.IEEE Communications Magazine\.Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p2.1)\.
- \[66\]C\. Huang, G\. Chen, P\. Xiao, Z\. Han, and R\. Tafazolli\(2026\)Large artificial intelligence models for future wireless communications\.arXiv preprint arXiv:2601\.06906\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p1.1)\.
- \[67\]R\. Huang\(2026\)Edge\-centric generative ai: a survey on efficient inference for large language models in resource\-constrained environments\.Journal of Computer and Communications14\(4\),pp\. 238–253\.Cited by:[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p3.1)\.
- \[68\]A\. Hussain, U\. E\. Farwa, A\. Sikandar, and K\. Hee\-Cheol\(2026\)The rise of foundation models: opportunities, technology, applications, challenges, recent trends, and future directions\.Applied System Innovation9\(2\),pp\. 35\.Cited by:[§VIII\-A](https://arxiv.org/html/2608.14694#S8.SS1.p1.1)\.
- \[69\]N\. Islam and S\. Shin\(2024\)Deep learning in physical layer: review on data driven end\-to\-end communication systems and their enabling semantic applications\.IEEE Open Journal of the Communications Society5,pp\. 4207–4240\.Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p2.1)\.
- \[70\]M\. Jagielskiet al\.\(2021\)Manipulating machine learning: poisoning attacks and countermeasures\.InIEEE Symposium on Security and Privacy,Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p2.1)\.
- \[71\]S\. Javaid, B\. He, G\. Nauryzbayev, and N\. Saeed\(2026\)Neuro\-symbolic AI for UAVs\-based communication networks\.IEEE Communications Standards Magazine\.External Links:[Document](https://dx.doi.org/10.1109/MCOMSTD.2026.3661529)Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p1.1)\.
- \[72\]S\. Javaid, N\. Khan, A\. Alwarafy, and N\. Saeed\(2026\)AGI and LLM\-driven spectrum intelligence in future wireless networks\.IEEE Wireless Communications33\(1\),pp\. 224–236\.External Links:[Document](https://dx.doi.org/10.1109/MWC.2025.3600789)Cited by:[§VIII\-E](https://arxiv.org/html/2608.14694#S8.SS5.p4.1)\.
- \[73\]S\. Javaid and N\. Saeed\(2026\)The post\-electromagnetic era: a vision for wireless communication beyond 6G\.Array29,pp\. 100714\.External Links:[Document](https://dx.doi.org/10.1016/j.array.2026.100714)Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p3.1)\.
- \[74\]F\. Jiang, C\. Pan, L\. Dong,et al\.\(2026\)A comprehensive survey of large ai models for future communications\.IEEE Communications Surveys & Tutorials\.Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p2.1)\.
- \[75\]F\. Jiang, C\. Pan, L\. Dong, K\. Wang, M\. Debbah, D\. Niyato, and Z\. Han\(2026\)A comprehensive survey of large ai models for future communications: foundations, applications and challenges\.IEEE Communications Surveys & Tutorials\.Cited by:[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1)\.
- \[76\]F\. Jiang, C\. Pan, L\. Dong, K\. Wang, M\. Debbah, D\. Niyato, and Z\. Han\(2026\)A comprehensive survey of large ai models for future communications: foundations, applications, and challenges\.IEEE Communications Surveys & Tutorials28\(\),pp\. 4731–4764\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2026.3660844)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p1.1)\.
- \[77\]J\. Jiang, Y\. Gao, X\. Wu, and S\. Xu\(2025\)Towards channel foundation models \(cfms\): motivations, methodologies and opportunities\.arXiv preprint arXiv:2507\.13637\.Cited by:[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p2.1)\.
- \[78\]J\. Jiang, W\. Yu, Y\. Li, Y\. Gao, and S\. Xu\(2025\)A mimo wireless channel foundation model via cir\-csi consistency\.In2025 IEEE International Conference on Machine Learning for Communication and Networking \(ICMLCN\),pp\. 1–6\.Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p3.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p5.1)\.
- \[79\]L\. Jing, T\. Yang, H\. Zhang, Y\. Shi, C\. Zhang, and B\. Zhang\(2026\)Signal compression for wireless communication and sensing: a general approach utilizing pretrained wireless foundation models\.IEEE Transactions on Mobile Computing\.Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p1.1)\.
- \[80\]L\. Jing and Y\. Tian\(2021\)Self\-supervised visual feature learning with deep neural networks: a survey\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(11\),pp\. 4037–4058\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2020.2992393)Cited by:[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p1.1)\.
- \[81\]G\. E\. Karniadakis, I\. G\. Kevrekidis, L\. Lu,et al\.\(2021\)Physics\-informed machine learning\.Nature Reviews Physics3,pp\. 422–440\.External Links:[Document](https://dx.doi.org/10.1038/s42254-021-00314-5)Cited by:[§VIII\-B](https://arxiv.org/html/2608.14694#S8.SS2.p2.1)\.
- \[82\]S\. M\. Kay\(1993\)Fundamentals of statistical signal processing, volume i: estimation theory\.Prentice Hall\.Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p2.1)\.
- \[83\]D\. Khan, M\. M\. Alam, M\. A\. Ouameur, A\. Ullah, I\. U\. Din, M\. Bagaa, D\. Massicotte, and B\. Razmpoosh\(2026\)Machine learning\-aided rf circuit and antenna design: a review on emerging trends and techniques\.IEEE Journal of Selected Topics in Electromagnetics, Antennas and Propagation2\(\),pp\. 131–150\.External Links:[Document](https://dx.doi.org/10.1109/JSTEAP.2026.3700094)Cited by:[§VIII\-A](https://arxiv.org/html/2608.14694#S8.SS1.p4.1)\.
- \[84\]F\. M\. A\. Khan, M\. Hallaq, H\. Abou\-Zeid, O\. Erak, O\. Waqar, S\. A\. Hassan, O\. Alhussein, and E\. Hossain\(2026\)Model compression for sustainable ai in xg wireless networks: recent advances, challenges, and future directions\.IEEE Communications Surveys & Tutorials\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p5.1)\.
- \[85\]N\. Khan, A\. Alwarafy, M\. Hayajneh, F\. Sallabi, and H\. El\-Sayed\(2025\)Digital twin\-based multi\-uav ris network for user fairness and network load balancing using multi\-tasking drl\.In2025 IEEE International Mediterranean Conference on Communications and Networking \(MeditCom\),Vol\.,pp\. 24–29\.External Links:[Document](https://dx.doi.org/10.1109/MeditCom64437.2025.11104471)Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p2.1)\.
- \[86\]N\. Khan, A\. Alwarafy, S\. Mumtaz, M\. Ur\-Rehman, I\. Politis, and N\. Saeed\(2026\-07\-17\)Semantic communication and large language models for agi based resource allocation in future wireless networks\.Discover Internet of Things\.External Links:[Document](https://dx.doi.org/10.1007/s43926-026-00430-7),[Link](https://doi.org/10.1007/s43926-026-00430-7),ISSN 2731\-7503Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p1.1)\.
- \[87\]N\. Khan, A\. Alwarafy, N\. Saeed, Z\. Mehmood, F\. Sallabi, and A\. Ahmad\(2026\)Multi\-relay UAV\-Assisted STAR\-RIS networks for efficient resource allocation: a digital twin\-driven mtdrl approach\.IEEE Open Journal of Vehicular Technology7\(\),pp\. 1151–1167\.External Links:[Document](https://dx.doi.org/10.1109/OJVT.2026.3672597)Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p1.1)\.
- \[88\]W\. U\. Khan, M\. Adil, J\. Malik, C\. K\. Sheemar, S\. Chatzinotas, S\. A\. Alqahtani, and K\. Yahya\(2026\)Multi\-modal foundation models for space\-air\-ground integrated 6g and beyond networks: a survey and tutorial\.IEEE Open Journal of the Communications Society7\(\),pp\. 4623–4659\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2026.3687570)Cited by:[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1)\.
- \[89\]K\. Khurshid, S\. Javaid, and N\. Saeed\(2026\)A communication\-centric 6G–LLM architecture for scalable tactical autonomous defense vehicle networks\.IEEE Network\.Note:Early AccessCited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p2.1)\.
- \[90\]K\. Khurshid and N\. Saeed\(2026\)Performance analysis of hybrid SOR\-MMSE detection for RIS\-assisted cell\-free networks\.IEEE Access14,pp\. 15447–15461\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2026.3658444)Cited by:[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p1.1)\.
- \[91\]J\. Kim, H\. V\. Poor,et al\.\(2023\)Security of deep learning\-based wireless communications: challenges and opportunities\.IEEE Communications Surveys & Tutorials25\(4\),pp\. 2558–2593\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p1.1)\.
- \[92\]A\. Klautauet al\.\(2019\)The raymobtime dataset: enabling deep learning for wireless communications\.IEEE Access7,pp\. 35379–35391\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p2.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.3.3.1.1.1)\.
- \[93\]P\. Kyösti, J\. Meinilä, L\. Hentilä, X\. Zhao, T\. Jämsä, C\. Schneider, M\. Narandžić, M\. Milojević, A\. Hong, J\. Ylitalo, V\. Holappa, M\. Alatossava, R\. Bultitude, Y\. de Jong, and T\. Rautiainen\(2007\-09\)WINNER ii channel models\.Deliverable D1\.1\.2, Version 1\.2IST\-4\-027756 WINNER II Project\.Note:Updated February 2008Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p2.1)\.
- \[94\]Y\. LeCun, Y\. Bengio, and G\. Hinton\(2015\)Deep learning\.Nature521\(7553\),pp\. 436–444\.External Links:[Document](https://dx.doi.org/10.1038/nature14539)Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p2.1)\.
- \[95\]W\. Lee and J\. Park\(2026\)LLM\-empowered resource allocation in wireless communications systems\.IEEE Access14\(\),pp\. 15260–15272\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2026.3655801)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p3.1)\.
- \[96\]H\. Lei, D\. Meng, K\.\-H\. Park, N\. Saeed, and G\. Pan\(2025\)DRL\-based resource allocation for aerial iot systems with no\-fly zones\.IEEE Transactions on Aerospace and Electronic Systems61\(6\),pp\. 17892–17907\.External Links:[Document](https://dx.doi.org/10.1109/TAES.2025.3606915)Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p1.1)\.
- \[97\]B\. Lester, R\. Al\-Rfou, and N\. Constant\(2021\)The power of scale for parameter\-efficient prompt tuning\.EMNLP\.Cited by:[§V\-C](https://arxiv.org/html/2608.14694#S5.SS3.p2.1),[§V\-D](https://arxiv.org/html/2608.14694#S5.SS4.p1.1)\.
- \[98\]X\. L\. Li and P\. Liang\(2021\)Prefix\-tuning: optimizing continuous prompts for generation\.ACL\.Cited by:[§V\-D](https://arxiv.org/html/2608.14694#S5.SS4.p1.1)\.
- \[99\]L\. Liang, H\. Ye, Y\. Sheng,et al\.\(2026\)Large language models for wireless communications: from adaptation to autonomy\.IEEE Communications Magazine\.Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p2.1)\.
- \[100\]L\. Liang, J\. Guo, J\. Zhang, C\. Chae, L\. Lu, S\. Xu, O\. A\. Dobre, S\. Jin, and G\. Y\. Li\(2026\)Foundation models for wireless communications: from phy intelligence to network autonomy\.arXiv preprint arXiv:2606\.06239\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p1.1)\.
- \[101\]L\. Liang, H\. Ye, Y\. Sheng, O\. Wang, J\. Wang, S\. Jin, and G\. Y\. Li\(2026\)Large language models for wireless communications: from adaptation to autonomy\.IEEE Communications Magazine\(\),pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/MCOM.001.2500476)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1)\.
- \[102\]J\. Lin, C\. Gan, and S\. Han\(2020\)Dynamic neural networks: a survey\.InIEEE Transactions on Pattern Analysis and Machine Intelligence,Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p3.1)\.
- \[103\]B\. Liu, S\. Gao, X\. Liu, X\. Cheng, and L\. Yang\(2025\)WiFo: wireless foundation model for channel prediction\.Science China Information Sciences68\(6\),pp\. 162302:1–162302:14\.External Links:[Document](https://dx.doi.org/10.1007/s11432-025-4349-0)Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p4.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.6.6.1.1.1)\.
- \[104\]F\. Liu, C\. Masouros, A\. Petropulu,et al\.\(2024\)Integrated sensing and communications: toward dual\-functional wireless networks for 6g and beyond\.IEEE Journal on Selected Areas in Communications42\(1\),pp\. 1–24\.Cited by:[§VI\-D](https://arxiv.org/html/2608.14694#S6.SS4.p1.1),[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p2.1)\.
- \[105\]J\. Liu, T\. Oyedare, and J\. Park\(2021\)Detecting out\-of\-distribution data in wireless communications applications of deep learning\.IEEE Transactions on Wireless Communications20\(10\),pp\. 6971–6983\.External Links:[Document](https://dx.doi.org/10.1109/TWC.2021.3089135)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p2.1),[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p2.1)\.
- \[106\]J\. Liu, T\. Oyedare, and J\. Park\(2021\)Detecting out\-of\-distribution data in wireless communications applications of deep learning\.IEEE Transactions on Wireless Communications\(\),pp\. 1–13\.External Links:[Document](https://dx.doi.org/)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p2.1)\.
- \[107\]P\. Liuet al\.\(2023\)Pre\-train, prompt, and predict: a systematic survey of prompting methods in natural language processing\.ACM Computing Surveys55\(9\)\.Cited by:[§V\-D](https://arxiv.org/html/2608.14694#S5.SS4.p1.1)\.
- \[108\]P\. Liu, X\. Ren, P\. Wang, H\. Yuan, Z\. Hao, G\. Chen, C\. Xu, D\. Ni, and S\. Cai\(2025\)An efficient graph\-transformer operator for learning physical dynamics with manifolds embedding\.arXiv preprint arXiv:2512\.10227\.Cited by:[§IV](https://arxiv.org/html/2608.14694#S4.p1.1)\.
- \[109\]X\. Liu, L\. Huang, X\. Mei, N\. Saeed, F\. Wang, Y\. Zhang, X\. Ma, and C\. Weng\(2025\)Multi\-strategy fusion enhanced channel estimation algorithm based on deep learning\.Ain Shams Engineering Journal16\(7\),pp\. 103416\.External Links:[Document](https://dx.doi.org/10.1016/j.asej.2025.103416)Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p1.1)\.
- \[110\]X\. Liuet al\.\(2023\)Self\-supervised learning: generative or contrastive\.IEEE Transactions on Knowledge and Data Engineering35\(1\),pp\. 857–876\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2021.3090866)Cited by:[§V\-A](https://arxiv.org/html/2608.14694#S5.SS1.p1.1)\.
- \[111\]X\. Liu, S\. Gao, B\. Liu, X\. Cheng, and L\. Yang\(2026\)WiFo\-misac: a wireless foundation model for multimodal sensing and communication integration via synesthesia of machines \(som\)\.arXiv preprint arXiv:2604\.18255\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.11.11.1.1.1)\.
- \[112\]Y\. Liu, T\. Chang, M\. Hong, Z\. Wu, A\. M\. So, E\. A\. Jorswieck, and W\. Yu\(2024\)A survey of recent advances in optimization methods for wireless communications\.IEEE Journal on Selected Areas in Communications42\(11\),pp\. 2992–3031\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p1.1)\.
- \[113\]Y\. Liu, M\. Chen,et al\.\(2023\)Graph neural networks for wireless communications: from theory to practice\.IEEE Communications Surveys & Tutorials25\(3\),pp\. 1833–1863\.Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p3.1)\.
- \[114\]H\. Luo, Y\. Yan, Y\. Bian, W\. Feng, R\. Zhang, Y\. Liu, J\. Wang, G\. Sun, D\. Niyato, H\. Yu,et al\.\(2026\)Ai reasoning for wireless communications and networking: a survey and perspectives\.ACM Computing Surveys58\(13\),pp\. 1–38\.Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p2.1)\.
- \[115\]A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu\(2018\)Towards deep learning models resistant to adversarial attacks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p3.1)\.
- \[116\]Y\. Mao, W\. Xu, J\. Sang, and H\. Liu\(2026\)BioLAMR: a biomimetically inspired large language model adaptation framework for automatic modulation recognition\.Biomimetics11\(4\),pp\. 288\.Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p1.1)\.
- \[117\]O\. Mashaal and H\. Abou\-Zeid\(2025\)IQFM a wireless foundational model for i/q streams in ai\-native 6g\.External Links:2506\.06718Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p3.1),[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p2.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.8.8.1.1.1)\.
- \[118\]F\. Matera, M\. Settembre, A\. Vizzarri, F\. Mazzenga, J\. Llorca, A\. Tulino, A\. Detti, D\. Ronzani, and S\. Barbarossa\(2024\)Opportunities and challenges for the integration of ai in network evolution toward 6g\.In2024 AEIT International Annual Conference \(AEIT\),pp\. 1–6\.Cited by:[§VIII\-E](https://arxiv.org/html/2608.14694#S8.SS5.p2.1)\.
- \[119\]B\. McMahan, E\. Moore, D\. Ramage,et al\.\(2017\)Communication\-efficient learning of deep networks from decentralized data\.Proceedings of AISTATS\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p3.1),[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p2.1)\.
- \[120\]H\. Miao, J\. Zhang, P\. Tang, Q\. Zhen, J\. Meng, X\. Liu, E\. Liu, P\. Liu, L\. Tian, and G\. Liu\(2025\)6G new mid\-band/fr3 \(6–24 ghz\): channel measurement, characteristics and modeling\.IEEE Open Journal of the Communications Society6\(\),pp\. 9942–9960\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2025.3636972)Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p2.1)\.
- \[121\]A\. F\. Molisch\(2012\)Wireless communications\.Vol\.34,John Wiley & Sons\.Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p1.1)\.
- \[122\]M\. E\. Morocho\-Cayamcela, M\. Maier, and W\. Lim\(2020\)Breaking wireless propagation environmental uncertainty with deep learning\.IEEE Transactions on Wireless Communications19\(8\),pp\. 5075–5087\.Cited by:[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p2.1)\.
- \[123\]C\. T\. Nguyenet al\.\(2023\)Transfer learning for wireless networks: a comprehensive survey\.Journal of Network and Computer Applications214,pp\. 103634\.Cited by:[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p2.1),[§VIII\-A](https://arxiv.org/html/2608.14694#S8.SS1.p4.1)\.
- \[124\]C\. T\. Nguyen, N\. V\. Huynh, N\. H\. Chu, Y\. M\. Saputra, D\. T\. Hoang, D\. N\. Nguyen, Q\. Pham, D\. Niyato, E\. Dutkiewicz, and W\. Hwang\(2021\)Transfer learning for wireless networks: a comprehensive survey\.\(\),pp\. 1–39\.External Links:[Document](https://dx.doi.org/)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p3.1)\.
- \[125\]C\. T\. Nguyen, N\. V\. Huynh, N\. H\. Chu, Y\. M\. Saputra, D\. T\. Hoang, D\. N\. Nguyen, Q\. Pham, D\. Niyato, E\. Dutkiewicz, and W\. Hwang\(2023\)Transfer learning for wireless networks: a comprehensive survey\.Journal of Network and Computer Applications214,pp\. 103634\.External Links:[Document](https://dx.doi.org/10.1016/j.jnca.2022.103634)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p2.1)\.
- \[126\]H\. Noh, B\. Shim, and H\. J\. Yang\(2025\)Adaptive resource allocation optimization using large language models in dynamic wireless environments\.IEEE Transactions on Vehicular Technology74\(10\),pp\. 16630–16635\.External Links:[Document](https://dx.doi.org/10.1109/TVT.2025.3572440)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p3.1)\.
- \[127\]T\. J\. O’Shea, J\. Corgan, and T\. C\. Clancy\(2018\)Convolutional radio modulation recognition networks\.Engineering Applications of Neural Networks\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p3.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.4.4.1.1.1)\.
- \[128\]V\. Palhares, S\. Taner, and C\. Studer\(2025\)CSI2Vec: towards a universal csi feature representation for positioning and channel charting\.External Links:2506\.05237Cited by:[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p3.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.7.7.1.1.1)\.
- \[129\]G\. Pan, Y\. Gao, Y\. Gao, W\. Yu, Z\. Zhong, X\. Yang, X\. Guo, and S\. Xu\(2026\)AI\-driven wireless positioning: fundamentals, standards, state\-of\-the\-art, and challenges\.IEEE Communications Surveys & Tutorials28\(\),pp\. 4394–4428\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2025.3648577)Cited by:[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1)\.
- \[130\]G\. Pan, K\. Huang, H\. Chen, S\. Zhang, C\. Häger, and H\. Wymeersch\(2025\)Large wireless localization model \(lwlm\): a foundation model for positioning in 6g networks\.arXiv preprint arXiv:2505\.10134\.Cited by:[§III](https://arxiv.org/html/2608.14694#S3.p1.1)\.
- \[131\]S\. J\. Pan and Q\. Yang\(2010\)A survey on transfer learning\.IEEE Transactions on Knowledge and Data Engineering22\(10\),pp\. 1345–1359\.Cited by:[§V\-C](https://arxiv.org/html/2608.14694#S5.SS3.p1.1)\.
- \[132\]Y\. Pang and X\. Zhao\(2026\)Beyond pixel overlap: a framework for decomposing segmentation evaluation metrics\.arXiv preprint arXiv:2607\.00886\.Cited by:[§VII\-C](https://arxiv.org/html/2608.14694#S7.SS3.p1.1)\.
- \[133\]N\. Papernot, P\. McDaniel, A\. Sinha, and M\. Wellman\(2018\)SoK: security and privacy in machine learning\.IEEE European Symposium on Security and Privacy\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p1.1)\.
- \[134\]E\. Perenda, S\. Rajendran, G\. Bovet, M\. Zheleva, and S\. Pollin\(2023\)Contrastive learning with self\-reconstruction for channel\-resilient modulation classification\.InIEEE INFOCOM 2023\-IEEE Conference on Computer Communications,pp\. 1–10\.Cited by:[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p3.1)\.
- \[135\]M\. U\. F\. Qaisar, W\. Yuan, O\. Günlü, T\. Riihonen, Y\. Cui, L\. Zhang, N\. Gonzalez\-Prelcic, M\. Di Renzo, and Z\. Han\(2026\)The role of isac in 6g networks: enabling next\-generation wireless systems\.IEEE Transactions on Network Science and Engineering\.Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p3.1)\.
- \[136\]Z\. Qin, G\. Y\. Li, and H\. V\. Poor\(2022\)Semantic communications: principles and challenges\.IEEE Wireless Communications29\(2\),pp\. 72–79\.Cited by:[§VI\-E](https://arxiv.org/html/2608.14694#S6.SS5.p1.1)\.
- \[137\]K\. Qu, S\. Guo, J\. Ye, and N\. Saeed\(2024\)Near\-field integrated sensing and communication: performance analysis and beamforming design\.IEEE Open Journal of the Communications Society5,pp\. 6353–6369\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2024.3470844)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1)\.
- \[138\]M\. Rahmani, S\. Norouzi, J\. Chen, H\. Ahmadi, T\. Braun, K\. Chowdhury, and A\. G\. BURR\(2026\)Hybrid gnn\-centric architectures for ai\-native 6g wireless networks: a comprehensive survey\.IEEE Communications Surveys and Tutorials28,pp\. 5678–5712\.Cited by:[§II](https://arxiv.org/html/2608.14694#S2.p1.1)\.
- \[139\]T\. S\. Rappaportet al\.\(2015\)Millimeter wave wireless communications\.Pearson\.Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p1.1)\.
- \[140\]T\. Raviv and N\. Shlezinger\(2026\)WiMamba: linear\-scale wireless foundation model\.arXiv preprint arXiv:2603\.26367\.Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p4.1)\.
- \[141\]Remcom Inc\.\(2024\)Wireless insite: 3d wireless prediction software\.Note:https://www\.remcom\.com/wireless\-insiteCited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p2.1)\.
- \[142\]C\. Rudin\(2019\)Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead\.Nature Machine Intelligence1\(5\),pp\. 206–215\.External Links:[Document](https://dx.doi.org/10.1038/s42256-019-0048-x)Cited by:[§VIII\-B](https://arxiv.org/html/2608.14694#S8.SS2.p1.1)\.
- \[143\]H\. Sadia, H\. Iqbal, S\. Fawad Hussain, and N\. Saeed\(2025\)Signal detection in intelligent reflecting surface\-assisted NOMA network using LSTM model: a machine learning approach\.IEEE Open Journal of the Communications Society6,pp\. 29–40\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2024.3521008)Cited by:[§VI\-B](https://arxiv.org/html/2608.14694#S6.SS2.p2.1)\.
- \[144\]N\. Saeed and M\. A\. Khan\(2026\)Localization in wireless networks: technologies and applications\.John Wiley & Sons\.Cited by:[§VI\-D](https://arxiv.org/html/2608.14694#S6.SS4.p1.1)\.
- \[145\]S\. Sai, D\. Sharma, M\. S\. Peelam, V\. Chamola, M\. Guizani, and D\. Niyato\(2026\)Machine learning techniques for wi\-fi csi\-based recognition and sensing: a comprehensive review\.IEEE Internet of Things Journal\.Cited by:[§V\-B](https://arxiv.org/html/2608.14694#S5.SS2.p2.1)\.
- \[146\]M\. Salman, N\. Saeed, and K\. Khurshid\(2026\)Transformers for internet of things security and communications: a comprehensive survey\.IEEE Internet of Things Journal\(\),pp\. 1–1\.Cited by:[§VIII\-C](https://arxiv.org/html/2608.14694#S8.SS3.p1.1)\.
- \[147\]W\. Samek, G\. Montavon, A\. Vedaldi, L\. K\. Hansen, and K\. Müller\(2021\)Explainable ai: interpreting, explaining and visualizing deep learning\.Proceedings of the IEEE109\(3\),pp\. 247–278\.External Links:[Document](https://dx.doi.org/10.1109/JPROC.2021.3054486)Cited by:[§VIII\-B](https://arxiv.org/html/2608.14694#S8.SS2.p1.1)\.
- \[148\]M\. Shanmugam, S\. Balaraman, and K\. Ranganathan\(2026\)Improving channel equalization in cell\-free mimo networks using reconfigurable intelligent surfaces and in\-context learning\.Annals of Telecommunications81\(3\),pp\. 259–273\.Cited by:[§VI\-A](https://arxiv.org/html/2608.14694#S6.SS1.p3.1)\.
- \[149\]Y\. Shen, J\. Zhang, and K\. B\. Letaief\(2018\)Machine learning in wireless networks: from theory to applications\.IEEE Communications Magazine56\(12\),pp\. 158–164\.Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p1.1)\.
- \[150\]Y\. Sheng, J\. Wang, X\. Zhou, L\. Liang, H\. Ye, S\. Jin, and G\. Y\. Li\(2025\)A wireless foundation model for multi\-task prediction\.arXiv preprint arXiv:2507\.05938\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p3.1),[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1),[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p2.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.5.5.1.1.1)\.
- \[151\]Y\. Shi, T\. Yang, K\. Ma, L\. Jing, Y\. Wang, M\. Zheng, and L\. Sun\(2026\)A unified adaptive feature composition framework for multi\-task generalization in wireless foundation models\.arXiv preprint arXiv:2606\.10277\.Cited by:[§IV](https://arxiv.org/html/2608.14694#S4.p1.1)\.
- \[152\]N\. Shlezinger and Y\. C\. Eldar\(2023\)Model\-based deep learning\.Proceedings of the IEEE111\(5\),pp\. 465–490\.External Links:[Document](https://dx.doi.org/10.1109/JPROC.2023.3240693)Cited by:[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p1.1),[§VIII\-B](https://arxiv.org/html/2608.14694#S8.SS2.p2.1)\.
- \[153\]M\. Tan and Q\. V\. Le\(2019\)EfficientNet: rethinking model scaling for convolutional neural networks\.Proceedings of the International Conference on Machine Learning \(ICML\)\.Cited by:[§VIII\-D](https://arxiv.org/html/2608.14694#S8.SS4.p3.1)\.
- \[154\]R\. Tan, R\. Li, and Z\. Zhao\(2025\)LLM4MAC: an llm\-driven reinforcement learning framework for mac protocol emergence\.In2025 IEEE 26th International Workshop on Signal Processing and Artificial Intelligence for Wireless Communications \(SPAWC\),pp\. 1–5\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p3.1)\.
- \[155\]Y\. Tay, M\. Dehghani, D\. Bahri, and D\. Metzler\(2023\)Efficient transformers: a survey\.ACM Computing Surveys55\(6\),pp\. 109:1–109:28\.External Links:[Document](https://dx.doi.org/10.1145/3530811)Cited by:[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p4.1)\.
- \[156\]Y\. Tay, M\. Dehghani, D\. Bahri, and D\. Metzler\(2023\)Efficient transformers: a survey\.ACM Computing Surveys55\(6\),pp\. 109:1–109:28\.External Links:[Document](https://dx.doi.org/10.1145/3530811)Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p4.1)\.
- \[157\]Y\. E\. Tok, A\. G\. Toprak, S\. N\. Karahan, O\. B\. Mercan, H\. M\. Aydin, and M\. Altintas\(2026\)Artificial intelligence for next\-generation 6g technologies and networks\.Discover Networks2\(1\),pp\. 3\.Cited by:[§VIII\-E](https://arxiv.org/html/2608.14694#S8.SS5.p2.1)\.
- \[158\]D\. Tse and P\. Viswanath\(2005\)Fundamentals of wireless communication\.Cambridge University Press\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511807213)Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p1.1)\.
- \[159\]L\. Vandenberghe and S\. Boyd\(2004\)Convex optimization\.Vol\.1,Cambridge university press Cambridge\.Cited by:[§II\-A](https://arxiv.org/html/2608.14694#S2.SS1.p2.1)\.
- \[160\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p1.1)\.
- \[161\]D\. Wang, H\. Zhang, C\. Ren, and Y\. Ren\(2026\)FM\-based agentic slicing for 6g: a hierarchical provider\-consumer framework for qos optimization at the ai\-native edge\.IEEE Transactions on Network Science and Engineering\.Cited by:[§I\-A](https://arxiv.org/html/2608.14694#S1.SS1.p1.1)\.
- \[162\]X\. Wang, Y\. Pan, N\. Cheng, Ç\. Yapar, R\. Sun, Z\. Yin, C\. Zhou, W\. Xu, Y\. Zhang, J\. Zhang,et al\.\(2026\)A tutorial on learning\-based radio map construction: data, paradigms, and physics\-awareness\.arXiv preprint arXiv:2603\.17499\.Cited by:[§VIII\-A](https://arxiv.org/html/2608.14694#S8.SS1.p1.1)\.
- \[163\]Y\. Wang, O\. Wang, S\. Zhou, and G\. Y\. Li\(2026\)Hierarchical wireless foundation model for multi\-task optimization\.arXiv preprint arXiv:2607\.16877\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p5.1)\.
- \[164\]B\. Wei, R\. Jiang, R\. Zhang, Y\. Liu, D\. Niyato, Y\. Sun, Y\. Lu, Y\. Li, S\. Mao, C\. Yuen,et al\.\(2025\)Large language models for next\-generation wireless network management: a survey and tutorial\.arXiv preprint arXiv:2509\.05946\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1)\.
- \[165\]Z\. Wu, S\. Pan, F\. Chen, G\. Long, C\. Zhang, and P\. S\. Yu\(2021\)A comprehensive survey on graph neural networks\.IEEE Transactions on Neural Networks and Learning Systems32\(1\),pp\. 4–24\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2020.2978386)Cited by:[§III\-B](https://arxiv.org/html/2608.14694#S3.SS2.p3.1)\.
- \[166\]H\. Wymeersch, N\. Tervo, S\. Wanstedt,et al\.\(2025\)Cross\-layer integrated sensing and communication: a joint industrial and academic perspective\.IEEE Open Journal of the Communications Society6,pp\. 6966–7015\.Cited by:[§VI\-D](https://arxiv.org/html/2608.14694#S6.SS4.p1.1)\.
- \[167\]H\. Wymeersch, N\. Tervo, S\. Wanstedt, S\. Saleh, J\. Ahlendorf, O\. Akgul,et al\.\(2025\)Cross\-layer integrated sensing and communication: a joint industrial and academic perspective\.IEEE Open Journal of the Communications Society6\(\),pp\. 6966–7015\.External Links:[Document](https://dx.doi.org/10.1109/OJCOMS.2025.3595459)Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1)\.
- \[168\]H\. Xie, C\. Li, X\. Huang, E\. Tanghe, W\. Joseph, S\. Ni, and X\. Yuan\(2026\)Learning\-driven channel representation for wireless localization: from channel observations to location inference\.arXiv preprint arXiv:2607\.14938\.Cited by:[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p3.1)\.
- \[169\]S\. Xu, C\. Kurisummoottil Thomas, O\. Hashash, N\. Muralidhar, W\. Saad, and N\. Ramakrishnan\(2024\)Large multi\-modal models \(lmms\) as universal foundation models for ai\-native wireless systems\.IEEE Network38\(5\),pp\. 10–20\.External Links:[Document](https://dx.doi.org/10.1109/MNET.2024.3427313)Cited by:[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1)\.
- \[170\]B\. Yang, W\. Chen, J\. Cheng, and B\. Ai\(2026\)ComHymba: low\-complexity domain\-informed foundation model for wireless communications\.External Links:2605\.23468Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p4.1)\.
- \[171\]T\. Yang, P\. Zhang, M\. Zheng, Y\. Shi, L\. Jing, J\. Huang, and N\. Li\(2025\)WirelessGPT: a generative pre\-trained multi\-task learning framework for wireless communication\.IEEE Network39\(5\),pp\. 58–65\.External Links:[Document](https://dx.doi.org/10.1109/MNET.2025.3579496)Cited by:[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1)\.
- \[172\]T\. Yang, P\. Zhang, M\. Zheng, Y\. Shi, L\. Jing, J\. Huang, and N\. Li\(2025\)WirelessGPT: a generative pre\-trained multi\-task learning framework for wireless communication\.IEEE Network39\(5\),pp\. 58–65\.External Links:[Document](https://dx.doi.org/10.1109/MNET.2025.3579496)Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p4.1),[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p2.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p2.1),[§IV\-A](https://arxiv.org/html/2608.14694#S4.SS1.p3.1),[TABLE III](https://arxiv.org/html/2608.14694#S6.T3.1.1.2.2.1.1.1)\.
- \[173\]Z\. Yang, G\. Chi, C\. Wu, H\. Liu, Y\. Gao, Y\. Liu, Y\. C\. Eldar, J\. Xu, and T\. X\. Han\(2026\)Generative ai for wireless communication and sensing: toward unified foundation models\.IEEE Transactions on Communications\.Cited by:[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p1.1)\.
- \[174\]H\. Ye, G\. Y\. Li, and B\. H\. Juang\(2018\)Power of deep learning for channel estimation and signal detection in ofdm systems\.IEEE Wireless Communications Letters7\(1\),pp\. 114–117\.External Links:[Document](https://dx.doi.org/)Cited by:[§II\-B](https://arxiv.org/html/2608.14694#S2.SS2.p1.1)\.
- \[175\]K\. Ying, Z\. Gao, T\. Yang, J\. Zhang, X\. Cheng, T\. Q\. Quek, and H\. V\. Poor\(2026\)From specialist to large models: a paradigm evolution towards semantic\-aware mimo\.IEEE Communications Magazine\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p1.1)\.
- \[176\]L\. Yu, L\. Shi, J\. Zhang, Z\. Zhang, Y\. Zhang, and G\. Liu\(2025\)ChannelGPT: a large model toward real\-world channel foundation model for 6g environment intelligence communication\.IEEE Communications Magazine63\(10\),pp\. 68–74\.Cited by:[§VI\-D](https://arxiv.org/html/2608.14694#S6.SS4.p2.1)\.
- \[177\]W\. Yu, F\. Sohrabi, and T\. Jiang\(2023\)Role of deep learning in wireless communications\.IEEE BITS: The Information Theory Magazine\.Note:Early version available as arXiv:2210\.02596External Links:2210\.02596Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p1.1),[§I](https://arxiv.org/html/2608.14694#S1.p2.1)\.
- \[178\]W\. Yue, T\. Lai, Q\. Mao, Q\. Li, and D\. Camacho\(2026\)A review of federated learning under data heterogeneity\.Expert Systems43\(6\),pp\. e70271\.Cited by:[§VIII\-A](https://arxiv.org/html/2608.14694#S8.SS1.p2.1)\.
- \[179\]C\. Zhang, X\. Lyu, C\. Ren, S\. Liu, and Q\. Cui\(2026\)Adaptive 3d\-rope: physics\-aligned rotary positional encoding for wireless foundation models\.arXiv preprint arXiv:2605\.00968\.Cited by:[§IV\-B](https://arxiv.org/html/2608.14694#S4.SS2.p3.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p3.1)\.
- \[180\]H\. Zhanget al\.\(2024\)DeepSense 6g: a large\-scale multimodal dataset for ai\-native wireless networks\.IEEE Data Descriptions\.Cited by:[§VII\-A](https://arxiv.org/html/2608.14694#S7.SS1.p3.1),[TABLE IV](https://arxiv.org/html/2608.14694#S7.T4.1.5.5.1.1.1)\.
- \[181\]H\. Zhang, M\. Farzanullah, M\. Ghassemi, A\. Bin Sediq, A\. Afana, and M\. Erol\-Kantarci\(2026\)Multi\-modal data\-enhanced foundation models for prediction and control in wireless networks: a survey\.External Links:2601\.03181Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p4.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p1.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p5.1),[§IV\-C](https://arxiv.org/html/2608.14694#S4.SS3.p1.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1)\.
- \[182\]H\. Zhang, S\. Gao, and X\. Cheng\(2026\)Tiny\-wifo: a lightweight wireless foundation model for channel prediction via multi\-component adaptive knowledge distillation\.IEEE Wireless Communications Letters\.Cited by:[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1)\.
- \[183\]T\. Zheng, J\. Guo, L\. Dai, S\. Jin, and J\. Zhang\(2026\)Muse\-fm: multi\-task environment\-aware foundation model for wireless communications\.IEEE Transactions on Wireless Communications\.Cited by:[§III\-A](https://arxiv.org/html/2608.14694#S3.SS1.p3.1),[§III\-C](https://arxiv.org/html/2608.14694#S3.SS3.p5.1),[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1),[§IV\-D](https://arxiv.org/html/2608.14694#S4.SS4.p2.1)\.
- \[184\]C\. Zhou, Q\. Li, C\. Li, J\. Yu, Y\. Liu, G\. Wang, K\. Zhang, C\. Ji, Q\. Yan, L\. He,et al\.\(2025\)A comprehensive survey on pretrained foundation models: a history from bert to chatgpt\.International Journal of Machine Learning and Cybernetics16\(12\),pp\. 9851–9915\.Cited by:[§I](https://arxiv.org/html/2608.14694#S1.p3.1)\.
- \[185\]F\. Zhu, X\. Wang, X\. Li, M\. Zhang, Y\. Chen, C\. Huang, Z\. Yang, X\. Chen, Z\. Zhang, R\. Jin,et al\.\(2025\)Wireless large ai model: shaping the ai\-native future of 6g and beyond\.arXiv preprint arXiv:2504\.1465310\.Cited by:[§VI\-C](https://arxiv.org/html/2608.14694#S6.SS3.p1.1)\.
- \[186\]Z\. Zhu, J\. Zhang, D\. Yu, and X\. Li\(2026\)Text\-attributed graph augmented large language models for question answering\.InInternational Conference on Knowledge Science, Engineering and Management,pp\. 397–408\.Cited by:[§VII\-B](https://arxiv.org/html/2608.14694#S7.SS2.p3.1)\.
- \[187\]H\. Zou, Y\. Yang, L\. Bariah, Y\. Tian, Y\. Lu, B\. Wang, A\. Bara, B\. Mefgouda, H\. Liu, Y\. Tao,et al\.\(2026\)Telecom world models: unifying digital twins, foundation models, and predictive planning for 6g\.arXiv preprint arXiv:2604\.06882\.Cited by:[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p4.1),[§III\-D](https://arxiv.org/html/2608.14694#S3.SS4.p5.1)\.

Similar Articles

AI Native Games: A Survey and Roadmap

arXiv cs.AI

This survey paper defines AI-native games as those where runtime generative AI is constitutive of the core game loop, proposes a dual-axis G/N taxonomy, analyzes 53 games, and provides a roadmap for controllable generation, multimodal systems, and AI safety in game design.