Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions
Summary
Proposes a deep wireless physical neural network where multi-hop MIMO relays realize trainable linear transforms and power amplifier nonlinearities serve as activation functions, enabling over-the-air inference for image classification.
View Cached Full Text
Cached at: 07/22/26, 08:20 AM
# Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions
Source: [https://arxiv.org/html/2607.18354](https://arxiv.org/html/2607.18354)
Meng Hua, Itsik Bergel, , and Deniz GündüzThis work was supported by UKRI under the projects AI\-R \(EP/X030806/1\) and INFORMED\-AI \(EP/Y028732/1\), and by the SNS JU project 6G\-GOALS under the EU Horizon program \(Grant Agreement No\. 101139232\)\. M\. Hua and D\. Gündüz are with the Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, U\.K\. \(e\-mail: \{m\.hua,d\.gunduz\}@imperial\.ac\.uk\)\. I\. Bergel is with the Faculty of Engineering, Bar\-Ilan University, Ramat Gan 5290002, Israel \(e\-mail: itsik\.bergel@biu\.ac\.il\)\.
###### Abstract
Wireless physical neural networks \(WPNNs\) embed neural computation directly into analog hardware, offering lower energy consumption and latency than conventional digital implementations\. In this paper, we propose a deep WPNN in which nonlinear activations are realized by a multi\-hop multiple\-input multiple\-output \(MIMO\) relay network, in which each relay implements a trainable complex linear gain and bias, followed by the power amplifier’s intrinsic nonlinearity acting as an activation function\. The cascade of multiple relays therefore realizes an over\-the\-air fully connected network whose parameters can be trained end\-to\-end\. We develop two transceiver designs for different channel state information \(CSI\) availability scenarios: a least squares \(LS\)\-based scheme requiring only receiver\-side CSI, and a singular\-value\-decomposition \(SVD\)\-based scheme requiring both transmitter\-side and receiver\-side CSI\. Simulation results show that the proposed architecture enables accurate over\-the\-air inference for image classification\. In particular, the results highlight the advantage of exploiting hardware nonlinearity for enhanced inference capability\.
## IIntroduction
The remarkable success of artificial intelligence \(AI\) has been largely driven by deep learning models executed on graphics processing units \(GPUs\), whose massive parallelism and numerical precision enable efficient training and inference\. However, GPUs inherit the von Neumann architecture, in which the physical separation between memory and computation incurs substantial energy and latency costs, particularly for large\-scale or real\-time deployments\. The emerging concept of thephysical neural network \(PNN\)offers a promising alternative\[[14](https://arxiv.org/html/2607.18354#bib.bib15),[10](https://arxiv.org/html/2607.18354#bib.bib11),[7](https://arxiv.org/html/2607.18354#bib.bib8)\]: linear mappings and nonlinear activations are realized through controllable analog mechanisms, including electronic, optical, or wireless, whose intrinsic dynamics and nonlinearity enable in\-situ inference with significantly lower energy and delay\.
Recently, research on implementing wireless PNNs \(WPNNs\) over the air has gained increasing attention\[[4](https://arxiv.org/html/2607.18354#bib.bib5)\]\. First, physical substrates such as the phase shifts of reconfigurable intelligent surfaces \(RISs\) and the amplification coefficients of relays can be regarded as trainable neurons within the wireless medium\. Second, fundamental algebraic operations, such as matrix multiplication and summation, can be inherently realized over the air by exploiting the superposition property of the wireless multiple\-access channel\[[1](https://arxiv.org/html/2607.18354#bib.bib1)\], thereby significantly reducing computation energy consumption and processing latency\. Nevertheless, research on over\-the\-air WPNNs remains in its early stage, with only a limited number of studies reported to date\[[12](https://arxiv.org/html/2607.18354#bib.bib13),[5](https://arxiv.org/html/2607.18354#bib.bib6),[16](https://arxiv.org/html/2607.18354#bib.bib17),[8](https://arxiv.org/html/2607.18354#bib.bib9),[9](https://arxiv.org/html/2607.18354#bib.bib10),[15](https://arxiv.org/html/2607.18354#bib.bib16),[6](https://arxiv.org/html/2607.18354#bib.bib7),[11](https://arxiv.org/html/2607.18354#bib.bib12),[2](https://arxiv.org/html/2607.18354#bib.bib2),[13](https://arxiv.org/html/2607.18354#bib.bib14),[3](https://arxiv.org/html/2607.18354#bib.bib3)\]\. In particular,\[[12](https://arxiv.org/html/2607.18354#bib.bib13),[5](https://arxiv.org/html/2607.18354#bib.bib6),[16](https://arxiv.org/html/2607.18354#bib.bib17),[8](https://arxiv.org/html/2607.18354#bib.bib9),[9](https://arxiv.org/html/2607.18354#bib.bib10),[15](https://arxiv.org/html/2607.18354#bib.bib16),[6](https://arxiv.org/html/2607.18354#bib.bib7),[11](https://arxiv.org/html/2607.18354#bib.bib12)\]investigated the use of RISs as neural network neurons, where the adjustable phase shifts of RIS elements are treated as trainable parameters of the network\. For instance, in\[[11](https://arxiv.org/html/2607.18354#bib.bib12)\], the authors proposed leveraging RISs to realize one\-dimensional convolution operations by exploiting multipath propagation delays, wherein each channel impulse response acts as an individual finite impulse response filter that convolves with the transmitted signal to emulate a digital convolutional neural network\. Relays have also been explored as WPNN hardware in\[[2](https://arxiv.org/html/2607.18354#bib.bib2),[13](https://arxiv.org/html/2607.18354#bib.bib14),[3](https://arxiv.org/html/2607.18354#bib.bib3)\], where the relay amplification coefficients are interpreted as analog neuron weights\. When multiple relays are employed, the combined effects of wireless channels and amplification gains form an equivalent virtual multiple\-input multiple\-output \(MIMO\) system\. For example,\[[2](https://arxiv.org/html/2607.18354#bib.bib2)\]demonstrated that by exploiting both the amplification gain and the inherent nonlinearity of power amplifiers \(PAs\), significant communication performance gains can be achieved\. However, the aforementioned studies have not fully exploited the spatial degrees of freedom or the intrinsic nonlinearity of physical devices, which fundamentally limit their expressive capacity\.
In this paper, we study deep physical neural networks via a multi\-hop MIMO relay architecture, as illustrated in Fig\.[1](https://arxiv.org/html/2607.18354#S1.F1)\. Our main contributions are as follows\. First, we propose a multi\-hop MIMO relay architecture for realizing a deep WPNN, formally associating the trainable amplification, bias, and PA nonlinearity at each relay with the weight, bias, and activation function of one FC layer\. Second, we propose two transceiver schemes for different channel state information \(CSI\) availability scenarios: a least squares \(LS\)\-based scheme requiring only receiver\-side CSI \(CSIR\), and a singular\-value\-decomposition \(SVD\)\-based scheme requiring both transmitter\-side and receiver\-side CSI \(CSIR/T\)\. Both schemes support end\-to\-end training with standard backpropagation through the PA model\. Third, through simulations on the Fashion\-MNIST dataset, we show that, contrary to expectation, the SVD\-based scheme that decouples eigenmodes does not benefit from PA nonlinearity in deep cascades, whereas the LS\-based scheme exploits nonlinearity and scales gracefully with depth\.

Figure 1:The architecture of multi\-hop MIMO relay systems\.
## IISystem model
As illustrated in Fig\.[1](https://arxiv.org/html/2607.18354#S1.F1), we consider a multi\-hop MIMO relay system, where information is transmitted from a source to a user throughMMMIMO relaysR1,…,RM\{R\_\{1\},\\ldots,R\_\{M\}\}\. For notational convenience, we denote the source node and the user asR0R\_\{0\}andRM\+1R\_\{M\+1\}, respectively\. All nodes in the network are equipped with equal number of transmit and receive antennas\. Specifically, letNmN\_\{m\}denote the number of transmit or receive antennas at nodeRmR\_\{m\}, form∈\{0,…,M\+1\}m\\in\\left\\\{\{0,\\ldots,M\+1\}\\right\\\}\.
The objective of this relay network is to perform a generic inference task on the signal𝐒∈𝒮\\bf S\\in\{\\cal S\}available atR0R\_\{0\}, and conveyed toRM\+1R\_\{M\+1\}throughMMMIMO relays\. The task is represented by𝒯:𝒮→𝒴\\mathcal\{T\}:\\mathcal\{S\}\\rightarrow\\mathcal\{Y\}, where𝒴\\mathcal\{Y\}denotes the task\-specific output space\. A concrete example of an image classification task will be introduced in Section[IV](https://arxiv.org/html/2607.18354#S4)\. Let𝐇m∈ℂNm×Nm−1\{\{\\mathbf\{H\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times\{N\_\{m\-1\}\}\}\}denote the complex baseband equivalent MIMO channel from nodeRm−1\{\{R\_\{m\-1\}\}\}to nodeRm\{\{R\_\{m\}\}\}form∈\{1,…,M\+1\}m\\in\\left\\\{\{1,\\ldots,M\+1\}\\right\\\}\. We adopt a Rician fading model for all wireless links\. Specifically, the channel matrix of themm\-th hop,𝐇m\{\{\\bf\{H\}\}\_\{m\}\}, is expressed as
𝐇m=αm\(KK\+1𝐇m,LoS\+1K\+1𝐇m,NLoS\),\\displaystyle\{\{\\mathbf\{H\}\}\_\{m\}\}\{\\text\{ = \}\}\\sqrt\{\{\\alpha\_\{m\}\}\}\\left\(\{\\sqrt\{\\frac\{K\}\{\{K\+1\}\}\}\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{LoS\}\}\}\}\+\\sqrt\{\\frac\{1\}\{\{K\+1\}\}\}\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{NLoS\}\}\}\}\}\\right\),\(1\)whereαm\{\{\\alpha\_\{m\}\}\}captures the large\-scale fading of themm\-th hop, andKKrepresents the Rician factor\. The line\-of\-sight \(LoS\) component𝐇m,LoS\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{LoS\}\}\}\}is modeled as a deterministic rank\-one matrix and is given by
𝐇m,LoS=𝐚r\(θmr;Nm\)𝐚tH\(θm−1t;Nm−1\),\\displaystyle\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{LoS\}\}\}\}=\{\{\\mathbf\{a\}\}\_\{\\text\{r\}\}\}\\left\(\{\\theta\_\{m\}^\{\\text\{r\}\};\{N\_\{m\}\}\}\\right\)\{\\mathbf\{a\}\}\_\{\\text\{t\}\}^\{H\}\\left\(\{\\theta\_\{m\-1\}^\{\\text\{t\}\};\{N\_\{m\-1\}\}\}\\right\),\(2\)where𝐚r\(θmr;Nm\)\{\{\\mathbf\{a\}\}\_\{\\text\{r\}\}\}\\left\(\{\\theta\_\{m\}^\{\\text\{r\}\};\{N\_\{m\}\}\}\\right\)and𝐚t\(θm−1t;Nm−1\)\{\{\\mathbf\{a\}\}\_\{\\text\{t\}\}\}\\left\(\{\\theta\_\{m\-1\}^\{\\text\{t\}\};\{N\_\{m\-1\}\}\}\\right\)denote the receive and transmit array response vectors at nodesRmR\_\{m\}andRm−1R\_\{m\-1\}, respectively\. The angles of arrival \(AoA\) and departure \(AoD\) associated with the LoS path are denoted byθmr\{\\theta\_\{m\}^\{\\text\{r\}\}\}andθm−1t\{\\theta\_\{m\-1\}^\{\\text\{t\}\}\}, respectively, which are assumed to be independently and uniformly distributed over\[0,π\]\\left\[\{0,\\pi\}\\right\]\. For a uniform linear array withNNantennas, its array response is given by
𝐚c\(θ;N\)=\[1,ejπsinθ,…,ejπsinθ\(N−1\)\]T,c∈\{r,t\}\.\\displaystyle\{\{\\mathbf\{a\}\}\_\{\\text\{c\}\}\}\\left\(\{\\theta;N\}\\right\)=\{\\left\[\{1,\{e^\{j\\pi\\sin\\theta\}\},\\ldots,\{e^\{j\\pi\\sin\\theta\\left\(\{N\-1\}\\right\)\}\}\}\\right\]^\{T\}\},\{\\text\{c\}\}\\in\\left\\\{\{\{\\text\{r,t\}\}\}\\right\\\}\.\(3\)The non\-LoS \(NLoS\) component𝐇m,NLoS\{\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{NLoS\}\}\}\}\}models the small\-scale scattering, where each entry is independently distributed as\[𝐇m,NLoS\]i,j∼𝒞𝒩\(0,1\)\{\\left\[\{\{\{\\mathbf\{H\}\}\_\{m,\{\\text\{NLoS\}\}\}\}\}\\right\]\_\{i,j\}\}\\sim\{\\cal CN\}\\left\(\{0,1\}\\right\)\.
Let𝐗m∈ℂNm×L\{\{\\mathbf\{X\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times L\}\}and𝐘m∈ℂNm×L\{\{\\mathbf\{Y\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times L\}\}denote the transmitted and received symbols at nodeRmR\_\{m\}form∈\{0,1,…,M\+1\}m\\in\\left\\\{\{0,1,\\ldots,M\+1\}\\right\\\}, respectively, whereLLdenotes the number of transmitted symbols, which will be specified later\. We have
𝐘m=𝐇m𝐗m−1\+𝐍m,m∈\{1,…,M\+1\},\\displaystyle\{\{\\bf\{Y\}\}\_\{m\}\}=\{\{\\bf\{H\}\}\_\{m\}\}\{\{\\bf\{X\}\}\_\{m\-1\}\}\+\{\{\\bf\{N\}\}\_\{m\}\},~m\\in\\left\\\{\{1,\\ldots,M\+1\}\\right\\\},\(4\)where𝐍m∈ℂNm×L\{\{\\mathbf\{N\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times L\}\}denotes the additive white Gaussian noise, with each entry in𝐍m\{\{\\bf\{N\}\}\_\{m\}\}independently distributed as\[𝐍m\]i,j∼𝒞𝒩\(0,σ2\)\{\\left\[\{\{\{\\bf\{N\}\}\_\{m\}\}\}\\right\]\_\{i,j\}\}\\sim\{\\cal CN\}\\left\(\{0,\{\\sigma^\{2\}\}\}\\right\)\.
We consider two cases, namely CSIR MIMO and CSIR/T MIMO, depending on the availability of CSI at each node\.
### II\-ACSIR MIMO System
In a CSIR MIMO system, the CSI is available at the receiver for each hop, which allows it to equalize the transmitted signals\. According to the input–output relationship in \([4](https://arxiv.org/html/2607.18354#S2.E4)\), with the knowledge of𝐇m\{\{\\mathbf\{H\}\}\_\{m\}\},RmR\_\{m\}can perform receiver\-side signal processing to estimate the transmitted symbols\. We adopt the LS estimator to recover an estimate𝐗^m−1\{\{\{\\bf\{\\hat\{X\}\}\}\}\_\{m\-1\}\}of𝐗m−1\{\{\{\\bf\{X\}\}\}\_\{m\-1\}\}\. By applying a node\-specific processing function parameterized by𝜽m\{\{\{\\bm\{\\theta\}\}\_\{m\}\}\}, the transmitted signal atRmR\_\{m\}can be expressed as
𝐗m=g𝜽m\(𝐘m,𝐇m,σ2\),m∈\{1,…,M\+1\},\\displaystyle\{\{\\bf\{X\}\}\_\{m\}\}=\{g\_\{\{\{\\bm\{\\theta\}\}\_\{m\}\}\}\}\\left\(\{\{\{\{\\bf\{Y\}\}\}\_\{m\}\},\{\{\\bf\{H\}\}\_\{m\}\},\{\\sigma^\{2\}\}\}\\right\),~m\\in\\left\\\{\{1,\\ldots,M\+1\}\\right\\\},\(5\)whereg𝜽m\(⋅\)\{g\_\{\{\{\\bm\{\\theta\}\}\_\{m\}\}\}\}\\left\(\\cdot\\right\)denotes the signal processing function implemented atRmR\_\{m\}\.
### II\-BCSIR/T MIMO System
In a CSIR/T MIMO system, the CSI is available at both the transmitter and receiver for each hop\. Therefore, the encoding function parameterized by WPNN parametersϕm\{\{\{\\bm\{\\phi\}\}\_\{m\}\}\}atRmR\_\{m\}can be expressed as
𝐗m=fϕm\(𝐘m,𝐇m,𝐇m\+1,σ2\),m∈\{1,…,M\},\\displaystyle\\\!\\\!\\\!\{\{\\bf\{X\}\}\_\{m\}\}=\{f\_\{\{\{\\bm\{\\phi\}\}\_\{m\}\}\}\}\\left\(\{\{\{\{\\bf\{Y\}\}\}\_\{m\}\},\{\{\\bf\{H\}\}\_\{m\}\},\{\{\\bf\{H\}\}\_\{m\+1\}\},\{\\sigma^\{2\}\}\}\\right\),~m\\in\\left\\\{\{1,\\ldots,M\}\\right\\\},\(6\)wherefϕm\(⋅\)\{f\_\{\{\{\\bm\{\\phi\}\}\_\{m\}\}\}\}\\left\(\\cdot\\right\)denotes the signal processing function implemented atRmR\_\{m\}\.
The WPNN realized by the multi\-hop MIMO relay network is trained end\-to\-end to approximate the task mapping𝒯\\mathcal\{T\}\. Denoting by𝒯^𝚯\(𝐒\)\\widehat\{\\mathcal\{T\}\}\_\{\\bm\{\\Theta\}\}\(\{\\bf S\}\)the mapping from the input sample𝐒\\bf Sto the user’s decision, parameterized by all trainable physical\-layer parameters𝚯\\bm\{\\Theta\}\(amplification matrices, biases, transmit/receive processing, and final read\-out layer\), the objective is to minimize a task\-specific loss:min𝚯𝔼\(𝐒,y\)∼P𝐒,y\[ℒ\(𝒯^𝚯\(S\),y\)\]\\min\_\{\\bm\{\\Theta\}\}\\mathbb\{E\}\_\{\(\{\\bf S\},y\)\\sim P\_\{\{\\bf S\},y\}\}\[\{\\cal L\}\(\\widehat\{\\mathcal\{T\}\}\_\{\\bm\{\\Theta\}\}\(S\),y\)\], wherey∈𝒴y\\in\\mathcal\{Y\}is the ground\-truth label\.
## IIIRole of PA Nonlinearity in Multi\-Hop MIMO Relay Systems
To motivate the role of PA nonlinearity, we contrast linear and nonlinear PA regimes\.
Linear PA: Assume that the PA at each relay operates strictly in its linear region so that𝐗m=𝐖m𝐘m\{\{\\mathbf\{X\}\}\_\{m\}\}=\{\{\\mathbf\{W\}\}\_\{m\}\}\{\{\\mathbf\{Y\}\}\_\{m\}\}, where𝐖m∈ℂNm×Nm\{\{\\mathbf\{W\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times N\_\{m\}\}\}is the linear amplification matrix\. Substituting it into \([4](https://arxiv.org/html/2607.18354#S2.E4)\) recursively yields
𝐘M\+1=𝐇M\+1𝐖M𝐇M𝐖M−1⋯𝐇1𝐗0\+𝐍eff,\\displaystyle\{\{\\mathbf\{Y\}\}\_\{M\+1\}\}=\{\{\\mathbf\{H\}\}\_\{M\+1\}\}\{\{\\mathbf\{W\}\}\_\{M\}\}\{\{\\mathbf\{H\}\}\_\{M\}\}\{\{\\mathbf\{W\}\}\_\{M\-1\}\}\\cdots\{\{\\mathbf\{H\}\}\_\{1\}\}\{\{\\mathbf\{X\}\}\_\{0\}\}\+\{\{\\mathbf\{N\}\}\_\{\{\\text\{eff\}\}\}\},\(7\)where𝐍eff\{\{\\mathbf\{N\}\}\_\{\{\\text\{eff\}\}\}\}aggregates the per\-hop noise and is independent of𝐗0\{\\bf X\}\_\{0\}\. The end\-to\-end map𝐗0→𝐘M\+1\{\{\\mathbf\{X\}\}\_\{0\}\}\\to\{\{\\mathbf\{Y\}\}\_\{M\+1\}\}is therefore linear and collapses to a single composite FC layer for anyMM, fundamentally limiting the expressive capability of the WPNN\. Basically, the achievable function class is the set of complex linear maps of limited rankminmrank\(𝐇m𝐖m\)\\mathop\{\\min\}\\limits\_\{m\}\{\\text\{rank\}\}\\left\(\{\{\{\\mathbf\{H\}\}\_\{m\}\}\{\{\\mathbf\{W\}\}\_\{m\}\}\}\\right\)\.
Nonlinear PA: When the PA exhibits a nonlinear transfer characteristic, the relay output becomes
𝐗m=PA\(𝐖m𝐘m\),\\displaystyle\{\{\\mathbf\{X\}\}\_\{m\}\}=\{\\text\{PA\}\}\\left\(\{\{\{\\mathbf\{W\}\}\_\{m\}\}\{\{\\mathbf\{Y\}\}\_\{m\}\}\}\\right\),\(8\)wherePA\(⋅\)\{\\rm\{PA\}\}\\left\(\\cdot\\right\)denotes the element\-wise nonlinear input–output map, whose specific form will be detailed in Section[IV](https://arxiv.org/html/2607.18354#S4)\. Each relay hop realizes linear mixing followed by a pointwise nonlinear feature transformation, so theMM\-hop relay chain is equivalent to anMM\-layer FC network with implicit activations\. This breaks the limited rank bound of linear PA, making the achievable function class grow withMM\. This is the structural basis for interpreting the multi\-hop MIMO relay system as a deep WPNN\.
## IVDeep Learning Optimization with Nonlinear PA
In this section, we instantiate the general framework developed in Section[II](https://arxiv.org/html/2607.18354#S2)for an image classification task, while noting that the proposed design can be readily extended to other inference tasks\. We then present the training strategy and the corresponding loss function used to train the WPNN for this task\.
### IV\-ATransmission Architecture Design
#### IV\-A1CSIR Design
Let the input sample be an image𝐒∈ℝC×H×W\{\\bf S\}\\in\{\{\\mathbb\{R\}\}^\{C\\times H\\times W\}\}, whereCC,HH, andWWrepresent the number of color channels, height, and width, respectively, and the output space𝒴\\mathcal\{Y\}is the discrete set of image classes\. The source nodeR0R\_\{0\}first normalizes𝐒\\bf Sso that its pixel values lie in\[0,1\]\\left\[\{0,1\}\\right\]\. The normalized image is then vectorized and reshaped into a complex\-valued matrix𝐒c∈ℂN0×L\{\{\\bf\{S\}\}\_\{c\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{0\}\}\\times L\}\}withL=C×H×W2N0L=\\frac\{\{C\\times H\\times W\}\}\{\{2\{N\_\{0\}\}\}\}\. Then, the signal transmitted byR0R\_\{0\}is given by
𝐗0=PA\(𝐅0𝐒c\+𝐛0𝟏LT\),\\displaystyle\{\{\\mathbf\{X\}\}\_\{0\}\}=\{\\text\{PA\}\}\\left\(\{\{\{\\mathbf\{F\}\}\_\{0\}\}\{\{\\mathbf\{S\}\}\_\{c\}\}\+\{\{\\mathbf\{b\}\}\_\{0\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\}\}\\right\),\(9\)where𝐅0∈ℂN0×N0\{\{\\bf\{F\}\}\_\{0\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{0\}\}\\times\{N\_\{0\}\}\}\}and𝐛0∈ℂN0×1\{\{\\bf\{b\}\}\_\{0\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{0\}\}\\times 1\}\}denote the precoder matrix and the bias vector, respectively, and𝟏L\{\{\\mathbf\{1\}\}\_\{L\}\}is theLL\-length vector of all ones\. The bias vector𝐛0\{\{\\bf\{b\}\}\_\{0\}\}plays a role analogous to that in digital neural networks and can be practically realized by injecting a direct current offset, which is trainable\. The Rapp PA model, denoted byPA\(⋅\)\{\\rm\{PA\}\}\\left\(\\cdot\\right\), can be modeled as\[[13](https://arxiv.org/html/2607.18354#bib.bib14)\]
PA\(x\)=x\(1\+\(\|x\|/xsat\)2p\)1/\(2p\),\\displaystyle\{\\text\{PA\}\}\\left\(x\\right\)=\\frac\{x\}\{\{\{\{\\left\(\{1\+\{\{\\left\(\{\{\{\\left\|x\\right\|\}\\mathord\{\\left/\{\\vphantom\{\{\\left\|x\\right\|\}\{\{x\_\{\{\\text\{sat\}\}\}\}\}\}\}\\right\.\\kern\-1\.2pt\}\{\{x\_\{\{\\text\{sat\}\}\}\}\}\}\}\\right\)\}^\{2p\}\}\}\\right\)\}^\{\{1\\mathord\{\\left/\{\\vphantom\{1\{\\left\(\{2p\}\\right\)\}\}\}\\right\.\\kern\-1\.2pt\}\{\\left\(\{2p\}\\right\)\}\}\}\}\}\},\(10\)whereppandxsatx\_\{\\rm sat\}represent the PA parameters\. Following\[[2](https://arxiv.org/html/2607.18354#bib.bib2)\], we setp=2p=2andxsat=1x\_\{\\rm sat\}=1, under which the amplitude response of the nonlinear power amplifier exhibits a smooth saturation behavior similar to that of atanh\\mathrm\{tanh\}function\. Accordingly, the signal received at relayR1R\_\{1\}can be represented as
𝐘1=𝐇1𝐗0\+𝐍1\.\\displaystyle\{\{\\bf\{Y\}\}\_\{1\}\}=\{\{\\bf\{H\}\}\_\{1\}\}\{\{\\bf\{X\}\}\_\{0\}\}\+\{\{\\bf\{N\}\}\_\{1\}\}\.\(11\)Then, a LS MIMO estimator is employed to exploit the CSI to decouple the entangled signal𝐘1\{\{\\bf\{Y\}\}\_\{1\}\}as𝐗^0\{\{\{\\bf\{\\hat\{X\}\}\}\}\_\{0\}\}:
𝐗^0=𝐇1\+𝐘1=𝐇1\+\(𝐇1𝐗0\+𝐍1\),\\displaystyle\{\{\{\\bf\{\\hat\{X\}\}\}\}\_\{0\}\}=\{\\bf\{H\}\}\_\{1\}^\{\+\}\{\{\\bf\{Y\}\}\_\{1\}\}=\{\\bf\{H\}\}\_\{1\}^\{\+\}\\left\(\{\{\{\\bf\{H\}\}\_\{1\}\}\{\{\\bf\{X\}\}\_\{0\}\}\+\{\{\\bf\{N\}\}\_\{1\}\}\}\\right\),\(12\)where\(⋅\)\+\{\\left\(\\cdot\\right\)^\{\+\}\}denotes the the Moore–Penrose pseudo\-inverse\. Next, a trainable relay amplification matrix, denoted by𝐅1∈ℂN1×N0\{\{\\bf\{F\}\}\_\{1\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{1\}\}\\times N\_\{0\}\}\}, is applied to scale𝐗^0\{\{\{\\bf\{\\hat\{X\}\}\}\}\_\{0\}\}, and a trainable bias vector𝐛1∈ℂN1×1\{\{\\bf\{b\}\}\_\{1\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{1\}\}\\times 1\}\}is added, and the result passes through the nonlinear PA\. The output signal at relayR1R\_\{1\}can thus be expressed as
𝐗1=PA\(𝐅1𝐗^0\+𝐛1𝟏LT\)\.\\displaystyle\{\{\\mathbf\{X\}\}\_\{1\}\}=\{\\text\{PA\}\}\\left\(\{\{\{\\mathbf\{F\}\}\_\{1\}\}\{\{\{\\mathbf\{\\hat\{X\}\}\}\}\_\{0\}\}\+\{\{\\mathbf\{b\}\}\_\{1\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\}\}\\right\)\.\(13\)It can be seen that this nonlinear transformation plays the role of an activation function, enabling the relay to realize both amplification and nonlinear feature mapping over the air\. Therefore, the relay effectively performs one FC layer, where the amplification matrix𝐅1\{\{\{\\bf\{F\}\}\_\{1\}\}\}and the bias vector𝐛1\{\{\{\\bf\{b\}\}\_\{1\}\}\}correspond to the trainable weights and bias of a conventional digital neural network, respectively\.
At the subsequent relay nodesR2,…,RM\{R\_\{2\}\},\\ldots,\{R\_\{M\}\}, similar signal processing operations are performed, where each relay employs its own trainable amplification matrix and bias vector, followed by the inherent nonlinear amplifier characteristic\. Consequently, the entire multi\-hop relay chain can be viewed as a cascade of over\-the\-air FC layers, thereby forming a multi\-layer WPNN\. In this sense, the CSIR MIMO relay network inherently realizes a deep neural architecture in the analog domain, with each relay acting as one neural layer that jointly contributes to the end\-to\-end inference process\.
#### IV\-A2CSIR/T Design
Given𝐇m∈ℂNm×Nm−1\{\{\\mathbf\{H\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times\{N\_\{m\-1\}\}\}\}, we first decompose the channel matrix𝐇m\{\{\\bf\{H\}\}\_\{m\}\}by SVD, yielding𝐇m=𝐔m𝚺m𝐕mH\{\{\\bf\{H\}\}\_\{m\}\}=\{\{\\bf\{U\}\}\_\{m\}\}\{\{\\bf\{\\Sigma\}\}\_\{m\}\}\{\\bf\{V\}\}\_\{m\}^\{H\}, where𝐔m∈ℂNm×Nm\{\{\\mathbf\{U\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times\{N\_\{m\}\}\}\}and𝐕m∈ℂNm−1×Nm−1\{\{\\mathbf\{V\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\-1\}\}\\times\{N\_\{m\-1\}\}\}\}are unitary matrices, and𝚺m∈ℂNm×Nm−1\{\{\\bf\{\\Sigma\}\}\_\{m\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{m\}\}\\times\{N\_\{m\-1\}\}\}\}is a diagonal matrix whose singular values are real and sorted in a descending order\. For a CSIR/T MIMO system, the CSI can be leveraged at the transmitter side, and the output signal at sourceR0R\_\{0\}is given by
𝐗0=PA\(𝐕1𝐅0𝐒c\+𝐛0𝟏LT\)\.\\displaystyle\{\{\\mathbf\{X\}\}\_\{0\}\}=\{\\text\{PA\}\}\\left\(\{\{\{\\mathbf\{V\}\}\_\{1\}\}\{\{\\mathbf\{F\}\}\_\{0\}\}\{\{\\mathbf\{S\}\}\_\{c\}\}\+\{\{\\mathbf\{b\}\}\_\{0\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\}\}\\right\)\.\(14\)At relayR1R\_\{1\}, the received signal is first processed by the combiner𝐔1H\{\\bf\{U\}\}\_\{1\}^\{H\}to decouple the spatial streams according to the SVD structure of𝐇1\{\{\{\\bf\{H\}\}\_\{1\}\}\}\. The resulting signal is then scaled by the pseudo\-inverse of the singular\-value matrix𝚺1\+\{\\bf\{\\Sigma\}\}\_\{1\}^\{\+\}to normalize the power across the eigenmodes\. Subsequently, an amplification matrix𝐅1\{\{\{\\bf\{F\}\}\_\{1\}\}\}is applied to control the relay gain and adapt the transmitted power level\. Finally, the processed signal is multiplied by the right singular matrix𝐕2\{\{\{\\bf\{V\}\}\_\{2\}\}\}, which serves as a pre\-processing operation aligned with the channel𝐇2\{\{\{\\bf\{H\}\}\_\{2\}\}\}toward the next hop\. This process atR1R\_\{1\}can be written as
𝐗1=PA\(𝐕2𝐅1𝚺1\+𝐔1H𝐘1\+𝐛1𝟏LT\)\.\\displaystyle\{\{\\mathbf\{X\}\}\_\{1\}\}=\{\\text\{PA\}\}\\left\(\{\{\{\\mathbf\{V\}\}\_\{2\}\}\{\{\\mathbf\{F\}\}\_\{1\}\}\{\\mathbf\{\\Sigma\}\}\_\{1\}^\{\+\}\{\\mathbf\{U\}\}\_\{1\}^\{H\}\{\{\\mathbf\{Y\}\}\_\{1\}\}\+\{\{\\mathbf\{b\}\}\_\{1\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\}\}\\right\)\.\(15\)This sequential combination of receive combining, singular\-value equalization, amplification, and transmit pre\-processing can be extended to subsequent relays, yielding an end\-to\-end mapping whose parameters\{𝐅m,𝐛m\}m=1M\\left\\\{\{\{\{\\mathbf\{F\}\}\_\{m\}\},\{\{\\mathbf\{b\}\}\_\{m\}\}\}\\right\\\}\_\{m=1\}^\{M\}are jointly trained\.
### IV\-BCommunication Design
Based on Subsection[IV\-A](https://arxiv.org/html/2607.18354#S4.SS1), the signal received at the user after signal processing is given by
𝐗M\+1=\{𝐅M\+1𝐗^M\+𝐛M\+1𝟏LT,CSIR,𝐅M\+1𝚺M\+1\+𝐔M\+1H𝐘M\+1\+𝐛M\+1𝟏LT,CSIR/T,\\displaystyle\{\{\\mathbf\{X\}\}\_\{M\+1\}\}=\\left\\\{\{\\begin\{array\}\[\]\{\*\{20\}\{l\}\}\{\{\{\\mathbf\{F\}\}\_\{M\+1\}\}\{\{\{\\mathbf\{\\hat\{X\}\}\}\}\_\{M\}\}\+\{\{\\mathbf\{b\}\}\_\{M\+1\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\},\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\\qquad\\qquad\\qquad\{\\text\{CSIR\}\},\}\\\\ \{\{\{\\mathbf\{F\}\}\_\{M\+1\}\}\{\\mathbf\{\\Sigma\}\}\_\{M\+1\}^\{\+\}\{\\mathbf\{U\}\}\_\{M\+1\}^\{H\}\{\{\\mathbf\{Y\}\}\_\{M\+1\}\}\+\{\{\\mathbf\{b\}\}\_\{M\+1\}\}\{\\mathbf\{1\}\}\_\{L\}^\{T\},\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\kern 1\.0pt\}\{\\text\{CSIR/T\}\},\}\\end\{array\}\}\\right\.\(18\)where𝐅M\+1∈ℂNM\+1×NM\{\{\\mathbf\{F\}\}\_\{M\+1\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{M\+1\}\}\\times\{N\_\{M\}\}\}\}and𝐛M\+1∈ℂNM\+1×1\{\{\\bf\{b\}\}\_\{M\+1\}\}\\in\{\{\\mathbb\{C\}\}^\{\{N\_\{M\+1\}\}\\times 1\}\}denote the combiner and bias at user, respectively, and𝐗^M\{\{\{\\bf\{\\hat\{X\}\}\}\}\_\{M\}\}denotes the estimated signals transmitted fromRMR\_\{M\}, which is similarly defined in \([12](https://arxiv.org/html/2607.18354#S4.E12)\)\. After obtaining the processed signal𝐗M\+1\{\{\\bf\{X\}\}\_\{M\+1\}\}, it is first converted into its real\-valued representation by concatenating the real and imaginary parts\. The resulting real\-valued tensor is then flattened into a one\-dimensional feature vector, which is fed into a task\-specific read\-out layer to produce the final inference output\. For the image classification example, the read\-out layer consists of an FC layer followed by a softmax activation producing the class probability vector𝐩^\{\{\\mathbf\{\\hat\{p\}\}\}\}over theCCclasses\.
### IV\-CTraining Loss
For the image classification task, a standard cross\-entropy loss is adopted to train the WPNN, expressed as
ℒloss=−∑i=1Cpilog\(p^i\),\\displaystyle\{\{\\cal L\}\_\{\{\\rm\{loss\}\}\}\}=\-\\sum\\limits\_\{i=1\}^\{C\}\{\{p\_\{i\}\}\\log\\left\(\{\{\{\\hat\{p\}\}\_\{i\}\}\}\\right\)\},\(19\)whereCCdenotes the number of classes,pi\{\{p\_\{i\}\}\}is the one\-hot true label, andp^i\{\{\{\\hat\{p\}\}\_\{i\}\}\}is the predicted probability of theiith class\. For other tasks,ℒloss\{\{\\cal L\}\_\{\{\\rm\{loss\}\}\}\}would be replaced by the corresponding task\-specific loss function without changing the underlying architecture or the training procedure\.
## VNumerical results
We evaluate the proposed multi\-hop MIMO relay\-based WPNN in terms of classification accuracy on the Fashion\-MNIST dataset, which contains 60,000 training examples and 10,000 test examples across 10 categories, where each example is a28×2828\\times 28grayscale image\. The path loss between the source node and the user is normalized to one\. The relay nodes are uniformly placed along the line connecting the source node and the user, yieldingαm=\(M\+1\)2\{\\alpha\_\{m\}\}=\{\\left\(\{M\+1\}\\right\)^\{2\}\}\. The signal\-to\-noise ratio \(SNR\) is defined asSNR=10log101σ2\{\\rm\{SNR=10lo\}\}\{\{\\rm\{g\}\}\_\{10\}\}\\frac\{\{\{1\}\}\}\{\{\{\\sigma^\{2\}\}\}\}in dB\. Moreover, the average transmit power at the linear PA is set to11W\. Unless specified otherwise, we setN0=28N\_\{0\}=28,L=14L=14, andNm=32N\_\{m\}=32,m∈\{1,…,M\+1\}m\\in\\left\\\{\{1,\\ldots,M\+1\}\\right\\\}\. The proposed WPNN is trained in an end\-to\-end manner using the Adam optimizer with a learning rate of10−410^\{\-4\}and a batch size of 64\. During training, one independent channel realization is generated for each image transmission\.
The following schemes are considered:
- •Upper bound:This scheme employs an ideal digital neural network with the same layer dimensions and depth as the proposed WPNN\. Each relay\-associated physical layer is replaced by a perfect digital FC layer, without wireless\-channel distortion\. The nonlinear activation is implemented by the standardtanh\(⋅\)\\tanh\(\\cdot\)function\.
- •LS, LPA:This scheme corresponds to the CSIR case, requiring only receiver CSI for each hop\. All PAs are modeled as ideal linear amplifiers\.
- •LS, NPA:This scheme adopts the same LS\-based transceiver architecture as “LS, LPA”, except that the nonlinear Rapp PA model is employed\.
- •SVD, LPA:This scheme corresponds to the CSIR/T case, exploiting both transmitter\- and receiver\-side CSI for each hop\. All PAs are modeled as ideal linear amplifiers\.
- •SVD, NPA:This scheme adopts the same SVD\-based transceiver architecture as “SVD, LPA”, except that the nonlinear Rapp PA model is employed\.
- •Training\-testing PA mismatch \(TPM\):The WPNN is trained assuming ideal linear PAs but is evaluated using nonlinear PAs, without retraining or fine\-tuning\. This mismatch setting is considered for both the LS\- and SVD\-based transceiver architectures\.

Figure 2:Classification accuracy versus SNR\.Fig\.[2](https://arxiv.org/html/2607.18354#S5.F2)compares the classification accuracy of different transceiver schemes under varying SNRs forM=1M=1relay andK=−∞K=\-\\infty\(dB\)\. We can observe that for SNRs above0dB, the LS scheme with nonlinear PA outperforms linear PA due to its superior expressiveness\. In particular, at an SNR of2020dB, the performance of “LS, NPA” scheme is already close to the upper bound\. For the SVD\-based scheme, the linear PA consistently performs better since the nonlinearity prevents full eigenmode decoupling, leaving residual inter\-stream interference\. Moreover, the TPM schemes under both the LS\- and SVD\-based architectures exhibit clear performance losses relative to their nonlinear\-PA counterparts\. These losses become more pronounced as the SNR increases, since the inconsistency between the assumed and actual PA models becomes the dominant performance\-limiting factor when noise is weak\. These results illustrate that by appropriately tuning the hardware parameters at each node, the proposed WPNN can closely approximate the performance of a digital neural network\.

Figure 3:Classification accuracy versus number of relaysMM\.Fig\.[3](https://arxiv.org/html/2607.18354#S5.F3)illustrates the impact of the number of relay nodesMMon the classification accuracy for different transceiver schemes underSNR=20\{\\rm SNR\}=20dB andK=−∞K=\-\\infty\(dB\)\. For the “LS, NPA” scheme, the classification accuracy increases withMM, from 0\.8988 atM=1M=1to 0\.9258 atM=5M=5, and consistently outperforms its linear counterpart across allMM\. In contrast, the “SVD, NPA” scheme exhibits noticeable degradation asMMincreases, since the nonlinear PA distorts the SVD\-based beamforming structure and such distortion accumulates across multiple hops\. Also, for both linear PA schemes, the accuracy remains the same with increasingMMdue to the absence of nonlinear activation, which limits the expressive capability of the cascaded architecture\. It is further observed that the accuracy of TPM schemes under both the LS\- and SVD\-based architectures diminishes withMM\. This is due to the accumulation of errors caused by PA\-model mismatch over multiple hops\. We further consider two transmit\-power constraints for the TPM scheme\. Specifically, the “LS, TPM \(0\.7 W\)” scheme corresponds to the case in which the relays are trained with an average transmit power of 0\.7 W\. Interestingly, “LS, TPM \(0\.7 W\)” outperforms “LS, TPM” because the lower training power keeps the PAs closer to their linear operating region, thereby reducing the mismatch between the linear PA model assumed during training and the nonlinear PA behavior encountered during testing\.

Figure 4:Classification accuracy versus Rician factorKK\.Fig\.[4](https://arxiv.org/html/2607.18354#S5.F4)illustrates the impact of the Rician factorKKon the classification accuracy forM=1M=1underSNR=−20dB\\mathrm\{SNR\}=\-20~\\mathrm\{dB\}andSNR=10dB\\mathrm\{SNR\}=10~\\mathrm\{dB\}\. It can be observed that the classification accuracy generally degrades asKKincreases for both LS\- and SVD\-based schemes\. This is because lower\-rank channels destroy the spatial degrees\-of\-freedom on which the WPNN relies\. AtSNR=10dB\\mathrm\{SNR\}=10~\\mathrm\{dB\}, the accuracy still decreases withKK, but the drop is much less drastic since the noise is no longer the dominating impairment during per\-hop equalization\. For instance, even atK=20dBK=20~\\mathrm\{dB\}, the LS\- and SVD\-based schemes achieve accuracies of0\.69570\.6957and0\.77280\.7728, respectively\.
## VIConclusion
This paper has proposed a deep WPNN realized through a multi\-hop MIMO relay network, in which each relay implements a trainable linear precoding stage followed by the intrinsic nonlinear activation of its PA\. CascadingMMsuch relays yields an over\-the\-air multi\-layer FC network that unifies communication and computation over the same wireless infrastructure\. Two transceiver designs were developed: an LS\-based scheme requiring only receiver CSI, and an SVD\-based scheme exploiting joint transmitter–receiver CSI\. Three main findings emerged from our study: First, PA nonlinearity is a resource, rather than merely an impairment, for over\-the\-air computing\. The LS scheme with a nonlinear PA monotonically improves withMM\. Second, this improvement is architecture\-dependent: nonlinearity disrupts SVD eigenmode decoupling, so a linear\-PA design is preferable for CSIR/T systems\. Third, hardware\-model mismatch and increasing channel rank deficiency \(large RicianKK\) cause errors that compound over hops, motivating mismatch\-aware training\. Two natural extensions are: \(i\) more complex over\-the\-air architectures beyond FC layers, e\.g\., convolutional or attention\-based WPNNs; and \(ii\) robust, mismatch\- and CSI\-uncertainty\-aware end\-to\-end training for deployment under realistic hardware imperfections\.
## References
- \[1\]\(2020\-05\)Federated learning over wireless fading channels\.IEEE Trans\. Wireless Commun\.19\(5\),pp\. 3546–3557\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[2\]I\. Bergel\(2024\-Dec\.\)Non\-linear relay optimization using deep\-learning tools\.IEEE Trans\. Wireless Commun\.23\(12\),pp\. 19289–19301\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1),[§IV\-A1](https://arxiv.org/html/2607.18354#S4.SS1.SSS1.p1.23)\.
- \[3\]C\. Bian, M\. Hua, and D\. Gündüz\(2025\-Nov\.\)Over\-the\-air inference through analog computation over multi\-hop MIMO networks\.IEEE Wireless Commun\. Lett\.14\(11\),pp\. 3739–3743\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[4\]M\. Hua, I\. Bergel, T\. Girici, M\. Di Renzo, and D\. Gunduz\(2026\)Wireless physical neural networks \(WPNNs\): opportunities and challenges\.External Links:[Link](https://arxiv.org/abs/2602.14094.)Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[5\]M\. Hua, C\. Bian, H\. Wu, and D\. Gunduz\(2025\)Implementing neural networks over\-the\-air via reconfigurable intelligent surfaces\.IEEE Trans\. Wireless Commun\.25,pp\. 11562–11576\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[6\]M\. Hua, H\. Wu, and D\. Gündüz\(2026\-Mar\.\)CNNs in the air via reconfigurable intelligent surfaces\.IEEE Wireless Commun\. Lett\.15,pp\. 2124–2128\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[7\]R\. Iten, T\. Metger, H\. Wilming, L\. del Rio, and R\. Renner\(2020\-01\)Discovering physical concepts with neural networks\.Phys\. Rev\. Lett\.124,pp\. 010508\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p1.1)\.
- \[8\]C\. Liu, Q\. Ma, Z\. J\. Luo, Q\. R\. Hong, Q\. Xiao, H\. C\. Zhang, L\. Miao, W\. M\. Yu, Q\. Cheng, L\. Li,et al\.\(2022\-Feb\.\)A programmable diffractive deep neural network based on a digital\-coding metasurface array\.Nat\. Electron\.5\(2\),pp\. 113–122\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[9\]M\. Liu, J\. An, C\. Huang, and C\. Yuen\(2026\)Over\-the\-air ODE\-inspired neural network for dual task\-oriented semantic communications\.IEEE Trans\. Cogn\. Commun\. Netw\.12,pp\. 805–819\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[10\]A\. Momeniet al\.\(2025\)Training of physical neural networks\.Nature645\(8079\),pp\. 53–61\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p1.1)\.
- \[11\]G\. Sanchezet al\.\(2023\-Dec\.\)AirNN: over\-the\-air computation for neural networks via reconfigurable intelligent surfaces\.IEEE/ACM Tran\. Netw\.31\(6\),pp\. 2470–2482\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[12\]K\. Stylianopoulos, P\. Di Lorenzo, and G\. C\. Alexandropoulos\(2026\-Mar\.\)Over\-the\-air edge inference via end\-to\-end metasurfaces\-integrated artificial neural networks\.IEEE Trans\. Wireless Commun\.25,pp\. 13818–13834\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[13\]R\. Wang, Y\. Jiang, and W\. Zhang\(2022\-Apr\.\)Distributed learning for MIMO relay networks\.IEEE J\. Sel\. Topics Signal Process\.16\(3\),pp\. 343–357\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1),[§IV\-A1](https://arxiv.org/html/2607.18354#S4.SS1.SSS1.p1.17)\.
- \[14\]L\. G\. Wright, T\. Onodera, M\. M\. Stein, T\. Wang, D\. T\. Schachter, Z\. Hu, and P\. L\. McMahon\(2022\)Deep physical neural networks trained with backpropagation\.Nature601\(7894\),pp\. 549–555\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p1.1)\.
- \[15\]Y\. Yang, Z\. Zhang, Y\. Tian, Z\. Yang, R\. Jin, L\. Liu, and C\. Huang\(2024\)Realizing over\-the\-air neural networks in RIS\-assisted MIMO communication systems\.InIEEE WCNC, Dubai, UAE,pp\. 1–5\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.
- \[16\]J\. Zhang, H\. Chen, and D\. M\. Blough\(2024\)A radio\-frequency\-based 2\-D convolutional layer using transmissive intelligent surfaces\.InIEEE VTC, Washington, DC, USA,pp\. 1–7\.Cited by:[§I](https://arxiv.org/html/2607.18354#S1.p2.1)\.Similar Articles
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
This paper introduces Learned Relay Representations (Relay), a method that allows masked diffusion models to propagate latent information across denoising steps, overcoming the hard reset problem and improving performance-latency trade-offs. The method is shown to outperform standard supervised finetuning on coding tasks while reducing inference latency by up to 32%.
Building The Ph(ysical)AI Layer Of Machine Intelligence
Researchers at MIT Lincoln Laboratory propose 'principle-driven foundation models' that encode signal-theoretic physical principles (Fourier decomposition, energy conservation, symmetry) instead of learning statistical correlations from large paired datasets. Trained exclusively on RF data, their 1.99M parameter frozen encoder achieves 77.7% average accuracy across 15 diverse tasks spanning audio, images, text, and video without any fine-tuning on target domains.
PE-MHL: Physics-Encoded Modular Hybrid Layers for Scalable Learning of Complex Systems
This paper proposes PE-MHL, a Physics-Encoded Modular Hybrid Layer framework that incrementally refines a physics-based model with data-driven sub-models, providing theoretical convergence guarantees and outperforming monolithic networks on control benchmarks.
Low-power analogue neural networks with trainable nonlinear connections for continuous control
This paper presents low-power analogue neural networks that place trainable nonlinear functions on connections, inspired by Kolmogorov-Arnold networks, enabling efficient continuous control tasks with far fewer nodes and connections than multilayer perceptrons, demonstrated on hardware with projected microWatt power.
Phases in a class of associative memories via hidden neurons
The paper analyzes a class of associative memories with hidden neurons, deriving phase diagrams and storage capacities using replica methods and linking to softmax attention in transformers.