CRFCAN: A Complex-Valued Cross-Domain Residual Network for Joint Channel and Phase Noise Estimation in Sub-THz OFDM Systems
Summary
Proposes CRFCAN, a complex-valued residual network for joint channel and phase noise estimation in sub-THz OFDM systems, which achieves end-to-end recovery with physics-inspired structure and outperforms conventional and state-of-the-art deep learning models.
View Cached Full Text
Cached at: 09/14/26, 08:38 AM
# CRFCAN: A Complex-Valued Cross-Domain Residual Network for Joint Channel and Phase Noise Estimation in Sub-THz OFDM Systems
Source: [https://arxiv.org/html/2609.12244](https://arxiv.org/html/2609.12244)
Xiaodai Dong††thanks:The authors are with the Department of Electrical and Computer Engineering, University of Victoria, Victoria, BC V8P 5C2, Canada\.††thanks:Corresponding author: Xiaodai Dong \(email: xdong@ece\.uvic\.ca\)\.
###### Abstract
In sub\-terahertz \(sub\-THz\) communications, the coupling of ultra\-wide bandwidth and severe phase noise \(PN\) impairments renders conventional joint channel and PN estimation highly complex and computationally prohibitive\. To address this, we propose CRFCAN, a complex\-valued residual FFT convolutional attention network designed for joint channel and PN estimation\. Unlike existing deep learning schemes that rely on cascaded networks or hybrid frameworks combining neural networks with conventional iterative estimators, CRFCAN performs joint recovery in a truly end\-to\-end fashion through a physics\-inspired cross\-domain structure\. Specifically, Fast Fourier Transform \(FFT\) and inverse FFT modules are embedded within residual groups to enable iterative feature interaction across the time and frequency domains, thereby capturing both frequency\-selective fading and time\-varying phase distortions\. In addition, two dedicated residual blocks are introduced for complex feature extraction and multiplicative phase\-distortion modeling, respectively\. A physics\-aware PN output tail with soft normalization is further employed to improve estimation stability while preserving the physical characteristics of the effective PN process\. Simulation results demonstrate that CRFCAN significantly outperforms conventional algorithms and state\-of\-the\-art deep learning models in terms of normalized mean square error \(NMSE\) and bit error rate \(BER\)\. Notably, CRFCAN achieves superior performance with single\-shot, fixed\-complexity inference and generalizes well to unseen PN models without fine\-tuning, highlighting its robustness and practicality for sub\-THz receivers\.
###### Index Terms:
Attention mechanism, complex\-valued neural networks \(CVNNs\), joint channel and phase noise estimation, orthogonal frequency\-division multiplexing \(OFDM\), sub\-terahertz \(sub\-THz\) communications, physics\-aware deep learning\.
## IINTRODUCTION
The rapid proliferation of data\-intensive applications, such as ultra\-high\-definition video streaming and immersive extended reality, has driven growing interest in the sub\-terahertz \(sub\-THz\) band \(100–300 GHz\) as a key enabling technology for future 6G wireless networks\[[1](https://arxiv.org/html/2609.12244#bib.bib22)\]\. Operating in the sub\-THz spectrum provides access to unprecedented multi\-gigahertz bandwidths, thereby supporting terabit\-per\-second \(Tbps\) data rates\[[2](https://arxiv.org/html/2609.12244#bib.bib23)\]\. However, these gains come at the expense of pronounced hardware impairments, among which phase noise \(PN\) arising from high\-frequency local oscillators is particularly severe\[[3](https://arxiv.org/html/2609.12244#bib.bib24)\]\. In contrast to lower\-frequency systems, PN in sub\-THz communications can induce substantial common phase error \(CPE\) and inter\-carrier interference \(ICI\) in multicarrier transmissions, while coexisting with highly frequency\-selective fading channels\[[4](https://arxiv.org/html/2609.12244#bib.bib42)\]\. As a result, accurate joint estimation of the channel state information \(CSI\) and PN becomes essential for reliable system operation\. Nevertheless, conventional iterative estimation approaches often entail prohibitive computational complexity when applied to the massive signal dimensions associated with sub\-THz bandwidths\[[5](https://arxiv.org/html/2609.12244#bib.bib4)\]\.
To mitigate the effects of imperfect CSI and PN, a variety of estimation schemes have been developed\. Classical approaches, including least\-squares \(LS\)\[[6](https://arxiv.org/html/2609.12244#bib.bib1)\]and linear minimum mean\-square error \(LMMSE\)\[[7](https://arxiv.org/html/2609.12244#bib.bib3)\]estimators, typically require substantial pilot overhead or incur excessive computational complexity when scaled to sub\-THz bandwidths\[[5](https://arxiv.org/html/2609.12244#bib.bib4)\]\. Moreover, under the combined effects of highly frequency\-selective fading and strong PN at sub\-THz frequencies, these model\-based estimators often exhibit slow or stalled convergence, resulting in a pronounced error floor even with additional iterations\. To specifically address oscillator impairments, sophisticated statistical frameworks have been proposed, such as the joint estimation of channel, carrier frequency offset \(CFO\), and PN using expectation conditional maximization \(ECM\)\[[8](https://arxiv.org/html/2609.12244#bib.bib2)\]and extended Kalman filtering \(EKF\)\[[9](https://arxiv.org/html/2609.12244#bib.bib26),[10](https://arxiv.org/html/2609.12244#bib.bib27)\]\. While these methods provide a theoretical performance benchmark, their reliance on recursive filtering and iterative step\-by\-step optimizations introduces significant processing latency in high\-speed links\[[11](https://arxiv.org/html/2609.12244#bib.bib25),[12](https://arxiv.org/html/2609.12244#bib.bib28)\]\. Moreover, the Jacobian\-based linearizations required for tracking non\-linear PN trajectories can become computationally prohibitive as the signal dimension increases\[[8](https://arxiv.org/html/2609.12244#bib.bib2)\]\. Parallel to these efforts, compressed sensing \(CS\)\-based techniques\[[13](https://arxiv.org/html/2609.12244#bib.bib5),[14](https://arxiv.org/html/2609.12244#bib.bib40)\]leverage the inherent channel sparsity to reduce overhead\. However, phase noise \(PN\) introduces additive correlated perturbations to the sensing matrix\[[15](https://arxiv.org/html/2609.12244#bib.bib6)\], which severely degrades the recovery reliability of conventional pursuit algorithms\. To combat such distortions, advanced schemes like PN\-aware sparse Bayesian learning \(PNA\-SBL\)\[[15](https://arxiv.org/html/2609.12244#bib.bib6)\]have been developed; yet, their intensive iterative nature results in prohibitive latency, rendering them unsuitable for real\-time sub\-THz applications\.
More recently, deep learning \(DL\) has emerged as a promising alternative owing to its low online inference complexity\. Significant progress has been made in applying DL to channel estimation\[[16](https://arxiv.org/html/2609.12244#bib.bib29),[17](https://arxiv.org/html/2609.12244#bib.bib30)\]and phase noise compensation\[[18](https://arxiv.org/html/2609.12244#bib.bib31),[19](https://arxiv.org/html/2609.12244#bib.bib32)\]independently\. However, when it comes to joint CSI and PN estimation, most existing DL\-based approaches either employ cascaded network architectures\[[20](https://arxiv.org/html/2609.12244#bib.bib7)\]that fail to capture their intrinsic coupling, or integrate neural networks with conventional iterative or CS\-based algorithms, in which the network is used only for channel estimation while PN is handled by model\-based methods\[[21](https://arxiv.org/html/2609.12244#bib.bib8)\]\. From a learning perspective, such hybrid designs expose the network to incomplete system information, thereby limiting its ability to fully exploit the representational power of deep models\. Moreover, most existing approaches are built upon conventional real\-valued convolutional layers, which inherently overlook the fundamental phase relationships present in complex\-valued wireless signals\[[22](https://arxiv.org/html/2609.12244#bib.bib9)\]\. Consequently, most DL models operate as black boxes, lacking physics\-inspired domain knowledge to effectively accommodate the distinctive time–frequency characteristics of sub\-THz impairments\.
Despite the initial success of DL\-based estimators, several critical issues remain insufficiently addressed\. First, most existing models are built upon real\-valued convolutional operations that treat the in\-phase \(I\) and quadrature \(Q\) components as independent feature channels\[[23](https://arxiv.org/html/2609.12244#bib.bib33)\]\. Such an additive feature\-mapping paradigm is inherently mismatched with the multiplicative nature of PN, making it difficult for the network to explicitly model the geometric rotation in the complex plane\. Second, while PN manifests as a time\-varying phase rotation in the time domain and induces structured ICI in the frequency domain, many current DL architectures operate predominantly within a single domain\. This design choice limits their ability to disentangle the coupled effects of frequency\-selective fading and PN, thereby leading to suboptimal performance in wideband sub\-THz systems\[[22](https://arxiv.org/html/2609.12244#bib.bib9)\]\. Finally, in the absence of domain\-specific physical constraints, these black\-box models often suffer from limited generalization capability and provide little interpretability, which undermines their reliability under severe hardware impairments\[[24](https://arxiv.org/html/2609.12244#bib.bib34)\]\.
To bridge these gaps, we propose CRFCAN, a novel complex\-valued residual framework designed for joint CSI and PN estimation\. By integrating physically\-inspired modules with a cross\-domain architecture, our approach realizes a robust and low\-latency solution for sub\-THz communications\. The main contributions of this paper are summarized as follows:
- •Physically Inspired Complex\-Valued Rotation Residual Learning: We develop a specialized residual channel attention block, termed RCAB\-PhaseRotation, which explicitly models phase noise \(PN\) as a multiplicative rotation in the complex plane rather than as an additive feature perturbation\. By embedding Euler’s formulation into a complex\-valued residual learning framework, the proposed module enables fine\-grained phase compensation through an adaptive gating mechanism\. This physically consistent design aligns the network behavior with the underlying signal impairment model, thereby enhancing the stability and accuracy of phase recovery\.
- •Single\-Shot Cross\-Domain Estimation Architecture: We propose a cross\-domain architecture that tightly integrates FFT and IFFT operations into a complex\-valued deep network, enabling joint processing in both the time and frequency domains\. This design allows frequency\-selective fading to be captured in the frequency domain while simultaneously tracking time\-varying phase rotations in the time domain\. In contrast to iterative or cascaded schemes, the proposed framework performs joint CSI and PN estimation in a single forward pass, substantially reducing processing latency\.
- •End\-to\-End Joint Optimization with Physical Constraints: We present an end\-to\-end training paradigm that explicitly accounts for the intrinsic coupling between CSI and PN\. By incorporating domain\-specific physical constraints, including a soft\-normalization output tail, the network search space is regularized to respect the geometric structure of complex\-valued signals\. Extensive simulations demonstrate that the proposed CRFCAN effectively mitigates the error floors observed in conventional LMMSE\- and CS\-based estimators, achieving robust performance under severe sub\-THz hardware impairments\.
The remainder of this paper is organized as follows\. Section II introduces the sub\-THz system model and characterizes the phase noise \(PN\) impairment\. Section III presents the proposed CRFCAN architecture, with particular emphasis on the RCAB\-PhaseRotation module and the cross\-domain processing pipeline\. Section IV reports the simulation results and performance evaluation, including comparisons with state\-of\-the\-art methods and an analysis of computational complexity\. Finally, Section V concludes the paper and outlines potential directions for future research\.
Fig\. 1:Proposed sub\-THz OFDM system model with joint channel and phase noise impairments\.
## IISystem Model
As illustrated in Fig\.[1](https://arxiv.org/html/2609.12244#S1.F1), we consider a sub\-THz Orthogonal Frequency\-Division Multiplexing \(OFDM\) transmission system that incorporates analog\-domain pulse shaping and hardware impairments\. The corresponding signal processing chain is described as follows\.
### II\-AOFDM System Model
At the transmitter, information bits are first mapped onto complex\-valued constellation symbols, which are grouped into frequency\-domain OFDM blocks ofNcN\_\{c\}subcarriers, denoted by
𝐗=\[X0,X1,…,XNc−1\]T\.\\mathbf\{X\}=\[X\_\{0\},X\_\{1\},\\dots,X\_\{N\_\{c\}\-1\}\]^\{T\}\.\(1\)Each block is transformed into the time domain through anNcN\_\{c\}\-point IFFT, represented by the operator𝐅H\\mathbf\{F\}^\{H\}in Fig\.[1](https://arxiv.org/html/2609.12244#S1.F1)\. A cyclic prefix \(CP\) of lengthPPis then appended to mitigate inter\-symbol interference \(ISI\) caused by multipath propagation\. The resulting discrete\-time OFDM symbol can be expressed as
x\[n\]=1Nc∑k=0Nc−1Xkej2πnkNc,n=0,1,…,Nc−1,x\[n\]=\\frac\{1\}\{\\sqrt\{N\_\{c\}\}\}\\sum\_\{k=0\}^\{N\_\{c\}\-1\}X\_\{k\}e^\{j\\frac\{2\\pi nk\}\{N\_\{c\}\}\},\\quad n=0,1,\\dots,N\_\{c\}\-1,\(2\)whereXkX\_\{k\}denotes the data symbol transmitted on thekk\-th subcarrier\. After CP insertion, the symbol is extended toNc\+PN\_\{c\}\+Psamples\.
To emulate the analog\-domain transmission chain and enable pulse shaping, the CP\-appended signal is upsampled by a factor ofMMand subsequently filtered by a transmit pulse\-shaping filter, such as a root\-raised cosine \(RRC\) filter\. The resulting continuous\-time baseband signal is denoted byx\(t\)x\(t\)and is transmitted over the sub\-THz propagation channel\.
When propagating through the channel with continuous\-time impulse responseh\(t\)h\(t\), the signal is impaired by time\-varying PNϕ\(t\)\\phi\(t\)introduced by the local oscillator\. The received continuous\-time baseband signal at the antenna can be written as
y\(t\)=ejϕ\(t\)∫−∞∞h\(τ\)x\(t−τ\)𝑑τ\+w¯\(t\),y\(t\)=e^\{j\\phi\(t\)\}\\int\_\{\-\\infty\}^\{\\infty\}h\(\\tau\)\\,x\(t\-\\tau\)\\,d\\tau\+\\bar\{w\}\(t\),\(3\)wherew¯\(t\)\\bar\{w\}\(t\)denotes additive white Gaussian noise \(AWGN\) at the analog front end\.
At the receiver, as illustrated in Fig\.[1](https://arxiv.org/html/2609.12244#S1.F1), the signaly\(t\)y\(t\)is first passed through a receive filter matched to the transmit pulse\-shaping filter, followed by sampling with periodTsT\_\{s\}, downsampling, and CP removal\. Assuming that the CP lengthPPis sufficient to cover the maximum channel delay spread, the resulting discrete\-time received block ofNcN\_\{c\}samples can be expressed as
y\[n\]=p\[n\]∑r=0Nc−1x\[r\]h\[\(n−r\)Nc\]\+w\[n\],n=0,1,…,Nc−1,y\[n\]=p\[n\]\\sum\_\{r=0\}^\{N\_\{c\}\-1\}x\[r\]\\,h\[\(n\-r\)\_\{N\_\{c\}\}\]\+w\[n\],\\quad n=0,1,\\dots,N\_\{c\}\-1,\(4\)wherep\[n\]=a\[n\]ejϕ\[nTs\]p\[n\]=a\[n\]e^\{j\\phi\[nT\_\{s\}\]\}denotes an effective discrete\-time multiplicative impairment induced by phase noise after pulse shaping, matched filtering, and sampling\. Although this formulation introduces an additional amplitude degree of freedom compared to a strict unit\-modulus model, explicitly separatingp\[n\]p\[n\]from the propagation channel preserves the inherent structure of the channel and leads to more reliable joint channel and phase\-noise estimation\.
The time\-domain formulation in \([4](https://arxiv.org/html/2609.12244#S2.E4)\) highlights the multiplicative nature of PN and its coupling with the channel convolution\. Applying anNcN\_\{c\}\-point FFT toy\[n\]y\[n\]yields the frequency\-domain representation
Y\[k\]=∑m=0Nc−1X\[m\]H\[m\]P\[k−m\]\+W\[k\],Y\[k\]=\\sum\_\{m=0\}^\{N\_\{c\}\-1\}X\[m\]\\,H\[m\]\\,P\[k\-m\]\+W\[k\],\(5\)whereH\[m\]H\[m\]is the channel frequency response andP\[k\]P\[k\]characterizes the spectral spreading induced by PN, giving rise to CPE and ICI\. The resulting frequency\-domain observationsY\[k\]Y\[k\], together with time\-domain features, serve as the inputs to the proposedCRFCANfor joint CSI and PN estimation\.
### II\-BSub\-THz Channel Model
To evaluate the performance of the proposed system under realistic propagation conditions, we adopt the tapped delay line C \(TDL\-C\) channel model specified in the 3GPP TR 38\.901 technical report\[[25](https://arxiv.org/html/2609.12244#bib.bib10)\], which is defined for carrier frequencies ranging from 0\.5 to 100 GHz\. The TDL\-C profile is designed to characterize non\-line\-of\-sight \(NLOS\) propagation scenarios with moderate delay spreads and has been widely used for sub\-THz and millimeter\-wave evaluations\. Considering a single\-input single\-output \(SISO\) configuration, the continuous\-time channel impulse response \(CIR\)h\(τ,t\)h\(\\tau,t\)is modeled as a superposition ofLtapL\_\{\\text\{tap\}\}discrete propagation paths:
h\(τ,t\)=∑l=0Ltap−1al\(t\)δ\(τ−τl\),h\(\\tau,t\)=\\sum\_\{l=0\}^\{L\_\{\\text\{tap\}\}\-1\}a\_\{l\}\(t\)\\,\\delta\(\\tau\-\\tau\_\{l\}\),\(6\)whereal\(t\)a\_\{l\}\(t\)andτl\\tau\_\{l\}denote the time\-varying complex gain and propagation delay of thell\-th tap, respectively\.
Each tap coefficiental\(t\)a\_\{l\}\(t\)is modeled as a zero\-mean, wide\-sense stationary \(WSS\) narrowband complex Gaussian process, corresponding to Rayleigh fading\. The temporal variation ofal\(t\)a\_\{l\}\(t\)is governed by Doppler effects, with its power spectral density \(PSD\) following the classical Jakes spectrum\[[26](https://arxiv.org/html/2609.12244#bib.bib35)\]\. The maximum Doppler frequency is given byfd=vfc/cf\_\{d\}=vf\_\{c\}/c, wherevvdenotes the relative velocity between the transmitter and receiver,fcf\_\{c\}is the carrier frequency, andccis the speed of light\.
According to\[[25](https://arxiv.org/html/2609.12244#bib.bib10)\], the normalized tap delays\{τl,norm\}\\\{\\tau\_\{l,\\mathrm\{norm\}\}\\\}are scaled by the target root\-mean\-square \(RMS\) delay spreadστ\\sigma\_\{\\tau\}to accommodate different deployment scenarios, such as urban microcell or urban macrocell environments, i\.e\.,
τl=τl,normστ\.\\tau\_\{l\}=\\tau\_\{l,\\mathrm\{norm\}\}\\,\\sigma\_\{\\tau\}\.\(7\)The average power of each tap, defined asPl=𝔼\[\|al\(t\)\|2\]P\_\{l\}=\\mathbb\{E\}\\\!\\left\[\|a\_\{l\}\(t\)\|^\{2\}\\right\], follows the exponential power delay profile \(PDP\) specified by the TDL\-C model\.
In the discrete\-time system formulation presented in Subsection[II\-A](https://arxiv.org/html/2609.12244#S2.SS1), the equivalent discrete\-time channel impulse responseh\[n\]h\[n\]in \([4](https://arxiv.org/html/2609.12244#S2.E4)\) is obtained by convolving the physical channelh\(τ,t\)h\(\\tau,t\)with the combined transmit and receive pulse\-shaping filter responseg\(τ\)g\(\\tau\)and sampling at periodTsT\_\{s\}, yielding
h\[n\]=∑l=0Ltap−1al\(nTs\)g\(nTs−τl\)\.h\[n\]=\\sum\_\{l=0\}^\{L\_\{\\text\{tap\}\}\-1\}a\_\{l\}\(nT\_\{s\}\)\\,g\(nT\_\{s\}\-\\tau\_\{l\}\)\.\(8\)This representation consistently incorporates the effects of multipath fading, Doppler spread, and pulse\-shaping\-induced inter\-symbol interference \(ISI\) into the discrete\-time baseband model\.
Fig\. 2:Overall architecture of the proposed CRFCAN, consisting of a complex\-convolution front\-end, alternating IFFT/FFT residual groups for cross\-domain refinement, and a dual\-output tail for joint CSI and PN estimation\.
### II\-CPhase Noise Model
At sub\-THz frequencies, the stability of the local oscillator \(LO\) becomes a dominant performance\-limiting factor\. Following the reference model in 3GPP TR 38\.803\[[27](https://arxiv.org/html/2609.12244#bib.bib11)\]and the analytical framework in\[[28](https://arxiv.org/html/2609.12244#bib.bib12)\], the PN processϕ\(t\)\\phi\(t\)is characterized in the frequency domain by its single\-sideband \(SSB\) PSD, denoted asL\(f\)L\(f\)in dBc/Hz\. The relationship betweenL\(f\)L\(f\)and the linear\-scale PSDSϕ\(f\)S\_\{\\phi\}\(f\)is given by
L\(f\)=10log10\(Sϕ\(f\)\)\+Δcal,L\(f\)=10\\log\_\{10\}\\\!\\left\(S\_\{\\phi\}\(f\)\\right\)\+\\Delta\_\{\\text\{cal\}\},\(9\)whereSϕ\(f\)S\_\{\\phi\}\(f\)represents the PN PSD in linear scale \(rad2/Hz\)\. To align the generalized 3GPP model with practical hardware characteristics, an implementation\-specific calibration factorΔcal\\Delta\_\{\\text\{cal\}\}is introduced\. As emphasized in\[[28](https://arxiv.org/html/2609.12244#bib.bib12)\], such calibration is essential for bridging the gap between idealized analytical models and measurement\-driven observations\. Physically,Δcal\\Delta\_\{\\text\{cal\}\}captures the enhanced noise\-suppression capability enabled by optimized phase\-locked loop \(PLL\) designs, as well as the reduced thermal noise floor of practical sub\-THz front\-end components\.
The PN spectrum of a PLL\-based oscillator can be interpreted as the response of a stable linear filter driven by white Gaussian noise\. Consistent with the 3GPP specification, the linear\-scale PSDSϕ\(f\)S\_\{\\phi\}\(f\)is modeled as a transfer function characterizing the synthesizer behavior, which takes the following fractional product form:
Sϕ\(f\)=PSD0⋅∏n=1N\[1\+\(f/fz,n\)az,n\]∏m=1M\[1\+\(f/fp,m\)ap,m\],S\_\{\\phi\}\(f\)=\\text\{PSD\}\_\{0\}\\cdot\\frac\{\\prod\_\{n=1\}^\{N\}\\left\[1\+\(f/f\_\{z,n\}\)^\{a\_\{z,n\}\}\\right\]\}\{\\prod\_\{m=1\}^\{M\}\\left\[1\+\(f/f\_\{p,m\}\)^\{a\_\{p,m\}\}\\right\]\},\(10\)whereffdenotes the frequency offset from the carrier andPSD0\\text\{PSD\}\_\{0\}represents the white noise floor in linear scale\. In this formulation, the fractional product captures the interaction between the intrinsic noise of the free\-running oscillator and the tracking behavior of the PLL loop\. The parameter sets\{fz,n\}\\\{f\_\{z,n\}\\\}and\{fp,m\}\\\{f\_\{p,m\}\\\}correspond to the zero and pole frequencies, respectively\. Physically, the poles\{fp,m\}\\\{f\_\{p,m\}\\\}reflect the bandwidth limitations of the oscillator and the low\-pass characteristics of the loop filter, whereas the zeros\{fz,n\}\\\{f\_\{z,n\}\\\}represent compensation mechanisms introduced within the PLL to ensure loop stability\.
TABLE I:PLL\-based phase noise model parameters \(3GPP TR 38\.803 RAN4\[[27](https://arxiv.org/html/2609.12244#bib.bib11)\]\)The PN model specified in TR 38\.803 is defined with respect to a reference carrier frequencyfbase=29\.55f\_\{\\mathrm\{base\}\}=29\.55GHz\. When the system operates at a different carrier frequencyfcf\_\{c\}, the PN PSD must be appropriately scaled to account for the degradation of the LO during frequency up\-conversion\. Specifically, the white noise floor of the PN spectrum at the target carrier frequencyfcf\_\{c\}is adjusted as
PSD0\(fc\)=PSD0\(fbase\)\+20log10\(fcfbase\)\+ΔFoM,\\mathrm\{PSD\}\_\{0\}\(f\_\{c\}\)=\\mathrm\{PSD\}\_\{0\}\(f\_\{\\mathrm\{base\}\}\)\+20\\log\_\{10\}\\\!\\left\(\\frac\{f\_\{c\}\}\{f\_\{\\mathrm\{base\}\}\}\\right\)\+\\Delta\_\{\\mathrm\{FoM\}\},\(11\)where the term20log10\(fc/fbase\)20\\log\_\{10\}\(f\_\{c\}/f\_\{\\mathrm\{base\}\}\)accounts for the theoretical increase in phase noise due to frequency multiplication\. The additional termΔFoM\\Delta\_\{\\mathrm\{FoM\}\}captures the frequency\-dependent degradation of the oscillator figure of merit \(FoM\), which arises from reduced resonator quality factors and increased losses in frequency multiplication chains at higher carrier frequencies\.
Following common practice, this additional degradation is modeled as a logarithmic function of the frequency scaling factor,
ΔFoM=βlog10\(fcfbase\),\\Delta\_\{\\mathrm\{FoM\}\}=\\beta\\log\_\{10\}\\\!\\left\(\\frac\{f\_\{c\}\}\{f\_\{\\mathrm\{base\}\}\}\\right\),\(12\)whereβ\\betais a slope coefficient \(e\.g\.,99dB/decade\) that characterizes hardware\-specific performance loss beyond the ideal20log10\(⋅\)20\\log\_\{10\}\(\\cdot\)scaling\.
With the frequency\-scaled noise floorPSD0\(fc\)\\mathrm\{PSD\}\_\{0\}\(f\_\{c\}\), the linear\-scale PN PSDSϕ\(f\)S\_\{\\phi\}\(f\)in \([10](https://arxiv.org/html/2609.12244#S2.E10)\) is updated accordingly\. This procedure preserves the standardized multi\-segment slope characteristics defined in TR 38\.803, while accurately reflecting the increased phase instability associated with higher carrier frequencies\.
## IIIProposed Network
### III\-ANetwork Architecture
In this subsection, we present the overall architecture of the proposed CRFCAN for joint estimation ofH^\\hat\{H\}andp^\\hat\{p\}\. The network takes as input a structured complex\-valued tensor of size2×Nc×Nt2\\times N\_\{c\}\\times N\_\{t\}, whereNcN\_\{c\}andNtN\_\{t\}denote the numbers of subcarriers and consecutive OFDM symbols, respectively\. The two input channels correspond to the received signal and the designed pilot sequence, where the data positions in the pilot channel are filled with zeros\. A comb\-type pilot pattern\[[29](https://arxiv.org/html/2609.12244#bib.bib44)\]is employed across the frequency domain and is cyclically shifted along the time axis over consecutive OFDM symbols, as illustrated in Fig\.[3](https://arxiv.org/html/2609.12244#S3.F3)\.
As shown in Fig\.[2](https://arxiv.org/html/2609.12244#S2.F2), CRFCAN consists of three functional stages: 1\) a complex\-convolution\-based front\-end for high\-dimensional feature extraction; 2\) across\-domain backbonecomposed of alternating IFFT Residual Group and FFT Residual Group modules; and 3\) aphysics\-aware back\-endwith two dedicated tail modules\. The network outputs the estimated channel frequency responseH^\\hat\{H\}and the phase noise trajectoryp^\\hat\{p\}through these two tail modules, respectively\.
#### III\-A1Complex\-Valued Neural Network
In sub\-THz communications, signals naturally reside in the complex domain, where phase information plays a central role in characterizing hardware impairments such as PN and frequency\-selective fading\. Real\-valued neural networks \(RVNNs\) typically process the in\-phase and quadrature components as independent channels through concatenation or stacking\[[30](https://arxiv.org/html/2609.12244#bib.bib36)\], which increases the degrees of freedom and decouples the intrinsic correlations between the real and imaginary components\[[31](https://arxiv.org/html/2609.12244#bib.bib37),[32](https://arxiv.org/html/2609.12244#bib.bib38)\]\. In contrast, the proposed CRFCAN adopts a CVNN formulation with features represented as𝐳=𝐩\+j𝐪\\mathbf\{z\}=\\mathbf\{p\}\+j\\mathbf\{q\}\.
In complex\-valued convolutional layers, kernels are modeled as complex operators that obey complex multiplication\[[33](https://arxiv.org/html/2609.12244#bib.bib13)\]:
𝐖∗𝐬=\(𝐀∗𝐩−𝐁∗𝐪\)\+j\(𝐁∗𝐩\+𝐀∗𝐪\),\\mathbf\{W\}\\ast\\mathbf\{s\}=\(\\mathbf\{A\}\\ast\\mathbf\{p\}\-\\mathbf\{B\}\\ast\\mathbf\{q\}\)\+j\(\\mathbf\{B\}\\ast\\mathbf\{p\}\+\\mathbf\{A\}\\ast\\mathbf\{q\}\),\(13\)where𝐖=𝐀\+j𝐁\\mathbf\{W\}=\\mathbf\{A\}\+j\\mathbf\{B\}and𝐬=𝐩\+j𝐪\\mathbf\{s\}=\\mathbf\{p\}\+j\\mathbf\{q\}\. This structured coupling constrains the learnable parameters and reduces degrees of freedom compared to an unconstrained RVNN of comparable size\. Regarding nonlinear modeling, activation design in the complex domain is fundamentally constrained by Liouville’s theorem, which implies that any bounded entire complex\-differentiable function must be constant\[[34](https://arxiv.org/html/2609.12244#bib.bib14)\]\. We therefore employ the split\-type activation CReLU, defined as
ℂReLU\(z\)=ReLU\(ℜ\(z\)\)\+jReLU\(ℑ\(z\)\)\.\\mathbb\{C\}\\mathrm\{ReLU\}\(z\)=\\mathrm\{ReLU\}\(\\mathfrak\{R\}\(z\)\)\+j\\,\\mathrm\{ReLU\}\(\\mathfrak\{I\}\(z\)\)\.\(14\)
Since the CRFCAN architecture explicitly incorporates FFT/IFFT\-based cross\-domain transitions and multiplicative phase\-rotation modules, adopting a complex\-valued formulation is a structural requirement to preserve the analytical relationships between magnitude and phase across the time–frequency domains\.
Fig\. 3:Comb\-type pilot pattern with cyclic shift across consecutive OFDM symbols\.
#### III\-A2Cross\-Domain Signal Flow
The architectural core of CRFCAN is a cross\-domain signal reasoning framework designed to disentangle the intertwined impairments of CSI and PN\. The necessity of cross\-domain processing stems from the dual nature of sub\-THz distortions: PNϕ\[n\]\\phi\[n\]is a time\-varying process with strong temporal correlations, whereas the sub\-THz channel𝐇\\mathbf\{H\}exhibits frequency\-selective fading with a structured spectral response\. Moreover, the multiplicative PN in the time domain becomes a convolutional distortion in the frequency domain after the FFT, breaking subcarrier orthogonality and inducing ICI\. Therefore, joint CSI–PN estimation inherently benefits from cross\-domain signal reasoning\. Although joint time–frequency processing has shown promise in other fields\[[35](https://arxiv.org/html/2609.12244#bib.bib41)\], it has not been applied to joint CSI–PN estimation in wireless communications\.
To facilitate domain interaction, CRFCAN integrates FFT and IFFT layers as deterministic and information\-preserving mapping operators\. Within CRFCAN, the signal flow alternates between domains: theIFFT Residual Grouprefines features in the frequency domain and projects them to the time domain via an IFFT, while theFFT Residual Groupperforms time\-domain refinement followed by an FFT\-based mapping back to the frequency domain\.
All intermediate feature maps are organized on a time–frequency \(TF\) grid and represented as
𝐅∈ℂC×Nc×Nt,\\mathbf\{F\}\\in\\mathbb\{C\}^\{C\\times N\_\{c\}\\times N\_\{t\}\},\(15\)whereCCdenotes the number of feature channels\. We use the superscripts\(f\)\(f\)and\(t\)\(t\)to distinguish two equivalent TF representations with the same tensor size:𝐅\(f\)\\mathbf\{F\}^\{\(f\)\}is indexed by the subcarrier indexkk, whereas𝐅\(t\)\\mathbf\{F\}^\{\(t\)\}is indexed by the sample indexnnwithin each OFDM symbol\. In CRFCAN, FFT and IFFT are applied only along the subcarrier dimensionNcN\_\{c\}, independently for each feature channel and OFDM symbol:
𝐅\(t\)=ℐℱℱ𝒯\(𝐅\(f\)\),𝐅\(f\)=ℱℱ𝒯\(𝐅\(t\)\)\.\\mathbf\{F\}^\{\(t\)\}=\\mathcal\{IFFT\}\\\!\\left\(\\mathbf\{F\}^\{\(f\)\}\\right\),\\qquad\\mathbf\{F\}^\{\(f\)\}=\\mathcal\{FFT\}\\\!\\left\(\\mathbf\{F\}^\{\(t\)\}\\right\)\.\(16\)
A key innovation of the CRFCAN architecture is to formulate cross\-domain learning as a sequence of structured residual connections\. Instead of direct reconstruction in each domain, thenn\-th stage learns a complex\-valued residual mappingΔ𝐗𝒟\(n\)\\Delta\\mathbf\{X\}^\{\(n\)\}\_\{\\mathcal\{D\}\}, leading to
𝐗𝒟\(n\)=𝐗𝒟\(n−1\)\+αnΔ𝐗𝒟\(n\),𝒟∈\{time,frequency\},\\mathbf\{X\}^\{\(n\)\}\_\{\\mathcal\{D\}\}=\\mathbf\{X\}^\{\(n\-1\)\}\_\{\\mathcal\{D\}\}\+\\alpha\_\{n\}\\Delta\\mathbf\{X\}^\{\(n\)\}\_\{\\mathcal\{D\}\},\\quad\\mathcal\{D\}\\in\\\{\\text\{time\},\\text\{frequency\}\\\},\(17\)whereαn\\alpha\_\{n\}is a learnable residual scaling coefficient\. This design preserves identity connections within each domain and focuses each cross\-domain stage on incremental refinement, which helps stabilize optimization when stacking multiple FFT/IFFT projections\.
From an engineering perspective, the cross\-domain flow in CRFCAN is single\-shot and non\-iterative at runtime, with fixed and predictable computational complexity\. This property is attractive for wideband sub\-THz systems with large signal dimensionality, enabling low\-latency processing in practical receivers\.
### III\-BFFT\-Powered Cross\-Domain Residual Groups
The hierarchical backbone of CRFCAN consists of a series of FFT\-powered Residual Groups \(RGs\) as shown in Fig\.[2](https://arxiv.org/html/2609.12244#S2.F2)\. These RGs iteratively refine the joint estimation of CSI and PN by alternating feature representations between the time domain and frequency domain\.
#### III\-B1Motivation: Limitations of Depth\-Only Designs
Joint CSI–PN estimation involves strong coupling between frequency\-selective channel responses and temporally correlated phase noise\. Although increasing network depth via stacking residual blocks can enhance representation power\[[36](https://arxiv.org/html/2609.12244#bib.bib16)\], depth\-only scaling often yields diminishing returns and training instability in structured estimation tasks\. Similar behavior has been reported in image super\-resolution, where blindly increasing depth brings limited improvement despite substantially increased complexity\[[37](https://arxiv.org/html/2609.12244#bib.bib17)\]\. Motivated by the residual\-in\-residual design in Residual Channel Attention Network \(RCAN\)\[[38](https://arxiv.org/html/2609.12244#bib.bib15)\], we adopt the Residual Group \(RG\) as a fundamental unit to enable stable deep refinement through group\-level residual learning\.
Formally, let𝐅g−1\\mathbf\{F\}\_\{g\-1\}denote the input feature map to thegg\-th residual group\. The output𝐅g\\mathbf\{F\}\_\{g\}is defined as
𝐅g=𝒯g\(𝐅g−1\+ℋg\(𝐅g−1\)\),\\mathbf\{F\}\_\{g\}=\\mathcal\{T\}\_\{g\}\\\!\\left\(\\mathbf\{F\}\_\{g\-1\}\+\\mathcal\{H\}\_\{g\}\(\\mathbf\{F\}\_\{g\-1\}\)\\right\),\(18\)whereℋg\(⋅\)\\mathcal\{H\}\_\{g\}\(\\cdot\)denotes the intra\-group nonlinear transformation and𝒯g\(⋅\)\\mathcal\{T\}\_\{g\}\(\\cdot\)corresponds to anFFTFFToperator for FFT residual groups and anIFFTIFFToperator for IFFT residual groups\. This group\-level skip connection encourages each RG to focus on incremental refinement\.
#### III\-B2Intra\-Group Processing: Convolutional Blocks and Temporal Modeling
Each RG encapsulates specialized processing blocks for refining complex\-valued feature maps as shown in Fig\.[2](https://arxiv.org/html/2609.12244#S2.F2)\. We employ two variants of the Residual Channel Attention Block \(RCAB\):
- •CRCAB \(Complex\-RCAB\):A complex\-valued convolutional block with CReLU activation and channel attention for feature reweighting\.
- •PRCAB \(Phase\-Rotation RCAB\):A phase\-interaction block that learns I/Q coupling via channel attention and applies a learned complex\-valued phase rotation to the input features\.
#### III\-B3Asymmetric Block Configuration for IFFT and FFT Residual Groups
Recognizing the different physical signatures in the time and frequency domains, we adopt an asymmetric block configuration for the IFFT and FFT RGs\.
##### IFFT Residual Group Configuration
The IFFT\-RG performs feature refinement in the frequency domain and then projects the refined representation to the time domain via an IFFT operator\. In our implementation, it adopts a linear\-heavy layout:CRCAB–CRCAB–CRCAB–CRCAB–CRCAB, where the CRCABs emphasize additive and approximately linear correlations inherent to the frequency\-domain channel response\. This design is motivated by the fact that the sub\-THz channel exhibits structured spectral characteristics, such as frequency selectivity and inter\-subcarrier correlation, which are most effectively captured through successive linear refinement in the frequency domain\. By concentrating the modeling capacity on channel\-related features prior to domain transformation, the IFFT\-RG produces a spectrally coherent representation that is subsequently projected to the time domain for further joint processing\.
Formally, let𝐅g−1\(f\)\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(f\)\}\}denote the input feature map to thegg\-th IFFT\-RG in the frequency domain\. Denote theii\-th CRCAB within this group as the mapping𝒢i,g\(⋅\)\\mathcal\{G\}\_\{i,g\}\(\\cdot\),i=1,…,5i=1,\\ldots,5\. Then, the group output defined in \([18](https://arxiv.org/html/2609.12244#S3.E18)\) can be explicitly written as
𝐅g\(t\)\\displaystyle\\mathbf\{F\}\_\{g\}^\{\\mathrm\{\(t\)\}\}=ℐℱℱ𝒯g\(𝐅g−1\(f\)\+ℋg\(𝐅g−1\(f\)\)\)\\displaystyle=\\mathcal\{IFFT\}\_\{g\}\\\!\\left\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(f\)\}\}\+\\mathcal\{H\}\_\{g\}\\\!\\left\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(f\)\}\}\\right\)\\right\)\(19\)=ℐℱℱ𝒯g\(𝐅g−1\(f\)\+𝒢5,g\(𝒢4,g\(𝒢3,g\(𝒢2,g\(𝒢1,g\(𝐅g−1\(f\)\)\)\)\)\)\)\.\\displaystyle=\\mathcal\{IFFT\}\_\{g\}\\\!\\Big\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(f\)\}\}\+\\mathcal\{G\}\_\{5,g\}\\\!\\big\(\\mathcal\{G\}\_\{4,g\}\\\!\\big\(\\mathcal\{G\}\_\{3,g\}\\\!\\big\(\\mathcal\{G\}\_\{2,g\}\\\!\\big\(\\mathcal\{G\}\_\{1,g\}\\left\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(f\)\}\}\\right\)\\big\)\\big\)\\big\)\\big\)\\Big\)\.
##### FFT Residual Group Configuration
The FFT\-RG performs feature refinement in the time domain and then maps the refined representation back to the frequency domain via an FFT operator\. In this stage, the effective phase distortion observed after pulse shaping and matched filtering is no longer a pure unit\-modulus rotation; instead, it predominantly resides within an annular region around the unit circle\. As a result, both amplitude\-related distortions and phase\-rotational behaviors must be jointly modeled\.
Accordingly, we adopt the layoutCRCAB–CRCAB–PRCAB–CRCAB–CRCAB, where the surrounding CRCABs mainly capture amplitude variations and approximately linear correlations in the time\-domain features, while the central PRCAB provides a dedicated inductive bias by generating a feature\-dependent phase\-rotation offset that explicitly captures and modulates the phase information in the time\-domain representation\. This asymmetric composition enables the FFT\-RG to generate a refined time\-domain representation whose subsequent FFT projection yields more coherent spectral features for joint CSI–PN estimation\.
Formally, let𝒫3,g\(⋅\)\\mathcal\{P\}\_\{3,g\}\(\\cdot\), denote the process of PRCAB\. Then, the group output defined in \([18](https://arxiv.org/html/2609.12244#S3.E18)\) can be explicitly written as
𝐅g\(f\)\\displaystyle\\mathbf\{F\}\_\{g\}^\{\\mathrm\{\(f\)\}\}=ℱℱ𝒯g\(𝐅g−1\(t\)\+ℋg\(𝐅g−1\(t\)\)\)\\displaystyle=\\mathcal\{FFT\}\_\{g\}\\\!\\Big\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(t\)\}\}\+\\mathcal\{H\}\_\{g\}\\\!\\left\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(t\)\}\}\\right\)\\Big\)\(20\)=ℱℱ𝒯g\(𝐅g−1\(t\)\+𝒢5,g\(𝒢4,g\(𝒫3,g\(𝒢2,g\(𝒢1,g\(𝐅g−1\(t\)\)\)\)\)\)\)\.\\displaystyle=\\mathcal\{FFT\}\_\{g\}\\\!\\Big\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(t\)\}\}\+\{\\mathcal\{G\}\}\_\{5,g\}\\\!\\big\(\{\\mathcal\{G\}\}\_\{4,g\}\\\!\\big\(\\mathcal\{P\}\_\{3,g\}\\\!\\big\(\{\\mathcal\{G\}\}\_\{2,g\}\\\!\\big\(\{\\mathcal\{G\}\}\_\{1,g\}\\left\(\\mathbf\{F\}\_\{g\-1\}^\{\\mathrm\{\(t\)\}\}\\right\)\\big\)\\big\)\\big\)\\big\)\\Big\)\.
In summary, the FFT\-powered residual groups jointly integrate complex\-valued convolution for local correlation and domain\-mapping layers for physical consistency\.
### III\-CResidual Attention Blocks for Joint Modeling
#### III\-C1CRCAB
Fig\. 4:Complex Residual Channel Attention Block \(CRCAB\)\.We first introduce the complex residual channel attention block \(CRCAB\), which extends the residual channel attention design to complex\-valued feature maps\. Given the input feature𝐅g,b−1∈ℂC×Nc×Nt\\mathbf\{F\}\_\{g,b\-1\}\\in\\mathbb\{C\}^\{C\\times N\_\{c\}\\times N\_\{t\}\}, CRCAB learns an intermediate feature𝐔g,b∈ℂC×Nc×Nt\\mathbf\{U\}\_\{g,b\}\\in\\mathbb\{C\}^\{C\\times N\_\{c\}\\times N\_\{t\}\}through two stacked complex convolutions with a complex activation function, and applies channel\-wise attention to rescale𝐔g,b\\mathbf\{U\}\_\{g,b\}before the residual addition, as illustrated in Fig\.[4](https://arxiv.org/html/2609.12244#S3.F4)\.
Following the channel\-attention mechanism in RCAN\[[38](https://arxiv.org/html/2609.12244#bib.bib15)\], a channel descriptor is obtained by applying adaptive average pooling separately to the real and imaginary parts of the complex feature map and recombining them into a complex representation\. The resulting channel descriptor is then fed to a lightweight gating function to obtain channel\-wise weights\. Concretely, the channel descriptor𝐳g,b∈ℂC×1×1\\mathbf\{z\}\_\{g,b\}\\in\\mathbb\{C\}^\{C\\times 1\\times 1\}is obtained by complex adaptive average pooling applied channel\-wise\. Letc∈\{1,…,C\}c\\in\\\{1,\\ldots,C\\\}denote the channel index\. Then,
𝐳g,b\(c\)\\displaystyle\\mathbf\{z\}\_\{g,b\}\(c\)=1NcNt∑nc=1Nc∑nt=1Ntℜ\{𝐔g,b\(c,nc,nt\)\}\\displaystyle=\\frac\{1\}\{N\_\{c\}N\_\{t\}\}\\sum\_\{n\_\{c\}=1\}^\{N\_\{c\}\}\\sum\_\{n\_\{t\}=1\}^\{N\_\{t\}\}\\Re\\\!\\left\\\{\\mathbf\{U\}\_\{g,b\}\(c,n\_\{c\},n\_\{t\}\)\\right\\\}\(21\)\+j1NcNt∑nc=1Nc∑nt=1Ntℑ\{𝐔g,b\(c,nc,nt\)\},\\displaystyle\+\\,j\\,\\frac\{1\}\{N\_\{c\}N\_\{t\}\}\\sum\_\{n\_\{c\}=1\}^\{N\_\{c\}\}\\sum\_\{n\_\{t\}=1\}^\{N\_\{t\}\}\\Im\\\!\\left\\\{\\mathbf\{U\}\_\{g,b\}\(c,n\_\{c\},n\_\{t\}\)\\right\\\},∀c∈\{1,…,C\}\.\\displaystyle\\forall c\\in\\\{1,\\ldots,C\\\}\.The channel descriptor𝐳g,b\\mathbf\{z\}\_\{g,b\}is then processed by a lightweight gating network to generate channel\-wise attention weights\. Specifically, a bottleneck structure is employed to first project𝐳g,b\\mathbf\{z\}\_\{g,b\}to a lower\-dimensional space for cross\-channel interaction, followed by a non\-linear activation and an up\-projection back to the original channel dimension\. This design reduces computational complexity while preserving modeling flexibility\.
Accordingly, the channel attention weights𝐬g,b\\mathbf\{s\}\_\{g,b\}are computed via a learnable bottleneck gating network, where𝐖g,b\(d\)\\mathbf\{W\}^\{\(d\)\}\_\{g,b\}and𝐖g,b\(u\)\\mathbf\{W\}^\{\(u\)\}\_\{g,b\}are learnable channel\-projection operators implemented byℂ\\mathbb\{C\}Conv,δ\(⋅\)\\delta\(\\cdot\)denotes aℂ\\mathbb\{C\}ReLU activation function, andf\(⋅\)f\(\\cdot\)denotes theℂ\\mathbb\{C\}Sigmoid gating function, as
𝐬g,b=f\(𝐖g,b\(u\)δ\(𝐖g,b\(d\)𝐳g,b\)\),\\mathbf\{s\}\_\{g,b\}=f\\\!\\left\(\\mathbf\{W\}^\{\(u\)\}\_\{g,b\}\\,\\delta\\\!\\left\(\\mathbf\{W\}^\{\(d\)\}\_\{g,b\}\\,\\mathbf\{z\}\_\{g,b\}\\right\)\\right\),\(22\)Then the attention weights are used to rescale the residual feature in a channel\-wise manner, and the CRCAB output is given by
𝐅g,b=𝐅g,b−1\+\(𝐬g,b⊙𝐔g,b\),\\mathbf\{F\}\_\{g,b\}=\\mathbf\{F\}\_\{g,b\-1\}\+\\left\(\\mathbf\{s\}\_\{g,b\}\\odot\\mathbf\{U\}\_\{g,b\}\\right\),\(23\)where⊙\\odotdenotes channel\-wise multiplication\.
#### III\-C2PRCAB
Fig\. 5:Phase\-Rotation Residual Channel Attention Block \(PRCAB\)\.While CRCAB mainly performs linear and amplitude\-dominant feature refinement, PN manifests as a multiplicative distortion in the complex domain\. To explicitly account for this physical property, we propose the PRCAB, which extracts phase\-noise–aware features by modeling multiplicative phase rotations\. The overall architecture of PRCAB is illustrated in Fig\.[5](https://arxiv.org/html/2609.12244#S3.F5)\.
Unlike CRCAB, the phase\-related processing in PRCAB is primarily conducted in the real domain\. Specifically, the complex\-valued input feature𝐅g,b−1∈ℂC×Nc×Nt\\mathbf\{F\}\_\{g,b\-1\}\\in\\mathbb\{C\}^\{C\\times N\_\{c\}\\times N\_\{t\}\}is first mapped to a real\-valued representation by concatenating its real and imaginary parts along the channel dimension, yielding𝐅~g,b−1∈ℝ2C×Nc×Nt\\tilde\{\\mathbf\{F\}\}\_\{g,b\-1\}\\in\\mathbb\{R\}^\{2C\\times N\_\{c\}\\times N\_\{t\}\}\. This complex\-to\-real transformation allows the network to explicitly capture the relative variations between the real and imaginary components, thereby facilitating phase\-aware feature learning\. It is worth emphasizing that this real\-valued representation is introduced solely for phase analysis and regression\. After a lightweight real\-valued processing pipeline, consisting of two convolutional layers with an intermediate ReLU activation as illustrated in Fig\.[5](https://arxiv.org/html/2609.12244#S3.F5), a compact auxiliary feature𝐔~g,b∈ℝC×Nc×Nt\\tilde\{\\mathbf\{U\}\}\_\{g,b\}\\in\\mathbb\{R\}^\{C\\times N\_\{c\}\\times N\_\{t\}\}is obtained, while the dimensionality of the main real\-valued feature stream remains unchanged throughout the block\. The transformed feature then follows a CRCAB\-style channel attention pipeline composed of real\-valued convolutions and activations\. The resulting channel\-attended real feature is defined as
𝐅~CA=𝐬~g,b⊙𝐔~g,b,\\tilde\{\\mathbf\{F\}\}\_\{\\mathrm\{CA\}\}=\\tilde\{\\mathbf\{s\}\}\_\{g,b\}\\odot\\tilde\{\\mathbf\{U\}\}\_\{g,b\},\(24\)which can be interpreted as a high\-dimensional phase\-aware representation\. To estimate the phase\-related feature, the real\-valued channel\-attended representation𝐅~CA∈ℝC×H×W\\tilde\{\\mathbf\{F\}\}\_\{\\mathrm\{CA\}\}\\in\\mathbb\{R\}^\{C\\times H\\times W\}is fed into a lightweight convolutional projection to regress a phase estimate𝜽~g,b\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}\. Specifically, the phase feature is obtained as
𝜽~g,b=αtanh\(𝒲θ\(𝐅~CA\)\),\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}=\\alpha\\,\\tanh\\\!\\left\(\\mathcal\{W\}\_\{\\theta\}\\\!\\left\(\\tilde\{\\mathbf\{F\}\}\_\{\\mathrm\{CA\}\}\\right\)\\right\),\(25\)where𝒲θ\(⋅\)\\mathcal\{W\}\_\{\\theta\}\(\\cdot\)denotes a learnable real\-valued1×11\\times 1convolutional head, andα\\alphais a learnable parameter that controls the maximum allowable phase deviation\. The estimated phase𝜽~g,b\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}is then mapped to a unit\-modulus complex phasor as
𝐮g,b=cos\(𝜽~g,b\)\+jsin\(𝜽~g,b\)=exp\(j𝜽~g,b\)\.\\mathbf\{u\}\_\{g,b\}=\\cos\\\!\\left\(\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}\\right\)\+j\\sin\\\!\\left\(\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}\\right\)=\\exp\\\!\\left\(j\\tilde\{\\boldsymbol\{\\theta\}\}\_\{g,b\}\\right\)\.\(26\)
Instead of directly applying the full phase rotation, PRCAB adopts a residual formulation to ensure stable training\. Specifically, a phase\-rotated residual is computed as
𝐅Δ=𝐅g,b−1⊙𝐮g,b−𝐅g,b−1,\\mathbf\{F\}\_\{\\Delta\}=\\mathbf\{F\}\_\{g,b\-1\}\\odot\\mathbf\{u\}\_\{g,b\}\-\\mathbf\{F\}\_\{g,b\-1\},\(27\)and the block output is given by
𝐅g,b=𝐅g,b−1\+η𝐅Δ=\(1−η\)𝐅g,b−1\+η\(𝐅g,b−1⊙𝐮g,b\),\\mathbf\{F\}\_\{g,b\}=\\mathbf\{F\}\_\{g,b\-1\}\+\\eta\\,\\mathbf\{F\}\_\{\\Delta\}=\(1\-\\eta\)\\,\\mathbf\{F\}\_\{g,b\-1\}\+\\eta\\left\(\\mathbf\{F\}\_\{g,b\-1\}\\odot\\mathbf\{u\}\_\{g,b\}\\right\),\(28\)whereη∈\(0,1\)\\eta\\in\(0,1\)is a learnable scalar gate that controls the step size of the phase\-rotation update\.
### III\-DPhysically Constrained Dual Output
The backbone of CRFCAN consists ofGGFFT Residual Groups andGGIFFT Residual Groups, which are arranged in an alternating manner to enable progressive cross\-domain refinement\. For channel estimation, a lightweight complex\-valued convolutional tail is employed to project the refined features to a CSI estimation output\. Let𝐅G\(f\)\\mathbf\{F\}\_\{G\}^\{\(f\)\}and𝐅G−1\(f\)\\mathbf\{F\}\_\{G\-1\}^\{\(f\)\}denote the last two frequency\-domain feature representations produced by the proposed network\. For channel estimation, the fused frequency\-domain tail input is defined as the final frequency domain residual projection𝐅tail\(f\)=𝐅G\(f\)\+𝐅G−1\(f\)\\mathbf\{F\}\_\{\\mathrm\{tail\}\}^\{\(f\)\}=\\mathbf\{F\}\_\{G\}^\{\(f\)\}\+\\mathbf\{F\}\_\{G\-1\}^\{\(f\)\}, and the CSI estimate is obtained by a lightweight complex\-valued convolutional projection:
𝐇^=𝒯H\(𝐅tail\(f\)\)=𝐖H∗\(𝐅G\(f\)\+𝐅G−1\(f\)\),\\widehat\{\\mathbf\{H\}\}=\\mathcal\{T\}\_\{\\mathrm\{H\}\}\\\!\\left\(\\mathbf\{F\}\_\{tail\}^\{\(f\)\}\\right\)=\\mathbf\{W\}\_\{\\mathrm\{H\}\}\*\\left\(\\mathbf\{F\}\_\{G\}^\{\(f\)\}\+\\mathbf\{F\}\_\{G\-1\}^\{\(f\)\}\\right\),\(29\)where𝐖H\\mathbf\{W\}\_\{\\mathrm\{H\}\}denotes a sequence of complex\-valued convolutional kernels that progressively reduce the channel dimension fromCCto11\. This tail mainly aggregates high\-level representations and does not impose additional constraints\.
In accordance with the discrete\-time system model in \([4](https://arxiv.org/html/2609.12244#S2.E4)\), the PN estimation branch of CRFCAN directly estimates the complex\-valued multiplicative impairment coefficientp\[n\]p\[n\], rather than the phase angleϕ\[n\]\\phi\[n\]itself\. For PN estimation, the same tail fusion strategy is adopted to obtain an initial estimate, while operating on time\-domain features instead of frequency\-domain ones\. Specifically,
𝐙PN=𝒯PN\(𝐅G\(t\)\+𝐅G−1\(t\)\),\\mathbf\{Z\}\_\{\\mathrm\{PN\}\}=\\mathcal\{T\}\_\{\\mathrm\{PN\}\}\\\!\\left\(\\mathbf\{F\}\_\{G\}^\{\(t\)\}\+\\mathbf\{F\}\_\{G\-1\}^\{\(t\)\}\\right\),\(30\)where𝒯PN\(⋅\)\\mathcal\{T\}\_\{\\mathrm\{PN\}\}\(\\cdot\)shares the same architectural form as𝒯H\(⋅\)\\mathcal\{T\}\_\{\\mathrm\{H\}\}\(\\cdot\)but is dedicated to PN estimation\.
Unlike channel estimation, PN primarily manifests as phase rotations accompanied by limited and bounded amplitude fluctuations after pulse shaping and matched filtering\. Therefore, instead of enforcing a strict unit\-modulus constraint, a soft magnitude constraint is introduced to accommodate practical amplitude fluctuations while stabilizing network training\. Specifically, the final PN estimate is obtained as
𝐩^=𝐙PN\|𝐙PN\|\[μ\+6σ\(sigmoid\(\|𝐙PN\|−1\)−12\)\],\\widehat\{\\mathbf\{p\}\}=\\frac\{\\mathbf\{Z\}\_\{\\mathrm\{PN\}\}\}\{\\left\|\\mathbf\{Z\}\_\{\\mathrm\{PN\}\}\\right\|\}\\left\[\\mu\+6\\sigma\\left\(\\operatorname\{sigmoid\}\\\!\\left\(\\left\|\\mathbf\{Z\}\_\{\\mathrm\{PN\}\}\\right\|\-1\\right\)\-\\frac\{1\}\{2\}\\right\)\\right\],\(31\)whereμ\\muspecifies the nominal magnitude of the effective PN coefficient, andσ\\sigmacontrols the allowable magnitude variation\. The scaling factor66is chosen such that the soft constraint approximately covers a±3σ\\pm 3\\sigmarange aroundμ\\mu, which corresponds to a high\-probability operating region in practical systems\.
This soft magnitude constraint explicitly injects physical prior knowledge into the network by preserving the estimated phase information while gently regularizing the output magnitude\. In the considered joint channel and PN estimation task, the two impairments are strongly coupled, such that relying solely on data\-driven learning and loss\-based supervision is often insufficient to enforce physically consistent PN representations\. By softly constraining the PN magnitude, the effective degrees of freedom of the PN branch are reduced, which helps prevent excessive gradient absorption by PN and enables the dominant amplitude\-related gradients to be more reliably allocated to channel estimation\. As a result, this design mitigates gradient explosion or vanishing caused by extreme batch\-wise PN realizations and stabilizes joint optimization\. By applying a mild magnitude constraint only at the output stage, the proposed approach further stabilizes gradient propagation, mitigates extreme amplitude outliers as a secondary effect, and yields physically plausible PN estimates without restricting the expressive capacity of the network\.
## IVNumerical Results
### IV\-ASimulation Setup
All numerical results are obtained using Monte Carlo simulations with datasets generated offline inMATLAB\. Unless otherwise specified, the considered system follows the sub\-THz OFDM signal model described in Section II and the simulation parameters summarized in Table[II](https://arxiv.org/html/2609.12244#S4.T2)are used throughout the numerical evaluations\.
TABLE II:Simulation ParametersThe network input consists of the received frequency\-domain signal after matched filtering, downsampling, and FFT processing, together with the corresponding pilot pattern, forming a complex\-valued tensor of size2×Nc×Nt2\\times N\_\{c\}\\times N\_\{t\}\. The two channels represent the processed received signal and the pilot grid, respectively, where transmitted pilot symbols are placed at pilot positions and all remaining entries are set to zero\. The network jointly outputs the estimated channel frequency response𝐇^∈ℂNc×Nt\\widehat\{\\mathbf\{H\}\}\\in\\mathbb\{C\}^\{N\_\{c\}\\times N\_\{t\}\}and the effective multiplicative phase\-noise impairment𝐩^∈ℂNc×Nt\\widehat\{\\mathbf\{p\}\}\\in\\mathbb\{C\}^\{N\_\{c\}\\times N\_\{t\}\}\.
In the simulation, the wireless channel is assumed to remain constant over theNtN\_\{t\}consecutive OFDM symbols, reflecting a quasi\-static fading condition over the considered observation interval\. Phase noise is generated as a continuous\-time process and applied at the sample level prior to receiver processing, which leads to symbol\-dependent phase distortions after matched filtering and sampling\. For the considered 100 GHz setting, the generalized PN models are conservatively refined following the above calibration principle\. It is motivated by the fact that the 3GPP RAN1 Set1\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]profile serves as a generalized measurement\-informed reference within the 3GPP framework, abstracted from multiple practical oscillator phase\-noise characteristics and thus providing a suitable baseline for hardware\-oriented calibration around 100 GHz\. This calibration is further supported by the measured results reported in\[[39](https://arxiv.org/html/2609.12244#bib.bib45)\], where a 100\.8\-GHz CMOS synthesizer achieves phase\-noise levels of−81\.9\-81\.9dBc/Hz at 100 kHz offset,−93\.0\-93\.0dBc/Hz at 1 MHz offset, and−104\.8\-104\.8dBc/Hz at 10 MHz offset\. By jointly considering these measured phase\-noise levels together with the three PN models adopted in our simulations,Δcal=−10\\Delta\_\{\\text\{cal\}\}=\-10dB is selected as a conservative calibration factor and applied to all three PN models such that the resulting PSDs remain broadly consistent with this practical 100\-GHz hardware noise range, as illustrated in Fig\.[6](https://arxiv.org/html/2609.12244#S4.F6)\. Additive white Gaussian noise is independently generated and added to the received time\-domain signal, with noise samples being independent across time and OFDM symbols\.
### IV\-BTraining Methodology
The training and evaluation datasets are generated independently under the same system and impairment models, using different random seeds to avoid overlap in channel and phase\-noise realizations\. For training, samples are generated over anEb/N0E\_\{b\}/N\_\{0\}range of2020–3030dB, with30,00030\{,\}000samples per point, and are randomly divided into70%70\\%training and30%30\\%validation sets\. The testing dataset is generated separately over a widerEb/N0E\_\{b\}/N\_\{0\}range of00–3030dB, with5,0005\{,\}000samples per point\. The moderate\-to\-high trainingEb/N0E\_\{b\}/N\_\{0\}regime is adopted to emphasize the underlying channel and phase\-noise structures while retaining a limited amount of additive noise for regularization\. A similar high\-Eb/N0E\_\{b\}/N\_\{0\}training strategy has also been used in\[[17](https://arxiv.org/html/2609.12244#bib.bib30)\]for channel estimation\.
The proposed network is trained in an end\-to\-end manner using a joint supervised learning objective that simultaneously accounts for channel estimation and phase\-noise estimation\. Let𝐇^\\widehat\{\\mathbf\{H\}\}and𝐩^\\widehat\{\\mathbf\{p\}\}denote the estimated channel frequency response and multiplicative phase\-noise process, respectively, and let𝐇\\mathbf\{H\}and𝐩\\mathbf\{p\}be their corresponding ground truths\. For both branches, a complex\-valued mean squared error \(CMSE\) loss is adopted, defined as
ℒCMSE\(𝐗^,𝐗\)=𝔼\[\|ℜ\{𝐗^−𝐗\}\|2\+\|ℑ\{𝐗^−𝐗\}\|2\],\\mathcal\{L\}\_\{\\mathrm\{CMSE\}\}\(\\widehat\{\\mathbf\{X\}\},\\mathbf\{X\}\)=\\mathbb\{E\}\\\!\\left\[\\left\|\\Re\\\{\\widehat\{\\mathbf\{X\}\}\-\\mathbf\{X\}\\\}\\right\|^\{2\}\+\\left\|\\Im\\\{\\widehat\{\\mathbf\{X\}\}\-\\mathbf\{X\}\\\}\\right\|^\{2\}\\right\],\(32\)where𝐗∈\{𝐇,𝐩\}\\mathbf\{X\}\\in\\\{\\mathbf\{H\},\\mathbf\{p\}\\\}\. Applying the CMSE jointly to both the channel and the effective multiplicative PN coefficient ensures that magnitude and phase errors are penalized symmetrically, which is consistent with the complex\-valued output representation adopted throughout the network\.
To automatically balance the heterogeneous learning difficulties of channel and phase\-noise estimation, an uncertainty\-based weighting strategy is employed\[[42](https://arxiv.org/html/2609.12244#bib.bib18)\]\. Specifically, the overall training loss is formulated as
ℒjoint=ℒHexp\(σH\)\+ℒPNexp\(σPN\)\+12\(σH\+σPN\),\\mathcal\{L\}\_\{\\mathrm\{joint\}\}=\\frac\{\\mathcal\{L\}\_\{\\mathrm\{H\}\}\}\{\\exp\(\\sigma\_\{\\mathrm\{H\}\}\)\}\+\\frac\{\\mathcal\{L\}\_\{\\mathrm\{PN\}\}\}\{\\exp\(\\sigma\_\{\\mathrm\{PN\}\}\)\}\+\\frac\{1\}\{2\}\\bigl\(\\sigma\_\{\\mathrm\{H\}\}\+\\sigma\_\{\\mathrm\{PN\}\}\\bigr\),\(33\)whereℒH\\mathcal\{L\}\_\{\\mathrm\{H\}\}andℒPN\\mathcal\{L\}\_\{\\mathrm\{PN\}\}denote the CMSE losses for channel and phase\-noise estimation, respectively, andσH\\sigma\_\{\\mathrm\{H\}\}andσPN\\sigma\_\{\\mathrm\{PN\}\}are learnable scalar parameters representing the task\-dependent log\-variance\. Under the strong coupling betweenHHandpp, the adopted adaptive weighting balances the two estimation objectives during training and improves the stability of joint optimization without requiring manual loss\-weight tuning\.
Training is organized into successive cycles indexed bykk, where each cycle spansEkE\_\{k\}epochs\. Within thekk\-th cycle, the learning rate is controlled by a warm\-up phase followed by a cosine annealing phase\[[43](https://arxiv.org/html/2609.12244#bib.bib20)\]\. Specifically, leteedenote the epoch index within the current cycle, with0≤e<Ek0\\leq e<E\_\{k\}\. The learning rate is defined as
η\(e\)\\displaystyle\\eta\(e\)=ηmax\(k\)\[ηr\(k\)\+12\(1−ηr\(k\)\)\\displaystyle=\\eta\_\{\\max\}^\{\(k\)\}\\Bigg\[\\eta\_\{r\}^\{\(k\)\}\+\\frac\{1\}\{2\}\\bigl\(1\-\\eta\_\{r\}^\{\(k\)\}\\bigr\)\(34\)×\(1\+cos\(πe−EwEk−Ew\)\)\],e≥Ew\.\\displaystyle\\times\\left\(1\+\\cos\\\!\\left\(\\pi\\frac\{e\-E\_\{\\mathrm\{w\}\}\}\{E\_\{k\}\-E\_\{\\mathrm\{w\}\}\}\\right\)\\right\)\\Bigg\],\\quad e\\geq E\_\{\\mathrm\{w\}\}\.whereEwE\_\{\\mathrm\{w\}\}denotes the warm\-up duration within each cycle andηr\(k\)=ηmin\(k\)/ηmax\(k\)\\eta\_\{r\}^\{\(k\)\}=\\eta\_\{\\min\}^\{\(k\)\}/\\eta\_\{\\max\}^\{\(k\)\}\. During the warm\-up phase0≤e<Ew0\\leq e<E\_\{\\mathrm\{w\}\}, the learning rate is linearly increased fromηmin\(k\)\\eta\_\{\\min\}^\{\(k\)\}toηmax\(k\)\\eta\_\{\\max\}^\{\(k\)\}\. Across cycles, both the maximum and minimum learning rates are progressively decayed to facilitate coarse\-to\-fine optimization\. This cyclic warm\-up and annealing strategy allows the optimizer to periodically escape shallow local minima while gradually refining the solution, which is particularly beneficial for stabilizing joint learning in the presence of strong task coupling\[[44](https://arxiv.org/html/2609.12244#bib.bib19)\]\.
### IV\-CBenchmark Schemes
To comprehensively evaluate the proposed joint channel and phase\-noise estimation framework, we compare it with representative baseline methods spanning classical signal processing and deep learning paradigms\. Specifically, we consider three representative baselines: 1\) a conventional iterative LS\-based estimator; 2\) a cascaded multi\-network learning framework; and 3\) an end\-to\-end neural receiver\. These baselines collectively represent conventional model\-driven signal processing, model\-assisted learning, and fully data\-driven neural receivers to joint estimation\.
#### IV\-C1Iterative LS\-Based Estimator
As a representative model\-based benchmark, we implement the iterative joint channel and phase\-noise compensation scheme proposed in\[[5](https://arxiv.org/html/2609.12244#bib.bib4)\]\. The method alternates between channel and phase\-noise estimation under a least\-squares criterion\. Let𝐀\\mathbf\{A\}denote the circulant convolution matrix constructed from the discrete\-time phase\-noise sequence𝐩\\mathbf\{p\}\. The channel and phase\-noise estimates are iteratively updated as
𝐡^\(i\+1\)=argmin𝐡‖𝐲−𝐀p\(i\)\(i\)𝐗𝐡\(i\)^‖2,\\hat\{\\mathbf\{h\}\}^\{\(i\+1\)\}=\\arg\\min\_\{\\mathbf\{h\}\}\\left\\\|\\mathbf\{y\}\-\\mathbf\{A\}^\{\(i\)\}\_\{p^\{\(i\)\}\}\\mathbf\{X\}\\hat\{\\mathbf\{h\}^\{\(i\)\}\}\\right\\\|^\{2\},\(35\)𝐩^\(i\+1\)=argmin𝐩‖𝐲−𝐀p\(i\)\(i\)𝐗𝐡^\(i\+1\)‖2\.\\hat\{\\mathbf\{p\}\}^\{\(i\+1\)\}=\\arg\\min\_\{\\mathbf\{p\}\}\\left\\\|\\mathbf\{y\}\-\\mathbf\{A\}^\{\(i\)\}\_\{p^\{\(i\)\}\}\\mathbf\{X\}\\hat\{\\mathbf\{h\}\}^\{\(i\+1\)\}\\right\\\|^\{2\}\.\(36\)To reduce the number of unknowns, the phase\-noise process is parameterized by a limited set of time\-domain samples and reconstructed via DFT\-based low\-pass interpolation, while the channel is modeled with a truncated impulse response\. When applied to highly frequency\-selective sub\-THz channels and rapidly varying PN, such simplified interpolation models may struggle to accurately capture the coupled distortions\. Moreover, the alternating least\-squares refinement requires iterative optimization, resulting in non\-negligible computational complexity\.
#### IV\-C2Cascaded Multi\-Network Baseline
We further consider the deep learning framework proposed in\[[20](https://arxiv.org/html/2609.12244#bib.bib7)\], which employs three cascaded neural networks for joint channel estimation, phase\-noise compensation, and data detection\. Specifically, a fully connected deep neural network, termed ChDNN, is first used to refine the initial nonlinear least\-squares channel estimate obtained from the block pilot\. The refined channel estimate is then used for equalization, based on which a least\-squares phase\-noise estimate is computed for each payload symbol and further refined by a second fully connected network, termed PnDNN\. Finally, the detected data symbols are reshaped into a two\-dimensional format and processed by a residual convolutional neural network, termed DnCNN, for data denoising\. This architecture follows a model\-assisted learning paradigm, where conventional estimators provide the initial inputs and the neural networks act as nonlinear refiners in a staged manner\.
#### IV\-C3End\-to\-End Neural Transceiver
We further consider the end\-to\-end neural receiver proposed in\[[22](https://arxiv.org/html/2609.12244#bib.bib9)\], which performs joint compensation of channel distortion and phase noise using a fully convolutional residual architecture\. The receiver, termed DeepSRX, processes the received signal block together with pilot\-derived side information and directly outputs soft bit estimates\. Channel and phase\-noise effects are implicitly mitigated within the neural network without explicitly estimating the channel response or the phase\-noise parameters\.
For a fair comparison, all benchmark schemes are evaluated under the same system configuration described in Section II\. The proposed CRFCAN uses only comb\-type pilots, which is the same pilot setting adopted by the end\-to\-end neural receiver\[[22](https://arxiv.org/html/2609.12244#bib.bib9)\]\. In contrast, the classical iterative estimator\[[5](https://arxiv.org/html/2609.12244#bib.bib4)\]and the cascaded multi\-network scheme\[[20](https://arxiv.org/html/2609.12244#bib.bib7)\]additionally require block\-type pilots for channel estimation\. As a result, CRFCAN operates with a lower overall pilot density than the conventional and cascaded benchmarks, while still delivering better performance\. All methods operate under identical channel realizations, phase\-noise models, modulation formats, andEb/N0E\_\{b\}/N\_\{0\}settings\. No additional prior statistical information is provided beyond what is explicitly assumed in each respective framework\.
### IV\-DOverall Performance Evaluation
Fig\. 6:Power spectral density \(PSD\) of the three PN models considered in this work\.In order to benchmark joint CSI–PN estimation under representative and physically distinct oscillator impairments, we consider three PN models with different spectral characteristics\. Fig\.[6](https://arxiv.org/html/2609.12244#S4.F6)depicts the corresponding power spectral density \(PSD\) of the generated PN processes, which serves as a compact characterization of their temporal correlation and effective impairment bandwidth\.
\(a\)PN model RAN4\[[27](https://arxiv.org/html/2609.12244#bib.bib11)\]\(b\)PN model RAN1 Set1\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]\(c\)PN model RAN1 Set2\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]
Fig\. 7:H and PN estimation error comparison under three PN models\.The proposed CRFCAN is trained only on the 3GPP TR38\.803 RAN4\[[27](https://arxiv.org/html/2609.12244#bib.bib11)\]dataset\. The trained model is then directly evaluated on 3GPP TR38\.803 RAN4, 3GPP TR38\.808 RAN1 Set1 and Set2\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]without any finetuning, thereby constituting a strict cross\-model generalization test\. In Fig\.[7](https://arxiv.org/html/2609.12244#S4.F7)\(a\)–\(c\), the left y\-axis reports the NMSE of the channel estimateH^\\hat\{H\}, while the right y\-axis reports the MSE of the PN estimate\. AsEb/N0E\_\{b\}/N\_\{0\}increases, both metrics exhibit a clear and synchronized decrease across all three PN models, indicating that CRFCAN effectively leverages improved observation quality and enhances the coupled H and PN estimates simultaneously rather than trading one for the other\. In the training\-domain case \(RAN4\), CRFCAN consistently achieves lower errors and better suppresses the high\-Eb/N0E\_\{b\}/N\_\{0\}error floor\. More importantly, under the unseen RAN1 Set1 and Set2 conditions, CRFCAN maintains the same decreasing trend and avoids catastrophic degradation, suggesting that it does not overfit to a specific PN statistics but learns a physically consistent refinement mechanism that generalizes to different PN spectra\. We attribute this cross\-model generalization capability to two architectural properties\. First, the FFT/IFFT\-powered cross\-domain backbone operates on the structural relationship between time\-domain phase rotations and frequency\-domain spectral spreading, which is governed by the Fourier transform and holds regardless of the specific PN spectral shape\. Consequently, the network learns a domain\-transformation–based refinement mechanism rather than memorizing the statistical profile of a particular PN model\. Second, the soft\-normalization constraint in the PN output tail restricts the estimated magnitude to a narrow annular region around the unit circle, which implicitly regularizes the solution space and prevents the network from overfitting to the amplitude statistics of the training PN model\.
\(a\)PN model RAN4\[[27](https://arxiv.org/html/2609.12244#bib.bib11)\]\(b\)PN model RAN1 Set1\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]\(c\)PN model RAN1 Set2\[[41](https://arxiv.org/html/2609.12244#bib.bib21)\]
Fig\. 8:System BER under three PN models\.Fig\.[8](https://arxiv.org/html/2609.12244#S4.F8)\(a\)–\(c\) compares the BER performance under three PN models\. Across all three PN models, CRFCAN consistently achieves the lowest BER among the practical schemes, demonstrating robust performance under different PN spectra\. The Perfect H only baseline exhibits a clear high\-Eb/N0E\_\{b\}/N\_\{0\}error floor, indicating that residual PN\-induced CPE/ICI remains the dominant impairment even with perfect CSI\. In contrast, CRFCAN consistently lowers the BER floor and achieves the best performance among all practical schemes, demonstrating more effective mitigation of PN\-induced distortions\. Notably, CRFCAN can even outperform Perfect H only at highEb/N0E\_\{b\}/N\_\{0\}, confirming that the proposed network provides explicit PN compensation beyond channel equalization\. Compared with the multi\-stage Cascade \(3 nets\) baseline, CRFCAN delivers further gains, suggesting that single\-shot cross\-domain refinement better avoids error propagation across stages\. DeepSRX\[[22](https://arxiv.org/html/2609.12244#bib.bib9)\]is included by adopting its residual\-network architecture, but its performance degrades under the joint CSI–PN setting, implying that a plain residual design is insufficient for this task\.
### IV\-ERobustness Analysis
Fig\. 9:EVM of CRFCAN versusEb/N0E\_\{b\}/N\_\{0\}for 4QAM and 16QAM under three PN modelsFig\.[9](https://arxiv.org/html/2609.12244#S4.F9)reports the EVM\[[45](https://arxiv.org/html/2609.12244#bib.bib39)\]performance of CRFCAN versusEb/N0E\_\{b\}/N\_\{0\}for 4QAM and 16QAM under three PN models\. The 4/16QAM Perfect curves correspond to oracle\-aided EVM computed with perfect knowledge of both the channel𝐇^\\widehat\{\\mathbf\{H\}\}and the phase noise𝐩^\\widehat\{\\mathbf\{p\}\}\. For both modulation formats, the EVM decreases consistently asEb/N0E\_\{b\}/N\_\{0\}increases, showing that the proposed receiver can effectively exploit improved observation quality\. In addition, 16QAM exhibits a noticeably higher EVM floor than 4QAM in the medium\-to\-highEb/N0E\_\{b\}/N\_\{0\}regime, which is consistent with the higher sensitivity of higher\-order modulation to residual PN\-induced CPE/ICI\. This behavior suggests that, once the thermal noise impact diminishes, the remaining performance is primarily limited by residual PN/ICI distortion rather than thermal noise\.
### IV\-FAblation Study
Fig\. 10:BER comparison of CRFCAN variants on the RAN4 dataset\. Each variant removes or replaces a single component while keeping the rest unchanged\.Before examining the contribution of individual modules, it should be noted that the adopted depth configuration, consisting of three pairs of residual groups with five residual blocks in each residual group, was selected based on preliminary comparisons over several depth configurations\. The final setting was retained as a practical balance between estimation performance and model complexity, and is therefore used in all subsequent ablation experiments\. To quantify the contribution of each proposed component, we conduct an ablation study by systematically removing or replacing individual modules while keeping the remaining architecture unchanged\. Five ablation variants are considered:
- •w/o FFT/IFFT: The FFT and IFFT domain\-transition operators within each residual group are replaced with identity mappings, reducing the cross\-domain backbone to a single\-domain convolutional network\.
- •w/o PRCAB: The phase\-rotation residual channel attention block \(PRCAB\) in each FFT\-RG is replaced with a standard CRCAB, removing the dedicated multiplicative phase\-modeling capability\.
- •w/o PN Tail: The physics\-aware soft\-normalization PN output tail is replaced with the same unconstrained convolutional tail used for channel estimation\.
- •w/o FFT/IFFT \+ PRCAB \+ PN Tail: All three components above are simultaneously removed\.
- •w/o Uncertainty Weighting: The uncertainty\-based adaptive loss weighting is replaced with a direct summation of the channel and PN CMSE losses, i\.e\.,ℒjoint=ℒH\+ℒPN\\mathcal\{L\}\_\{\\mathrm\{joint\}\}=\\mathcal\{L\}\_\{\\mathrm\{H\}\}\+\\mathcal\{L\}\_\{\\mathrm\{PN\}\}\.
Fig\.[10](https://arxiv.org/html/2609.12244#S4.F10)reports the BER performance of all variants on the RAN4 dataset\. At lowEb/N0E\_\{b\}/N\_\{0\}, all variants perform similarly since AWGN dominates\. AsEb/N0E\_\{b\}/N\_\{0\}increases and PN\-induced distortions become the limiting factor, clear performance gaps emerge\. Removing the FFT/IFFT domain transitions causes the most significant degradation, with its high\-Eb/N0E\_\{b\}/N\_\{0\}error floor approaching that of the combined ablation variant \(w/o All Three\)\. This confirms that the cross\-domain backbone is the most critical architectural component\. Replacing PRCAB with CRCAB also leads to a noticeable BER increase at highEb/N0E\_\{b\}/N\_\{0\}, indicating that the explicit multiplicative phase\-rotation modeling provides a meaningful inductive bias beyond what additive convolutional refinement can achieve\. Removing the PN tail results in a comparatively smaller degradation, suggesting that the soft\-normalization constraint acts as a useful regularizer rather than a dominant performance driver\. Replacing the uncertainty\-based loss weighting with equal\-weight summation also degrades performance, confirming that adaptive loss balancing is beneficial for joint optimization under heterogeneous task difficulties\.
Notably, only the full CRFCAN and the w/o PN Tail variant achieve BER below the Perfect H only baseline at highEb/N0E\_\{b\}/N\_\{0\}, indicating that the combination of cross\-domain processing and phase\-aware modeling is necessary to provide effective PN compensation beyond channel equalization\. When all three architectural components are removed simultaneously, the network reduces to a plain complex\-valued residual network and fails to surpass the Perfect H only bound, further corroborating the necessity of the proposed design\.
### IV\-GComputational Complexity
The computational cost of CRFCAN is dominated by the complex\-valued convolutional operations within the residual groups\. For an OFDM frame ofNcN\_\{c\}subcarriers andNtN\_\{t\}symbols, the convolutional cost scales as
𝒞conv=𝒪\(G⋅B⋅C2⋅K2⋅Nc⋅Nt\),\\mathcal\{C\}\_\{\\mathrm\{conv\}\}=\\mathcal\{O\}\\\!\\left\(G\\cdot B\\cdot C^\{2\}\\cdot K^\{2\}\\cdot N\_\{c\}\\cdot N\_\{t\}\\right\),\(37\)whereGGdenotes the number of residual group pairs,BBthe number of attention blocks per group,CCthe feature channel dimension, andKKthe convolutional kernel size\. The FFT/IFFT domain transitions contribute
𝒞FFT=𝒪\(G⋅C⋅Nt⋅NclogNc\),\\mathcal\{C\}\_\{\\mathrm\{FFT\}\}=\\mathcal\{O\}\\\!\\left\(G\\cdot C\\cdot N\_\{t\}\\cdot N\_\{c\}\\log N\_\{c\}\\right\),\(38\)which is negligible relative to𝒞conv\\mathcal\{C\}\_\{\\mathrm\{conv\}\}\. With the default configuration \(G=3G\{=\}3,B=5B\{=\}5,C=64C\{=\}64,K=3K\{=\}3,Nc=64N\_\{c\}\{=\}64,Nt=8N\_\{t\}\{=\}8\), the network contains approximately4\.7×1064\.7\\times 10^\{6\}real\-valued parameters\.
A key property of CRFCAN is that its computational cost is fixed and deterministic, independent of the channel realization or PN severity\. This contrasts with iterative model\-based approaches\[[5](https://arxiv.org/html/2609.12244#bib.bib4)\]and multi\-stage cascaded schemes\[[20](https://arxiv.org/html/2609.12244#bib.bib7)\], whose runtime depends on convergence behavior or requires sequential execution of multiple networks with intermediate processing steps\.
## VConclusion
This paper proposed CRFCAN for joint CSI and PN estimation in sub\-THz OFDM systems\. CRFCAN is built upon a complex\-valued and cross\-domain learning framework, where FFT/IFFT\-powered residual groups alternately refine time\- and frequency\-domain representations and phase\-aware modules explicitly enhance the modeling of PN\-induced distortions\. Simulation results under three PN models demonstrate that CRFCAN achieves consistently lower H estimation NMSE and PN estimation MSE, and further translates these gains into improved BER and EVM with a significantly reduced high\-Eb/N0E\_\{b\}/N\_\{0\}error floor\. Ablation studies confirm the necessity of the FFT/IFFT\-based cross\-domain backbone and the attention\-based phase modeling\. With single\-shot inference, cross\-model generalization capability, and end\-to\-end training simplicity, CRFCAN provides an effective and practical receiver solution for wideband sub\-THz systems with oscillator impairments\.
## References
- \[1\]ITU\-R\(2023\)Framework and overall objectives of the future development of IMT for 2030 and beyond\.Recommendation ITU\-R M\.2160\-0International Telecommunication Union,Geneva, Switzerland\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p1.1)\.
- \[2\]W\. Jiang, Q\. Zhou, J\. He, M\. A\. Habibi, S\. Melnyk, M\. El\-Absi, B\. Han, M\. Di Renzo, H\. D\. Schotten, F\. Luo,et al\.\(2024\)Terahertz communications and sensing for 6g and beyond: a comprehensive review\.IEEE Communications Surveys & Tutorials26\(4\),pp\. 2326–2381\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p1.1)\.
- \[3\]R\. Chen, B\. Yan, and M\. F\. Chang\(2025\)A review of circuits and systems for advanced sub\-thz transceivers in wireless communication\.Electronics14\(5\),pp\. 861\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p1.1)\.
- \[4\]A\. G\. Armada\(2001\)Understanding the effects of phase noise in orthogonal frequency division multiplexing \(ofdm\)\.IEEE transactions on broadcasting47\(2\),pp\. 153–159\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p1.1)\.
- \[5\]Q\. Zou, A\. Tarighat, and A\. H\. Sayed\(2007\)Compensation of phase noise in ofdm wireless systems\.IEEE transactions on signal processing55\(11\),pp\. 5407–5424\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p1.1),[§I](https://arxiv.org/html/2609.12244#S1.p2.1),[§IV\-C1](https://arxiv.org/html/2609.12244#S4.SS3.SSS1.p1.1),[§IV\-C3](https://arxiv.org/html/2609.12244#S4.SS3.SSS3.p2.1),[§IV\-G](https://arxiv.org/html/2609.12244#S4.SS7.p2.1)\.
- \[6\]J\. Van De Beek, O\. Edfors, M\. Sandell, S\. K\. Wilson, and P\. O\. Borjesson\(1995\)On channel estimation in ofdm systems\.In1995 IEEE 45th Vehicular Technology Conference\. Countdown to the Wireless Twenty\-First Century,Vol\.2,pp\. 815–819\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[7\]O\. Edfors, M\. Sandell, J\. Van de Beek, S\. K\. Wilson, and P\. O\. Borjesson\(1998\)OFDM channel estimation by singular value decomposition\.IEEE Transactions on communications46\(7\),pp\. 931–939\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[8\]O\. H\. Salim, A\. A\. Nasir, H\. Mehrpouyan, W\. Xiang, S\. Durrani, and R\. A\. Kennedy\(2014\)Channel, phase noise, and frequency offset in ofdm systems: joint estimation, data detection, and hybrid cramer\-rao lower bound\.IEEE Transactions on Communications62\(9\),pp\. 3311–3325\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[9\]D\. Petrovic, W\. Rave, and G\. Fettweis\(2003\)Phase noise suppression in ofdm using a kalman filter\.InProc\. WPMC,Vol\.176\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[10\]H\. Mehrpouyan, A\. A\. Nasir, S\. D\. Blostein, T\. Eriksson, G\. K\. Karagiannidis, and T\. Svensson\(2012\)Joint estimation of channel and oscillator phase noise in mimo systems\.IEEE Transactions on Signal Processing60\(9\),pp\. 4790–4807\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[11\]H\. Ju, H\. Zhang, L\. Li, X\. Li, and B\. Dong\(2024\)A comparative study of deep learning and iterative algorithms for joint channel estimation and signal detection in ofdm systems\.Signal Processing223,pp\. 109554\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[12\]T\. T\. Nguyen, S\. T\. Le, M\. Wuilpart, T\. Yakusheva, and P\. Mégret\(2017\)Simplified extended kalman filter phase noise estimation for co\-ofdm transmissions\.Optics Express25\(22\),pp\. 27247–27261\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[13\]C\. R\. Berger, Z\. Wang, J\. Huang, and S\. Zhou\(2010\)Application of compressive sensing to sparse channel estimation\.IEEE Communications Magazine48\(11\),pp\. 164–174\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[14\]J\. Meng, W\. Yin, Y\. Li, N\. T\. Nguyen, and Z\. Han\(2011\)Compressive sensing based high\-resolution channel estimation for ofdm system\.IEEE Journal of Selected Topics in Signal Processing6\(1\),pp\. 15–25\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[15\]R\. Zhang, B\. Shim, and H\. Zhao\(2020\)Downlink compressive channel estimation with phase noise in massive mimo systems\.IEEE transactions on communications68\(9\),pp\. 5534–5548\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p2.1)\.
- \[16\]X\. Yi and C\. Zhong\(2020\)Deep learning for joint channel estimation and signal detection in ofdm systems\.IEEE Communications Letters24\(12\),pp\. 2780–2784\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1)\.
- \[17\]H\. He, C\. Wen, S\. Jin, and G\. Y\. Li\(2018\)Deep learning\-based channel estimation for beamspace mmwave massive mimo systems\.IEEE Wireless Communications Letters7\(5\),pp\. 852–855\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1),[§IV\-B](https://arxiv.org/html/2609.12244#S4.SS2.p1.1)\.
- \[18\]P\. Neshaastegaran and M\. Jian\(2024\)Deep learning\-assisted phase noise mitigation for high\-order modulation with minimal overhead\.In2024 IEEE International Conference on Communications Workshops \(ICC Workshops\),pp\. 1152–1158\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1)\.
- \[19\]K\. Seo, J\. Jeon, G\. Lee, J\. Lee, and S\. Noh\(2025\)Deep sequential feature learning for phase noise compensation in sub\-thz systems\.IEEE Transactions on Communications\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1)\.
- \[20\]A\. Mohammadian, C\. Tellambura, and G\. Y\. Li\(2021\)Deep learning\-based phase noise compensation in multicarrier systems\.IEEE Wireless Communications Letters10\(10\),pp\. 2110–2114\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1),[§IV\-C2](https://arxiv.org/html/2609.12244#S4.SS3.SSS2.p1.1),[§IV\-C3](https://arxiv.org/html/2609.12244#S4.SS3.SSS3.p2.1),[§IV\-G](https://arxiv.org/html/2609.12244#S4.SS7.p2.1)\.
- \[21\]S\. R\. Mattu and A\. Chockalingam\(2022\)Learning\-based channel estimation and phase noise compensation in doubly\-selective channels\.IEEE Communications Letters26\(5\),pp\. 1052–1056\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1)\.
- \[22\]D\. Marasinghe, L\. H\. Nguyen, N\. Rajatheva, and M\. Latva\-Aho\(2025\)Phase noise resilient neural transceivers for high data\-rate sub\-thz links\.IEEE Wireless Communications Letters\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p3.1),[§I](https://arxiv.org/html/2609.12244#S1.p4.1),[§IV\-C3](https://arxiv.org/html/2609.12244#S4.SS3.SSS3.p1.1),[§IV\-C3](https://arxiv.org/html/2609.12244#S4.SS3.SSS3.p2.1),[§IV\-D](https://arxiv.org/html/2609.12244#S4.SS4.p3.1)\.
- \[23\]L\. Li, H\. Chen, H\. Chang, and L\. Liu\(2019\)Deep residual learning meets ofdm channel estimation\.IEEE Wireless Communications Letters9\(5\),pp\. 615–618\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p4.1)\.
- \[24\]H\. He, S\. Jin, C\. Wen, F\. Gao, G\. Y\. Li, and Z\. Xu\(2019\)Model\-driven deep learning for physical layer communications\.IEEE Wireless Communications26\(5\),pp\. 77–83\.Cited by:[§I](https://arxiv.org/html/2609.12244#S1.p4.1)\.
- \[25\]3GPP\(2025\)Study on channel model for frequencies from 0\.5 to 100 GHz\.Technical reportTechnical ReportTR 38\.901,3rd Generation Partnership Project \(3GPP\)\.Cited by:[§II\-B](https://arxiv.org/html/2609.12244#S2.SS2.p1.1),[§II\-B](https://arxiv.org/html/2609.12244#S2.SS2.p3.1)\.
- \[26\]W\. C\. Jakes\(1974\)Microwave mobile communications\.Wiley,New York\.Cited by:[§II\-B](https://arxiv.org/html/2609.12244#S2.SS2.p2.1)\.
- \[27\]3GPP\(2024\)Study on new radio access technology: radio frequency \(rf\) and co\-existence aspects\.Technical reportTechnical ReportTR 38\.803,3rd Generation Partnership Project \(3GPP\)\.Cited by:[§II\-C](https://arxiv.org/html/2609.12244#S2.SS3.p1.1),[TABLE I](https://arxiv.org/html/2609.12244#S2.T1),[TABLE I](https://arxiv.org/html/2609.12244#S2.T1.4),[7\(a\)](https://arxiv.org/html/2609.12244#S4.F7.sf1),[7\(a\)](https://arxiv.org/html/2609.12244#S4.F7.sf1.4),[8\(a\)](https://arxiv.org/html/2609.12244#S4.F8.sf1),[8\(a\)](https://arxiv.org/html/2609.12244#S4.F8.sf1.4),[§IV\-D](https://arxiv.org/html/2609.12244#S4.SS4.p2.1)\.
- \[28\]A\. Piemontese, G\. Colavolpe, and T\. Eriksson\(2022\)A new analytical model of phase noise in communication systems\.In2022 IEEE Wireless Communications and Networking Conference \(WCNC\),pp\. 926–931\.Cited by:[§II\-C](https://arxiv.org/html/2609.12244#S2.SS3.p1.1),[§II\-C](https://arxiv.org/html/2609.12244#S2.SS3.p1.2)\.
- \[29\]S\. Coleri, M\. Ergen, A\. Puri, and A\. Bahai\(2002\)Channel estimation techniques based on pilot arrangement in ofdm systems\.IEEE Transactions on broadcasting48\(3\),pp\. 223–229\.Cited by:[§III\-A](https://arxiv.org/html/2609.12244#S3.SS1.p1.1)\.
- \[30\]P\. Dong, H\. Zhang, G\. Y\. Li, I\. S\. Gaspar, and N\. NaderiAlizadeh\(2019\)Deep cnn\-based channel estimation for mmwave massive mimo systems\.IEEE Journal of Selected Topics in Signal Processing13\(5\),pp\. 989–1000\.Cited by:[§III\-A1](https://arxiv.org/html/2609.12244#S3.SS1.SSS1.p1.1)\.
- \[31\]Y\. Leng, Q\. Lin, L\. Yung, J\. Lei, Y\. Li, and Y\. Wu\(2025\)Unveiling the power of complex\-valued transformers in wireless communications\.IEEE Transactions on Communications74,pp\. 612–627\.Cited by:[§III\-A1](https://arxiv.org/html/2609.12244#S3.SS1.SSS1.p1.1)\.
- \[32\]J\. Chen, W\. Wong, B\. Hamdaoui, A\. Elmaghbub, K\. Sivanesan, R\. Dorrance, and L\. L\. Yang\(2022\)An analysis of complex\-valued cnns for rf data\-driven wireless device classification\.InICC 2022\-IEEE International Conference on Communications,pp\. 4318–4323\.Cited by:[§III\-A1](https://arxiv.org/html/2609.12244#S3.SS1.SSS1.p1.1)\.
- \[33\]C\. Trabelsi, O\. Bilaniuk, Y\. Zhang, D\. Serdyuk, S\. Subramanian, J\. F\. Santos, S\. Mehri, N\. Rostamzadeh, Y\. Bengio, and C\. J\. Pal\(2017\)Deep complex networks\.arXiv preprint arXiv:1705\.09792\.Cited by:[§III\-A1](https://arxiv.org/html/2609.12244#S3.SS1.SSS1.p2.1)\.
- \[34\]J\. Bassey, L\. Qian, and X\. Li\(2021\)A survey of complex\-valued neural networks\.arXiv preprint arXiv:2101\.12249\.Cited by:[§III\-A1](https://arxiv.org/html/2609.12244#S3.SS1.SSS1.p2.2)\.
- \[35\]C\. Tang, C\. Luo, Z\. Zhao, W\. Xie, and W\. Zeng\(2021\)Joint time\-frequency and time domain learning for speech enhancement\.InProceedings of the twenty\-ninth international conference on international joint conferences on artificial intelligence,pp\. 3816–3822\.Cited by:[§III\-A2](https://arxiv.org/html/2609.12244#S3.SS1.SSS2.p1.1)\.
- \[36\]K\. He, X\. Zhang, S\. Ren, and J\. Sun\(2016\)Deep residual learning for image recognition\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 770–778\.Cited by:[§III\-B1](https://arxiv.org/html/2609.12244#S3.SS2.SSS1.p1.1)\.
- \[37\]B\. Lim, S\. Son, H\. Kim, S\. Nah, and K\. Mu Lee\(2017\)Enhanced deep residual networks for single image super\-resolution\.InProceedings of the IEEE conference on computer vision and pattern recognition workshops,pp\. 136–144\.Cited by:[§III\-B1](https://arxiv.org/html/2609.12244#S3.SS2.SSS1.p1.1)\.
- \[38\]Y\. Zhang, K\. Li, K\. Li, L\. Wang, B\. Zhong, and Y\. Fu\(2018\)Image super\-resolution using very deep residual channel attention networks\.InProceedings of the European conference on computer vision \(ECCV\),pp\. 286–301\.Cited by:[§III\-B1](https://arxiv.org/html/2609.12244#S3.SS2.SSS1.p1.1),[§III\-C1](https://arxiv.org/html/2609.12244#S3.SS3.SSS1.p2.1)\.
- \[39\]X\. Liu and H\. C\. Luong\(2019\)A fully integrated 0\.27\-thz injection\-locked frequency synthesizer with frequency\-tracking loop in 65\-nm cmos\.IEEE Journal of Solid\-State Circuits55\(4\),pp\. 1051–1063\.Cited by:[§IV\-A](https://arxiv.org/html/2609.12244#S4.SS1.p3.1),[TABLE II](https://arxiv.org/html/2609.12244#S4.T2.5.15.2)\.
- \[40\]I\. Loshchilov and F\. Hutter\(2017\)Decoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101\.Cited by:[TABLE II](https://arxiv.org/html/2609.12244#S4.T2.5.17.2)\.
- \[41\]3GPP\(2021\)Study on supporting NR from 52\.6 GHz to 71 GHz\.Technical ReportTechnical ReportTR 38\.808,3rd Generation Partnership Project \(3GPP\)\.Note:Release 17, V17\.0\.0External Links:[Link](https://www.3gpp.org/dynareport/38808.htm)Cited by:[7\(b\)](https://arxiv.org/html/2609.12244#S4.F7.sf2),[7\(b\)](https://arxiv.org/html/2609.12244#S4.F7.sf2.4),[7\(c\)](https://arxiv.org/html/2609.12244#S4.F7.sf3),[7\(c\)](https://arxiv.org/html/2609.12244#S4.F7.sf3.4),[8\(b\)](https://arxiv.org/html/2609.12244#S4.F8.sf2),[8\(b\)](https://arxiv.org/html/2609.12244#S4.F8.sf2.4),[8\(c\)](https://arxiv.org/html/2609.12244#S4.F8.sf3),[8\(c\)](https://arxiv.org/html/2609.12244#S4.F8.sf3.4),[§IV\-A](https://arxiv.org/html/2609.12244#S4.SS1.p3.1),[§IV\-D](https://arxiv.org/html/2609.12244#S4.SS4.p2.1)\.
- \[42\]A\. Kendall, Y\. Gal, and R\. Cipolla\(2018\)Multi\-task learning using uncertainty to weigh losses for scene geometry and semantics\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 7482–7491\.Cited by:[§IV\-B](https://arxiv.org/html/2609.12244#S4.SS2.p3.1)\.
- \[43\]P\. Goyal, P\. Dollár, R\. Girshick, P\. Noordhuis, L\. Wesolowski, A\. Kyrola, A\. Tulloch, Y\. Jia, and K\. He\(2017\)Accurate, large minibatch sgd: training imagenet in 1 hour\.arXiv preprint arXiv:1706\.02677\.Cited by:[§IV\-B](https://arxiv.org/html/2609.12244#S4.SS2.p4.1)\.
- \[44\]I\. Loshchilov and F\. Hutter\(2016\)Sgdr: stochastic gradient descent with warm restarts\.arXiv preprint arXiv:1608\.03983\.Cited by:[§IV\-B](https://arxiv.org/html/2609.12244#S4.SS2.p4.2)\.
- \[45\]3GPP\(2024\)NR; User Equipment \(UE\) radio transmission and reception; Part 1: Range 1 Standalone\.Technical SpecificationTechnical ReportTS 38\.101\-1,3rd Generation Partnership Project \(3GPP\)\.Note:Release 18, V18\.8\.0External Links:[Link](https://www.3gpp.org/dynareport/38101-1.htm)Cited by:[§IV\-E](https://arxiv.org/html/2609.12244#S4.SS5.p1.1)\.Similar Articles
Neural Network-Assisted CLEAN for Channel Modeling in Low-SNR Regimes
This paper proposes NN-CLEAN, a hybrid framework that embeds a multi-head residual network into the iterative CLEAN extraction loop for efficient multipath parameter estimation in low-SNR channel modeling. It matches traditional Grid-Search CLEAN accuracy while greatly reducing computational complexity and enabling parallelization for real-time MIMO systems.
Frequency Domain Reservoir Computing
This paper introduces FRESCO, an Echo State Network architecture operating entirely in the frequency domain to achieve O(N) complexity for dense recurrent updates, matching state-of-the-art performance on benchmarks while reducing computational costs.
Dualformer: Efficient Feature Extractor for Complex-valued Blind Communication Signal Analysis
This paper proposes Dualformer, a dual-channel neural network architecture based on transformers, designed for efficient feature extraction from complex-valued signals in blind communication analysis tasks such as automatic modulation recognition, signal scheme recognition, and signal structure parsing. Extensive experiments show consistent performance improvements over existing methods.
Radio-Frequency Convolutional Neural Networks
Radio-frequency convolutional neural networks (RF-CNNs) repurpose existing wireless communication hardware for efficient AI inference on edge devices, demonstrating deep CNN performance with significant energy savings.
Learning to Resolve Neutron Resonances with Fully Convolutional Neural Networks
This preliminary study applies a fully convolutional neural network to automatically detect neutron resonances in transmission spectra, achieving ~93% classification accuracy but failing to generalize to unseen isotopes. The authors suggest future work with larger datasets and physics-informed features.