一种用于联合发射预编码和STAR-RIS系数优化的粒子群辅助梯度元学习算法
摘要
本文提出一种粒子群辅助的梯度元学习算法,用于多用户无线系统中联合优化发射预编码和STAR-RIS系数,相比传统方法实现了加权和速率的提升。
arXiv:2609.29150v1 Announce Type: new
Abstract: This paper investigates the joint optimization of the transmit precoder and the transmission/reflection coefficients of a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) to maximize the weighted sum rate (WSR) in a multi-user downlink. We propose a particle-swarm-assisted gradient meta-learning (PSA-GML) algorithm for this non-convex problem. The original problem is first equivalently transformed via an amplitude-split parameterization and a collapsed precoder representation, which automatically satisfy the energy-conservation constraint and reduce the search dimension. Particle swarm optimization (PSO) then performs a global search over the STAR-RIS coefficients to yield a high-quality, initialization-robust warm start, with the transmit precoder obtained in closed form. Departing from conventional alternating optimization (AO), a coordinate-wise long short-term memory (LSTM) meta-optimizer trained by first-order gradient meta-learning further refines the coefficients and precoder jointly, learning per-coordinate adaptive update rules from data. The meta-optimizer is trained offline and applied to unseen channels without further adaptation. Numerical results show that PSA-GML attains an 11.06 bits/s/Hz WSR at 10 dB with N=32 elements and K=4 users, exceeding AO by 13.1% (and by 6.2% even with multiple random restarts) and the random-phase scheme by 35.1%. In the interference-limited regime it reaches 83.9% of the hand-designed Adam refinement without manual hyper-parameter tuning, and it transfers zero-shot across regimes, indicating that the learned update rule captures the intrinsic WSR landscape structure.
查看缓存全文
缓存时间: 2026/09/25 09:45
# A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization
Source: [https://arxiv.org/html/2609.29150](https://arxiv.org/html/2609.29150)
Kang ZhouAffiliation:School of Artificial Intelligence, Mianyang City College, Mianyang 621000, China zhoukang\.sicnu@foxmail\.comAffiliation:
###### Abstract
This paper investigates the joint optimization of the transmit precoder and the transmission/reflection coefficients of a simultaneously transmitting and reflecting reconfigurable intelligent surface \(STAR\-RIS\) to maximize the weighted sum rate \(WSR\) in a multi\-user downlink system\. We propose a particle\-swarm\-assisted gradient meta\-learning \(PSA\-GML\) algorithm to solve this non\-convex problem\. Specifically, the original problem is first equivalently transformed into a tractable form by an amplitude\-split parameterization and a collapsed precoder representation, which automatically satisfy the energy\-conservation constraint and reduce the search dimension\. Then, particle swarm optimization \(PSO\) performs a global search over the STAR\-RIS coefficients to produce a high\-quality warm start that is robust to initialization, with the transmit precoder obtained in closed form\. Departing from conventional alternating optimization \(AO\), whose outcome is sensitive to the starting point, a coordinate\-wise long short\-term memory \(LSTM\) meta\-optimizer trained by first\-order gradient meta\-learning is further employed to jointly refine the STAR\-RIS coefficients and the transmit precoder, learning per\-coordinate adaptive update rules from data\. Finally, the meta\-optimizer is trained offline over multiple channel realizations and applied to unseen channels without further adaptation\. Numerical results show that PSA\-GML attains an 11\.06 bits/s/Hz WSR at a transmit SNR of 10 dB withN=32N=32elements andK=4K=4users, exceeding the conventional alternating optimization \(AO\) by 13\.1%—and by 6\.2% even when AO is further equipped with multiple random restarts—and the random\-phase scheme by 35\.1%, with robustness against initialization\. Moreover, in the interference\-limited regime the learned optimizer attains a final WSR of83\.9%83\.9\\%of that reached by the hand\-designed Adam refinement without any manual hyper\-parameter tuning, and it transfers zero\-shot across operating regimes, indicating that the learned update rule captures the intrinsic structure of the WSR landscape rather than the specifics of the training scenario\.
###### Index Terms:
STAR\-RIS, weighted sum rate, joint optimization, particle swarm optimization, gradient meta\-learning\.
## IIntroduction
The transition of wireless networks from 5G to 6G is propelled by use cases such as ultra\-dense deployment, massive machine\-type communication \(mMTC\), and ultra\-reliable low\-latency communication \(URLLC\), which impose growing demands for spectral efficiency, connectivity, and quality of service \(QoS\)\[[1](https://arxiv.org/html/2609.29150#bib.bib1),[2](https://arxiv.org/html/2609.29150#bib.bib2),[3](https://arxiv.org/html/2609.29150#bib.bib3)\]\. Massive multiple\-input multiple\-output \(MIMO\) improves spectral efficiency by deploying many antennas but incurs high hardware cost and energy consumption\[[4](https://arxiv.org/html/2609.29150#bib.bib4)\], motivating more cost\-effective solutions\. Recently, reconfigurable intelligent surfaces \(RISs\) have emerged as a promising technology that reconfigures the wireless propagation environment through a large number of low\-cost passive reflecting elements\[[5](https://arxiv.org/html/2609.29150#bib.bib5),[6](https://arxiv.org/html/2609.29150#bib.bib6),[28](https://arxiv.org/html/2609.29150#bib.bib28)\]\. By adaptively tuning the phase shifts of the reflecting elements, an RIS can create favorable signal paths and suppress interference, thereby improving both the spectral efficiency and the energy efficiency of wireless networks at a low hardware cost\[[7](https://arxiv.org/html/2609.29150#bib.bib7),[27](https://arxiv.org/html/2609.29150#bib.bib27)\]\. However, a conventional RIS can only reflect the incident signal, which means that the transmitter and the receiver must be located on the same side of the surface\. To overcome this limitation and achieve full\-space360∘360^\{\\circ\}coverage, the simultaneously transmitting and reflecting RIS \(STAR\-RIS\) has been proposed\[[8](https://arxiv.org/html/2609.29150#bib.bib8),[12](https://arxiv.org/html/2609.29150#bib.bib12)\]\. Different from a conventional RIS, each element of a STAR\-RIS can split the incident signal into a transmitted part and a reflected part, so that users on both sides of the surface can be served simultaneously\[[12](https://arxiv.org/html/2609.29150#bib.bib12)\]\. Among the various operating protocols, the energy\-splitting \(ES\) mode is the most general one, in which the transmission and reflection amplitudes are independently controlled subject to an energy\-conservation constraint\[[8](https://arxiv.org/html/2609.29150#bib.bib8)\]\.
For RIS\-aided systems, the joint optimization of the transmit precoder and the passive coefficients is essential to fully exploit the achievable performance\. To this end, a variety of optimization approaches have been proposed, including semidefinite relaxation \(SDR\)\[[5](https://arxiv.org/html/2609.29150#bib.bib5)\], alternating optimization \(AO\)\[[9](https://arxiv.org/html/2609.29150#bib.bib9)\], the fractional programming method\[[10](https://arxiv.org/html/2609.29150#bib.bib10)\], the weighted minimum mean\-square\-error \(WMMSE\) method\[[11](https://arxiv.org/html/2609.29150#bib.bib11)\], and the discrete phase\-shift design\[[29](https://arxiv.org/html/2609.29150#bib.bib29)\]\. Although these methods achieve satisfactory performance, they generally entail high computational complexity or converge to locally optimal solutions\. For STAR\-RIS\-aided systems, the joint transmission/reflection coefficient and beamforming design has been investigated in\[[12](https://arxiv.org/html/2609.29150#bib.bib12),[13](https://arxiv.org/html/2609.29150#bib.bib13)\]\. Joint beamforming and coefficient design for STAR\-RIS assisted NOMA systems and coverage characterization of STAR\-RIS networks were studied in\[[14](https://arxiv.org/html/2609.29150#bib.bib14),[15](https://arxiv.org/html/2609.29150#bib.bib15)\]\. Meanwhile, meta\-heuristic algorithms such as particle swarm optimization \(PSO\)\[[16](https://arxiv.org/html/2609.29150#bib.bib16)\]and genetic algorithms have been applied to RIS phase\-shift design owing to their global search capability\[[17](https://arxiv.org/html/2609.29150#bib.bib17)\]\. Meta\-heuristic methods, however, do not exploit the gradient information available in the problem\.
More recently, learning\-based methods have been developed to learn optimization strategies from data\. Deep unfolding networks\[[18](https://arxiv.org/html/2609.29150#bib.bib18),[19](https://arxiv.org/html/2609.29150#bib.bib19)\]unroll iterative optimization algorithms into layer\-wise neural architectures and have achieved notable success in signal processing\. As a related line, learning\-to\-optimize \(L2O\) methods train a recurrent meta\-optimizer that maps gradients to parameter updates, thereby discovering a problem\-specific update rule\[[20](https://arxiv.org/html/2609.29150#bib.bib20),[21](https://arxiv.org/html/2609.29150#bib.bib21)\]\. Moreover, model\-based deep learning\[[22](https://arxiv.org/html/2609.29150#bib.bib22)\]and meta\-learning\[[23](https://arxiv.org/html/2609.29150#bib.bib23),[24](https://arxiv.org/html/2609.29150#bib.bib24)\]have been exploited to improve both the adaptability and the robustness of wireless optimization algorithms\. In particular, the gradient meta\-learning joint optimization \(GML\-JO\) approach in\[[25](https://arxiv.org/html/2609.29150#bib.bib25)\]delegates the two blocks of optimization variables to a pair of learned networks and reports reduced sensitivity to the starting point\. Nevertheless, GML\-JO initializes the optimization from random points, which can restrict the achievable WSR on a sharply peaked landscape\.
This paper studies a STAR\-RIS aided multi\-user downlink and formulates its WSR maximization problem, which is non\-convex and intractable because the transmit precoder, the transmission/reflection amplitudes, and the phase shifts are tightly coupled under the unit\-modulus and energy\-conservation constraints\. To tackle this problem, we develop a particle\-swarm\-assisted gradient meta\-learning \(PSA\-GML\) algorithm\. Unlike AO methods, which are sensitive to the initialization, the proposed algorithm first employs PSO to perform a global search over the STAR\-RIS coefficients, producing a high\-quality warm start that is robust to initialization\. Then, building on the gradient meta\-learning paradigm of\[[25](https://arxiv.org/html/2609.29150#bib.bib25)\]but departing from it, a gradient meta\-learning stage employs a lightweight coordinate\-wise LSTM meta\-optimizer trained by first\-order gradient meta\-learning to jointly refine the STAR\-RIS coefficients and the transmit precoder\. The contributions of this paper are summarized as follows:
- •We formulate the WSR maximization of a STAR\-RIS aided multi\-user downlink, where the transmit precoder and the STAR\-RIS transmission/reflection coefficients are jointly optimized\. Through an amplitude\-split parameterization and a collapsed precoder representation, we perform an equivalent transformation of the original non\-convex problem, which automatically satisfies the energy\-conservation constraint and reduces the search dimension\.
- •We then develop the PSA\-GML algorithm\. Especially, particle swarm optimization performs a global search over the STAR\-RIS coefficients to produce a high\-quality warm start that is robust to initialization, and a coordinate\-wise LSTM meta\-optimizer trained by first\-order gradient meta\-learning learns per\-coordinate adaptive update rules to jointly refine the STAR\-RIS coefficients and the precoder\.
- •The designed meta\-optimizer is trained offline over multiple channel realizations by unrolling the update recursion and minimizing the negative WSR averaged over the unrolled steps, which is applied to unseen channels without further adaptation, thereby avoiding the initialization sensitivity of conventional gradient\-based methods\.
Numerical experiments show that PSA\-GML reaches an 11\.06 bits/s/Hz WSR, delivering a 13\.1% performance enhancement over the benchmark AO method \(6\.2% over a multiple\-random\-restart AO\) and a 35\.1% enhancement over the random\-phase scheme\. Moreover, PSA\-GML is largely insensitive to the starting point, and in the interference\-limited regime the learned optimizer attains a final WSR of4\.734\.73bits/s/Hz—83\.9%83\.9\\%of the5\.645\.64bits/s/Hz reached by the hand\-designed Adam refinement—while capturing about31\.6%31\.6\\%of the WSR gain that Adam realizes over the PSO\-only baseline\.
The remainder of this paper proceeds as follows\. Section II describes the system model and the problem formulation, and Section III derives the equivalent transformations of the original problem\. Section IV develops the proposed PSA\-GML algorithm, Section V reports the numerical results, and Section VI draws the conclusion\.
Notations:Throughout the paper, italic letters denote scalars, bold lowercase letters denote vectors, and bold uppercase letters denote matrices;diag\(𝐚\)\\textrm\{diag\}\(\\mathbf\{a\}\)constructs a diagonal matrix from the entries of𝐚\\mathbf\{a\};\(⋅\)H\(\\cdot\)^\{H\},\(⋅\)T\(\\cdot\)^\{T\}, and∥⋅∥\\\|\\cdot\\\|stand for the conjugate transpose, transpose, and Euclidean norm, respectively;𝒞𝒩\(μ,σ2\)\\mathcal\{CN\}\(\\mu,\\sigma^\{2\}\)is the complex Gaussian distribution with meanμ\\muand varianceσ2\\sigma^\{2\}; and\|⋅\|\|\\cdot\|denotes the modulus of a complex scalar\.
## IISystem Model and Problem Formulation
As shown in Fig\.[1](https://arxiv.org/html/2609.29150#S2.F1), we consider a STAR\-RIS aided downlink communication system, where a base station \(BS\) equipped withNtN\_\{t\}antennas servesKKsingle\-antenna users with the aid of a STAR\-RIS withNNelements\. TheKKusers are partitioned intoKtK\_\{t\}transmission users andKrK\_\{r\}reflection users located on opposite sides of the STAR\-RIS, withK=Kt\+KrK=K\_\{t\}\+K\_\{r\}\. The direct links between the BS and the users are assumed to be severely blocked, so that the BS–user communication is established exclusively through the STAR\-RIS\. In the energy\-splitting \(ES\) mode, the incident signal at each element is split into a transmitted part and a reflected part\[[8](https://arxiv.org/html/2609.29150#bib.bib8)\], and the transmission and reflection amplitudes are independently adjustable subject to an energy\-conservation constraint, which is more general than the mode\-switching \(MS\) and time\-switching \(TS\) protocols\. Let𝜽t=\[θt,1,…,θt,N\]T\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\}=\[\\theta\_\{t,1\},\\dots,\\theta\_\{t,N\}\]^\{T\}and𝜽r=\[θr,1,…,θr,N\]T\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}=\[\\theta\_\{r,1\},\\dots,\\theta\_\{r,N\}\]^\{T\}denote the transmission and reflection phase\-shift vectors, and let𝜷t=\[βt,1,…,βt,N\]T\\mbox\{\\boldmath\{$\\beta$\}\}\_\{t\}=\[\\beta\_\{t,1\},\\dots,\\beta\_\{t,N\}\]^\{T\}and𝜷r=\[βr,1,…,βr,N\]T\\mbox\{\\boldmath\{$\\beta$\}\}\_\{r\}=\[\\beta\_\{r,1\},\\dots,\\beta\_\{r,N\}\]^\{T\}denote the corresponding amplitude coefficients\. The transmission and reflection coefficient vectors are then given by
𝐯t=𝜷t⊙ej𝜽t,𝐯r=𝜷r⊙ej𝜽r,\\mathbf\{v\}\_\{t\}=\\mbox\{\\boldmath\{$\\beta$\}\}\_\{t\}\\odot e^\{j\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\}\},\\quad\\mathbf\{v\}\_\{r\}=\\mbox\{\\boldmath\{$\\beta$\}\}\_\{r\}\\odot e^\{j\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\},\(1\)where⊙\\odotdenotes the element\-wise product\. The amplitude coefficients satisfy the energy\-conservation constraint
βt,n2\+βr,n2=1,∀n=1,…,N\.\\beta\_\{t,n\}^\{2\}\+\\beta\_\{r,n\}^\{2\}=1,\\quad\\forall n=1,\\dots,N\.\(2\)
Let𝐆∈ℂN×Nt\\mathbf\{G\}\\in\\mathbb\{C\}^\{N\\times N\_\{t\}\}denote the BS\-to\-RIS channel, and let𝐇t∈ℂKt×N\\mathbf\{H\}\_\{t\}\\in\\mathbb\{C\}^\{K\_\{t\}\\times N\}and𝐇r∈ℂKr×N\\mathbf\{H\}\_\{r\}\\in\\mathbb\{C\}^\{K\_\{r\}\\times N\}denote the RIS\-to\-transmission\-user and RIS\-to\-reflection\-user channels, respectively\. Both the i\.i\.d\. Rayleigh and the Saleh\-Valenzuela \(SV\) channel models are considered\[[26](https://arxiv.org/html/2609.29150#bib.bib26)\]\. For the SV channel, the BS\-to\-RIS channel is modeled as
𝐆=NNtNcNray∑c=1Nc∑ℓ=1Nrayαc,ℓ𝐚R\(φc,ℓ\)𝐚TH\(ψc,ℓ\),\\mathbf\{G\}=\\sqrt\{\\frac\{NN\_\{t\}\}\{N\_\{c\}N\_\{\\mathrm\{ray\}\}\}\}\\sum\_\{c=1\}^\{N\_\{c\}\}\\sum\_\{\\ell=1\}^\{N\_\{\\mathrm\{ray\}\}\}\\alpha\_\{c,\\ell\}\\,\\mathbf\{a\}\_\{R\}\(\\varphi\_\{c,\\ell\}\)\\mathbf\{a\}\_\{T\}^\{H\}\(\\psi\_\{c,\\ell\}\),\(3\)whereNcN\_\{c\}andNrayN\_\{\\mathrm\{ray\}\}denote the numbers of clusters and rays per cluster, respectively,αc,ℓ∼𝒞𝒩\(0,1\)\\alpha\_\{c,\\ell\}\\sim\\mathcal\{CN\}\(0,1\)is the complex path gain,φc,ℓ\\varphi\_\{c,\\ell\}andψc,ℓ\\psi\_\{c,\\ell\}are the angle of arrival \(AoA\) at the STAR\-RIS and the angle of departure \(AoD\) at the BS, respectively, and𝐚R∈ℂN\\mathbf\{a\}\_\{R\}\\in\\mathbb\{C\}^\{N\}and𝐚T∈ℂNt\\mathbf\{a\}\_\{T\}\\in\\mathbb\{C\}^\{N\_\{t\}\}are the steering vectors of theNN\-element andNtN\_\{t\}\-element half\-wavelength\-spaced uniform linear arrays \(ULAs\), respectively, with
𝐚M\(φ\)=1M\[1,ejπsinφ,…,ejπ\(M−1\)sinφ\]T\.\\mathbf\{a\}\_\{M\}\(\\varphi\)=\\frac\{1\}\{\\sqrt\{M\}\}\\big\[1,e^\{j\\pi\\sin\\varphi\},\\dots,e^\{j\\pi\(M\-1\)\\sin\\varphi\}\\big\]^\{T\}\.\(4\)The RIS\-to\-user channel admits the analogous SV form, i\.e\., the channel between the STAR\-RIS and thekk\-th user is given by
𝐡k=NNcNray∑c=1Nc∑ℓ=1Nrayαc,ℓ,k𝐚N\(φc,ℓ,k\),\\mathbf\{h\}\_\{k\}=\\sqrt\{\\frac\{N\}\{N\_\{c\}N\_\{\\mathrm\{ray\}\}\}\}\\sum\_\{c=1\}^\{N\_\{c\}\}\\sum\_\{\\ell=1\}^\{N\_\{\\mathrm\{ray\}\}\}\\alpha\_\{c,\\ell,k\}\\,\\mathbf\{a\}\_\{N\}\(\\varphi\_\{c,\\ell,k\}\),\(5\)where𝐡kH\\mathbf\{h\}\_\{k\}^\{H\}is thekk\-th row of𝐇t\\mathbf\{H\}\_\{t\}or𝐇r\\mathbf\{H\}\_\{r\}\. For the i\.i\.d\. Rayleigh channel, all entries of𝐆\\mathbf\{G\},𝐇t\\mathbf\{H\}\_\{t\}, and𝐇r\\mathbf\{H\}\_\{r\}are drawn independently from𝒞𝒩\(0,1\)\\mathcal\{CN\}\(0,1\)\. All channels are normalized to unit variance, with the large\-scale path loss absorbed into the transmit SNR\. The effective channel from the BS to thekk\-th user is thekk\-th row of the composite channel matrix
𝐀=\[𝐇tdiag\(𝐯t\)𝐆𝐇rdiag\(𝐯r\)𝐆\]∈ℂK×Nt,\\mathbf\{A\}=\\begin\{bmatrix\}\\mathbf\{H\}\_\{t\}\\textrm\{diag\}\(\\mathbf\{v\}\_\{t\}\)\\mathbf\{G\}\\\\ \\mathbf\{H\}\_\{r\}\\textrm\{diag\}\(\\mathbf\{v\}\_\{r\}\)\\mathbf\{G\}\\end\{bmatrix\}\\in\\mathbb\{C\}^\{K\\times N\_\{t\}\},\(6\)where the firstKtK\_\{t\}rows correspond to the transmission users and the remainingKrK\_\{r\}rows correspond to the reflection users\. Let𝐰k∈ℂNt\\mathbf\{w\}\_\{k\}\\in\\mathbb\{C\}^\{N\_\{t\}\}denote the transmit precoding vector for thekk\-th user\. The received signal at thekk\-th user is
yk=𝐚kH𝐰ksk\+∑j≠k𝐚kH𝐰jsj\+nk,y\_\{k\}=\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{k\}s\_\{k\}\+\\sum\_\{j\\neq k\}\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{j\}s\_\{j\}\+n\_\{k\},\(7\)where𝐚kH\\mathbf\{a\}\_\{k\}^\{H\}is thekk\-th row of𝐀\\mathbf\{A\},sks\_\{k\}is the data symbol for thekk\-th user with𝔼\[\|sk\|2\]=1\\mathbb\{E\}\[\|s\_\{k\}\|^\{2\}\]=1, andnk∼𝒞𝒩\(0,σ2\)n\_\{k\}\\sim\\mathcal\{CN\}\(0,\\sigma^\{2\}\)is the additive white Gaussian noise\. Accordingly, the signal\-to\-interference\-plus\-noise ratio \(SINR\) of thekk\-th user is
γk=\|𝐚kH𝐰k\|2∑j≠k\|𝐚kH𝐰j\|2\+σ2\.\\gamma\_\{k\}=\\frac\{\|\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{k\}\|^\{2\}\}\{\\sum\_\{j\\neq k\}\|\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{j\}\|^\{2\}\+\\sigma^\{2\}\}\.\(8\)The WSR is then defined as
R\(𝐖,𝚽\)=∑k=1Kωklog2\(1\+γk\),R\(\\mathbf\{W\},\\mbox\{\\boldmath\{$\\Phi$\}\}\)=\\sum\_\{k=1\}^\{K\}\\omega\_\{k\}\\log\_\{2\}\(1\+\\gamma\_\{k\}\),\(9\)whereωk\>0\\omega\_\{k\}\>0is the priority weight of thekk\-th user,𝐖=\[𝐰1,…,𝐰K\]\\mathbf\{W\}=\[\\mathbf\{w\}\_\{1\},\\dots,\\mathbf\{w\}\_\{K\}\]collects the precoding vectors, and𝚽\\Phicollects all STAR\-RIS coefficients\. Our objective is to maximize the WSR subject to the transmit\-power and energy\-conservation constraints, which is formulated as
max𝐖,𝜷t,𝜷r,𝜽t,𝜽rR\(𝐖,𝚽\)\\displaystyle\\max\_\{\\mathbf\{W\},\\mbox\{\\boldmath\{$\\beta$\}\}\_\{t\},\\mbox\{\\boldmath\{$\\beta$\}\}\_\{r\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\}\\;R\(\\mathbf\{W\},\\mbox\{\\boldmath\{$\\Phi$\}\}\)\(10a\)s\.t\.∑k=1K‖𝐰k‖2≤Pmax,\\displaystyle\\text\{s\.t\.\}\\;\\;\\sum\_\{k=1\}^\{K\}\\\|\\mathbf\{w\}\_\{k\}\\\|^\{2\}\\leq P\_\{\\max\},\(10b\)βt,n2\+βr,n2=1,βt,n,βr,n≥0,∀n,\\displaystyle\\qquad\\beta\_\{t,n\}^\{2\}\+\\beta\_\{r,n\}^\{2\}=1,\\ \\beta\_\{t,n\},\\beta\_\{r,n\}\\geq 0,\\ \\forall n,\(10c\)θt,n,θr,n∈\[0,2π\),∀n\.\\displaystyle\\qquad\\theta\_\{t,n\},\\theta\_\{r,n\}\\in\[0,2\\pi\),\\ \\forall n\.\(10d\)
Problem \([10](https://arxiv.org/html/2609.29150#S2.E10)\) is non\-convex due to the fractional form of the SINR \([8](https://arxiv.org/html/2609.29150#S2.E8)\), the unit energy constraint \([10c](https://arxiv.org/html/2609.29150#S2.E10.3)\), and the strong coupling between the precoder and the STAR\-RIS coefficients, which renders conventional convex optimization inapplicable\. To gain further insight, consider the Lagrangian associated with the power constraint \([10b](https://arxiv.org/html/2609.29150#S2.E10.2)\):
ℒ\(𝐖,𝚽,μ\)=−R\(𝐖,𝚽\)\+μ\(∑k‖𝐰k‖2−Pmax\),\\mathcal\{L\}\(\\mathbf\{W\},\\mbox\{\\boldmath\{$\\Phi$\}\},\\mu\)=\-R\(\\mathbf\{W\},\\mbox\{\\boldmath\{$\\Phi$\}\}\)\+\\mu\\Big\(\\textstyle\\sum\_\{k\}\\\|\\mathbf\{w\}\_\{k\}\\\|^\{2\}\-P\_\{\\max\}\\Big\),\(11\)whereμ≥0\\mu\\geq 0is the dual variable\. The stationarity condition∂ℒ/∂𝐰k=𝟎\\partial\\mathcal\{L\}/\\partial\\mathbf\{w\}\_\{k\}=\\mathbf\{0\}couples all precoders through the interference terms of \([8](https://arxiv.org/html/2609.29150#S2.E8)\), so no closed\-form precoder exists in general, and the dual gap is generally nonzero owing to the non\-convexity\. We therefore avoid Lagrangian dualization and instead enforce the constraints exactly by reparameterization, as detailed in Section III\.
BSNtN\_\{t\}ObstacleSTAR\-RISNNelementsR1R\_\{1\}R2R\_\{2\}T1T\_\{1\}T2T\_\{2\}𝐆\\mathbf\{G\}𝐡t,k\\mathbf\{h\}\_\{t,k\}𝐡r,k\\mathbf\{h\}\_\{r,k\}BlockedReflection regionTransmission region
Fig\. 1:Downlink STAR\-RIS aided multi\-user communication system, where the BS serves the transmission usersT1,T2T\_\{1\},T\_\{2\}and the reflection usersR1,R2R\_\{1\},R\_\{2\}exclusively through the STAR\-RIS, since the direct BS–user links are blocked by the obstacle\.
## IIIEquivalent Transformations of the Original Problem
In this section, we transform the original non\-convex problem \([10](https://arxiv.org/html/2609.29150#S2.E10)\) into a sequence of equivalent tractable forms, which lays the foundation for the proposed PSA\-GML algorithm\.
### III\-AAmplitude\-Split Parameterization
The energy\-conservation constraint \([10c](https://arxiv.org/html/2609.29150#S2.E10.3)\) couples the transmission and reflection amplitudes of each element\. The following lemma shows that this constraint can be absorbed by a smooth amplitude\-split parameterization\.
###### Lemma 1
For any feasible amplitude pair\(βt,n,βr,n\)\(\\beta\_\{t,n\},\\beta\_\{r,n\}\)satisfyingβt,n2\+βr,n2=1\\beta\_\{t,n\}^\{2\}\+\\beta\_\{r,n\}^\{2\}=1withβt,n,βr,n≥0\\beta\_\{t,n\},\\beta\_\{r,n\}\\geq 0, there exists a unique angleφn∈\[0,π/2\]\\varphi\_\{n\}\\in\[0,\\pi/2\]such thatβt,n=cosφn\\beta\_\{t,n\}=\\cos\\varphi\_\{n\}andβr,n=sinφn\\beta\_\{r,n\}=\\sin\\varphi\_\{n\}, and vice versa\. Hence, the constraint \([10c](https://arxiv.org/html/2609.29150#S2.E10.3)\) is equivalent to the amplitude\-split parameterization
βt,n=cosφn,βr,n=sinφn,φn∈\[0,π/2\],\\beta\_\{t,n\}=\\cos\\varphi\_\{n\},\\quad\\beta\_\{r,n\}=\\sin\\varphi\_\{n\},\\quad\\varphi\_\{n\}\\in\[0,\\pi/2\],\(12\)which absorbs the energy\-conservation constraint automatically\.
###### Proof:
Sinceβt,n2\+βr,n2=1\\beta\_\{t,n\}^\{2\}\+\\beta\_\{r,n\}^\{2\}=1andβt,n,βr,n≥0\\beta\_\{t,n\},\\beta\_\{r,n\}\\geq 0, the point\(βt,n,βr,n\)\(\\beta\_\{t,n\},\\beta\_\{r,n\}\)lies on the unit circle in the first quadrant\. Therefore, there exists a unique angleφn=arctan\(βr,n/βt,n\)∈\[0,π/2\]\\varphi\_\{n\}=\\arctan\(\\beta\_\{r,n\}/\\beta\_\{t,n\}\)\\in\[0,\\pi/2\]\(withφn=π/2\\varphi\_\{n\}=\\pi/2whenβt,n=0\\beta\_\{t,n\}=0\) satisfyingβt,n=cosφn\\beta\_\{t,n\}=\\cos\\varphi\_\{n\}andβr,n=sinφn\\beta\_\{r,n\}=\\sin\\varphi\_\{n\}\. Conversely, for anyφn∈\[0,π/2\]\\varphi\_\{n\}\\in\[0,\\pi/2\], the identitycos2φn\+sin2φn=1\\cos^\{2\}\\varphi\_\{n\}\+\\sin^\{2\}\\varphi\_\{n\}=1holds, so the resulting pair is feasible\.■\\blacksquare∎
With Lemma[1](https://arxiv.org/html/2609.29150#Thmlemma1), the optimization variables reduce to the amplitude\-split angles𝝋=\[φ1,…,φN\]T\\mbox\{\\boldmath\{$\\varphi$\}\}=\[\\varphi\_\{1\},\\dots,\\varphi\_\{N\}\]^\{T\}, the phase shifts𝜽t,𝜽r\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}, and the precoder𝐖\\mathbf\{W\}, and problem \([10](https://arxiv.org/html/2609.29150#S2.E10)\) is equivalently reformulated as
max𝐖,𝝋,𝜽t,𝜽rR\(𝐖,𝚽\)\\displaystyle\\max\_\{\\mathbf\{W\},\\mbox\{\\boldmath\{$\\varphi$\}\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\}\\;R\(\\mathbf\{W\},\\mbox\{\\boldmath\{$\\Phi$\}\}\)\(13a\)s\.t\.∑k=1K‖𝐰k‖2≤Pmax,\\displaystyle\\text\{s\.t\.\}\\;\\;\\sum\_\{k=1\}^\{K\}\\\|\\mathbf\{w\}\_\{k\}\\\|^\{2\}\\leq P\_\{\\max\},\(13b\)φn∈\[0,π/2\],θt,n,θr,n∈\[0,2π\),∀n\.\\displaystyle\\qquad\\varphi\_\{n\}\\in\[0,\\pi/2\],\\ \\theta\_\{t,n\},\\theta\_\{r,n\}\\in\[0,2\\pi\),\\ \\forall n\.\(13c\)
### III\-BClosed\-Form Zero\-Forcing Precoder
For a fixed set of STAR\-RIS coefficients, the zero\-forcing \(ZF\) precoder admits a closed form and nulls all inter\-user interference, as stated below\.
###### Proposition 1
For a fixed composite channel𝐀∈ℂK×Nt\\mathbf\{A\}\\in\\mathbb\{C\}^\{K\\times N\_\{t\}\}with full row rank \(which holds whenK≤NtK\\leq N\_\{t\}\), the zero\-forcing \(ZF\) precoder
𝐖ZF=𝐀H\(𝐀𝐀H\)−1,\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}=\\mathbf\{A\}^\{H\}\(\\mathbf\{A\}\\mathbf\{A\}^\{H\}\)^\{\-1\},\(14\)satisfies𝐚kH𝐰j=0\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{j\}=0for allj≠kj\\neq k, and reduces the SINR \([8](https://arxiv.org/html/2609.29150#S2.E8)\) to the interference\-free formγk=\|𝐚kH𝐰k\|2/σ2\\gamma\_\{k\}=\|\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{k\}\|^\{2\}/\\sigma^\{2\}\.
###### Proof:
Substituting \([14](https://arxiv.org/html/2609.29150#S3.E14)\) into the composite channel yields𝐀𝐖ZF=𝐀𝐀H\(𝐀𝐀H\)−1=𝐈K\\mathbf\{A\}\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}=\\mathbf\{A\}\\mathbf\{A\}^\{H\}\(\\mathbf\{A\}\\mathbf\{A\}^\{H\}\)^\{\-1\}=\\mathbf\{I\}\_\{K\}, where𝐈K\\mathbf\{I\}\_\{K\}is theK×KK\\times Kidentity matrix\. Consequently, the\(k,j\)\(k,j\)\-th entry satisfies𝐚kH𝐰j=\[𝐀𝐖ZF\]k,j=δk,j\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{j\}=\[\\mathbf\{A\}\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}\]\_\{k,j\}=\\delta\_\{k,j\}, which vanishes forj≠kj\\neq kand equals one forj=kj=k\. Substituting into \([8](https://arxiv.org/html/2609.29150#S2.E8)\) givesγk=\|𝐚kH𝐰k\|2/σ2\\gamma\_\{k\}=\|\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{k\}\|^\{2\}/\\sigma^\{2\}\.■\\blacksquare∎Note that in practical use the ZF precoder is power\-normalized to satisfy the power constraint \([10b](https://arxiv.org/html/2609.29150#S2.E10.2)\), i\.e\.,𝐖ZF←Pmax𝐖ZF/‖𝐖ZF‖F\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}\\leftarrow\\sqrt\{P\_\{\\max\}\}\\,\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}/\\\|\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}\\\|\_\{F\}, which scales the useful signal𝐚kH𝐰k\\mathbf\{a\}\_\{k\}^\{H\}\\mathbf\{w\}\_\{k\}uniformly and preserves the zero\-forcing property\.
### III\-CCollapsed Precoder Representation
To jointly optimize the STAR\-RIS coefficients and the precoder with a reduced search dimension, we adopt the collapsed precoder representation\.
###### Lemma 2
For a fixed composite channel𝐀∈ℂK×Nt\\mathbf\{A\}\\in\\mathbb\{C\}^\{K\\times N\_\{t\}\}with full row rank, it is without loss of optimality to restrict the transmit precoder to the collapsed form
𝐖=𝐀H𝐗,\\mathbf\{W\}=\\mathbf\{A\}^\{H\}\\mathbf\{X\},\(15\)where𝐗∈ℂK×K\\mathbf\{X\}\\in\\mathbb\{C\}^\{K\\times K\}is a collapsed variable, since the null\-space component of𝐖\\mathbf\{W\}does not affect the SINR and only increases the transmit power\.
###### Proof:
Any precoder admits the orthogonal decomposition𝐖=𝐀H𝐗\+𝐖⟂\\mathbf\{W\}=\\mathbf\{A\}^\{H\}\\mathbf\{X\}\+\\mathbf\{W\}\_\{\\perp\}with𝐖⟂∈null\(𝐀\)\\mathbf\{W\}\_\{\\perp\}\\in\\mathrm\{null\}\(\\mathbf\{A\}\), i\.e\.,𝐀𝐖⟂=𝟎\\mathbf\{A\}\\mathbf\{W\}\_\{\\perp\}=\\mathbf\{0\}\. Then𝐀𝐖=𝐀𝐀H𝐗\\mathbf\{A\}\\mathbf\{W\}=\\mathbf\{A\}\\mathbf\{A\}^\{H\}\\mathbf\{X\}, so the received signal \([7](https://arxiv.org/html/2609.29150#S2.E7)\) and the SINR \([8](https://arxiv.org/html/2609.29150#S2.E8)\) depend on𝐗\\mathbf\{X\}alone and are independent of𝐖⟂\\mathbf\{W\}\_\{\\perp\}\. On the other hand, since𝐀H𝐗∈col\(𝐀H\)\\mathbf\{A\}^\{H\}\\mathbf\{X\}\\in\\mathrm\{col\}\(\\mathbf\{A\}^\{H\}\)is orthogonal tonull\(𝐀\)\\mathrm\{null\}\(\\mathbf\{A\}\), the transmit power satisfies‖𝐖‖F2=‖𝐀H𝐗‖F2\+‖𝐖⟂‖F2\\\|\\mathbf\{W\}\\\|\_\{F\}^\{2\}=\\\|\\mathbf\{A\}^\{H\}\\mathbf\{X\}\\\|\_\{F\}^\{2\}\+\\\|\\mathbf\{W\}\_\{\\perp\}\\\|\_\{F\}^\{2\}, which is minimized by𝐖⟂=𝟎\\mathbf\{W\}\_\{\\perp\}=\\mathbf\{0\}\. Therefore, for any feasible𝐖\\mathbf\{W\}, the collapsed form \([15](https://arxiv.org/html/2609.29150#S3.E15)\) achieves the same WSR with no larger transmit power\.■\\blacksquare∎
With Lemma[2](https://arxiv.org/html/2609.29150#Thmlemma2), problem \([13](https://arxiv.org/html/2609.29150#S3.E13)\) is further reformulated as
max𝐗,𝝋,𝜽t,𝜽rR\(𝐀H𝐗,𝚽\)\\displaystyle\\max\_\{\\mathbf\{X\},\\mbox\{\\boldmath\{$\\varphi$\}\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\}\\;R\(\\mathbf\{A\}^\{H\}\\mathbf\{X\},\\mbox\{\\boldmath\{$\\Phi$\}\}\)\(16a\)s\.t\.‖𝐀H𝐗‖F2≤Pmax,\\displaystyle\\text\{s\.t\.\}\\;\\;\\\|\\mathbf\{A\}^\{H\}\\mathbf\{X\}\\\|\_\{F\}^\{2\}\\leq P\_\{\\max\},\(16b\)φn∈\[0,π/2\],θt,n,θr,n∈\[0,2π\),∀n\.\\displaystyle\\qquad\\varphi\_\{n\}\\in\[0,\\pi/2\],\\ \\theta\_\{t,n\},\\theta\_\{r,n\}\\in\[0,2\\pi\),\\ \\forall n\.\(16c\)
### III\-DUnconstrained Reparameterization
To enable end\-to\-end gradient propagation, the bounded amplitude\-split angles are reparameterized through the sigmoid function\.
###### Proposition 2
The box constraintφn∈\[0,π/2\]\\varphi\_\{n\}\\in\[0,\\pi/2\]is enforced by the unconstrained reparameterization
φn=π2ς\(φ~n\)=π/21\+e−φ~n,\\varphi\_\{n\}=\\frac\{\\pi\}\{2\}\\varsigma\(\\tilde\{\\varphi\}\_\{n\}\)=\\frac\{\\pi/2\}\{1\+e^\{\-\\tilde\{\\varphi\}\_\{n\}\}\},\(17\)whereφ~n∈ℝ\\tilde\{\\varphi\}\_\{n\}\\in\\mathbb\{R\}is an unconstrained variable andς\(⋅\)\\varsigma\(\\cdot\)denotes the sigmoid function\. With \([17](https://arxiv.org/html/2609.29150#S3.E17)\), the composite channel \([6](https://arxiv.org/html/2609.29150#S2.E6)\) becomes fully differentiable with respect to𝛗~\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\},𝛉t\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\}, and𝛉r\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\.
###### Proof:
The sigmoid functionς\(x\)=1/\(1\+e−x\)\\varsigma\(x\)=1/\(1\+e^\{\-x\}\)is a strictly increasing bijection fromℝ\\mathbb\{R\}onto the open interval\(0,1\)\(0,1\), soπ2ς\(φ~n\)\\frac\{\\pi\}\{2\}\\varsigma\(\\tilde\{\\varphi\}\_\{n\}\)is a bijection fromℝ\\mathbb\{R\}onto\(0,π/2\)\(0,\\pi/2\)\. Hence, every feasibleφn∈\[0,π/2\]\\varphi\_\{n\}\\in\[0,\\pi/2\]is reachable \(the endpoints00andπ/2\\pi/2are approached in the limit\|φ~n\|→∞\|\\tilde\{\\varphi\}\_\{n\}\|\\to\\infty\), and everyφ~n\\tilde\{\\varphi\}\_\{n\}produces a feasible angle\. Since the map is smooth, the composite channel \([6](https://arxiv.org/html/2609.29150#S2.E6)\) is differentiable in𝝋~\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\}\.■\\blacksquare∎Note that the phase shifts𝜽t\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\}and𝜽r\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}require no reparameterization, since they enter the composite channel \([6](https://arxiv.org/html/2609.29150#S2.E6)\) only throughejθt,ne^\{j\\theta\_\{t,n\}\}andejθr,ne^\{j\\theta\_\{r,n\}\}, which are2π2\\pi\-periodic; any real\-valued phase is therefore equivalent to one in\[0,2π\)\[0,2\\pi\), so the phases can be optimized directly as unconstrained variables\.
## IVProposed PSA\-GML Algorithm
In this section, we present the proposed PSA\-GML algorithm, which combines global exploration and local refinement\. The key observation is that the non\-convex WSR landscape of \([16](https://arxiv.org/html/2609.29150#S3.E16)\) contains many local optima, so a purely gradient\-based method is sensitive to initialization, whereas a purely heuristic method does not exploit gradient information\. The proposed algorithm therefore consists of two stages: a PSO\-based global warm start \(Stage 1\), followed by a gradient meta\-learning refinement \(Stage 2\)\.
cd\(s−1\)c\_\{d\}^\{\(s\-1\)\}⊙\\odot\+\+cdsc\_\{d\}^\{s\}tanh\\tanh⊙\\odothdsh\_\{d\}^\{s\}f=ςf=\\varsigmai=ςi=\\varsigmac~=tanh\\tilde\{c\}=\\tanho=ςo=\\varsigma⊙\\odotgdg\_\{d\}hd\(s−1\)h\_\{d\}^\{\(s\-1\)\}gradienthidden\[gd;hd\(s−1\)\]\[\\,g\_\{d\};\\ h\_\{d\}^\{\(s\-1\)\}\\,\]Δξd\\Delta\\xi\_\{d\}readoutcell statecdc\_\{d\}carriedhidden statehdh\_\{d\}carried overSSstepsFig\. 2:Neural\-network structure of the coordinate\-wise LSTM learned optimizerℳψ\\mathcal\{M\}\_\{\\psi\}\. For each coordinatedd, the preprocessed gradientgdg\_\{d\}and the recurrent hidden statehd\(s−1\)h\_\{d\}^\{\(s\-1\)\}are concatenated and mapped through four gates—forgetff, inputii, candidatec~\\tilde\{c\}, and outputoo—via weights𝐖∗∈ℝH×1\\mathbf\{W\}\_\{\*\}\\in\\mathbb\{R\}^\{H\\times 1\}\(ongdg\_\{d\}\) and𝐔∗∈ℝH×H\\mathbf\{U\}\_\{\*\}\\in\\mathbb\{R\}^\{H\\times H\}\(onhdh\_\{d\}\)\. The gates update the cell statecdc\_\{d\}and the hidden statehdh\_\{d\}, which are carried across theSSrefinement steps \(dashed loops\); a readout layer then produces the adaptive per\-coordinate updateΔξd=ν\(𝐚Thds\+b\)\\Delta\\xi\_\{d\}=\\nu\\,\(\\mathbf\{a\}^\{T\}h\_\{d\}^\{s\}\+b\)\.### IV\-APSO\-Based Global Warm Start
In the first stage, PSO performs a global search over the STAR\-RIS coefficients\. Each particle encodes a candidate solution
𝐱=\[𝝋;𝜽t;𝜽r\]∈𝒮,\\mathbf\{x\}=\[\\mbox\{\\boldmath\{$\\varphi$\}\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\]\\in\\mathcal\{S\},\(18\)where the search space is the box𝒮=\[0,π/2\]N×\[0,2π\)N×\[0,2π\)N\\mathcal\{S\}=\[0,\\pi/2\]^\{N\}\\times\[0,2\\pi\)^\{N\}\\times\[0,2\\pi\)^\{N\}, so that the amplitude\-split angles and the phase shifts respect \([12](https://arxiv.org/html/2609.29150#S3.E12)\) and \([10d](https://arxiv.org/html/2609.29150#S2.E10.4)\) by construction\. The fitness of a particle is the WSR achieved by the closed\-form ZF precoder \([14](https://arxiv.org/html/2609.29150#S3.E14)\) evaluated at the composite channel𝐀\(𝐱\)\\mathbf\{A\}\(\\mathbf\{x\}\):
f\(𝐱\)=R\(𝐖ZF\(𝐱\),𝐱\)=∑k=1Kωklog2\(1\+γk\(𝐱\)\),f\(\\mathbf\{x\}\)=R\\big\(\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}\(\\mathbf\{x\}\),\\mathbf\{x\}\\big\)=\\sum\_\{k=1\}^\{K\}\\omega\_\{k\}\\log\_\{2\}\\\!\\big\(1\+\\gamma\_\{k\}\(\\mathbf\{x\}\)\\big\),\(19\)whereγk\(𝐱\)\\gamma\_\{k\}\(\\mathbf\{x\}\)is obtained by substituting𝐀\(𝐱\)\\mathbf\{A\}\(\\mathbf\{x\}\)and the normalized ZF precoder into \([8](https://arxiv.org/html/2609.29150#S2.E8)\)\. This keeps the fitness consistent with the final objective \([16](https://arxiv.org/html/2609.29150#S3.E16)\) while avoiding an inner precoder optimization\.
Let𝐱i\(t\)\\mathbf\{x\}\_\{i\}^\{\(t\)\}and𝐯i\(t\)\\mathbf\{v\}\_\{i\}^\{\(t\)\}denote the position and velocity of theii\-th particle at iterationtt, and let𝐩i\\mathbf\{p\}\_\{i\}and𝐠\\mathbf\{g\}denote the personal best and global best positions, respectively\. The velocity and position are updated according to
𝐯i\(t\+1\)\\displaystyle\\mathbf\{v\}\_\{i\}^\{\(t\+1\)\}=ω𝐯i\(t\)\+c1r1\(𝐩i−𝐱i\(t\)\)\+c2r2\(𝐠−𝐱i\(t\)\),\\displaystyle=\\omega\\mathbf\{v\}\_\{i\}^\{\(t\)\}\+c\_\{1\}r\_\{1\}\(\\mathbf\{p\}\_\{i\}\-\\mathbf\{x\}\_\{i\}^\{\(t\)\}\)\+c\_\{2\}r\_\{2\}\(\\mathbf\{g\}\-\\mathbf\{x\}\_\{i\}^\{\(t\)\}\),\(20\)𝐱i\(t\+1\)\\displaystyle\\mathbf\{x\}\_\{i\}^\{\(t\+1\)\}=Π𝒮\(𝐱i\(t\)\+𝐯i\(t\+1\)\),\\displaystyle=\\Pi\_\{\\mathcal\{S\}\}\\\!\\big\(\\mathbf\{x\}\_\{i\}^\{\(t\)\}\+\\mathbf\{v\}\_\{i\}^\{\(t\+1\)\}\\big\),\(21\)whereω\\omegais the inertia weight,c1c\_\{1\}andc2c\_\{2\}are the acceleration coefficients,r1,r2∼𝒰\(0,1\)r\_\{1\},r\_\{2\}\\sim\\mathcal\{U\}\(0,1\)are random numbers, andΠ𝒮\(⋅\)\\Pi\_\{\\mathcal\{S\}\}\(\\cdot\)denotes the element\-wise projection \(clipping\) onto the box𝒮\\mathcal\{S\}\. The personal and global best positions are updated as
𝐩i\\displaystyle\\mathbf\{p\}\_\{i\}←𝐱i\(t\+1\)iff\(𝐱i\(t\+1\)\)\>f\(𝐩i\),\\displaystyle\\leftarrow\\mathbf\{x\}\_\{i\}^\{\(t\+1\)\}\\ \\text\{if\}\\ \\ f\\big\(\\mathbf\{x\}\_\{i\}^\{\(t\+1\)\}\\big\)\>f\(\\mathbf\{p\}\_\{i\}\),\(22\)𝐠\\displaystyle\\mathbf\{g\}←argmaxif\(𝐩i\),\\displaystyle\\leftarrow\\arg\\max\_\{i\}f\(\\mathbf\{p\}\_\{i\}\),\(23\)so that the global\-best fitness is never degraded, i\.e\.,f\(𝐠\(t\+1\)\)≥f\(𝐠\(t\)\)f\(\\mathbf\{g\}^\{\(t\+1\)\}\)\\geq f\(\\mathbf\{g\}^\{\(t\)\}\)for alltt\. AfterIPSOI\_\{\\mathrm\{PSO\}\}iterations, the global best𝐠\\mathbf\{g\}is taken as the warm\-start point for the refinement stage\. Since PSO explores the entire feasible region𝒮\\mathcal\{S\}, the resulting warm start is largely insensitive to the random initialization, which addresses the initialization sensitivity of gradient\-based methods\.
### IV\-BGradient Meta\-Learning Refinement
Starting from the PSO warm start, the second stage refines the solution through a gradient meta\-learning backbone, i\.e\., a learned optimizer that maps the gradient of the objective \([9](https://arxiv.org/html/2609.29150#S2.E9)\) to an adaptive update\. At the warm start,𝐗\(0\)=\(𝐀𝐀H\)−1\\mathbf\{X\}^\{\(0\)\}=\(\\mathbf\{A\}\\mathbf\{A\}^\{H\}\)^\{\-1\}reproduces exactly the ZF precoder of Stage 1, so the two stages share the same objective at initialization; as𝐗\\mathbf\{X\}evolves, the collapsed precoder \([15](https://arxiv.org/html/2609.29150#S3.E15)\) becomes a strict generalization of the ZF solution\. With the collapsed precoder \([15](https://arxiv.org/html/2609.29150#S3.E15)\) and the unconstrained reparameterization \([17](https://arxiv.org/html/2609.29150#S3.E17)\), the optimization reduces to updating the vector
𝝃=\[𝝋~;𝜽t;𝜽r;vec\(𝐗\)\]∈ℝD,\\mbox\{\\boldmath\{$\\xi$\}\}=\[\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\};\\mathrm\{vec\}\(\\mathbf\{X\}\)\]\\in\\mathbb\{R\}^\{D\},\(24\)jointly, withD=3N\+2K2D=3N\+2K^\{2\}, where the complex precoder𝐗\\mathbf\{X\}is represented by its real and imaginary parts\. At each refinement step, the variables are reconstructed, the precoder is normalized to satisfy the power constraint \([10b](https://arxiv.org/html/2609.29150#S2.E10.2)\) by the differentiable scaling𝐖←𝐖Pmax/∑k‖𝐰k‖2\\mathbf\{W\}\\leftarrow\\mathbf\{W\}\\sqrt\{P\_\{\\max\}/\\sum\_\{k\}\\\|\\mathbf\{w\}\_\{k\}\\\|^\{2\}\}, the WSR \([9](https://arxiv.org/html/2609.29150#S2.E9)\) is computed, and the gradient∇𝝃R\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}Ris obtained by automatic differentiation through this scaling\.
Since each entry of the composite channelAk,m=∑nHk,nvnGn,mA\_\{k,m\}=\\sum\_\{n\}H\_\{k,n\}v\_\{n\}G\_\{n,m\}is a smooth function of the coefficients,RRis differentiable in𝝃\\xi\. Differentiating \([9](https://arxiv.org/html/2609.29150#S2.E9)\) and \([8](https://arxiv.org/html/2609.29150#S2.E8)\) yields the scalar derivative
∂R∂γk=ωkln2\(1\+γk\),\\frac\{\\partial R\}\{\\partial\\gamma\_\{k\}\}=\\frac\{\\omega\_\{k\}\}\{\\ln 2\\,\(1\+\\gamma\_\{k\}\)\},\(25\)and the chain rule
∇𝝃R=∑k=1K∂R∂γk∇𝝃γk\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}R=\\sum\_\{k=1\}^\{K\}\\frac\{\\partial R\}\{\\partial\\gamma\_\{k\}\}\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}\\gamma\_\{k\}\(26\)closes the gradient, where eachγk\\gamma\_\{k\}is a smooth ratio of quadratics in𝐀\\mathbf\{A\}and𝐖\\mathbf\{W\}\. The dependence of the composite channel on the coefficients is rank\-one in each element, i\.e\.,
∂Ak,m∂φn=Hk,n∂vn∂φnGn,m,∂Ak,m∂θn=jHk,nvnGn,m,\\frac\{\\partial A\_\{k,m\}\}\{\\partial\\varphi\_\{n\}\}=H\_\{k,n\}\\frac\{\\partial v\_\{n\}\}\{\\partial\\varphi\_\{n\}\}G\_\{n,m\},\\quad\\frac\{\\partial A\_\{k,m\}\}\{\\partial\\theta\_\{n\}\}=j\\,H\_\{k,n\}v\_\{n\}G\_\{n,m\},\(27\)with∂vt,n/∂φn=−sinφnejθt,n\\partial v\_\{t,n\}/\\partial\\varphi\_\{n\}=\-\\sin\\varphi\_\{n\}\\,e^\{j\\theta\_\{t,n\}\}and∂vr,n/∂φn=cosφnejθr,n\\partial v\_\{r,n\}/\\partial\\varphi\_\{n\}=\\cos\\varphi\_\{n\}\\,e^\{j\\theta\_\{r,n\}\}for the transmission and reflection coefficients, respectively, and∂Ak,m/∂θn=jHk,nvt,nGn,m\\partial A\_\{k,m\}/\\partial\\theta\_\{n\}=j\\,H\_\{k,n\}v\_\{t,n\}G\_\{n,m\}fork≤Ktk\\leq K\_\{t\}orjHk,nvr,nGn,mj\\,H\_\{k,n\}v\_\{r,n\}G\_\{n,m\}fork\>Ktk\>K\_\{t\}, which renders the backward pass fully explicit\.
Instead of a hand\-designed optimizer, the variables are updated by a*learned optimizer*:
𝝃\(s\+1\)=𝝃\(s\)\+ℳψ\(∇𝝃R\(𝝃\(s\)\)\),\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\+1\)\}=\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\+\\mathcal\{M\}\_\{\\psi\}\\\!\\left\(\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}R\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\)\\right\),\(28\)whereℳψ\(⋅\)\\mathcal\{M\}\_\{\\psi\}\(\\cdot\)is a coordinate\-wise long short\-term memory \(LSTM\) meta\-optimizer parameterized byψ\\psi\[[20](https://arxiv.org/html/2609.29150#bib.bib20)\], whose neural\-network structure is illustrated in Fig\.[2](https://arxiv.org/html/2609.29150#S4.F2)\. In a coordinate\-wise LSTM, the same cell is applied independently to every scalar coordinate of the gradient, and the hidden state carried by each coordinate stores its own momentum\-like memory, thereby playing the role of the per\-coordinate adaptive statistics that make Adam robust\. To account for the distinct scales of the STAR\-RIS coefficients and the precoder, two separate LSTMs are employed, i\.e\.,ℳc\\mathcal\{M\}\_\{c\}for\[𝝋~;𝜽t;𝜽r\]\[\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\};\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}\]andℳx\\mathcal\{M\}\_\{x\}forvec\(𝐗\)\\mathrm\{vec\}\(\\mathbf\{X\}\), and the gradient is preprocessed by the symmetric logarithmic mapping
gd=sgn\(∇ξdR\)log\(1\+\|∇ξdR\|\),g\_\{d\}=\\mathrm\{sgn\}\\\!\\left\(\\nabla\_\{\\xi\_\{d\}\}R\\right\)\\,\\log\\\!\\left\(1\+\\left\|\\nabla\_\{\\xi\_\{d\}\}R\\right\|\\right\),\(29\)to handle its wide dynamic range\. Concretely, for thedd\-th optimization coordinate with preprocessed gradientgdg\_\{d\}, the meta\-optimizer maintains hidden and cell state vectors𝐡d,𝐜d∈ℝH\\mathbf\{h\}\_\{d\},\\mathbf\{c\}\_\{d\}\\in\\mathbb\{R\}^\{H\}\(with hidden sizeHH\) and updates them according to
𝐟d\\displaystyle\\mathbf\{f\}\_\{d\}=ς\(𝐖fgd\+𝐔f𝐡d\+𝐛f\),\\displaystyle=\\varsigma\(\\mathbf\{W\}\_\{f\}g\_\{d\}\+\\mathbf\{U\}\_\{f\}\\mathbf\{h\}\_\{d\}\+\\mathbf\{b\}\_\{f\}\),\(30\)𝐢d\\displaystyle\\mathbf\{i\}\_\{d\}=ς\(𝐖igd\+𝐔i𝐡d\+𝐛i\),\\displaystyle=\\varsigma\(\\mathbf\{W\}\_\{i\}g\_\{d\}\+\\mathbf\{U\}\_\{i\}\\mathbf\{h\}\_\{d\}\+\\mathbf\{b\}\_\{i\}\),𝐜~d\\displaystyle\\tilde\{\\mathbf\{c\}\}\_\{d\}=tanh\(𝐖cgd\+𝐔c𝐡d\+𝐛c\),\\displaystyle=\\tanh\(\\mathbf\{W\}\_\{c\}g\_\{d\}\+\\mathbf\{U\}\_\{c\}\\mathbf\{h\}\_\{d\}\+\\mathbf\{b\}\_\{c\}\),𝐨d\\displaystyle\\mathbf\{o\}\_\{d\}=ς\(𝐖ogd\+𝐔o𝐡d\+𝐛o\),\\displaystyle=\\varsigma\(\\mathbf\{W\}\_\{o\}g\_\{d\}\+\\mathbf\{U\}\_\{o\}\\mathbf\{h\}\_\{d\}\+\\mathbf\{b\}\_\{o\}\),𝐜d\\displaystyle\\mathbf\{c\}\_\{d\}←𝐟d⊙𝐜d\+𝐢d⊙𝐜~d,\\displaystyle\\leftarrow\\mathbf\{f\}\_\{d\}\\odot\\mathbf\{c\}\_\{d\}\+\\mathbf\{i\}\_\{d\}\\odot\\tilde\{\\mathbf\{c\}\}\_\{d\},𝐡d\\displaystyle\\mathbf\{h\}\_\{d\}←𝐨d⊙tanh\(𝐜d\),\\displaystyle\\leftarrow\\mathbf\{o\}\_\{d\}\\odot\\tanh\(\\mathbf\{c\}\_\{d\}\),and outputs the coordinate updateΔξd=ν\(𝐚T𝐡d\+b\)\\Delta\\xi\_\{d\}=\\nu\\,\(\\mathbf\{a\}^\{T\}\\mathbf\{h\}\_\{d\}\+b\)with output scaleν\\nu, where𝐚∈ℝH\\mathbf\{a\}\\in\\mathbb\{R\}^\{H\}andbbare the readout weight and bias\. The parameters\{𝐖∗,𝐔∗,𝐛∗,𝐚,b\}\\\{\\mathbf\{W\}\_\{\\ast\},\\mathbf\{U\}\_\{\\ast\},\\mathbf\{b\}\_\{\\ast\},\\mathbf\{a\},b\\\}are shared across all coordinates, so the per\-coordinate memory resides solely in the recurrent state\(𝐡d,𝐜d\)\(\\mathbf\{h\}\_\{d\},\\mathbf\{c\}\_\{d\}\), which mimics the momentum and running statistics of a hand\-designed optimizer such as Adam\[[20](https://arxiv.org/html/2609.29150#bib.bib20)\]\.
The meta\-optimizer is trained*offline*over a set of channel realizations by unrolling the recursion \([28](https://arxiv.org/html/2609.29150#S4.E28)\) forTTsteps and minimizing the negative average WSR:
minψ1\|𝒞\|∑c∈𝒞\[−1T∑s=1TRc\(𝝃\(s\)\)\],\\min\_\{\\psi\}\\;\\frac\{1\}\{\|\\mathcal\{C\}\|\}\\sum\_\{c\\in\\mathcal\{C\}\}\\left\[\-\\frac\{1\}\{T\}\\sum\_\{s=1\}^\{T\}R\_\{c\}\\\!\\left\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\\right\)\\right\],\(31\)where𝒞\\mathcal\{C\}is the training set of channel realizations and𝝃\(0\)\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(0\)\}is the PSO warm start of channelcc\. The average loss in \([31](https://arxiv.org/html/2609.29150#S4.E31)\) encourages both fast convergence and a high final WSR\. Notably, since the WSR landscape at the warm start is sharply peaked—the ZF precoder drives the inter\-user interference nearly to zero, so any perturbation sharply degrades the SINR—a second\-order backpropagation through \([28](https://arxiv.org/html/2609.29150#S4.E28)\) is numerically unstable\. We therefore detach the gradient∇𝝃R\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}Rfrom the computational graph during meta\-training, which corresponds to a first\-order \(truncated\-backpropagation\) approximation in the spirit of first\-order MAML \(FOMAML\)\[[23](https://arxiv.org/html/2609.29150#bib.bib23)\]and renders the training stable while the LSTM still learns per\-coordinate adaptive update rules\. Fig\.[3](https://arxiv.org/html/2609.29150#S4.F3)summarizes the complete offline meta\-training procedure\.
Offline meta\-training: gradient meta\-learning \(EEepochs\)training channel set\{c∈𝒞\}\\\{c\\in\\mathcal\{C\}\\\}\(minibatch ofBB\)PSO warm start→𝝃c\(0\)\\rightarrow\\mbox\{\\boldmath\{$\\xi$\}\}\_\{c\}^\{\(0\)\}\(Stage 1\)coordinate\-wise LSTMℳψ\\mathcal\{M\}\_\{\\psi\}— the learned optimizer \(parametersψ\\psi\)𝝃\(s\)\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}forward \+ gradientRc\(𝝃\(s\)\),∇𝝃RR\_\{c\}\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\),\\ \\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}Rpreprocessg=sgn\(∇R\)log\(1\+\|∇R\|\)g=\\mathrm\{sgn\}\(\\nabla R\)\\log\(1\{\+\}\|\\nabla R\|\)LSTM update𝝃\(s\+1\)=𝝃\(s\)\+Δ𝝃\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\+1\)\}=\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\+\\Delta\\mbox\{\\boldmath\{$\\xi$\}\}𝝃\(s\+1\)\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\+1\)\}unrollTTstepsL=−1T∑s=1TRc\(𝝃\(s\)\)L=\-\\frac\{1\}\{T\}\\sum\_\{s=1\}^\{T\}R\_\{c\}\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\), averaged over the minibatchtrajectory WSRfirst\-order update: detach∇𝝃R\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}R;ψ←Adam\(∇ψL\)\\psi\\leftarrow\\mathrm\{Adam\}\(\\nabla\_\{\\psi\}L\)repeatEEepochstrained optimizerℳψ∗\\mathcal\{M\}\_\{\\psi^\{\*\}\}\(deploy:SS\-step refinement on new channels\)afterEEepochsFig\. 3:Offline meta\-training of the coordinate\-wise LSTM learned optimizerℳψ\\mathcal\{M\}\_\{\\psi\}\.*Inner loop:*for each training channel, the PSO warm start𝝃c\(0\)\\mbox\{\\boldmath\{$\\xi$\}\}\_\{c\}^\{\(0\)\}is unrolled forTTsteps through the LSTM, producing a WSR trajectory\.*Outer loop:*the meta\-loss−1T∑sRc\(𝝃\(s\)\)\-\\frac\{1\}\{T\}\\sum\_\{s\}R\_\{c\}\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\)is averaged over a minibatch and minimized by a first\-order \(FOMAML\) update of the LSTM parametersψ\\psi, repeated forEEepochs; the trained optimizerℳψ∗\\mathcal\{M\}\_\{\\psi^\{\*\}\}is then deployed forSS\-step refinement on new channels\.
### IV\-CProposed PSA\-GML Algorithm
The overall procedure of the proposed PSA\-GML algorithm is summarized in Algorithm 1\.
Algorithm 1Proposed PSA\-GML Algorithm1:procedurePSA\-GML\(
𝐆,𝐇t,𝐇r,\{ωk\},σ2,Pmax\\mathbf\{G\},\\mathbf\{H\}\_\{t\},\\mathbf\{H\}\_\{r\},\\\{\\omega\_\{k\}\\\},\\sigma^\{2\},P\_\{\\max\}\)
2:Initialize
PPparticles
\{𝐱i\(0\)\}\\\{\\mathbf\{x\}\_\{i\}^\{\(0\)\}\\\}and velocities
\{𝐯i\(0\)\}\\\{\\mathbf\{v\}\_\{i\}^\{\(0\)\}\\\}randomly in
𝒮\\mathcal\{S\}
3:Initialize personal best
𝐩i←𝐱i\(0\)\\mathbf\{p\}\_\{i\}\\leftarrow\\mathbf\{x\}\_\{i\}^\{\(0\)\}and global best
𝐠\\mathbf\{g\}
4:for
t←1t\\leftarrow 1to
IPSOI\_\{\\mathrm\{PSO\}\}do
5:foreach particle
iido
6:Construct
𝐀\(𝐱i\)\\mathbf\{A\}\(\\mathbf\{x\}\_\{i\}\)via \([6](https://arxiv.org/html/2609.29150#S2.E6)\) and
𝐖ZF\\mathbf\{W\}\_\{\\mathrm\{ZF\}\}via \([14](https://arxiv.org/html/2609.29150#S3.E14)\)
7:Compute fitness
f\(𝐱i\)f\(\\mathbf\{x\}\_\{i\}\)via \([19](https://arxiv.org/html/2609.29150#S4.E19)\)
8:Update
𝐩i\\mathbf\{p\}\_\{i\}and
𝐠\\mathbf\{g\}via \([22](https://arxiv.org/html/2609.29150#S4.E22)\)–\([23](https://arxiv.org/html/2609.29150#S4.E23)\)
9:foreach particle
iido
10:Update
𝐯i\\mathbf\{v\}\_\{i\}and
𝐱i\\mathbf\{x\}\_\{i\}via \([20](https://arxiv.org/html/2609.29150#S4.E20)\)–\([21](https://arxiv.org/html/2609.29150#S4.E21)\)
11:Initialize
𝝋\(0\)←𝐠\\mbox\{\\boldmath\{$\\varphi$\}\}^\{\(0\)\}\\leftarrow\\mathbf\{g\}; map to unconstrained
𝝋~\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\}via \([17](https://arxiv.org/html/2609.29150#S3.E17)\); set
𝐗\(0\)←\(𝐀𝐀H\)−1\\mathbf\{X\}^\{\(0\)\}\\leftarrow\(\\mathbf\{A\}\\mathbf\{A\}^\{H\}\)^\{\-1\}
12:Initialize
Rbest←−∞R\_\{\\mathrm\{best\}\}\\leftarrow\-\\infty
13:for
s←1s\\leftarrow 1to
SSdo
14:Reconstruct
𝝋=π2ς\(𝝋~\)\\mbox\{\\boldmath\{$\\varphi$\}\}=\\frac\{\\pi\}\{2\}\\varsigma\(\\tilde\{\\mbox\{\\boldmath\{$\\varphi$\}\}\}\); form
𝐀\\mathbf\{A\}and
𝐖=𝐀H𝐗\\mathbf\{W\}=\\mathbf\{A\}^\{H\}\\mathbf\{X\}; normalize
𝐖\\mathbf\{W\}to satisfy \([10b](https://arxiv.org/html/2609.29150#S2.E10.2)\)
15:Compute
RRvia \([9](https://arxiv.org/html/2609.29150#S2.E9)\) and
∇𝝃R\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}Rvia \([25](https://arxiv.org/html/2609.29150#S4.E25)\)–\([27](https://arxiv.org/html/2609.29150#S4.E27)\)
16:if
R\>RbestR\>R\_\{\\mathrm\{best\}\}then
Rbest←RR\_\{\\mathrm\{best\}\}\\leftarrow R,
𝝃best←𝝃\\mbox\{\\boldmath\{$\\xi$\}\}\_\{\\mathrm\{best\}\}\\leftarrow\\mbox\{\\boldmath\{$\\xi$\}\}
17:Update
𝝃←𝝃\+ℳψ\(∇𝝃R\)\\mbox\{\\boldmath\{$\\xi$\}\}\\leftarrow\\mbox\{\\boldmath\{$\\xi$\}\}\+\\mathcal\{M\}\_\{\\psi\}\\\!\\left\(\\nabla\_\{\\mbox\{\\boldmath\{$\\xi$\}\}\}R\\right\)via \([28](https://arxiv.org/html/2609.29150#S4.E28)\)
18:return
𝐖∗\\mathbf\{W\}^\{\\ast\},
\(𝝋∗,𝜽t∗,𝜽r∗\)\(\\mbox\{\\boldmath\{$\\varphi$\}\}^\{\\ast\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{t\}^\{\\ast\},\\mbox\{\\boldmath\{$\\theta$\}\}\_\{r\}^\{\\ast\}\)reconstructed from
𝝃best\\mbox\{\\boldmath\{$\\xi$\}\}\_\{\\mathrm\{best\}\}
### IV\-DConvergence and Complexity Analysis
We first establish a monotone non\-decreasing property of the proposed algorithm, which is a weaker guarantee than convergence to a stationary point but ensures that the returned solution is no worse than the warm start\.
###### Proposition 3
The proposed PSA\-GML algorithm produces a WSR no smaller than the WSR of its PSO warm start\.
###### Proof:
In Stage 1, the global best is updated only when a strictly larger fitness is found, i\.e\.,f\(𝐠\(t\+1\)\)=max\{f\(𝐠\(t\)\),maxif\(𝐱i\(t\+1\)\)\}≥f\(𝐠\(t\)\)f\(\\mathbf\{g\}^\{\(t\+1\)\}\)=\\max\\big\\\{f\(\\mathbf\{g\}^\{\(t\)\}\),\\max\_\{i\}f\(\\mathbf\{x\}\_\{i\}^\{\(t\+1\)\}\)\\big\\\}\\geq f\(\\mathbf\{g\}^\{\(t\)\}\)\. In Stage 2, the refinement is initialized at the warm start—the coefficients are taken from𝐠\\mathbf\{g\}and the collapsed precoder is the ZF solution𝐗\(0\)=\(𝐀𝐀H\)−1\\mathbf\{X\}^\{\(0\)\}=\(\\mathbf\{A\}\\mathbf\{A\}^\{H\}\)^\{\-1\}, so thatR\(𝝃\(0\)\)=f\(𝐠\(IPSO\)\)R\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(0\)\}\)=f\(\\mathbf\{g\}^\{\(I\_\{\\mathrm\{PSO\}\}\)\}\)—and the best iterate is retained along the trajectory, givingmaxsR\(𝝃\(s\)\)≥R\(𝝃\(0\)\)\\max\_\{s\}R\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(s\)\}\)\\geq R\(\\mbox\{\\boldmath\{$\\xi$\}\}^\{\(0\)\}\)\. Combining the two stages completes the proof\.■\\blacksquare∎
We then analyze the per\-evaluation computational complexity of the proposed algorithm and the AO baseline\. For the PSO stage, each particle evaluation involves constructing the composite channel𝐀\\mathbf\{A\}in \([6](https://arxiv.org/html/2609.29150#S2.E6)\), computing the ZF precoder in \([14](https://arxiv.org/html/2609.29150#S3.E14)\), and evaluating the WSR in \([9](https://arxiv.org/html/2609.29150#S2.E9)\), which costs𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)dominated by the matrix inversion\. WithPPparticles andIPSOI\_\{\\mathrm\{PSO\}\}iterations, the total complexity of Stage 1 is𝒪\(PIPSO\(NNtK\+K3\)\)\\mathcal\{O\}\(PI\_\{\\mathrm\{PSO\}\}\(NN\_\{t\}K\+K^\{3\}\)\)\. For Stage 2, each unrolled step involves the forward WSR evaluation𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)together with the coordinate\-wise LSTM update𝒪\(DH2\)\\mathcal\{O\}\(DH^\{2\}\)with hidden sizeHH, giving𝒪\(S\(NNtK\+K3\+DH2\)\)\\mathcal\{O\}\(S\(NN\_\{t\}K\+K^\{3\}\+DH^\{2\}\)\)in total\. We note that the LSTM term is in fact the dominant per\-step cost—withN=32N=32,Nt=32N\_\{t\}=32,K=4K=4, andH=32H=32, one hasD=3N\+2K2=128D=3N\+2K^\{2\}=128andDH2≈1\.3×105DH^\{2\}\\approx 1\.3\\times 10^\{5\}, whereasNNtK\+K3≈4\.1×103NN\_\{t\}K\+K^\{3\}\\approx 4\.1\\times 10^\{3\}\. In contrast, each iteration of the AO baseline alternately updates the STAR\-RIS coefficients by a gradient\-ascent step and the precoder by the regularized ZF solution\[[9](https://arxiv.org/html/2609.29150#bib.bib9)\], both of which share the same order𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)dominated by theK×KK\\times Kmatrix inversion\. Hence, the per\-iteration complexity of the proposed refinement and that of the AO baseline are of the same order, and the practical computational advantage of PSA\-GML stems from the parallelism of thePPparticles in Stage 1 and from its initialization robustness, which removes the need for the multiple random restarts that AO would otherwise require, rather than from a lower asymptotic per\-iteration cost\. Summarizing, the overall complexity of the proposed algorithm is
𝒞PSA\-GML=𝒪\(CLOSE\\displaystyle\\mathcal\{C\}\_\{\\mathrm\{PSA\\text\{\-\}GML\}\}=\\mathcal\{O\}\\\!\\big\(PIPSO\(NNtK\+K3\)\\displaystyle PI\_\{\\mathrm\{PSO\}\}\(NN\_\{t\}K\+K^\{3\}\)OPEN\+S\(NNtK\+K3\+DH2\)\),\\displaystyle\+S\(NN\_\{t\}K\+K^\{3\}\+DH^\{2\}\)\\big\),\(32\)whereD=3N\+2K2D=3N\+2K^\{2\}andHHis the LSTM hidden size\.
## VSimulation Results
In this section, we evaluate the performance of the proposed PSA\-GML algorithm in STAR\-RIS aided multi\-user downlink systems through simulation experiments\.
### V\-AParameter Setting and Baseline
Unless otherwise stated, the BS is equipped withNt=32N\_\{t\}=32antennas and the STAR\-RIS hasN=32N=32elements servingKt=2K\_\{t\}=2transmission users andKr=2K\_\{r\}=2reflection users\. The noise power is set toσ2=1\\sigma^\{2\}=1, and the transmit SNR is1010dB\. Both i\.i\.d\. Rayleigh channels and Saleh\-Valenzuela \(SV\) channels withNc=3N\_\{c\}=3clusters andNray=8N\_\{\\mathrm\{ray\}\}=8rays are considered\. The PSO parameters are set asP=20P=20,IPSO=50I\_\{\\mathrm\{PSO\}\}=50,ω=0\.7\\omega=0\.7, andc1=c2=1\.5c\_\{1\}=c\_\{2\}=1\.5, and the refinement stage usesS=300S=300steps of the learned optimizer with hidden sizeH=32H=32and output scaleν=0\.05\\nu=0\.05, trained offline over120120channel realizations via first\-order gradient meta\-learning before deployment\. All results are averaged over independent channel realizations, with1010realizations for the SNR and user sweeps,88for the element sweep,1616for the interference\-limited study, and2020for the convergence behavior\.
We compare the proposed PSA\-GML algorithm with the following baselines:
- •Random phase \+ ZF:The STAR\-RIS coefficients are randomly generated, and the precoder is the closed\-form ZF solution \([14](https://arxiv.org/html/2609.29150#S3.E14)\)\.
- •AO/BCD:An alternating optimization \(block coordinate descent\) baseline that alternately updates the STAR\-RIS coefficients by gradient ascent and the precoder by the regularized ZF solution until convergence\[[9](https://arxiv.org/html/2609.29150#bib.bib9)\]\. It is initialized from random phases and iterated until the WSR increment falls below a small tolerance or a maximum number of iterationsNAON\_\{\\mathrm\{AO\}\}is reached\. Because the AO outcome is sensitive to its starting point, in addition to the single\-run AO depicted in the figures, we also report the best result over2020random restarts, whose total online budget is no smaller than that of the proposed algorithm, to ensure a fair comparison\.
- •PSO only:The PSO stage of the proposed algorithm without the refinement stage\.
- •Refine only:The gradient refinement stage initialized from random coefficients rather than the PSO warm start\.
### V\-BConvergence Behavior
Fig\.[4](https://arxiv.org/html/2609.29150#S5.F4)shows the convergence behavior of the proposed PSA\-GML algorithm and the baseline methods\. It is observed that the proposed algorithm converges rapidly to a high WSR within a small number of refinement steps and starts from a higher initial WSR than the refine\-only baseline\. In contrast, the refine\-only baseline, which is initialized from random coefficients, converges to a substantially lower WSR\.
Fig\. 4:Convergence behavior of the proposed PSA\-GML algorithm and the baselines\.
### V\-CPerformance versus SNR
Fig\.[5](https://arxiv.org/html/2609.29150#S5.F5)illustrates the WSR of all methods versus the transmit SNR\. As the SNR increases, the WSR of all methods increases\. In the low\-SNR regime, the performance gap among the algorithms is relatively small, while in the high\-SNR regime, the proposed PSA\-GML algorithm achieves a more pronounced gain over the baselines\. At SNR=10=10dB, the proposed algorithm attains11\.0611\.06bits/s/Hz versus9\.789\.78bits/s/Hz for a single AO run and10\.4210\.42bits/s/Hz for the AO baseline with the best of2020random restarts, i\.e\., a13\.1%13\.1\\%and a6\.2%6\.2\\%advantage, respectively, which confirms that the gain over AO is not an artifact of an unfavorable AO initialization\.
Fig\. 5:WSR performance comparison versus the transmit SNR\.
### V\-DLearned Optimizer in the Interference\-Limited Regime
To isolate the contribution of the meta\-learning component, we evaluate Stage 2 in an interference\-limited scenario with a low transmit SNR of00dB andKt=Kr=4K\_\{t\}=K\_\{r\}=4\(i\.e\.,K=8K=8\) users\. The meta\-optimizer is trained offline over120120channel realizations with aT=8T=8\-step unroll and hidden sizeH=32H=32, and is evaluated on held\-out channels with300300refinement steps, the same budget as the Adam reference\. A short unroll \(T=8T=8\) is used during training to keep the computational graph shallow and the backpropagation numerically stable, whereas the deployed optimizer is unrolled forS=300S=300steps; this is valid because the coordinate\-wise LSTM learns a per\-step update rule that is applied identically at every step and can therefore be safely unrolled beyond the training horizon\.
Table[I](https://arxiv.org/html/2609.29150#S5.T1)compares \(i\) the PSO warm start alone, and PSO followed by \(ii\) a plain SGD refinement, \(iii\) the learned optimizer, and \(iv\) the hand\-designed Adam refinement, all with the same300300\-step budget, and Fig\.[6](https://arxiv.org/html/2609.29150#S5.F6)depicts the corresponding average WSR trajectories\. Three observations are in order\. First, under Adam the refinement stage improves the WSR by about31%31\\%relative to the PSO\-only baseline\. Second, the WSR of the SGD refinement \(4\.324\.32bits/s/Hz\) is nearly identical to that of the PSO\-only baseline \(4\.314\.31bits/s/Hz\)\. Third, the learned optimizer, which is learned entirely from data, reaches83\.9%83\.9\\%of the Adam reference\. This is notable because the hand\-designed Adam here serves as a performance ceiling whose per\-coordinate adaptive statistics are manually tuned, whereas the learned optimizer approaches this ceiling without any manual hyper\-parameter tuning and transfers directly to unseen channels\.
Table[I](https://arxiv.org/html/2609.29150#S5.T1)also reports the same ablation in the main regime \(SNR=10=10dB,K=4K=4\), and the two central observations carry over unchanged\. First, plain SGD again fails to improve over the PSO warm start \(10\.8510\.85bits/s/Hz for both\), confirming that per\-coordinate adaptivity—rather than the mere availability of gradient information—is the essential ingredient on the sharply peaked WSR landscape\. Second, the learned optimizer attains11\.0611\.06bits/s/Hz, i\.e\.,99\.4%99\.4\\%of the11\.1211\.12bits/s/Hz Adam reference, essentially closing the gap to the hand\-designed ceiling\. The refinement gain is smaller in the main regime \(\+2\.0%\+2\.0\\%over PSO\) than in the interference\-limited regime \(\+31%\+31\\%over PSO\), which is a direct consequence of the higher quality of the PSO warm start at a moderate load and SNR: when the warm start already lies close to a locally optimal solution, the headroom left for any refinement optimizer is inherently small\.
Beyond generalization to unseen channel realizations within the training distribution, the learned update rule exhibits a zero\-shot transfer property across operating regimes\. Applying the meta\-optimizer trained exclusively in the main regime \(SNR=10=10dB,K=4K=4\) zero\-shot to this interference\-limited setting—without any fine\-tuning—yields an average WSR of4\.854\.85bits/s/Hz, i\.e\., approximately86%86\\%of the Adam reference and on par with the dedicated model in Table[I](https://arxiv.org/html/2609.29150#S5.T1)\. This contrasts sharply with a hand\-designed optimizer such as Adam, whose step size and momentum statistics must be re\-tuned per scenario\. The observed zero\-shot transfer therefore indicates that the LSTM learns a per\-coordinate update rule intrinsic to the WSR landscape rather than to any particular operating point\.
TABLE I:WSR \(bits/s/Hz\) of the PSO warm start followed by different refinement optimizers, in the main regime \(SNR=10=10dB,K=4K=4\) and the interference\-limited regime \(SNR=0=0dB,K=8K=8\)MethodMainIntfPSO only10\.854\.31PSO \+ SGD refinement10\.854\.32PSO \+ learned optimizer11\.064\.73PSO \+ Adam refinement11\.125\.64Fig\. 6:Average WSR convergence of the PSO warm start followed by different refinement optimizers \(SGD, the learned optimizer, and Adam\) in the interference\-limited regime\.
### V\-EImpact of the Number of Elements and Users
Figs\.[7](https://arxiv.org/html/2609.29150#S5.F7)and[8](https://arxiv.org/html/2609.29150#S5.F8)evaluate the WSR versus the number of STAR\-RIS elementsNNand the number of usersKK, respectively\. It is observed that the WSR of all methods grows withNN, and the proposed algorithm attains the highest WSR across the entire range ofNN, with its relative advantage over the baselines most pronounced at small\-to\-moderateNN\. Similarly, the proposed algorithm attains the highest sum rate for all considered values ofKK\.
Fig\. 7:WSR performance comparison versus the number of STAR\-RIS elementsNN\.Fig\. 8:Sum\-rate performance comparison versus the number of usersKK\.
### V\-FWarm\-Start Gain Analysis
Fig\.[9](https://arxiv.org/html/2609.29150#S5.F9)investigates the gain of the PSO warm start, defined as the relative WSR improvement of the proposed algorithm over the refine\-only baseline, versusNNunder both i\.i\.d\. Rayleigh and SV channels\. Two observations are noteworthy\. First, the warm\-start gain is substantial under the correlated SV channel, reaching20%20\\%or more for all considered values ofNNand peaking at37\.4%37\.4\\%atN=16N=16\. Second, the gain is much smaller under the i\.i\.d\. Gaussian channel \(around10%10\\%forN≤64N\\leq 64, declining to6\.9%6\.9\\%atN=128N=128\), and it decreases monotonically withNNunder the correlated SV channel \(from37\.4%37\.4\\%to20\.0%20\.0\\%\)\.
Fig\. 9:Gain of the PSO warm start versusNNunder i\.i\.d\. Rayleigh and Saleh\-Valenzuela channels\.
### V\-GRuntime and Complexity
Fig\.[10](https://arxiv.org/html/2609.29150#S5.F10)compares the average runtime of the proposed algorithm and the baselines\. The runtime of the proposed PSA\-GML algorithm is dominated by theS=300S=300refinement steps of the learned optimizer \(about0\.90\.9s per realization atN=32N=32,K=4K=4\) and exceeds that of the AO baseline \(about0\.40\.4s\) in the current Python/Torch prototype, whose unrolled LSTM update carries a non\-negligible per\-step overhead\. Nevertheless, as shown in Table[II](https://arxiv.org/html/2609.29150#S5.T2), the per\-iteration complexity of the proposed algorithm and that of the AO baseline are of the same order𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\), and the larger observed runtime is a constant\-factor effect of the unrolled LSTM in the current Python/Torch prototype: theSSrefinement steps each carry a per\-step overhead𝒪\(DH2\)\\mathcal\{O\}\(DH^\{2\}\)that is not yet optimized in the prototype, and this overhead is amenable to batching and hardware acceleration\. The offline training of the meta\-optimizer overNcN\_\{c\}channel realizations adds a one\-time cost, while the online inference only involves the PSO warm start andSSrefinement steps\.
Fig\. 10:Average runtime comparison of the proposed algorithm and the baselines\.TABLE II:Computational Complexity ComparisonAOProposed PSA\-GMLPer\-Iteration Complexity𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)𝒪\(NNtK\+K3\)\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)Refinement Per\-Step Overhead–𝒪\(DH2\)\\mathcal\{O\}\(DH^\{2\}\)Offline Training–Nc\[PIPSO𝒪\(NNtK\+K3\)\+S\(𝒪\(NNtK\+K3\)\+𝒪\(DH2\)\)\]N\_\{c\}\\left\[PI\_\{\\mathrm\{PSO\}\}\\,\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)\+S\\left\(\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)\+\\mathcal\{O\}\(DH^\{2\}\)\\right\)\\right\]Online Inference: PSO stage–PIPSO𝒪\(NNtK\+K3\)PI\_\{\\mathrm\{PSO\}\}\\,\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)Online Inference: refinement / AO iterationsNAO𝒪\(NNtK\+K3\)N\_\{\\mathrm\{AO\}\}\\,\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)S\[𝒪\(NNtK\+K3\)\+𝒪\(DH2\)\]S\\left\[\\mathcal\{O\}\(NN\_\{t\}K\+K^\{3\}\)\+\\mathcal\{O\}\(DH^\{2\}\)\\right\]
### V\-HOptimized Coefficients
Fig\.[11](https://arxiv.org/html/2609.29150#S5.F11)visualizes the optimized amplitude\-split and phase\-shift coefficients obtained by the proposed algorithm for one channel realization\. It is observed that the transmission and reflection powersβt,n2\\beta\_\{t,n\}^\{2\}andβr,n2\\beta\_\{r,n\}^\{2\}are adaptively balanced across elements, and the phase shifts are distributed over\[0,2π\)\[0,2\\pi\)\.
Fig\. 11:Optimized amplitude\-split and phase\-shift coefficients of the STAR\-RIS\.
## VIConclusion
This paper addressed WSR maximization for a STAR\-RIS aided multi\-user downlink and introduced the particle\-swarm\-assisted gradient meta\-learning \(PSA\-GML\) algorithm, which combines a PSO global warm start over the STAR\-RIS coefficients with a coordinate\-wise LSTM meta\-optimizer trained by first\-order gradient meta\-learning to jointly refine the coefficients and the transmit precoder\. The experiments show that PSA\-GML reached an 11\.06 bits/s/Hz WSR, delivering a 13\.1% performance enhancement over the benchmark AO method \(6\.2% over a multiple\-random\-restart AO\) and a 35\.1% enhancement over the random\-phase scheme, with robustness against initialization and a per\-iteration complexity of the same order as the AO baseline\. Notably, the benefit of the PSO warm start reached 20% or more under correlated Saleh\-Valenzuela channels, and the learned optimizer reached 83\.9% of the hand\-designed Adam refinement in the interference\-limited regime\.
Future extensions include imperfect channel state information and non\-line\-of\-sight propagation to further stress\-test the algorithmic robustness, meta\-learning of the PSO hyper\-parameters to speed up convergence, and generalization to multi\-antenna users and multiple STAR\-RISs\.
## References
- \[1\]E\. G\. Larsson, O\. Edfors, F\. Tufvesson, and T\. L\. Marzetta, “Massive MIMO for next generation wireless systems,”*IEEE Commun\. Mag\.*, vol\. 52, no\. 2, pp\. 186–195, Feb\. 2014\.
- \[2\]M\. Shafi*et al\.*, “5G: A tutorial overview of standards, trials, challenges, deployment, and practice,”*IEEE J\. Sel\. Areas Commun\.*, vol\. 35, no\. 6, pp\. 1201–1221, Jun\. 2017\.
- \[3\]C\.\-X\. Wang*et al\.*, “On the road to 6G: Visions, requirements, key technologies, and testbeds,”*IEEE Commun\. Surveys Tuts\.*, vol\. 25, no\. 2, pp\. 905–974, 2023\.
- \[4\]T\. L\. Marzetta, “Noncooperative cellular wireless with unlimited numbers of base station antennas,”*IEEE Trans\. Wireless Commun\.*, vol\. 9, no\. 11, pp\. 3590–3600, Nov\. 2010\.
- \[5\]Q\. Wu and R\. Zhang, “Intelligent reflecting surface enhanced wireless network via joint active and passive beamforming,”*IEEE Trans\. Wireless Commun\.*, vol\. 18, no\. 11, pp\. 5394–5409, Nov\. 2019\.
- \[6\]E\. Basar, M\. Di Renzo, J\. de Rosny, M\. Debbah, M\.\-S\. Alouini, and R\. Zhang, “Wireless communications through reconfigurable intelligent surfaces,”*IEEE Access*, vol\. 7, pp\. 116753–116773, 2019\.
- \[7\]Ö\. Özdogan, E\. Björnson, and E\. G\. Larsson, “Intelligent reflecting surfaces: Physics, propagation, and pathloss modeling,”*IEEE Wireless Commun\. Lett\.*, vol\. 9, no\. 5, pp\. 581–585, May 2020\.
- \[8\]Y\. Liu*et al\.*, “STAR: Simultaneous transmission and reflection for 360∘coverage by intelligent surfaces,”*IEEE Wireless Commun\.*, vol\. 28, no\. 6, pp\. 102–109, Dec\. 2021\.
- \[9\]H\. Guo, Y\.\-C\. Liang, J\. Chen, and E\. G\. Larsson, “Weighted sum\-rate maximization for reconfigurable intelligent surface aided wireless networks,”*IEEE Trans\. Wireless Commun\.*, vol\. 19, no\. 5, pp\. 3064–3076, May 2020\.
- \[10\]K\. Shen and W\. Yu, “Fractional programming for communication systems—Part I: Power control and beamforming,”*IEEE Trans\. Signal Process\.*, vol\. 66, no\. 10, pp\. 2616–2630, May 2018\.
- \[11\]Q\. Shi, M\. Razaviyayn, Z\.\-Q\. Luo, and C\. He, “An iteratively weighted MMSE approach to distributed sum\-utility maximization for a MIMO interfering broadcast channel,”*IEEE Trans\. Signal Process\.*, vol\. 59, no\. 9, pp\. 4331–4340, Sep\. 2011\.
- \[12\]X\. Mu, Y\. Liu, L\. Guo, J\. Lin, and R\. Schober, “Simultaneously transmitting and reflecting \(STAR\) RIS aided wireless communications,”*IEEE Trans\. Wireless Commun\.*, vol\. 21, no\. 5, pp\. 3083–3098, May 2022\.
- \[13\]J\. Xu, Y\. Liu, X\. Mu, and O\. A\. Dobre, “STAR\-RISs: Simultaneous transmitting and reflecting reconfigurable intelligent surfaces,”*IEEE Commun\. Lett\.*, vol\. 25, no\. 9, pp\. 3134–3138, Sep\. 2021\.
- \[14\]J\. Zuo, Y\. Liu, Z\. Ding, L\. Song, and H\. V\. Poor, “Joint design for simultaneously transmitting and reflecting \(STAR\) RIS assisted NOMA systems,”*IEEE Trans\. Wireless Commun\.*, vol\. 22, no\. 1, pp\. 611–626, Jan\. 2023\.
- \[15\]C\. Wu, Y\. Liu, X\. Mu, X\. Gu, and O\. A\. Dobre, “Coverage characterization of STAR\-RIS networks: NOMA and OMA,”*IEEE Commun\. Lett\.*, vol\. 25, no\. 9, pp\. 3036–3040, Sep\. 2021\.
- \[16\]J\. Kennedy and R\. Eberhart, “Particle swarm optimization,” in*Proc\. IEEE Int\. Conf\. Neural Networks \(ICNN\)*, 1995, pp\. 1942–1948\.
- \[17\]J\. C\. Bansal, P\. K\. Singh, M\. Saraswat, A\. Verma, S\. S\. Jadon, and A\. Abraham, “Inertia weight strategies in particle swarm optimization,” in*Proc\. World Congr\. Nature Biologically Inspired Comput\.*, 2011, pp\. 633–640\.
- \[18\]K\. Gregor and Y\. LeCun, “Learning fast approximations of sparse coding,” in*Proc\. Int\. Conf\. Mach\. Learn\. \(ICML\)*, 2010, pp\. 399–406\.
- \[19\]V\. Monga, Y\. Li, and Y\. C\. Eldar, “Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,”*IEEE Signal Process\. Mag\.*, vol\. 38, no\. 2, pp\. 18–44, Mar\. 2021\.
- \[20\]M\. Andrychowicz, M\. Denil, S\. Gómez Colmenarejo, M\. W\. Hoffman, D\. Pfau, T\. Schaul, B\. Shillingford, and N\. de Freitas, “Learning to learn by gradient descent by gradient descent,” in*Proc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\)*, 2016, pp\. 3981–3989\.
- \[21\]T\. Chen, X\. Chen, W\. Chen, H\. Heaton, J\. Liu, Z\. Wang, and W\. Yin, “Learning to optimize: A primer and a benchmark,”*J\. Mach\. Learn\. Res\.*, vol\. 23, pp\. 1–59, 2022\.
- \[22\]A\. Zappone, M\. Di Renzo, and M\. Debbah, “Wireless networks design in the era of deep learning: Model\-based, AI\-based, or both?”*IEEE Trans\. Commun\.*, vol\. 67, no\. 10, pp\. 7331–7376, Oct\. 2019\.
- \[23\]C\. Finn, P\. Abbeel, and S\. Levine, “Model\-agnostic meta\-learning for fast adaptation of deep networks,” in*Proc\. Int\. Conf\. Mach\. Learn\. \(ICML\)*, 2017, pp\. 1126–1135\.
- \[24\]T\. Hospedales, A\. Antoniou, P\. Micaelli, and A\. Storkey, “Meta\-learning in neural networks: A survey,”*IEEE Trans\. Pattern Anal\. Mach\. Intell\.*, vol\. 44, no\. 9, pp\. 5149–5169, Sep\. 2022\.
- \[25\]K\. Zhou, W\. Zhou, D\. Cai, X\. Lei, Y\. Xu, Z\. Ding, and P\. Fan, “A gradient meta\-learning joint optimization for beamforming and antenna position in pinching\-antenna systems,”*IEEE Trans\. Commun\.*, accepted\.
- \[26\]A\. A\. M\. Saleh and R\. Valenzuela, “A statistical model for indoor multipath propagation,”*IEEE J\. Sel\. Areas Commun\.*, vol\. 5, no\. 2, pp\. 128–137, Feb\. 1987\.
- \[27\]C\. Huang, A\. Zappone, G\. C\. Alexandropoulos, M\. Debbah, and C\. Yuen, “Reconfigurable intelligent surfaces for energy efficiency in wireless communication,”*IEEE Trans\. Wireless Commun\.*, vol\. 18, no\. 8, pp\. 4157–4170, Aug\. 2019\.
- \[28\]Y\. Liu, X\. Liu, X\. Mu, T\. Hou, J\. Xu, M\. Di Renzo, and N\. Al\-Dhahir, “Reconfigurable intelligent surfaces: Principles and opportunities,”*IEEE Commun\. Surveys Tuts\.*, vol\. 23, no\. 3, pp\. 1546–1577, 2021\.
- \[29\]Q\. Wu and R\. Zhang, “Beamforming optimization for wireless network aided by intelligent reflecting surface with discrete phase shifts,”*IEEE Trans\. Commun\.*, vol\. 68, no\. 3, pp\. 1838–1851, Mar\. 2020\.相似文章
结合梯度下降的自适应混合粒子群优化
本文提出自适应混合粒子群优化(AHPSO),利用群体多样性的sigmoid函数在搜索过程中自动调节梯度影响。结果表明,在特定问题类别上,它优于标准粒子群优化,并能与CMA-ES相媲美,但优势并非普遍适用。
多目标优化中梯度聚合的统一框架
本文提出了一个多目标优化中梯度聚合的统一理论框架,建立了收敛到帕累托平稳性的速率。作者引入了一个充分对齐条件,并展示了其在现有算法和新算法(如 capped MGDA)中的应用。
单一学习率不够:LoRA微调的自适应各向异性学习率
本文提出了一种用于LoRA微调的自适应各向异性学习率模型,以解决模块内异质性问题,提升性能并在各基准测试中提高秩容量利用率。
用于深度强化学习与模型预测控制之间共享控制权限的Composite-Gradient Learning
本文提出了一种复合梯度学习方法,该方法整合了深度强化学习和模型预测控制,用于自主系统中的共享控制权限,并在交通网络上进行评估,显示在强交互下有适度效益。
联合仿射谱整形:超越仅权重Muon的权重与偏置更新耦合
本文提出联合仿射谱整形(JRI),将仅权重的谱优化器(如Muon)扩展到仿射层中联合更新权重和偏置,在BERT-mini IMDb分类任务上显示出小而一致的准确率提升。