Neural operator learning for collision-aware trajectory planning of spacecraft swarms

arXiv cs.LG Papers

Summary

This paper introduces a permutation-equivariant neural operator that maps spacecraft, target, and debris distributions to collision-aware trajectories for entire swarms in one forward pass, trained with self-supervised physics objectives and adversarial threats. It generalizes zero-shot from 10 to 1,000 spacecraft amid dense debris, matching optimal-control accuracy while reducing proximity.

arXiv:2608.00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and learning-based planners rarely transfer across swarm sizes or debris densities. Here we introduce a permutation-equivariant neural operator that maps distributions of spacecraft, targets and debris to collision-aware trajectories for an entire swarm in a single forward pass, paired with a batched Gauss-Newton finish that enforces exact orbital dynamics. The operator is trained without optimal-trajectory labels, combining self-supervised physics objectives with adversarial threats generated against its own rollouts. Trained on ten spacecraft, it generalizes zero-shot to swarms of 1,000 amid more than 11,000 catalogued objects, matching a per-agent optimal-control solver's accuracy, evading worst-case threats that a debris-blind baseline cannot, and reducing proximity within the swarm several-fold. Physics-grounded operator learning thus offers a fast, scalable alternative to optimal control for crowded orbits.
Original Article
View Cached Full Text

Cached at: 08/04/26, 07:40 AM

# Neural operator learning for collision-aware trajectory planning of spacecraft swarms
Source: [https://arxiv.org/html/2608.00320](https://arxiv.org/html/2608.00320)
\[1,2\]\\fnmSidhdharth D\.\\surSikka

\[1\]\\orgdivSchool of Aeronautics and Astronautics,\\orgnamePurdue University,\\orgaddress\\cityWest Lafayette,\\stateIN,\\countryUSA

2\]\\orgnameManifold Research Group,\\orgaddress\\citySan Francisco,\\stateCA,\\countryUSA

3\]\\orgdivDepartment of Mathematics,\\orgnamePurdue University,\\orgaddress\\cityWest Lafayette,\\stateIN,\\countryUSA

4\]\\orgnameIndependent Researcher,\\orgaddress\\cityFremont,\\stateCA,\\countryUSA

###### Abstract

Autonomous spacecraft swarms must plan fuel\-efficient, collision\-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and learning\-based planners rarely transfer across swarm sizes or debris densities\. Here we introduce a permutation\-equivariant neural operator that maps distributions of spacecraft, targets and debris to collision\-aware trajectories for an entire swarm in a single forward pass, paired with a batched Gauss–Newton finish that enforces exact orbital dynamics\. The operator is trained without optimal\-trajectory labels, combining self\-supervised physics objectives with adversarial threats generated against its own rollouts\. Trained on ten spacecraft, it generalizes zero\-shot to swarms of 1,000 amid more than 11,000 catalogued objects, matching a per\-agent optimal\-control solver’s accuracy, evading worst\-case threats that a debris\-blind baseline cannot, and reducing proximity within the swarm several\-fold\. Physics\-grounded operator learning thus offers a fast, scalable alternative to optimal control for crowded orbits\.

###### keywords:

Neural operators, Operator learning, Distributional control, Multi\-agent systems, Collision\-aware planning, Spacecraft swarms, Orbital autonomy

Low Earth Orbit \(LEO\) is becoming increasingly crowded as satellite constellations expand and debris accumulates\. Roughly 16,000 active satellites and an estimated 140 million debris fragments now occupy near\-Earth space\[[1](https://arxiv.org/html/2608.00320#bib.bib1)\]\. Between December 2025 and May 2026 alone, SpaceX’s Starlink constellation executed more than 207,000 automated collision\-avoidance maneuvers, over three times its rate a year earlier, averaging more than 40 maneuvers per satellite per year, nearly one per week\[[2](https://arxiv.org/html/2608.00320#bib.bib2)\]\. As orbital congestion increases, trajectory planning becomes a persistent, large\-scale coordination problem rather than an occasional corrective maneuver\. Frequent avoidance maneuvers reduce mission lifetime, consume fuel, and create coordination burdens that grow with constellation size\.

Classical approaches to multi\-agent orbital maneuvering rely primarily on centralized optimization: mixed\-integer linear programs offer safety guarantees\[[3](https://arxiv.org/html/2608.00320#bib.bib3)\], linearized and distributed convex optimization improves tractability\[[4](https://arxiv.org/html/2608.00320#bib.bib4)\], and model predictive and receding\-horizon schemes handle constraints in closed loop\[[5](https://arxiv.org/html/2608.00320#bib.bib5),[6](https://arxiv.org/html/2608.00320#bib.bib6)\], but all must re\-solve programs whose collision constraints multiply with the number of agents and debris objects\. Reactive schemes such as velocity obstacles\[[7](https://arxiv.org/html/2608.00320#bib.bib7)\], artificial potential fields\[[8](https://arxiv.org/html/2608.00320#bib.bib8)\], and deployed rule\-based screening avoid this cost but grow conservative and prone to mutual conflict in dense traffic\[[2](https://arxiv.org/html/2608.00320#bib.bib2)\], while heuristic\[[9](https://arxiv.org/html/2608.00320#bib.bib9)\], reinforcement\-learning\[[10](https://arxiv.org/html/2608.00320#bib.bib10)\], and decision\-theoretic\[[11](https://arxiv.org/html/2608.00320#bib.bib11)\]planners, including learned swarm\-navigation policies\[[12](https://arxiv.org/html/2608.00320#bib.bib12)\], are flexible but seldom transfer across swarm sizes and debris densities\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/summary_handmade.png)Figure 1:Neural operator framework for collision\-aware swarm planning\.The spacecraft swarm and debris field are represented as distributions over orbital elements\. Separate attention pathways encode target alignment and debris avoidance, while a physics\-informed baseline \(green dashed\) plus learned residual \(blue solid\) generates dynamically grounded, collision\-aware trajectories for the full swarm\. As a final step, the resulting debris\-avoiding trajectories are refined by a batched Gauss–Newton finish that closes each one onto exact two\-body dynamics\.A key difficulty is structural\.Collision constraints couple each spacecraft to other spacecraft and debris objects through pairwise interactions, causing the number of safety relationships to grow rapidly with swarm size and debris density\. Rather than treating a swarm as a collection of individual spacecraft, one can instead model it as a probability distribution evolving under controlled dynamics\. Mean Field Games \(MFGs\) formalize this perspective by characterizing the limit of infinitely many interacting agents through coupled Hamilton–Jacobi–Bellman and Fokker–Planck equations\[[13](https://arxiv.org/html/2608.00320#bib.bib13)\]\. While this representation scales more naturally with population size and has shown promise in large\-scale robotic systems\[[14](https://arxiv.org/html/2608.00320#bib.bib14)\], solving the associated partial differential equations remains computationally demanding in high\-dimensional physical domains, even with dedicated machine\-learning solvers\[[15](https://arxiv.org/html/2608.00320#bib.bib15),[16](https://arxiv.org/html/2608.00320#bib.bib16),[17](https://arxiv.org/html/2608.00320#bib.bib17)\]\.

Learning offers a way to amortize this cost without requiring labeled optimal solutions: physics\-informed networks embed governing equations directly in the training objective\[[18](https://arxiv.org/html/2608.00320#bib.bib18),[19](https://arxiv.org/html/2608.00320#bib.bib19)\], and differentiable simulation trains control policies end\-to\-end through the system dynamics\[[20](https://arxiv.org/html/2608.00320#bib.bib20)\]\. Operator learning extends this paradigm from individual problem instances to families of them: neural operators learn mappings between functional inputs and outputs, enabling amortized inference across problem instances\[[21](https://arxiv.org/html/2608.00320#bib.bib21),[22](https://arxiv.org/html/2608.00320#bib.bib22),[23](https://arxiv.org/html/2608.00320#bib.bib23),[24](https://arxiv.org/html/2608.00320#bib.bib24),[25](https://arxiv.org/html/2608.00320#bib.bib25),[26](https://arxiv.org/html/2608.00320#bib.bib26)\]\. Distribution\-driven control methods have demonstrated that differentiable particle simulations can be used to train neural operators that map initial distributions to target configurations without explicitly solving MFG equations\[[27](https://arxiv.org/html/2608.00320#bib.bib27),[28](https://arxiv.org/html/2608.00320#bib.bib28)\]\.

Here we introduce a two\-stage planner for collision\-aware trajectory planning of spacecraft swarms in dense debris fields: a self\-supervised, permutation\-equivariant, time\-conditioned neural operator, followed by a lightweight per\-agent Gauss–Newton finish that closes each predicted trajectory onto exact two\-body dynamics\. The operator learns distribution\-to\-trajectory maps for spacecraft evolving under Keplerian orbital dynamics, conditions jointly on initial orbital distributions, target configurations, debris fields, and mission duration, and returns trajectories for the entire swarm through a single forward inference step\. Unlike supervised operator\-learning approaches that require precomputed optimal trajectories or numerical solution data, the operator is trained directly from physics\-informed objectives, including terminal accuracy, fuel\-effort surrogates, and closest\-point\-of\-approach penalties, with no trajectory\-optimization labels\. To expose the model to informative close\-approach regimes, we also generate adversarial debris by placing passive objects along nominal debris\-unaware rollouts, creating training and evaluation cases in which collision avoidance requires conditioning on the debris field\.

This distributional viewpoint also explains why the method can generalize across swarm sizes: the operator acts on empirical samples from the underlying spacecraft, target, and debris distributions, applying the same permutation\-equivariant interaction rule across arbitrary numbers of sample points\. At deployment, collision\-aware planning is reduced to a single forward inference pass over theNNspacecraft andMMdebris objects, followed by a Gauss–Newton finish that is batched across the swarm; together they avoid the repeated solution of large nonlinear or mixed\-integer trajectory\-optimization programs as the swarm and debris field change\. Our method is summarized in Fig\.[1](https://arxiv.org/html/2608.00320#S0.F1)and in Supplementary Video 1\.

We evaluate the method on real ephemeris data, spanning both ambient debris fields and adversarial close\-approach scenarios\. The planner produces low\-cost, collision\-aware, dynamically consistent trajectories and extrapolates stably to swarms ofN=1000N=1000, providing a fast inference\-time alternative to conventional optimal control in crowded orbital environments\.

## Results

We evaluate whether the learned neural operator can \(i\) generate low\-cost orbital transfers, \(ii\) avoid close approaches in dense debris fields, and \(iii\) generalize beyond the training distribution in both swarm size and debris density\. The operator is trained on short\-duration missions of 1–12 hours, with agent countsN∈\{1,…,10\}N\\in\\\{1,\\dots,10\\\}and one adversarially placed debris object per agent\. We report both interpolation performance within this regime and extrapolation to swarms of up toN=1000N=1000agents amid the full\>11,000\{\>\}11\{,\}000\-object catalog\.

We evaluate the planner over a2×22\\times 2family of test conditions in which every scenario draws itsNNspacecraft initial states from the real Two\-Line Element \(TLE\) catalog, so no reported result depends on synthetic starting geometry\. The*maneuver*axis sets the retargeting magnitude: a*minor*maneuver perturbs each orbital element by exactly1%1\\%of its scale \(random sign per element\) and represents station\-keeping or fine retargeting, while a*major*maneuver perturbs each element by exactly10%10\\%and represents rapid\-response or defensive retargeting; fixing the maneuver magnitude makes fuel costs directly comparable across trials within a class\. The*debris*axis sets the threat construction: a*debris*scenario surrounds the swarm with ambient catalog objects, the routine collision\-avoidance regime, whereas an*adversarial*scenario places one worst\-case object on each method’s own debris\-unaware predicted path, so a planner that does not condition on the debris field is struck unless it actively deviates\. These two axes span the method’s intended dual use: passive debris avoidance under realistic catalog density, and worst\-case defensive pursuit\-evasion against a threat that targets the nominal trajectory\. Proximity is reported as a per\-spacecraft rate, the percentage of planned maneuvers that pass within100100m of another agent or debris object, and all reported performance values are medians over500500Monte Carlo trials per cell\.

### Adversarial and debris\-field trajectory planning

Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)reports terminal accuracy, fuel cost, and per\-spacecraft proximity across the four scenarios and swarm sizes\. We compare the operator\-warm Gauss–Newton finish \(GNw\), which closes each operator rollout onto exact two\-body dynamics, with the same finish cold\-started without the operator seed \(GNc\)\. GNc reaches the same target orbit but lacks the operator’s learned collision avoidance, so the GNw–GNc gap isolates what the operator contributes\. Both run batched across the swarm at every size, includingN=1000N=1000, where a per\-agent nonlinear program is intractable\.

The debris scenarios are the realistic operating regime: real spacecraft initial conditions together with the actual catalogued debris field, the setting a swarm would face in today’s crowded low Earth orbit\. Here the operator\-warm finish keeps close approaches rare, at or below0\.28%0\.28\\%of maneuvers within100100m, while the debris\-blind GNc approaches within100100m on22–3\.5%3\.5\\%of maneuvers, an order of magnitude more, at a fuel cost within ten percent of GNw’s\. This behavior holds far beyond theN≤10N\\leq 10training range, degrading gradually rather than abruptly out toN=1000N=1000controlled spacecraft planned amid the full\>11,000\{\>\}11\{,\}000\-object catalog, with terminal error held at10−410^\{\-4\}–10−2%10^\{\-2\}\\%throughout\.

The same learned avoidance extends to an adversarial setting\. An adversarial object is one positioned directly on a spacecraft’s intended path, so that failing to deviate means near\-certain collision, whether a fragment on an unusually high\-conjunction trajectory or a hostile satellite maneuvering to intercept in a contested scenario\. We construct this worst case by seeding one such object on each method’s own debris\-unaware path\. A planner blind to it is struck almost every time \(GNc,99\.399\.3–99\.8%99\.8\\%of maneuvers at all sizes\), whereas the operator\-warm finish clears it on essentially every maneuver \(at most0\.21%0\.21\\%atN=1000N=1000, essentially none of which involves the threat itself\) at comparable fuel and accuracy\. Nominal debris avoidance and defensive evasion are therefore one capability, driven by the same conditioning on the surrounding object field and exercised here against a deliberately harder threat\.

Table 1:Terminal accuracy, fuel cost, and collision behavior across debris and adversarial environments
In the single\-agent setting \(N=1N=1\) we additionally solve the full optimal\-control problem with IPOPT\[[29](https://arxiv.org/html/2608.00320#bib.bib29)\]as a reference: a nonlinear program over the controlled two\-body dynamics that minimizes fuel while driving the agent to its target orbit and enforcing hard debris\-avoidance constraints\. It is tractable only when both the agent and debris counts are small, so we report it atN=1N=1in the adversarial scenarios, where each agent faces a single worst\-case object; against the full catalog it is intractable even atN=1N=1, as it imposes a separate collision constraint at every debris object and time step\. On the real single\-agent transfers it attains0\.0012%0\.0012\\%terminal error atΔ​v=0\.582\\Delta v=0\.582km/s for the minor maneuver and0\.045%0\.045\\%at5\.525\.52km/s for the major maneuver\. The operator\-warm finish GNw approaches this optimum, reaching comparable terminal accuracy \(Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\) at a fuel cost within2%2\\%on the minor maneuver and within11%11\\%on the major maneuver, while remaining batched and scalable to the swarm sizes where IPOPT cannot run\.

IPOPT serves as a quality benchmark rather than a scalable baseline: a per\-agent nonlinear program becomes prohibitively expensive asNNgrows, whereas the operator is evaluated once per scenario for the entire swarm\. This asymmetry is reflected in Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1): although the operator is trained only onN≤10N\\leq 10, the finished trajectories retain bounded terminal errors and low proximity\-violation rates atN=100N=100andN=1000N=1000, where direct nonlinear optimization is no longer practical as a routine online planner\.

Figure[2](https://arxiv.org/html/2608.00320#Sx1.F2)compares learned single\-agent transfers with nonlinear optimal\-control solutions and illustrates the effect of debris conditioning in an adversarial setting\. The learned trajectories recover the optimized transfer geometry for both minor and major maneuvers\. When adversarial debris is provided to the operator, the predicted trajectory changes to preserve separation above the100100m threshold; when the same debris information is omitted, multiple close approaches are predicted\. The debris input thus directly shapes the trajectory toward avoidance\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/short_minor_traj.png)

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/short_major_traj.png)

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/dist_vs_time_nominal_pd0_color.png)

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/dist_vs_time_avoid_pdadv_color.png)

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/reltrajlog_nominal_pd0_color.png)

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/reltrajlog_avoid_pdadv_color.png)

Figure 2:Learned transfer geometry and debris\-conditioned avoidance\.Top:Learned single\-agent transfers closely match nonlinear optimal\-control solutions for minor and major maneuvers\.Middle:Under identical initial conditions, omitting debris information produces predicted close approaches, whereas debris conditioning keeps separation above the100100m safety threshold\.Bottom:Log\-scaled relative\-trajectory spheres show the same effect geometrically: the nominal rollout intersects the debris\-centered danger region, while debris conditioning reshapes the trajectory to clear it\.Although the operator is trained only onN≤10N\\leq 10, the finished terminal error remains tightly bounded when extrapolating toN=100N=100andN=1000N=1000across all four scenarios \(Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\), indicating stable degradation rather than abrupt failure far beyond the training distribution\.

All scenarios draw their initial conditions, and the debris scenarios their debris fields, from a publicly available Two\-Line Element \(TLE\) catalog downloaded from Space\-Track \([https://www\.space\-track\.org](https://www.space-track.org/)\) on March 24, 2025, which contains over11,00011\{,\}000resident space objects\. Each TLE is propagated to its epoch using the SGP4 orbit propagation model\[[30](https://arxiv.org/html/2608.00320#bib.bib30)\]and converted to Cartesian state vectors\. The finished median terminal error stays in the10−310^\{\-3\}–10−2%10^\{\-2\}\\%range across swarm sizes spanning three orders of magnitude, and theN=1000N=1000debris case plans1,0001\{,\}000controlled spacecraft amid the full catalog of more than11,00011\{,\}000objects, far outside the training distribution; collision behavior is reported in Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\.

### Dynamic feasibility via a Gauss–Newton finish

The operator outputs element\-space trajectories that are not, by construction, hard\-constrained to be the integral of a physical control sequence under exact two\-body dynamics\. We close each rollout with a per\-agent Gauss–Newton finish: a single\-shooting step over the control sequence, so every iterate is dynamically exact by construction, whose residual matches the orbit’s conserved angular\-momentum and eccentricity vectors, a phase\-free, non\-singular terminal target, under a fuel regularizer\. Warm\-started from the operator \(GNw\) the finish inherits the operator’s collision\-aware geometry; cold\-started \(GNc\) it reaches the same target orbit without that geometry\. Because the per\-agent solve reduces to a small fixed\-size linear system, the finish is batched across the swarm and runs atN=1000N=1000, where a per\-agent nonlinear program is infeasible; its wall\-clock cost is set by the trajectory length rather than by the swarm or debris count, so it stays near\-constant as the swarm grows \([Runtime and scaling](https://arxiv.org/html/2608.00320#Sx1.SSx5)\)\. The full construction, including the conserved\-vector residual and its batched solution, is derived in[Methods](https://arxiv.org/html/2608.00320#Sx3)\.

GNw drives terminal error to10−310^\{\-3\}–10−2%10^\{\-2\}\\%\(Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\), one to two orders below the operator’s raw element\-space output and comparable to the single\-agent optimum, at a fuel cost close to that single\-agent optimum \(Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\), and the accuracy is bounded across three orders of magnitude inNN\.

### Collision avoidance

Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)reports medians, in this section we report collision behavior with greater granularity for three quantities across the debris and adversarial scenarios: the operator’s raw output \(ML\), the operator\-warm finish \(GNw\), and the debris\-blind cold finish \(GNc\)\. The metric is the per\-spacecraft rate at100100m\. The contrast is decisive and holds at every swarm size\. Against the adversarial threat the debris\-blind GNc is struck on almost every maneuver \(99\.399\.3–99\.8%99\.8\\%\), whereas the operator clears it essentially always \(at most0\.21%0\.21\\%\); under ambient catalog debris the same ordering holds at lower levels, with GNc approaching a debris object an order of magnitude more often than GNw\. Throughout, GNw tracks ML to within a few tenths of a percent: the dynamics\-closing finish preserves the operator’s learned avoidance rather than eroding it, and the large GNw–GNc gap is precisely the avoidance value the operator supplies through the warm start\.

Splitting the residual by pair type shows that avoidance of*external*objects is effectively complete\. In the adversarial scenarios GNw’s agent–debris rate is at most0\.002%0\.002\\%at every swarm size \(1010events across500,000500\{,\}000maneuvers in the worst cell\), and under catalog debris it stays at or below0\.06%0\.06\\%at100100m and0\.02%0\.02\\%at the tighter5050m radius, against3\.2%3\.2\\%\(at100100m\) and12%12\\%\(at a conservative500500m screen\) for GNc\. What remains for ML and GNw is almost entirely*agent–agent*, appearing only once the swarm is dense \(N≥100N\\geq 100\) and reflecting the swarm’s own internal density\. This component is directly shaped by training: the scenario construction deliberately forces sustained agent–agent conflict \(clustered starts and converging targets, Methods\) under a dedicated agent–agent penalty, and the resulting operator reduces swarm\-internal proximity three\- to eight\-fold relative to the avoidance\-free GNc \(0\.070\.07–0\.22%0\.22\\%versus0\.510\.51–0\.67%0\.67\\%atN=1000N=1000across the four scenarios\)\. Further improving agent–agent deconfliction, through both training design and architecture design, is a direction for future work\. \(A per\-*trial*rate, which flags an entire trial if any single agent comes close, is naturally higher and grows with swarm size, so we report the per\-spacecraft rate as the true per\-vehicle risk\.\)

Figure[3](https://arxiv.org/html/2608.00320#Sx1.F3)shows the distribution of each spacecraft’s closest approach for all three methods: under the adversarial scenarios GNc collapses onto the threat \(∼10\{\\sim\}10m\) while the bulk of the ML and GNw mass sits kilometers above the100100m threshold, and for the operator methods the only sub\-threshold mass is a thin swarm\-internal tail in the densest cells, consistent with the sub\-0\.3%0\.3\\%per\-spacecraft rates in Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/minsep_violins.png)Figure 3:Distribution of each spacecraft’s closest approach \(500500Monte Carlo trials per cell; one sample per spacecraft per trial\)\.Each cell shows three violins per swarm size, for the operator output \(ML, green\), the operator\-warm finish \(GNw, red\), and the cold debris\-blind finish \(GNc, blue\); black bars mark medians and the dotted and dashed lines mark the100100m and500500m thresholds\. GNc sits lowest in every cell and collapses onto the threat under the adversarial scenarios, whereas the ML and GNw mass sits well above the threshold at every swarm size; their only sub\-threshold mass is a thin swarm\-internal \(agent–agent\) tail in the densest cells, matching the per\-spacecraft proximity rates in Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\.
### Grid\-independent proximity scoring

Every separation reported in this paper is measured with a closest\-point\-of\-approach \(CPA\) refinement of the sampled trajectory defined in equation \([7](https://arxiv.org/html/2608.00320#Sx3.E7)\): within each rollout interval we solve analytically for the instant of minimum separation, rather than reading off the smallest distance at the grid points\. This matters because two objects usually pass closest*between*samples, so a grid\-only scorer both overstates how far apart they get and under\-counts the close approaches that fall between samples, by an amount that depends on how finely the trajectory happens to be discretized\.

Figure[4](https://arxiv.org/html/2608.00320#Sx1.F4)makes the effect concrete\. Re\-scoring the densest \(N=1000N=1000\) cells while sweeping the grid from coarse to fine, the grid\-sampled minimum overstates the true clearance by roughly1\.51\.5–2×2\\timesand drifts with resolution, so its apparent safety margin is largely an artifact of the step count\. The CPA measure, by contrast, is essentially flat across the sweep, recovering the same physical closest\-approach distance at every resolution\. The proximity rates reported throughout are therefore both faithful, reflecting the distances objects actually reach, and grid\-independent, so they cannot be made to look safer by integrating on a coarser grid\. This behavior is representative: the same flatness holds across all four scenarios and every swarm size\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/cpa_ablation.png)Figure 4:CPA scoring is faithful and grid\-independent\.Median minimum separation for the densest debris cell \(N=1000N=1000, minor maneuver\) as the scoring grid is swept from coarse to fine \(integration stepsKK; log scale\)\. The grid\-sampled minimum \(orange, dashed\) overstates the true clearance and drifts with resolution, only approaching the correct value as the grid is refined, whereas the within\-interval CPA measure \(blue, solid\) used throughout is flat across the sweep, recovering the same physical closest\-approach distance at every resolution\. The shaded band is clearance that the grid scorer reports but that does not physically exist; the vertical line marks the grid used elsewhere in the paper\. The same behavior holds across all scenarios and swarm sizes\.
### Runtime and scaling

We measured wall\-clock time for both stages on the same workstation used for the Monte Carlo evaluations \(a single NVIDIA RTX 2080 Ti, 11 GB\)\. Both the operator rollout and the Gauss–Newton finish use a constant physical timestepΔ​t=120\\Delta t=120s, so the number of integration steps scales with mission duration \(K≈T/Δ​tK\\approx T/\\Delta t, from≈31\{\\approx\}31for a 1\-hour transfer to≈361\{\\approx\}361for the full 12\-hour horizon\) rather than with swarm size\. Figure[5](https://arxiv.org/html/2608.00320#Sx1.F5)reports the median time per\(N,M\)\(N,M\)cell at a representative 6\-hour horizon \(K≈181K\\approx 181\)\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/Figures/runtime_heatmap.png)Figure 5:Runtime scaling of the two pipeline stages\.Median wall\-clock time \(s\) on a single NVIDIA RTX 2080 Ti as a function of swarm sizeNNand debris countMM\.Left:neural\-operator inference, which grows only mildly \(0\.60\.6–2\.92\.9s\)\.Right:the Gauss–Newton finish, near\-uniform at≈35\\approx\\\!35s regardless ofNNorMM\(all cells span only3434–3636s\), since its cost is set by the per\-iteration rollout length rather than the agent or debris count\. Both panels use a zero\-based color scale, so the finish reads as a near\-flat band\. Neither stage scales combinatorially\.Operator inference stays below 3 seconds throughout, rising only from0\.70\.7s atN=1N=1to2\.92\.9s atN=1000N=1000withM=1000M=1000debris\. The Gauss–Newton finish reduces to a batched, fixed\-size linear solve per agent \([Methods](https://arxiv.org/html/2608.00320#Sx3)\), so its cost is set by the sequential RK4 rollouts in each iteration rather than byNNorMM: it runs in≈35\{\\approx\}35s and is essentially flat in swarm size \(34\.934\.9s atN=1N=1,36\.436\.4s atN=1000N=1000\)\. The finish dominates, so the full pipeline replans the entire swarm in under a minute at the 6\-hour horizon and roughly twice that at 12 hours; neither stage scales combinatorially with agent or debris count\. The single\-agent IPOPT solve, by comparison, costs≈5\.7\{\\approx\}5\.7s per agent and does not amortize across the swarm, serving as a quality reference atN=1N=1rather than a scalable planner\.

### Duration generalization

The operator is trained on mission durations up to 12 hours and is not expected to extrapolate reliably to substantially longer horizons in a single rollout\. The measured inference latency suggests that the operator could be embedded in a closed\-loop receding\-horizon controller, repeatedly querying the model with updated spacecraft and debris states while remaining within the temporal support of the training distribution\. We do not evaluate this mode in the present work\. All reported metrics are from single\-shot rollouts on the trained horizon; long\-horizon receding\-horizon execution is therefore a deployment hypothesis rather than a measured result\.

## Discussion

This work demonstrated that collision\-aware trajectory planning for an entire spacecraft swarm can be amortized into a single forward pass of a permutation\-equivariant neural operator, trained without optimal\-trajectory labels from self\-supervised physics objectives and adversarial threats generated against the model’s own rollouts\. Where classical planners re\-solve a nonlinear program per agent and per scenario, the operator absorbs that cost at training time: trained on ten spacecraft, it transfers zero\-shot to swarms two orders of magnitude larger amid the full catalogued debris field, with rare proximity violations and stable terminal accuracy, suggesting that it captures structure in the underlying dynamics rather than memorizing fixed configurations\.

Two design choices carry the result and are not specific to astrodynamics\. The first is the division of labor between learning and numerics: the operator supplies collision\-aware geometry in one batched inference, and the Gauss–Newton finish closes it onto exact dynamics, so the learned component is never asked to guarantee physics and the numerical component is never asked to see debris\. The second is that the training distribution, as much as the model, determines what is learned: effective avoidance emerged from deliberately constructed scenarios, crossing\-orbit threats with well\-conditioned avoidance gradients and clustered, converging swarm geometries that force sustained agent–agent conflict\.

Several limitations remain\. For very small maneuvers the finishedΔ​v\\Delta vis mildly suboptimal, because the smooth quadratic control surrogate biases the warm start away from the sharp impulse\-like profiles that minimize fuel for near\-zero terminal change, and accuracy degrades outside the trained 12\-hour horizon, consistent with operator\-learning models under distribution shift\. The interaction penalties are soft and the finish carries no collision term, so the method offers no worst\-case collision\-avoidance guarantee\. At the largest swarms the residual proximity is almost entirely agent–agent \(Table[1](https://arxiv.org/html/2608.00320#Sx1.T1.fig1)\); training on conflicting geometries under a dedicated agent–agent penalty reduces it several\-fold relative to the avoidance\-free baseline, and further improvement, through both training design and architecture design, is a direction for future work\. Both training and evaluation assume deterministic two\-body Keplerian motion, withoutJ2J\_\{2\}, atmospheric drag, solar\-radiation pressure, third\-body perturbations, state\-estimation uncertainty or thrust execution error, effects that matter operationally because catalog uncertainty can be comparable to the100100–500500m thresholds considered here\. Finally, extrapolation beyond the training distribution is observed empirically but not theoretically characterized\.

As orbits grow more congested, planning methods whose cost scales with a single batched inference rather than with the number of pairwise constraints will become a prerequisite for swarm autonomy, in routine collision avoidance and in the adversarial regime alike\. The recipe demonstrated here, self\-supervised physics losses, adversarial scenario generation and a certified numerical finish, is not specific to orbital mechanics, and offers a template for scalable, collision\-aware multi\-agent planning wherever swarms must move through contested, cluttered environments\.

## Methods

### Problem setup and notation

We consider a swarm ofNNcontrolled spacecraft operating in the presence ofMMunpowered debris objects over a fixed horizonTT\. Each spacecraft state is represented in Keplerian orbital elements,

𝒙​\(t\)=\[aειΩων\]⊤∈ℝ6,\\boldsymbol\{x\}\(t\)=\\begin\{bmatrix\}a&\\varepsilon&\\iota&\\Omega&\\omega&\\nu\\end\{bmatrix\}^\{\\top\}\\in\\mathbb\{R\}^\{6\},whereaais the semi\-major axis,ε\\varepsilonis the eccentricity,ι\\iotais the inclination,Ω\\Omegais the right ascension of the ascending node,ω\\omegais the argument of periapsis, andν\\nuis the true anomaly\. Each debris object is represented similarly as𝒅​\(t\)∈ℝ6\\boldsymbol\{d\}\(t\)\\in\\mathbb\{R\}^\{6\}\. Spacecraft apply control accelerations in the Radial–Transverse–Normal \(RTN\) frame,

𝒖​\(t\)=\[uRuSuW\]⊤∈ℝ3,\\boldsymbol\{u\}\(t\)=\\begin\{bmatrix\}u\_\{\\mathrm\{R\}\}&u\_\{\\mathrm\{S\}\}&u\_\{\\mathrm\{W\}\}\\end\{bmatrix\}^\{\\top\}\\in\\mathbb\{R\}^\{3\},whereuRu\_\{\\mathrm\{R\}\},uSu\_\{\\mathrm\{S\}\}, anduWu\_\{\\mathrm\{W\}\}denote the radial, along\-track, and cross\-track acceleration components, respectively\. The target specification constrains only the first five orbital elements\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\), allowing free phasing in true anomalyν\\nu\.

### Orbital dynamics

Spacecraft orbital element dynamics are described by the Gauss Variational Equations \(GVE\),

a˙=2​a2h​\(ε​sin⁡\(ν\)​uR\+pr​uS\),\\displaystyle\\dot\{a\}=\\frac\{2a^\{2\}\}\{h\}\\\!\\left\(\\varepsilon\\sin\(\\nu\)\\,u\_\{\\mathrm\{R\}\}\+\\frac\{p\}\{r\}\\,u\_\{\\mathrm\{S\}\}\\right\),\(1a\)ε˙=1h​\(p​sin⁡\(ν\)​uR\+\(\(p\+r\)​cos⁡\(ν\)\+r​ε\)​uS\),\\displaystyle\\dot\{\\varepsilon\}=\\frac\{1\}\{h\}\\\!\\left\(p\\sin\(\\nu\)\\,u\_\{\\mathrm\{R\}\}\+\\big\(\(p\+r\)\\cos\(\\nu\)\+r\\varepsilon\\big\)\\,u\_\{\\mathrm\{S\}\}\\right\),\(1b\)ι˙=r​cos⁡\(θ\)h​uW,\\displaystyle\\dot\{\\iota\}=\\frac\{r\\cos\(\\theta\)\}\{h\}\\,u\_\{\\mathrm\{W\}\},\(1c\)Ω˙=r​sin⁡\(θ\)h​sin⁡\(ι\)​uW,\\displaystyle\\dot\{\\Omega\}=\\frac\{r\\sin\(\\theta\)\}\{h\\sin\(\\iota\)\}\\,u\_\{\\mathrm\{W\}\},\(1d\)ω˙=1h​ε​\(−p​cos⁡\(ν\)​uR\+\(p\+r\)​sin⁡\(ν\)​uS\)−r​sin⁡\(θ\)h​cot⁡\(ι\)​uW,\\displaystyle\\dot\{\\omega\}=\\frac\{1\}\{h\\varepsilon\}\\\!\\left\(\-p\\cos\(\\nu\)\\,u\_\{\\mathrm\{R\}\}\+\(p\+r\)\\sin\(\\nu\)\\,u\_\{\\mathrm\{S\}\}\\right\)\-\\frac\{r\\sin\(\\theta\)\}\{h\}\\cot\(\\iota\)\\,u\_\{\\mathrm\{W\}\},\(1e\)ν˙=hr2\+1h​ε​\(p​cos⁡\(ν\)​uR−\(p\+r\)​sin⁡\(ν\)​uS\),\\displaystyle\\dot\{\\nu\}=\\frac\{h\}\{r^\{2\}\}\+\\frac\{1\}\{h\\varepsilon\}\\\!\\left\(p\\cos\(\\nu\)\\,u\_\{\\mathrm\{R\}\}\-\(p\+r\)\\sin\(\\nu\)\\,u\_\{\\mathrm\{S\}\}\\right\),\(1f\)with

p=a​\(1−ε2\),r=p1\+ε​cos⁡\(ν\),h=μ​p,θ=ω\+ν\.p=a\(1\-\\varepsilon^\{2\}\),\\quad r=\\frac\{p\}\{1\+\\varepsilon\\cos\(\\nu\)\},\\quad h=\\sqrt\{\\mu\\,p\},\\quad\\theta=\\omega\+\\nu\.\(2\)
For compactness, we also write the controlled dynamics as

𝒙˙=𝒇​\(𝒙\)\+𝒈​\(𝒙\)​𝒖,\\dot\{\\boldsymbol\{x\}\}=\\boldsymbol\{f\}\(\\boldsymbol\{x\}\)\+\\boldsymbol\{g\}\(\\boldsymbol\{x\}\)\\boldsymbol\{u\},\(3\)where𝒇\\boldsymbol\{f\}captures the uncontrolled evolution and𝒈\\boldsymbol\{g\}collects the control\-affine coefficients implied by equation \([1](https://arxiv.org/html/2608.00320#Sx3.E1)\)\. Debris are modeled as unpowered \(𝒖≡𝟎\\boldsymbol\{u\}\\equiv\\boldsymbol\{0\}\) and propagated under the corresponding uncontrolled dynamics𝒅˙=𝒇​\(𝒅\)\\boldsymbol\{\\dot\{d\}\}=\\boldsymbol\{f\}\(\\boldsymbol\{d\}\)\.

### Distributional formulation

LetP0P\_\{0\}denote the initial spacecraft distribution overℝ6\\mathbb\{R\}^\{6\},P1P\_\{1\}the target distribution overℝ5\\mathbb\{R\}^\{5\}\(for the first five elements\), andPdP\_\{d\}the initial debris distribution overℝ6\\mathbb\{R\}^\{6\}\. We seek a trajectory mapF​\(𝒙,t\)F\(\\boldsymbol\{x\},t\)whose pushforward defines the time\-varying swarm distribution,

Pt=F​\(⋅,t\)\#​P0\.P\_\{t\}=F\(\\cdot,t\)\_\{\\\#\}P\_\{0\}\.\(4\)Rather than solving a large coupled optimal control problem directly for each\(P0,Pd,P1,T\)\(P\_\{0\},P\_\{d\},P\_\{1\},T\)instance, we learn an operator that amortizes this mapping across scenarios\.

### Neural trajectory operator

We define a time\-conditioned neural operator

𝒢θ:\(P0,Pd,P1,t,T\)↦F​\(⋅,t\)\#​P0,\\mathcal\{G\}\_\{\\theta\}:\(P\_\{0\},P\_\{d\},P\_\{1\},t,T\)\\mapsto F\(\\cdot,t\)\_\{\\\#\}P\_\{0\},\(5\)implemented as a permutation\-equivariant transformer with multi\-head attention\. Our architecture builds upon the operator\-learning framework introduced in Huang et al\.\[[27](https://arxiv.org/html/2608.00320#bib.bib27)\], which uses sampling\-invariant, permutation\-equivariant attention blocks to learn solution operators for mean\-field games from samples of the initial and terminal distributions\.

In contrast, our modified network introduces two key differences tailored to the spacecraft\-swarm navigation setting: first, we include dual cross\-attention streams in parallel, one between agents and the target orbits, and one between agents and the debris set, rather than a unified attention block over all inputs\. This separation enables dedicated feature pathways for target alignment and debris avoidance\. We maintain a shallow projection head that concatenates fused agent features, target\-attention features, debris\-attention features, and the original lifted queries, to output per\-agent orbital\-element predictions\. Detail on the composition of the operator is presented in Fig\.[6](https://arxiv.org/html/2608.00320#Sx3.F6)\.

![Refer to caption](https://arxiv.org/html/2608.00320v1/x1.png)Figure 6:Architecture of solution operator network, whereX0X\_\{0\},X1X\_\{1\}, andXdX\_\{d\}are sequence of spacecraft/debris locations with arbitrary sequence length, MLP stands for Multi\-layer Perceptron feed\-forward network, and MHCA stands for Multi\-head Cross Attention network\.
### Terminal loss

Given samples\{𝒙0,i\}i=1N∼P0\\\{\\boldsymbol\{x\}\_\{0,i\}\\\}\_\{i=1\}^\{N\}\\sim P\_\{0\}and associated target samples\{𝒙des,i\}i=1N∼P1\\\{\\boldsymbol\{x\}\_\{\\mathrm\{des\},i\}\\\}\_\{i=1\}^\{N\}\\sim P\_\{1\}, the operator produces terminal states𝒙i​\(T\)=F​\(𝒙0,i,T\)\\boldsymbol\{x\}\_\{i\}\(T\)=F\(\\boldsymbol\{x\}\_\{0,i\},T\)\. We penalize mismatch in the first five elements using

𝒯​\(F​\(⋅,T\)\#​P0;P1\)=1N​∑i=1N‖Π5​\(𝒙i​\(T\)\)−𝒙des,i‖22,\\mathcal\{T\}\\\!\\left\(F\(\\cdot,T\)\_\{\\\#\}P\_\{0\};P\_\{1\}\\right\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\\\|\\Pi\_\{5\}\\\!\\big\(\\boldsymbol\{x\}\_\{i\}\(T\)\\big\)\-\\boldsymbol\{x\}\_\{\\mathrm\{des\},i\}\\right\\\|\_\{2\}^\{2\},\(6\)whereΠ5​\(⋅\)\\Pi\_\{5\}\(\\cdot\)projects onto\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\)\.

### Interaction penalties

Collision avoidance is enforced using a closest\-point\-of\-approach \(CPA\) penalty evaluated in Cartesian space\. A state𝒙\\boldsymbol\{x\}expressed in Keplerian elements is converted to its Earth\-centered inertial \(ECI\) position–velocity state through the standard Keplerian element\-to\-Cartesian mapΓ\\Gamma, written𝝌=Γ​\(𝒙\)=\(𝒑,𝒗\)\\boldsymbol\{\\chi\}=\\Gamma\(\\boldsymbol\{x\}\)=\(\\boldsymbol\{p\},\\boldsymbol\{v\}\), and this conversion is applied to every spacecraft and debris state at every timestep\. Debris are propagated as unpowered objects by advancing only the true anomalyν\\nuunder two\-body Keplerian motion, keeping\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\)fixed\.

Given a spacecraft state𝝌s=\(𝒑s,𝒗s\)\\boldsymbol\{\\chi\}\_\{s\}=\(\\boldsymbol\{p\}\_\{s\},\\boldsymbol\{v\}\_\{s\}\)and a debris state𝝌d=\(𝒑d,𝒗d\)\\boldsymbol\{\\chi\}\_\{d\}=\(\\boldsymbol\{p\}\_\{d\},\\boldsymbol\{v\}\_\{d\}\)at the start of a discrete interval of durationΔ​t\\Delta t, we compute the relative position and velocity

Δ​𝒑=𝒑s−𝒑d,Δ​𝒗=𝒗s−𝒗d\.\\Delta\\boldsymbol\{p\}=\\boldsymbol\{p\}\_\{s\}\-\\boldsymbol\{p\}\_\{d\},\\qquad\\Delta\\boldsymbol\{v\}=\\boldsymbol\{v\}\_\{s\}\-\\boldsymbol\{v\}\_\{d\}\.The CPA time is

tcpa​\(𝝌s,𝝌d\)=clip​\(−Δ​𝒑⊤​Δ​𝒗‖Δ​𝒗‖22\+ϵ,0,Δ​t\),t\_\{\\mathrm\{cpa\}\}\(\\boldsymbol\{\\chi\}\_\{s\},\\boldsymbol\{\\chi\}\_\{d\}\)=\\mathrm\{clip\}\\\!\\left\(\-\\frac\{\\Delta\\boldsymbol\{p\}^\{\\top\}\\Delta\\boldsymbol\{v\}\}\{\\\|\\Delta\\boldsymbol\{v\}\\\|\_\{2\}^\{2\}\+\\epsilon\},\\,0,\\,\\Delta t\\right\),and the CPA separation is

dcpa​\(𝝌s,𝝌d\)=‖Δ​𝒑\+Δ​𝒗​tcpa​\(𝝌s,𝝌d\)‖2\.d\_\{\\mathrm\{cpa\}\}\(\\boldsymbol\{\\chi\}\_\{s\},\\boldsymbol\{\\chi\}\_\{d\}\)=\\left\\\|\\Delta\\boldsymbol\{p\}\+\\Delta\\boldsymbol\{v\}\\,t\_\{\\mathrm\{cpa\}\}\(\\boldsymbol\{\\chi\}\_\{s\},\\boldsymbol\{\\chi\}\_\{d\}\)\\right\\\|\_\{2\}\.\(7\)We penalize violations of a pair\-type\-dependent safety radius using a quadratic hinge, evaluated at every interval along the horizon,

𝒥CPA​\(Ps,Pd\)=\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{CPA\}\}\(P\_\{s\},P\_\{d\}\)=\{\}𝔼𝒙s∼Ps,𝒙d∼Pd​\(κ​max⁡\{0,rd−dcpa​\(Γ​\(𝒙s\),Γ​\(𝒙d\)\)\}\)2\\displaystyle\\mathbb\{E\}\_\{\\boldsymbol\{x\}\_\{s\}\\sim P\_\{s\},\\,\\boldsymbol\{x\}\_\{d\}\\sim P\_\{d\}\}\\left\(\\kappa\\max\\\!\\left\\\{0,\\;r\_\{d\}\-d\_\{\\mathrm\{cpa\}\}\\\!\\left\(\\Gamma\(\\boldsymbol\{x\}\_\{s\}\),\\Gamma\(\\boldsymbol\{x\}\_\{d\}\)\\right\)\\right\\\}\\right\)^\{2\}\+ws​𝔼𝒙s,𝒙s′∼Ps​\(κ​max⁡\{0,rs−dcpa​\(Γ​\(𝒙s\),Γ​\(𝒙s′\)\)\}\)2,\\displaystyle\+w\_\{s\}\\,\\mathbb\{E\}\_\{\\boldsymbol\{x\}\_\{s\},\\,\\boldsymbol\{x\}^\{\\prime\}\_\{s\}\\sim P\_\{s\}\}\\left\(\\kappa\\max\\\!\\left\\\{0,\\;r\_\{s\}\-d\_\{\\mathrm\{cpa\}\}\\\!\\left\(\\Gamma\(\\boldsymbol\{x\}\_\{s\}\),\\Gamma\(\\boldsymbol\{x\}^\{\\prime\}\_\{s\}\)\\right\)\\right\\\}\\right\)^\{2\},\(8\)wherePsP\_\{s\}andPdP\_\{d\}are the spacecraft and debris populations, and the second expectation excludes the self\-pair𝒙s′=𝒙s\\boldsymbol\{x\}^\{\\prime\}\_\{s\}=\\boldsymbol\{x\}\_\{s\}\. The spacecraft–debris safety radiusrd=1r\_\{d\}=1km applies at all times; the spacecraft–spacecraft safety radiusrs=100r\_\{s\}=100m applies only fort≥0\.2​Tt\\geq 0\.2\\,Tand its term carries the weightws=10w\_\{s\}=10;κ\\kappais a scaling constant\. The spacecraft–spacecraft radius matches the100100m threshold used in evaluation, and the grace window over the first20%20\\%of the transfer exempts the deliberately conflicting clustered start \(see Scenario Sampling\), which is unavoidable by construction, while still penalizing any conflict the swarm has not resolved by mid\-transfer, including the converging arrival\. This yields a smooth, differentiable surrogate that concentrates gradient signal near predicted close\-approach events\.

##### Adversarial debris generation\.

To stress\-test avoidance behavior during training, every debris object is generated adversarially as a*crossing\-orbit*threat to the model’s*nominal*rollout, the trajectory produced when the model is prompted with an empty debris set\. Concretely, for a given\(P0,P1,T\)\(P\_\{0\},P\_\{1\},T\), we first perform a rollout withPd=∅P\_\{d\}=\\varnothingto obtain each agent’s nominal trajectory\. For each debris object we select an agent and a random hit timethitt\_\{\\mathrm\{hit\}\}in the latter half of the transfer, and take the agent’s nominal ECI state\(𝒑,𝒗\)\(\\boldsymbol\{p\},\\boldsymbol\{v\}\)atthitt\_\{\\mathrm\{hit\}\}\. The debris velocity is the agent’s velocity rotated by a random angle drawn from\[20∘,75∘\]\[20^\{\\circ\},75^\{\\circ\}\]about a random axis perpendicular to it; the rotation preserves speed, so the debris orbit remains bound, but crosses the agent’s path with a substantial relative velocity, as in a genuine conjunction, rather than trailing it co\-orbitally\. The debris position is offset from the agent’s by a sub\-safety\-radius near\-miss distance \(0\.350\.35–0\.65​rd0\.65\\,r\_\{d\}, random direction\), the perturbed state is converted back to orbital elements, and, because the debris propagation model advances onlyν\\nu, the initial true anomalyν0\\nu\_\{0\}is chosen so that the object reaches this state atthitt\_\{\\mathrm\{hit\}\}under Keplerian motion\. The near\-miss offset is essential: a threat placed at exact position–velocity coincidence produces a closest\-approach distance of zero, at which the CPA penalty’s gradient with respect to position vanishes identically, so the model receives a large penalty but no direction in which to evade\. The crossing near\-miss instead yields a well\-conditioned avoidance gradient on every sample\.

### Fuel\-cost surrogate from Gauss variational dynamics

To encourage fuel\-efficient transfers, we penalize the magnitude of the control acceleration implied by the Gauss Variational Equations \(GVE\)\. The element\-rate dynamics take the control\-affine form𝒙˙=𝒇​\(𝒙\)\+𝒈​\(𝒙\)​𝒖\\dot\{\\boldsymbol\{x\}\}=\\boldsymbol\{f\}\(\\boldsymbol\{x\}\)\+\\boldsymbol\{g\}\(\\boldsymbol\{x\}\)\\,\\boldsymbol\{u\}, where𝒙∈ℝ6\\boldsymbol\{x\}\\in\\mathbb\{R\}^\{6\}is the orbital\-element state,𝒖∈ℝ3\\boldsymbol\{u\}\\in\\mathbb\{R\}^\{3\}is the control acceleration in the rotating RTN frame,𝒇∈ℝ6\\boldsymbol\{f\}\\in\\mathbb\{R\}^\{6\}is the uncontrolled Keplerian rate \(nonzero only in theν\\nucomponent\), and𝒈​\(𝒙\)∈ℝ6×3\\boldsymbol\{g\}\(\\boldsymbol\{x\}\)\\in\\mathbb\{R\}^\{6\\times 3\}collects the GVE control\-affine coefficients\. Given a predicted trajectory𝒙​\(t\)\\boldsymbol\{x\}\(t\), we estimate𝒙˙​\(t\)\\dot\{\\boldsymbol\{x\}\}\(t\)by central finite differences over the discretized rollout and infer the corresponding control via the Moore–Penrose pseudoinverse𝒈†\\boldsymbol\{g\}^\{\\dagger\}:

𝒖^​\(t\)=𝒈​\(𝒙​\(t\)\)†​\(𝒙˙​\(t\)−𝒇​\(𝒙​\(t\)\)\)\.\\hat\{\\boldsymbol\{u\}\}\(t\)=\\boldsymbol\{g\}\(\\boldsymbol\{x\}\(t\)\)^\{\\dagger\}\\left\(\\dot\{\\boldsymbol\{x\}\}\(t\)\-\\boldsymbol\{f\}\(\\boldsymbol\{x\}\(t\)\)\\right\)\.We then define the fuel\-cost surrogate as the time integral of squared control magnitude,

ℱ=∫0T‖𝒖^​\(t\)‖22​𝑑t,\\mathcal\{F\}=\\int\_\{0\}^\{T\}\\\|\\hat\{\\boldsymbol\{u\}\}\(t\)\\\|\_\{2\}^\{2\}\\,dt,\(9\)implemented in discrete time using the rollout grid\. This term does not require solving an optimal control problem during training, yet it preserves the physical meaning of minimizing RTN control effort under the element\-rate dynamics\.

### Trajectory parameterization with a biased baseline and learned residual

We represent each spacecraft trajectory in orbital elements as a smooth baseline transfer augmented by a learned residual\. Letτ=t/T∈\[0,1\]\\tau=t/T\\in\[0,1\]denote normalized time, and let𝒙​\(t\)=\[a,ε,ι,Ω,ω,ν\]⊤\\boldsymbol\{x\}\(t\)=\[a,\\varepsilon,\\iota,\\Omega,\\omega,\\nu\]^\{\\top\}\. For the first five “slow” elements\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\), we define a deterministic baseline𝒙¯1:5​\(τ\)\\bar\{\\boldsymbol\{x\}\}\_\{1:5\}\(\\tau\)that interpolates between the initial and target values, and we learn an additive bias \(residual\)Δ​𝒙1:5​\(τ\)\\Delta\\boldsymbol\{x\}\_\{1:5\}\(\\tau\):

𝒙1:5​\(τ\)=𝒙¯1:5​\(τ\)\+Δ​𝒙1:5​\(τ\)\.\\boldsymbol\{x\}\_\{1:5\}\(\\tau\)=\\bar\{\\boldsymbol\{x\}\}\_\{1:5\}\(\\tau\)\+\\Delta\\boldsymbol\{x\}\_\{1:5\}\(\\tau\)\.The baseline uses linear interpolation for\(a,ε,ι\)\(a,\\varepsilon,\\iota\)and wrapped interpolation for angular elements\(Ω,ω\)\(\\Omega,\\omega\)via

angdiff​\(α,β\)=atan2​\(sin⁡\(α−β\),cos⁡\(α−β\)\),\\mathrm\{angdiff\}\(\\alpha,\\beta\)=\\mathrm\{atan2\}\(\\sin\(\\alpha\-\\beta\),\\cos\(\\alpha\-\\beta\)\),so that angle differences remain on the principal branch\. The resulting slow\-element sequence is projected by clamping physical bounds \(e\.g\.,a\>0a\>0,0≤ε<10\\leq\\varepsilon<1,0<ι<π0<\\iota<\\pi\) and wrapping angles to\[0,2​π\)\[0,2\\pi\)\.

The true anomalyν\\nuis then propagated forward in time using a physically grounded Keplerian rate with an additional learned correction\. Specifically, at discrete timestkt\_\{k\}with stepsΔ​tk\\Delta t\_\{k\}, we update

νk\+1=νk\+Δ​tk​\(ν˙kep​\(ak,εk,νk\)\+Δ​ν˙k\),\\nu\_\{k\+1\}=\\nu\_\{k\}\+\\Delta t\_\{k\}\\Big\(\\dot\{\\nu\}\_\{\\mathrm\{kep\}\}\(a\_\{k\},\\varepsilon\_\{k\},\\nu\_\{k\}\)\+\\Delta\\dot\{\\nu\}\_\{k\}\\Big\),whereν˙kep\\dot\{\\nu\}\_\{\\mathrm\{kep\}\}is the two\-body Keplerian rate andΔ​ν˙k\\Delta\\dot\{\\nu\}\_\{k\}is the network output evaluated at the normalized timeτk=tk/T∈\[0,1\]\\tau\_\{k\}=t\_\{k\}/T\\in\[0,1\]\. This parameterization biases the model toward feasible transfers while allowing the learned residual to represent nontrivial maneuver geometry and phasing behavior\.

### Training objective

For each sampled scenario\(P0,Pd,P1,T\)\(P\_\{0\},P\_\{d\},P\_\{1\},T\), the underlying planning task can be viewed as a finite\-horizon collision\-aware optimal control problem\. Given spacecraft samples\{𝒙0,i\}i=1N∼P0\\\{\\boldsymbol\{x\}\_\{0,i\}\\\}\_\{i=1\}^\{N\}\\sim P\_\{0\}, target samples\{𝒙des,i\}i=1N∼P1\\\{\\boldsymbol\{x\}\_\{\\mathrm\{des\},i\}\\\}\_\{i=1\}^\{N\}\\sim P\_\{1\}, and debris samples\{𝒅0,j\}j=1M∼Pd\\\{\\boldsymbol\{d\}\_\{0,j\}\\\}\_\{j=1\}^\{M\}\\sim P\_\{d\}, the ideal per\-instance problem is to find trajectories and controls that minimize fuel expenditure while reaching the target distribution and avoiding close approaches:

min\{𝒙i,𝒖i\}i=1N\\displaystyle\\min\_\{\\\{\\boldsymbol\{x\}\_\{i\},\\boldsymbol\{u\}\_\{i\}\\\}\_\{i=1\}^\{N\}\}∑i=1N∫0T‖𝒖i​\(t\)‖22​𝑑t\+λT​∑i=1N‖Π5​\(𝒙i​\(T\)\)−Π5​\(𝒙des,i\)‖22\+λI​𝒥CPA\\displaystyle\\sum\_\{i=1\}^\{N\}\\int\_\{0\}^\{T\}\\\|\\boldsymbol\{u\}\_\{i\}\(t\)\\\|\_\{2\}^\{2\}\\,dt\+\\lambda\_\{T\}\\sum\_\{i=1\}^\{N\}\\left\\\|\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{i\}\(T\)\\right\)\-\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{\\mathrm\{des\},i\}\\right\)\\right\\\|\_\{2\}^\{2\}\+\\lambda\_\{I\}\\,\\mathcal\{J\}\_\{\\mathrm\{CPA\}\}\(10\)s\.t\.\\displaystyle\\mathrm\{s\.t\.\}𝒙˙i​\(t\)=𝒇​\(𝒙i​\(t\)\)\+𝒈​\(𝒙i​\(t\)\)​𝒖i​\(t\),𝒙i​\(0\)=𝒙0,i,\\displaystyle\\dot\{\\boldsymbol\{x\}\}\_\{i\}\(t\)=\\boldsymbol\{f\}\(\\boldsymbol\{x\}\_\{i\}\(t\)\)\+\\boldsymbol\{g\}\(\\boldsymbol\{x\}\_\{i\}\(t\)\)\\boldsymbol\{u\}\_\{i\}\(t\),\\qquad\\boldsymbol\{x\}\_\{i\}\(0\)=\\boldsymbol\{x\}\_\{0,i\},𝒅˙j​\(t\)=𝒇​\(𝒅j​\(t\)\),𝒅j​\(0\)=𝒅0,j,\\displaystyle\\dot\{\\boldsymbol\{d\}\}\_\{j\}\(t\)=\\boldsymbol\{f\}\(\\boldsymbol\{d\}\_\{j\}\(t\)\),\\qquad\\boldsymbol\{d\}\_\{j\}\(0\)=\\boldsymbol\{d\}\_\{0,j\},whereΠ5\\Pi\_\{5\}projects onto the first five orbital elements\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\)and𝒥CPA\\mathcal\{J\}\_\{\\mathrm\{CPA\}\}denotes the closest\-point\-of\-approach penalty, evaluated separately over spacecraft–debris and spacecraft–spacecraft pairs with their respective safety radii \(see[Interaction penalties](https://arxiv.org/html/2608.00320#Sx3.SSx6)\)\. This formulation expresses the desired collision\-aware planning problem for a single scenario, but solving equation \([10](https://arxiv.org/html/2608.00320#Sx3.E10)\) repeatedly would require expensive numerical optimization and would make supervised training dependent on a large library of precomputed optimal trajectories\.

Instead, we train𝒢θ\\mathcal\{G\}\_\{\\theta\}as an amortized solution operator over a distribution of scenarios\. At each training instance, we sample\(P0,Pd,P1,T\)\(P\_\{0\},P\_\{d\},P\_\{1\},T\), construct a debris set with an adversarial fraction, and generate trajectories using the biased baseline plus learned residual parameterization, namely

𝒙θ​\(t;𝒙0\)=𝒙¯​\(t\)\+Δ​𝒙θ​\(t;𝒙0\),\\boldsymbol\{x\}\_\{\\theta\}\(t;\\boldsymbol\{x\}\_\{0\}\)=\\bar\{\\boldsymbol\{x\}\}\(t\)\+\\Delta\\boldsymbol\{x\}\_\{\\theta\}\(t;\\boldsymbol\{x\}\_\{0\}\),\(11\)where𝒙¯\\bar\{\\boldsymbol\{x\}\}is the deterministic baseline of[Trajectory parameterization with a biased baseline and learned residual](https://arxiv.org/html/2608.00320#Sx3.SSx8)andΔ​𝒙θ\\Delta\\boldsymbol\{x\}\_\{\\theta\}is the network residual, conditioned on the full scenario\(P0,Pd,P1,T\)\(P\_\{0\},P\_\{d\},P\_\{1\},T\)\. The rollout realizes the trajectory map of equation \([5](https://arxiv.org/html/2608.00320#Sx3.E5)\) sample\-wise:𝒢θ​\(P0,Pd,P1,t,T\)=𝒙θ​\(t;⋅\)\#​P0\\mathcal\{G\}\_\{\\theta\}\(P\_\{0\},P\_\{d\},P\_\{1\},t,T\)=\\boldsymbol\{x\}\_\{\\theta\}\(t;\\cdot\)\_\{\\\#\}P\_\{0\}\. The network parameters are optimized by minimizing the expected self\-supervised loss

minθ⁡𝔼\(P0,Pd,P1,T\)∼𝒟​\[λf​ℱθ\+λI​ℐθ\+λT​𝒯θ\+λS​𝒮θ\],\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{\(P\_\{0\},P\_\{d\},P\_\{1\},T\)\\sim\\mathcal\{D\}\}\\left\[\\lambda\_\{f\}\\mathcal\{F\}\_\{\\theta\}\+\\lambda\_\{I\}\\mathcal\{I\}\_\{\\theta\}\+\\lambda\_\{T\}\\mathcal\{T\}\_\{\\theta\}\+\\lambda\_\{S\}\\mathcal\{S\}\_\{\\theta\}\\right\],\(12\)where𝒟\\mathcal\{D\}is the training distribution over planning scenarios\. The four terms are defined as follows:

ℱθ=𝔼𝒙0∼P0​∫0T‖𝒈​\(𝒙θ​\(t;𝒙0\)\)†​\(𝒙˙θ​\(t;𝒙0\)−𝒇​\(𝒙θ​\(t;𝒙0\)\)\)‖22​𝑑t\\mathcal\{F\}\_\{\\theta\}=\\mathbb\{E\}\_\{\\boldsymbol\{x\}\_\{0\}\\sim P\_\{0\}\}\\int\_\{0\}^\{T\}\\left\\\|\\boldsymbol\{g\}\(\\boldsymbol\{x\}\_\{\\theta\}\(t;\\boldsymbol\{x\}\_\{0\}\)\)^\{\\dagger\}\\left\(\\dot\{\\boldsymbol\{x\}\}\_\{\\theta\}\(t;\\boldsymbol\{x\}\_\{0\}\)\-\\boldsymbol\{f\}\(\\boldsymbol\{x\}\_\{\\theta\}\(t;\\boldsymbol\{x\}\_\{0\}\)\)\\right\)\\right\\\|\_\{2\}^\{2\}\\,dt\(13\)is the fuel cost, the distributional form of equation \([9](https://arxiv.org/html/2608.00320#Sx3.E9)\); and

ℐθ=𝒥CPA​\(𝒙θ​\(t;⋅\)\#​P0,Pd\)\\mathcal\{I\}\_\{\\theta\}=\\mathcal\{J\}\_\{\\mathrm\{CPA\}\}\\left\(\\boldsymbol\{x\}\_\{\\theta\}\(t;\\cdot\)\_\{\\\#\}P\_\{0\},\\;P\_\{d\}\\right\)\(14\)is the closest\-point\-of\-approach interaction penalty of equation \([8](https://arxiv.org/html/2608.00320#Sx3.E8)\), evaluated on the pushforward ofP0P\_\{0\}under the rollout along the horizon; and

𝒯θ=𝔼\(𝒙0,𝒙des\)∼\(P0,P1\)​‖Π5​\(𝒙θ​\(T;𝒙0\)\)−Π5​\(𝒙des\)‖22\\mathcal\{T\}\_\{\\theta\}=\\mathbb\{E\}\_\{\(\\boldsymbol\{x\}\_\{0\},\\boldsymbol\{x\}\_\{\\mathrm\{des\}\}\)\\sim\(P\_\{0\},P\_\{1\}\)\}\\left\\\|\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{\\theta\}\(T;\\boldsymbol\{x\}\_\{0\}\)\\right\)\-\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{\\mathrm\{des\}\}\\right\)\\right\\\|\_\{2\}^\{2\}\(15\)is the terminal loss of equation \([6](https://arxiv.org/html/2608.00320#Sx3.E6)\), where the expectation is over paired samples, each spacecraft with its own assigned target; and

𝒮θ=𝔼𝒙0∼P0​‖Π5​\(𝒙θ​\(0;𝒙0\)\)−Π5​\(𝒙0\)‖22\\mathcal\{S\}\_\{\\theta\}=\\mathbb\{E\}\_\{\\boldsymbol\{x\}\_\{0\}\\sim P\_\{0\}\}\\left\\\|\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{\\theta\}\(0;\\boldsymbol\{x\}\_\{0\}\)\\right\)\-\\Pi\_\{5\}\\\!\\left\(\\boldsymbol\{x\}\_\{0\}\\right\)\\right\\\|\_\{2\}^\{2\}\(16\)is the initial loss, which anchors the rollout atτ=0\\tau=0to the sampled initial state; in both anchoring terms the projectionΠ5\\Pi\_\{5\}excludes the true anomaly\. This objective trains the operator without ground\-truth optimal trajectories or numerically generated trajectory labels, while preserving the structure of the per\-instance optimal control problem in equation \([10](https://arxiv.org/html/2608.00320#Sx3.E10)\)\. The specific choice of the weights\(λf,λI,λT,λS\)\(\\lambda\_\{f\},\\lambda\_\{I\},\\lambda\_\{T\},\\lambda\_\{S\}\)is discussed in Table[5](https://arxiv.org/html/2608.00320#Sx3.T5)\.

### Dynamic\-feasibility finish via Gauss–Newton terminal targeting

The operator’s element\-space rollout is not, by construction, the integral of a physical control sequence under exact two\-body dynamics\. We finish each rollout, per agent, with a single\-shooting Gauss–Newton \(GN\) step with Levenberg–Marquardt damping\. The control sequence𝒖0:T−1\\boldsymbol\{u\}\_\{0:T\-1\}is the only decision variable; the state is the exact RK4 rollout𝒙k\+1=ΦΔ​tRK4​\(𝒙k,𝒖k\)\\boldsymbol\{x\}\_\{k\+1\}=\\Phi^\{\\mathrm\{RK4\}\}\_\{\\Delta t\}\(\\boldsymbol\{x\}\_\{k\},\\boldsymbol\{u\}\_\{k\}\)from the fixed initial state, so every iterate is dynamically feasible and there are no dynamics constraints\.

We target the orbit through its conserved vectors rather than its element angles: the specific angular momentum𝒉=𝒓×𝒗\\boldsymbol\{h\}=\\boldsymbol\{r\}\\times\\boldsymbol\{v\}and the eccentricity vector𝒆=\(𝒗×𝒉\)/μ−𝒓/‖𝒓‖\\boldsymbol\{e\}=\(\\boldsymbol\{v\}\\times\\boldsymbol\{h\}\)/\\mu\-\\boldsymbol\{r\}/\\\|\\boldsymbol\{r\}\\\|\. Both are invariant along an orbit, so the residual is phase\-free, and smooth in\(𝒓,𝒗\)\(\\boldsymbol\{r\},\\boldsymbol\{v\}\), so the Jacobian stays well\-defined for the near\-circular and near\-equatorial orbits where element\-angle residuals are singular\. With targets\(𝒉⋆,𝒆⋆\)\(\\boldsymbol\{h\}^\{\\star\},\\boldsymbol\{e\}^\{\\star\}\)from the goal orbitP1P\_\{1\}, the residual and objective are

𝝆​\(𝒖\)=\[\(𝒉​\(T\)−𝒉⋆\)/sh;\(𝒆​\(T\)−𝒆⋆\)/se\]∈ℝ6,min𝒖⁡‖𝝆​\(𝒖\)‖22\+λf​‖𝒖‖22\.\\boldsymbol\{\\rho\}\(\\boldsymbol\{u\}\)=\\big\[\(\\boldsymbol\{h\}\(T\)\-\\boldsymbol\{h\}^\{\\star\}\)/s\_\{h\};\\;\(\\boldsymbol\{e\}\(T\)\-\\boldsymbol\{e\}^\{\\star\}\)/s\_\{e\}\\big\]\\in\\mathbb\{R\}^\{6\},\\qquad\\min\_\{\\boldsymbol\{u\}\}~\\\|\\boldsymbol\{\\rho\}\(\\boldsymbol\{u\}\)\\\|\_\{2\}^\{2\}\+\\lambda\_\{f\}\\\|\\boldsymbol\{u\}\\\|\_\{2\}^\{2\}\.\(17\)Each iteration forms the Jacobian𝑱=∂𝝆/∂𝒖∈ℝ6×3​T\\boldsymbol\{J\}=\\partial\\boldsymbol\{\\rho\}/\\partial\\boldsymbol\{u\}\\in\\mathbb\{R\}^\{6\\times 3T\}by automatic differentiation through the RK4 rollout, six vector–Jacobian products, one per residual component, sharing a single retained backward graph, and takes the damped Gauss–Newton step

Δ​𝒖=−\(𝑱⊤​𝑱\+\(λf\+μ\)​𝑰\)−1​\(𝑱⊤​𝝆\+λf​𝒖\),\\Delta\\boldsymbol\{u\}=\-\\big\(\\boldsymbol\{J\}^\{\\top\}\\boldsymbol\{J\}\+\(\\lambda\_\{f\}\+\\mu\)\\boldsymbol\{I\}\\big\)^\{\-1\}\\big\(\\boldsymbol\{J\}^\{\\top\}\\boldsymbol\{\\rho\}\+\\lambda\_\{f\}\\boldsymbol\{u\}\\big\),\(18\)whereμ≥0\\mu\\geq 0is the Levenberg–Marquardt damping\. The decision vector𝒖\\boldsymbol\{u\}has dimension3​T3T\(hundreds to thousands\), but the residual has only six components, so we never form the3​T×3​T3T\\times 3Tsystem\. Writingα=λf\+μ\\alpha=\\lambda\_\{f\}\+\\muand𝒃=−\(𝑱⊤​𝝆\+λf​𝒖\)\\boldsymbol\{b\}=\-\(\\boldsymbol\{J\}^\{\\top\}\\boldsymbol\{\\rho\}\+\\lambda\_\{f\}\\boldsymbol\{u\}\), the Woodbury identity collapses equation \([18](https://arxiv.org/html/2608.00320#Sx3.E18)\) to a single6×66\\times 6solve,

Δ​𝒖=1α​\(𝒃−𝑱⊤​\(α​𝑰6\+𝑱​𝑱⊤\)−1​𝑱​𝒃\),\\Delta\\boldsymbol\{u\}=\\tfrac\{1\}\{\\alpha\}\\Big\(\\boldsymbol\{b\}\-\\boldsymbol\{J\}^\{\\top\}\\big\(\\alpha\\boldsymbol\{I\}\_\{6\}\+\\boldsymbol\{J\}\\boldsymbol\{J\}^\{\\top\}\\big\)^\{\-1\}\\boldsymbol\{J\}\\,\\boldsymbol\{b\}\\Big\),\(19\)whose cost is independent of the horizon lengthTTand identical for every agent\. The swarm is therefore solved as one batched stack of6×66\\times 6systems, the property that lets the finish run atN=1000N=1000, where a per\-agent nonlinear program is intractable\. The6×66\\times 6inner solve is carried out in double precision to absorb the1/α1/\\alphacancellation when the fuel term is active \(λf\>0\\lambda\_\{f\}\>0\); for the pure terminal target \(λf=0\\lambda\_\{f\}=0\) the step reduces to the numerically benign minimum\-norm formΔ​𝒖=−𝑱⊤​\(𝑱​𝑱⊤\+μ​𝑰6\)−1​𝝆\\Delta\\boldsymbol\{u\}=\-\\boldsymbol\{J\}^\{\\top\}\(\\boldsymbol\{J\}\\boldsymbol\{J\}^\{\\top\}\+\\mu\\boldsymbol\{I\}\_\{6\}\)^\{\-1\}\\boldsymbol\{\\rho\}\.

The damping is adapted by a*per\-agent*trust region\. At each iteration the trial step is accepted for every agent whose combined objective‖𝝆‖22\+λf​‖𝒖‖22\\\|\\boldsymbol\{\\rho\}\\\|\_\{2\}^\{2\}\+\\lambda\_\{f\}\\\|\\boldsymbol\{u\}\\\|\_\{2\}^\{2\}decreases and rejected otherwise;μ\\muis scaled down by0\.50\.5on acceptance and up by44on rejection, clamped to a fixed range\. Because acceptance is decided independently per agent inside the batched iteration, a stiff agent that needs heavy damping does not stall the rest of the swarm\. We run at most2525iterations and stop early once the terminal residual falls below10−710^\{\-7\}\. The same conserved\-vector targeting defines the single\-agent IPOPT reference\.

We report two variants\. GNw warm\-starts𝒖\\boldsymbol\{u\}from the operator rollout via an analytical RK4 inverse, per interval, inverting the constant input\-to\-state mapB=\[\(Δ​t2/2\)​I3;Δ​t​I3\]B=\[\(\\Delta t^\{2\}/2\)I\_\{3\};\\,\\Delta t\\,I\_\{3\}\]on the zero\-control residual, so the finish begins from the operator’s collision\-aware geometry; GNc cold\-starts from𝒖=0\\boldsymbol\{u\}=0\. The fuel regularizerλf\\lambda\_\{f\}\(a small default;λf=0\\lambda\_\{f\}=0recovers the pure minimum\-norm step\) trades a slight amount of terminal accuracy for lower control energy, and is added as the residual rows in equation \([18](https://arxiv.org/html/2608.00320#Sx3.E18)\) so the solve stays6×66\\times 6\. The finish carries no collision term: avoidance is supplied entirely by the collision\-aware operator seed, and the minimum\-norm character of the step perturbs that trajectory only as much as exact dynamics require\. All reportedΔ​v\\Delta vand terminal\-error values are computed on the finished, dynamically\-exact trajectory\.

### Architecture and training details

This subsection gives the exact layer structure of the operator𝒢θ\\mathcal\{G\}\_\{\\theta\}\(equation \([5](https://arxiv.org/html/2608.00320#Sx3.E5)\)\), its initialization, and the training hyperparameters\.

#### Input feature lifting

Each spacecraft or debris state𝒙=\[a,ε,ι,Ω,ω,ν\]⊤∈ℝ6\\boldsymbol\{x\}=\[a,\\,\\varepsilon,\\,\\iota,\\,\\Omega,\\,\\omega,\\,\\nu\]^\{\\top\}\\in\\mathbb\{R\}^\{6\}is expressed in non\-dimensional units using the scale factors

Lc=6\.37×106​m,Vc=μE/Lc,Tc=Lc/Vc≈803​s\.L\_\{c\}=6\.37\\times 10^\{6\}\\ \\text\{m\},\\quad V\_\{c\}=\\sqrt\{\\mu\_\{E\}/L\_\{c\}\},\\quad T\_\{c\}=L\_\{c\}/V\_\{c\}\\approx 803\\ \\text\{s\}\.\(20\)Before encoding, every point inP0P\_\{0\},P1P\_\{1\}, andPdP\_\{d\}is lifted fromℝ6\\mathbb\{R\}^\{6\}toℝ8\\mathbb\{R\}^\{8\}by appending the normalized timeτ∈\[0,1\]\\tau\\in\[0,1\]and the non\-dimensional mission durationTnd=T/TcT\_\{\\mathrm\{nd\}\}=T/T\_\{c\}:

𝒙~\(j\)=\[𝒙\(j\),τ,Tnd\]∈ℝ8\.\\tilde\{\\boldsymbol\{x\}\}^\{\(j\)\}=\\bigl\[\\boldsymbol\{x\}^\{\(j\)\},\\;\\tau,\\;T\_\{\\mathrm\{nd\}\}\\bigr\]\\in\\mathbb\{R\}^\{8\}\.\(21\)This conditioning is applied identically across all three input sets, yielding batched tensors of shape\(B,N,8\)\(B,N,8\),\(B,N,8\)\(B,N,8\), and\(B,M,8\)\(B,M,8\)for spacecraft, target, and debris, respectively\.

#### Network architecture

The operator is structured in three sequential stages\.

##### Per\-Point Encoders

Three independent two\-layer pointwise MLPs, implemented as1×11\{\\times\}1convolutions over the particle dimension to preserve permutation equivariance,project each input stream into a shared latent space of widthhh:

ϕ:ℝ8→ℝh,ϕ\(⋅\)=σ\(W2σ\(W1⋅\+b1\)\+b2\),\\phi:\\;\\mathbb\{R\}^\{8\}\\to\\mathbb\{R\}^\{h\},\\quad\\phi\(\\cdot\)=\\sigma\\bigl\(W\_\{2\}\\,\\sigma\(W\_\{1\}\\,\\cdot\+b\_\{1\}\)\+b\_\{2\}\\bigr\),\(22\)whereσ\\sigmadenotes GELU and dropout \(ratepp\) is applied after the first linear map\. Separate parameter setsϕ0\\phi\_\{0\},ϕ1\\phi\_\{1\},ϕd\\phi\_\{d\}are maintained for the spacecraft, target, and debris streams, producing encoded representations

𝐇0∈ℝB×h×N,𝐇1∈ℝB×h×N,𝐇d∈ℝB×h×M\.\\mathbf\{H\}\_\{0\}\\in\\mathbb\{R\}^\{B\\times h\\times N\},\\quad\\mathbf\{H\}\_\{1\}\\in\\mathbb\{R\}^\{B\\times h\\times N\},\\quad\\mathbf\{H\}\_\{d\}\\in\\mathbb\{R\}^\{B\\times h\\times M\}\.\(23\)

##### Dual Cross\-Attention Stack

The operator employs two parallel cross\-attention streams\. This appendix gives the precise per\-layer update\. At each layerℓ=1,…,L\\ell=1,\\ldots,L, the spacecraft encodings𝐗q\(ℓ−1\)∈ℝN×B×h\\mathbf\{X\}\_\{q\}^\{\(\\ell\-1\)\}\\in\\mathbb\{R\}^\{N\\times B\\times h\}serve as queries into both streams:

𝐘1\(ℓ\)\\displaystyle\\mathbf\{Y\}\_\{1\}^\{\(\\ell\)\}=MHA1​\(𝐗q\(ℓ−1\),𝐇1,𝐇1\),\\displaystyle=\\mathrm\{MHA\}\_\{1\}\\\!\\bigl\(\\mathbf\{X\}\_\{q\}^\{\(\\ell\-1\)\},\\,\\mathbf\{H\}\_\{1\},\\,\\mathbf\{H\}\_\{1\}\\bigr\),\(24\)𝐘d\(ℓ\)\\displaystyle\\mathbf\{Y\}\_\{d\}^\{\(\\ell\)\}=MHAd​\(𝐗q\(ℓ−1\),𝐇d,𝐇d\),\\displaystyle=\\mathrm\{MHA\}\_\{d\}\\\!\\bigl\(\\mathbf\{X\}\_\{q\}^\{\(\\ell\-1\)\},\\,\\mathbf\{H\}\_\{d\},\\,\\mathbf\{H\}\_\{d\}\\bigr\),\(25\)𝐗q\(ℓ\)\\displaystyle\\mathbf\{X\}\_\{q\}^\{\(\\ell\)\}=𝐗q\(ℓ−1\)\+𝐘1\(ℓ\)\+𝐘d\(ℓ\),\\displaystyle=\\mathbf\{X\}\_\{q\}^\{\(\\ell\-1\)\}\+\\mathbf\{Y\}\_\{1\}^\{\(\\ell\)\}\+\\mathbf\{Y\}\_\{d\}^\{\(\\ell\)\},\(26\)where eachMHA\\mathrm\{MHA\}block is a standard scaled dot\-product multi\-head attention withhheadsh\_\{\\mathrm\{heads\}\}heads, head dimensionh/hheadsh/h\_\{\\mathrm\{heads\}\}, and no projection biases\. The two streams share no parameters\. The residual accumulation in the last line propagates both goal\-directed and avoidance signals simultaneously through the depth of the stack\.

##### Final Decoder

AfterLLattention layers, features from the accumulated query𝐗q\(L\)\\mathbf\{X\}\_\{q\}^\{\(L\)\}, the final target\-attention output𝐘1\(L\)\\mathbf\{Y\}\_\{1\}^\{\(L\)\}, the final debris\-attention output𝐘d\(L\)\\mathbf\{Y\}\_\{d\}^\{\(L\)\}, and the raw lifted inputs are concatenated per point:

𝐙=\[𝐗q\(L\),𝐘1\(L\),𝐘d\(L\),P~0,P~1\]∈ℝB×\(3​h\+16\)×N,\\mathbf\{Z\}=\\bigl\[\\mathbf\{X\}\_\{q\}^\{\(L\)\},\\;\\mathbf\{Y\}\_\{1\}^\{\(L\)\},\\;\\mathbf\{Y\}\_\{d\}^\{\(L\)\},\\;\\tilde\{P\}\_\{0\},\\;\\tilde\{P\}\_\{1\}\\bigr\]\\in\\mathbb\{R\}^\{B\\times\(3h\+16\)\\times N\},\(27\)whereP~0\\tilde\{P\}\_\{0\}andP~1\\tilde\{P\}\_\{1\}are the lifted \(dimension\-8\) inputs, and3​h\+2×8=3​h\+163h\+2\{\\times\}8=3h\+16\. A final two\-layer pointwise MLP projects this to the output:

Conv1d​\(3​h\+16→h\)→Dropout​\(p\)→GELU→Conv1d​\(h→7\)\.\\text\{Conv1d\}\(3h\{\+\}16\\to h\)\\to\\text\{Dropout\}\(p\)\\to\\text\{GELU\}\\to\\text\{Conv1d\}\(h\\to 7\)\.\(28\)The seven output channels per agent correspond to residuals on the five slow elements\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\)and one residual correction to the true\-anomaly rateΔ​ν˙\\Delta\\dot\{\\nu\}, consistent with the trajectory parameterization\.

##### Hyperparameters and Parameter Count

Table[2](https://arxiv.org/html/2608.00320#Sx3.T2)lists the architectural hyperparameters\. With these settings, the approximate parameter counts by component are given in Table[3](https://arxiv.org/html/2608.00320#Sx3.T3)\.

Table 2:Architectural hyperparameters of the neural operator\.Table 3:Approximate parameter counts by component \(h=1024h=1024,L=5L=5\)\.
##### Weight Initialization

All1×11\{\\times\}1convolutional layers are initialized with Kaiming uniform initialization \(a=5a=\\sqrt\{5\}, fan\-in mode\); any standalone linear layers use Xavier uniform\. All biases are initialized to zero\. The final output convolution is not rescaled, so initial outputs are of unit magnitude\.

#### Inference rollout

At inference, the operator is queried on a uniform grid ofK\+1K\{\+\}1time points,τk=k/K\\tau\_\{k\}=k/K, withKKdetermined by the physical time stepΔ​tphys=120​s\\Delta t\_\{\\mathrm\{phys\}\}=120\\ \\text\{s\}and the mission durationTT\. The slow\-element sequence is produced in a single batched forward pass by stacking allK\+1K\{\+\}1time queries, and the true anomalyν\\nuis then integrated sequentially using the recurrence in equation \([1](https://arxiv.org/html/2608.00320#Sx3.E1)\) augmented by the network’s residual correctionΔ​ν˙k\\Delta\\dot\{\\nu\}\_\{k\}\.

#### Training hyperparameters

##### Scenario Sampling

At each training step a single LEO reconfiguration scenario is drawn\. The number of controlled spacecraftN∈\{1,…,10\}N\\in\\\{1,\\dots,10\\\}is sampled uniformly at random\. Initial states are generated as a*dense cluster*: a cluster\-center orbit is drawn uniformly within the bounds given in Table[4](https://arxiv.org/html/2608.00320#Sx3.T4)\(subject to the feasibility constraint that the periapsis exceeds Earth\+\+100 km and the apoapsis lies below Earth\+\+2000 km\), converted to a Cartesian state, and each agent is placed uniformly within a5050m ball of the center in*both*position and a matched velocity spread, so pairwise separations span roughly0–100100m\. This manufactures genuine, unavoidable agent–agent conflict at everyN≥2N\\geq 2; LEO coverage comes from randomizing the cluster center across scenarios rather than from scattering agents, under which agent–agent conflict essentially never arises atN≤10N\\leq 10\.

Table 4:Orbital element sampling bounds used during training\.Targets are per\-agent but*converging*: one maneuver class is drawn per scenario \(minor or major, equal probability\), one deviation vector is sampled at that class’s magnitude, each of the five slow elements perturbed by exactly its class scale with an independent random sign and the terminal true anomaly left free, and the*same*deviation is applied to every agent’s ownP0P\_\{0\}\. Each agent therefore flies an exactly\-1%1\\%or exactly\-10%10\\%maneuver of its own orbit, while the targets form a cluster mirroring the start cluster: an unmodified \(non\-deconflicting\) rollout keeps the phase\-locked formation in conflict all the way to a co\-located arrival, and the escape, a phase\-staggered arrival, is free because the target constrains only the five slow elements\. The mission durationTTis drawn uniformly from\[1​hr,12​hr\]\[1\\ \\text\{hr\},12\\ \\text\{hr\}\]\. Debris consist of one adversarially generated crossing\-orbit object per agent \(M=NM=N\), using the construction described in[Interaction penalties](https://arxiv.org/html/2608.00320#Sx3.SSx6)\.

##### Loss Weights

The complete weighted objective is given by equation \([12](https://arxiv.org/html/2608.00320#Sx3.E12)\), with coefficients given in Table[5](https://arxiv.org/html/2608.00320#Sx3.T5)\. LEO feasibility is maintained by construction rather than by penalty: scenario sampling reflects any target perturbation that would leave the feasible set back toward its interior, and numerical guards \(see Numerical Stability\) protect against ill\-conditioned element configurations\. The terminal and initial losses𝒯\\mathcal\{T\}and𝒮\\mathcal\{S\}are normalized element\-wise by the scale vector\[1\.37,0\.1,π,2​π,2​π\]\[1\.37,\\;0\.1,\\;\\pi,\\;2\\pi,\\;2\\pi\]for\(a,ε,ι,Ω,ω\)\(a,\\varepsilon,\\iota,\\Omega,\\omega\)to account for their differing numerical ranges\.

Table 5:Loss term weights used during training\.
##### Optimizer

The network is trained with Adam using the hyperparameters in Table[6](https://arxiv.org/html/2608.00320#Sx3.T6)\. Gradients are clipped by globalℓ2\\ell\_\{2\}norm before each parameter update\.

Table 6:Optimizer and scheduler hyperparameters\.
##### Numerical Stability

Three mechanisms guard against training instabilities that arise from stiff orbital dynamics and adversarial debris generation\.

1. 1\.EMA spike guard\.An exponential moving average of the loss \(α=0\.99\\alpha=0\.99\) is maintained throughout training\. Any step whose loss exceedsmax⁡\(103​ℒ¯EMA,102\)\\max\(10^\{3\}\\,\\bar\{\\mathcal\{L\}\}\_\{\\mathrm\{EMA\}\},\\;10^\{2\}\)is skipped, and the learning rate is reduced by a factor of0\.90\.9before the next step\.
2. 2\.Non\-finite clamping\.All network outputs and intermediate Cartesian coordinate conversions \(orbital elements to inertial position–velocity for the CPA computation\) are passed throughnan\_to\_numwith zero replacement to prevent gradient corruption from ill\-conditioned element configurations\.
3. 3\.Per\-term gradient validation\.Each loss term is back\-propagated individually with graph retention before the full backward pass, and parameter gradients are checked for non\-finite values\. An infinite or NaN gradient raises a runtime exception and halts training, enabling targeted diagnosis rather than silent divergence\.

## Declarations

Funding This research received no external funding\. Conflict of interest / Competing interests The authors declare no competing interests\. Ethics approval and consent to participate Not applicable; this computational study involves no human participants, animals, or biological materials\. Consent for publication Not applicable\. Data availability The Two\-Line Element catalog used for initial conditions and debris fields is publicly available from Space\-Track \([https://www\.space\-track\.org](https://www.space-track.org/); snapshot of 24 March 2025\)\. The processed ephemeris file, trained model weights, and the Monte Carlo evaluation outputs supporting all tables and figures will be deposited in a public repository with a DOI upon publication, and are available to editors and reviewers on request during assessment\. Materials availability Not applicable\. Code availability The training and evaluation code used to generate all results in this study will be released in a DOI\-minting public repository upon publication, and is likewise available to editors and reviewers on request\. Use of artificial intelligence tools Large language model tools \(Anthropic Claude\) assisted with code development and manuscript editing; all methods, results, analyses, and conclusions were developed and verified by the authors\. No generative artificial intelligence was used to create any figure or image in this manuscript\. Author contributions S\.D\.S\. conceived the study, developed the methodology, implemented the experiments, performed the analysis, and wrote the manuscript\. R\.L\. and S\.M\. supervised the research, contributed key ideas, and revised the manuscript\. Z\.L\. and S\.G\. advised on technical aspects of implementation and contributed to manuscript drafting and revision\. All authors reviewed and approved the final manuscript\.

## References

- \\bibcommenthead
- European Space Agency Space Debris Office \[2026\]European Space Agency Space Debris Office: ESA’s Annual Space Environment Report\. Technical report, ESA/ESOC, Darmstadt, Germany \(2026\)
- Space\.com \[2026\]Space\.com: Every SpaceX Starlink satellite has to dodge a collision almost weekly, and experts fear the worst\. Space\.com\. Reporting SpaceX’s semi\-annual orbital\-safety filing to the FCC, covering December 2025–May 2026 \(2026\)
- Lee and Ho \[2023\]Lee, H\.W\., Ho, K\.: Regional constellation reconfiguration problem: Integer linear programming formulation and Lagrangian heuristic method\. Journal of Spacecraft and Rockets60\(6\), 1828–1845 \(2023\)
- Basu et al\. \[2023\]Basu, H\., Pedari, Y\., Almassalkhi, M\., Ossareh, H\.R\.: Computationally efficient collision\-free trajectory planning of satellite swarms under unmodeled orbital perturbations\. Journal of Guidance, Control, and Dynamics46\(8\), 1548–1563 \(2023\)
- Eren et al\. \[2017\]Eren, U\., Prach, A\., Koçer, B\.B\., Raković, S\.V\., Kayacan, E\., Açıkmeşe, B\.: Model predictive control in aerospace systems: Current state and opportunities\. Journal of Guidance, Control, and Dynamics40\(7\), 1541–1566 \(2017\)
- Chen et al\. \[2024\]Chen, R\., Dong, M\., Bai, Y\., Zhao, Y\., Chen, X\.: Trajectory planning and control of spacecraft avoiding dynamic debris swarm\. Aerospace Science and Technology151, 109273 \(2024\)
- van den Berg et al\. \[2011\]van den Berg, J\., Guy, S\.J\., Lin, M\., Manocha, D\.: Reciprocal n\-body collision avoidance\. In: Robotics Research\. Springer Tracts in Advanced Robotics, vol\. 70, pp\. 3–19\. Springer, Berlin, Heidelberg \(2011\)\.
- Pedari et al\. \[2023\]Pedari, Y\., Basu, H\., Ossareh, H\.R\.: A novel framework for trajectory planning and safe navigation of satellite swarms\. IFAC\-PapersOnLine56\(2\), 547–552 \(2023\)
- Jung and Chung \[2025\]Jung, I\., Chung, D\.: Genetic algorithm\-based approach for improving temporal resolution in constellation operation of national satellites\. International Journal of Aeronautical and Space Sciences26\(1\), 314–326 \(2025\)
- Xu et al\. \[2024\]Xu, L\., Zhang, G\., Qiu, S\., Cao, X\.: Reinforcement learning\-based multi\-impulse rendezvous approach for satellite constellation reconfiguration\. Acta Astronautica224, 325–337 \(2024\)
- Kuhl et al\. \[2025\]Kuhl, W\., Wang, J\., Eddy, D\., Kochenderfer, M\.J\.: Markov decision processes for satellite maneuver planning and collision avoidance\. In: 2025 IEEE Aerospace Conference, pp\. 1–9\. IEEE, Big Sky, MT \(2025\)\.
- An et al\. \[2026\]An, X\., Luo, S\., Zhang, H\., Yang, Q\., Ma, Y\., Wang, B\., Du, J\., Wang, Q\.: Autonomous navigation of intelligent microrobotic swarms in unknown environments\. Nature Machine Intelligence8\(6\), 955–968 \(2026\)
- Bensoussan et al\. \[2013\]Bensoussan, A\., Frehse, J\., Yam, P\.: Mean Field Games and Mean Field Type Control Theory\. SpringerBriefs in Mathematics\. Springer, New York \(2013\)\.
- Wang et al\. \[2022\]Wang, G\., Yao, W\., Zhang, X\., Li, Z\.: A mean\-field game control for large\-scale swarm formation flight in dense environments\. Sensors22\(14\), 5437 \(2022\)
- Guo et al\. \[2019\]Guo, X\., Hu, A\., Xu, R\., Zhang, J\.: Learning mean\-field games\. In: Advances in Neural Information Processing Systems, vol\. 32\. Curran Associates, Inc\., Red Hook, NY \(2019\)\.
- Ruthotto et al\. \[2020\]Ruthotto, L\., Osher, S\.J\., Li, W\., Nurbekyan, L\., Fung, S\.W\.: A machine learning framework for solving high\-dimensional mean field game and mean field control problems\. Proceedings of the National Academy of Sciences117\(17\), 9183–9193 \(2020\)
- Laurière et al\. \[2022\]Laurière, M\., Perrin, S\., Pérolat, J\., Girgin, S\., Muller, P\., Élie, R\., Geist, M\., Pietquin, O\.: Learning in Mean Field Games: A Survey\. Preprint at[https://arxiv\.org/abs/2205\.12944](https://arxiv.org/abs/2205.12944)\(2022\)
- Raissi et al\. \[2019\]Raissi, M\., Perdikaris, P\., Karniadakis, G\.E\.: Physics\-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\. Journal of Computational Physics378, 686–707 \(2019\)
- Karniadakis et al\. \[2021\]Karniadakis, G\.E\., Kevrekidis, I\.G\., Lu, L\., Perdikaris, P\., Wang, S\., Yang, L\.: Physics\-informed machine learning\. Nature Reviews Physics3\(6\), 422–440 \(2021\)
- Zhang et al\. \[2025\]Zhang, Y\., Hu, Y\., Song, Y\., Zou, D\., Lin, W\.: Learning vision\-based agile flight via differentiable physics\. Nature Machine Intelligence7\(6\), 954–966 \(2025\)
- Lu et al\. \[2021\]Lu, L\., Jin, P\., Pang, G\., Zhang, Z\., Karniadakis, G\.E\.: Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators\. Nature Machine Intelligence3\(3\), 218–229 \(2021\)
- Berner et al\. \[2026\]Berner, J\., Liu\-Schiaffini, M\., Kossaifi, J\., Duruisseaux, V\., Bonev, B\., Azizzadenesheli, K\., Anandkumar, A\.: Principled approaches for extending neural architectures to function spaces for operator learning\. Nature Machine Intelligence8, 1173–1181 \(2026\)
- Wang et al\. \[2021\]Wang, S\., Wang, H\., Perdikaris, P\.: Learning the solution operator of parametric partial differential equations with physics\-informed deeponets\. Science Advances7\(40\), 8605 \(2021\)
- Yang et al\. \[2023\]Yang, L\., Liu, S\., Meng, T\., Osher, S\.J\.: In\-context operator learning with data prompts for differential equation problems\. Proceedings of the National Academy of Sciences120\(39\), 2310142120 \(2023\)
- Xiao et al\. \[2025\]Xiao, P\., Zheng, M\., Jiao, A\., Yang, X\., Lu, L\.: Quantum DeepONet: Neural operators accelerated by quantum computing\. Quantum9, 1761 \(2025\)
- Xu et al\. \[2025\]Xu, W\., Han, J\., Lai, R\.: Self\-Supervised Amortized Neural Operators for Optimal Control: Scaling Laws and Applications\. Preprint at[https://arxiv\.org/abs/2512\.24897](https://arxiv.org/abs/2512.24897)\(2025\)
- Huang and Lai \[2025\]Huang, H\., Lai, R\.: Unsupervised solution operator learning for mean\-field games\. Journal of Computational Physics537, 114057 \(2025\)
- Cole et al\. \[2026\]Cole, F\., Wang, D\., Chen, Y\., Lu, Y\., Lai, R\.: In\-Context Operator Learning on the Space of Probability Measures\. Preprint at[https://arxiv\.org/abs/2601\.09979](https://arxiv.org/abs/2601.09979)\(2026\)
- Wächter and Biegler \[2006\]Wächter, A\., Biegler, L\.T\.: On the implementation of an interior\-point filter line\-search algorithm for large\-scale nonlinear programming\. Mathematical Programming106\(1\), 25–57 \(2006\)
- Vallado et al\. \[2006\]Vallado, D\.A\., Crawford, P\., Hujsak, R\., Kelso, T\.S\.: Revisiting spacetrack report \#3\. In: AIAA/AAS Astrodynamics Specialist Conference and Exhibit\. American Institute of Aeronautics and Astronautics, Keystone, CO \(2006\)\.

Similar Articles

Neural Operators for Immersed-Boundary Soft Swimmers Locomotion

arXiv cs.LG

The paper develops neural operator surrogates to predict hydrodynamic fields (velocity, vorticity, pressure) around immersed-boundary soft swimmers, achieving low global relative errors on held-out trajectories while identifying pressure accuracy and physical consistency as areas for further work.

An Exploration of Collision-based Enemy Morphology Generation

arXiv cs.AI

This paper explores three novel approaches for procedurally generating enemy morphologies (body plans and collision information) specifically conditioned on player collision interactions, finding all outperform an evolutionary baseline adapted from robotics.