Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations
Summary
This paper investigates using an 8-billion parameter LLM to improve autonomous cyber defense, then distills its policy into a lightweight 64,910-parameter RL agent, demonstrating feasibility across CybORG scenarios.
View Cached Full Text
Cached at: 08/03/26, 07:32 AM
# Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations
Source: [https://arxiv.org/html/2607.28826](https://arxiv.org/html/2607.28826)
11institutetext:Department of Electrical and Computer Engineering, Royal Military College of Canada, Kingston, Canada22institutetext:Department of Mathematics and Computer Science, Royal Military College of Canada, Kingston, Canada33institutetext:Defence Research and Development Canada, Ottawa, Canada44institutetext:Department of Computer and Software Engineering, Polytechnique Montreal, Montreal, Canada###### Abstract
Autonomous Cyber Operations \(ACO\) have become increasingly important for defending enterprise networks as cyber threats continue to evolve in scale and sophistication\. ACO applications commonly employ Reinforcement Learning \(RL\) agents to learn defensive behaviors through direct interaction with environments\. However, RL agents typically require extensive exploration during training, often resulting in unstable behavior and poor initial decision\-making before converging toward effective defense strategies\.
In this work, we investigate the use of a Large Language Model \(LLM\) to improve autonomous defensive decision\-making within an ACO environment\. Through prompt engineering rather than environment\-specific fine\-tuning, we demonstrate that an 8\-billion parameter LLM pretrained on cybersecurity data can outperform a baseline RL agent in a modified CybORG CAGE Challenge 2 environment\. We then propose an online policy distillation framework that transfers the LLM’s defensive policy into a lightweight RL agent containing only 64,910 parameters, reducing model size by several orders of magnitude while maintaining effective defensive capabilities\. This provides a potential pathway toward operationalizing the defensive capabilities of computationally expensive frontier cybersecurity models within lightweight, operationally deployable agents\.
To evaluate transferability, we additionally construct multiple CybORG scenarios ranging from 4 to 12 hosts and assess the feasibility of the proposed approach across varying network configurations\. Furthermore, we systematically evaluate multiple teacher\-guided RL stabilization strategies and observe that none consistently surpass the optimized teacher policy, suggesting potential policy\-alignment limitations between reward\-driven RL optimization and teacher\-guided defense strategies\.
Our results demonstrate the potential of using cybersecurity\-focused LLMs as sources of expertise for autonomous cyber defense, while policy distillation provides a practical path toward operationalizing the defensive capabilities of frontier cybersecurity models within computationally efficient and scalable agents\.
utonomous Cyber Operations \(ACO\); Policy Distillation; Teacher\-Guided Learning; Behavioral Cloning
## 1Introduction
Reinforcement Learning \(RL\) has emerged as a promising approach for countering continuously evolving adversarial cyber threats\. Unlike traditional Machine Learning \(ML\) techniques, which often require maintaining large labeled datasets that must continually be updated to remain effective against new attack strategies, RL enables agents to learn defensive behaviors through direct interaction with an environment, reducing dependence on manually curated datasets\. However, traditional RL approaches present several challenges in the cybersecurity domain\. In particular, RL agents often require long training times and must initially perform unfavorable actions in order to learn from their consequences\. In cybersecurity environments, where incorrect decisions may result in severe operational impacts such as the compromise of an enterprise network, this exploratory learning strategy can be problematic\.
Previous work has explored integrating a teacher into the RL pipeline, where an external entity with prior knowledge of the field initially guides the agent’s actions\. As training progresses and predefined criteria are met \(e\.g\., a fixed number of episodes\), the teacher’s influence is gradually reduced until the agent learns independently through environmental feedback\. While this approach can significantly accelerate learning and reduce harmful exploration, it introduces an important limitation: complex tasks like cybersecurity may contain multiple favorable policies\. Consequently, the teacher’s policy may not align with the policy ultimately reinforced through environmental rewards\. Even when both policies are effective, this misalignment may negatively impact training stability and overall agent performance\.
One potential solution is to focus directly on learning the teacher’s policy, rather than jointly optimizing through the teacher’s guidance and environmental feedback\. For this approach to be effective, the teacher itself must demonstrate strong defensive decision\-making capabilities\. Large Language Models \(LLMs\), particularly those pre\-trained on cybersecurity\-related data, provide a promising candidate for this teacher role due to their flexible input format, embedded domain knowledge, and reasoning capabilities\. However, LLMs typically contain significantly more parameters than lightweight RL agents, resulting in increased computational requirements and slower inference times\. Recent advances in autonomous cyber defense increasingly rely on these large frontier cybersecurity models capable of advanced contextual reasoning\. However, continuously deploying such models in operational environments remains computationally expensive due to their large parameter counts, increased inference latency, and infrastructure requirements\. Consequently, an important open problem is whether the defensive policies of these models can be distilled into compact and computationally efficient agents suitable for scalable and resource\-constrained deployments\.
In this work, we distill the policy of an LLM pretrained on cybersecurity\-related data into a lightweight RL agent and demonstrate that the resulting agent outperforms an RL agent trained solely through interaction with a modified version of the CybORG CAGE Challenge 2 environment\[[21](https://arxiv.org/html/2607.28826#bib.bib28)\]\. Rather than relying on environment\-specific fine\-tuning, we optimize the LLM’s defensive behavior through prompt engineering and subsequently transfer this policy into a lightweight RL agent\. Furthermore, we evaluate the transferability of the proposed approach across multiple CybORG scenarios ranging from 4 to 12 hosts\. We additionally investigate whether teacher\-guided RL agents can consistently surpass the optimized teacher policy when transitioning toward independent reinforcement learning\. The major contributions for this work are summarized as follows:
- •LLM\-to\-RL policy distillationWe propose a method to efficiently distill the knowledge of an 8 billion parameter cybersecurity\-focused LLM into a lightweight 64,910 parameter RL agent while maintaining comparable defensive performance\.
- •Transferability evaluation across network topologies\.We construct multiple simulated network environments and evaluate our proposed approach across varying CybORG scenarios to assess its transferability\.
- •Teacher\-guided RL stabilization analysis\.We systematically evaluate multiple teacher\-guided RL stabilization strategies and demonstrate that none consistently surpass the optimized teacher policy, highlighting potential policy\-alignment limitations between reward\-driven and teacher\-driven defense policies\.
The rest of this paper is organized as follows: Section[2](https://arxiv.org/html/2607.28826#S2)discusses related work in ACO, RL, and teacher\-guided learning approaches\. Section[3](https://arxiv.org/html/2607.28826#S3)presents the methodology used to implement the proposed solution\. Section[4](https://arxiv.org/html/2607.28826#S4)evaluates the performance and transferability of the proposed approach across multiple CybORG environments\. Finally, Section[5](https://arxiv.org/html/2607.28826#S5)concludes the paper and discusses future research directions\.
To support reproducibility, the source code, configurations, prompts, and experimental setup are available in the accompanying GitHub repository at https://github\.com/Poly\-AIvsAI/LLMDistillationACO\.
## 2Related Work
In this section, we discuss previous work in ACO, teacher\-guided RL, knowledge distillation and LLMs\. We then synthesize our findings to motivate the proposed approach and identify the research gap addressed in this work\.
### 2\.1ACO and RL
Researchers have increasingly explored RL for applications in ACO, where agents learn defense strategies by directly interacting with their respective environment\[[19](https://arxiv.org/html/2607.28826#bib.bib26)\]\. For RL to be effective in cybersecurity domains, the environment must both realistically model cyber operations and provide meaningful signals that enable the agent to converge towards effective policies\. CybORG satisfies these requirements, by providing a cybersecurity environment where blue agents can be trained to defend a simulated enterprise network\[[3](https://arxiv.org/html/2607.28826#bib.bib19),[8](https://arxiv.org/html/2607.28826#bib.bib29)\]\.
Previous work, such as that conducted by Palmer et al\., Wiebe et al\., and McDonald et al\. demonstrate the feasibility of training RL agents within the CybORG environment\[[13](https://arxiv.org/html/2607.28826#bib.bib11),[27](https://arxiv.org/html/2607.28826#bib.bib12),[11](https://arxiv.org/html/2607.28826#bib.bib13)\]\. However, these approaches typically initialize agents without prior domain knowledge, requiring them to learn entirely through environmental exploration\. Consequently, agents often exhibit poor initial performance and require substantial training time before converging toward effective defensive policies\. In cybersecurity, where incorrect actions may have a detrimental impact on security posture, this exploratory learning process can directly impede ACO’s practical adoption\.
### 2\.2Teacher\-Guided RL
Teacher\-guided RL has emerged as a promising approach for mitigating the inherent limitations associated with learning from scratch\. One common strategy involves using a teacher model to guide the student’s early learning process before gradually reducing the teacher’s influence over time\. For example, M\. Pfeiffer et al\., proposed generating a labeled dataset using the teacher to train the student prior to transitioning to conventional RL training\[[14](https://arxiv.org/html/2607.28826#bib.bib22)\]\. More recent work incorporated the teacher’s guidance directly within the RL environment through approaches such as reward shaping, feature space modification, action guidance, and auxiliary loss signals rather than through a separated pretraining phase\[[4](https://arxiv.org/html/2607.28826#bib.bib24),[26](https://arxiv.org/html/2607.28826#bib.bib23),[20](https://arxiv.org/html/2607.28826#bib.bib1)\]\.
These approaches demonstrated that teacher guidance can accelerate convergence and improve initial agent performance\. However, they still rely on environmental rewards as the primary optimization objective\. In complex domains such as cybersecurity, multiple viable defensive policies may exist for achieving favorable outcomes\. Consequently, the teacher’s policy may not align with the policy ultimately reinforced through environmental feedback\. Even when both policies are effective, this mismatch may require the student to partially unlearn the teacher’s guidance during later training stages, potentially destabilizing learning and reducing the effectiveness of teacher\-guided training\.
### 2\.3Knowledge Distillation
One potential approach for mitigating possible teacher\-policy misalignment is to derive the student’s training directly from the teacher rather than jointly optimizing through environmental rewards\. This concept is not novel in itself and has previously been explored through various forms of knowledge distillation and imitation learning across multiple domains\[[16](https://arxiv.org/html/2607.28826#bib.bib2),[15](https://arxiv.org/html/2607.28826#bib.bib3),[28](https://arxiv.org/html/2607.28826#bib.bib4),[1](https://arxiv.org/html/2607.28826#bib.bib5)\]\. Prior work like those done by Agarwal et al\. and Pozzi et\. al demonstrated that complex teacher models can effectively transfer knowledge into smaller and computationally efficient student models\[[1](https://arxiv.org/html/2607.28826#bib.bib5),[15](https://arxiv.org/html/2607.28826#bib.bib3)\]\.
However, these approaches typically do not focus on online interaction with dynamic RL environments\. This presents an opportunity to investigate whether policy distillation can be performed during active interaction with a cybersecurity environment while deriving learning signals directly from the teacher policy rather than environmental reward signals\. Under this paradigm, the environment continues providing state transitions and observations, while the teacher acts as the primary source of supervisory guidance\.
Because the student no longer directly optimizes against environmental rewards, the effectiveness of the overall approach becomes heavily dependent on the quality and consistency of the teacher policy\. Consequently, the teacher must demonstrate strong defensive decision\-making capabilities, ideally without requiring extensive environment\-specific fine\-tuning\.
### 2\.4LLMs
LLMs have shown significant promise in cybersecurity due to their ability to recognize complex patterns in text and generate contextually relevant responses\[[6](https://arxiv.org/html/2607.28826#bib.bib8),[2](https://arxiv.org/html/2607.28826#bib.bib9),[9](https://arxiv.org/html/2607.28826#bib.bib10)\]\. Their large parameter size and extensive pretraining enable them to encode substantial domain knowledge, making them promising candidates for acting as teachers in complex cybersecurity environments\.
However, directly deploying LLMs in operational settings presents several challenges\. Their large model sizes require substantial computational resources and typically introduce increased inference times, limitations that are particularly problematic in time\-sensitive cybersecurity environments\. This motivates investigating whether the capabilities of an LLM can be distilled into a significantly smaller and more efficient RL agent while preserving strong defensive performance\.
To the best of our knowledge, no existing LLM has been specifically trained for the CybORG environment\. One possible solution would be to follow prior teacher\-guided RL approaches where a generalized LLM initially guides training before the agent transitions to learning directly from environmental rewards\[[20](https://arxiv.org/html/2607.28826#bib.bib1),[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\. However, this may reintroduce the previously discussed policy\-alignment challenges\. Another option would involve fine\-tuning the LLM directly for CybORG; however, this process is computationally expensive and difficult to scale within rapidly evolving cybersecurity environments\.
Prompt engineering presents a computationally efficient alternative to fine\-tuning, where the underlying model parameters remain fixed while the input prompts are optimized to better align the model with the target environment\[[10](https://arxiv.org/html/2607.28826#bib.bib15)\]\. Previous work has already demonstrated that prompt engineering can substantially improve LLM performance without modifying model parameters\[[18](https://arxiv.org/html/2607.28826#bib.bib7)\]\. For example, Son et al\. improved Llama3’s TruthfulQA performance from 3\.2 to 13\.65 BLEU using Retrieval\-Augmented Generation \(RAG\), representing a 323% relative improvement\[[18](https://arxiv.org/html/2607.28826#bib.bib7)\]\.
### 2\.5Discussion
Prior work demonstrated the feasibility of applying RL to ACO environments such as CybORG; however, these approaches typically require agents to learn entirely through exploration, resulting in poor initial performance and extended convergence times\[[11](https://arxiv.org/html/2607.28826#bib.bib13),[13](https://arxiv.org/html/2607.28826#bib.bib11),[27](https://arxiv.org/html/2607.28826#bib.bib12)\]\. Teacher\-guided RL approaches helped mitigate some of these limitations, but continued reliance on environmental reward signals may introduce policy\-alignment challenges between the teacher and student\[[20](https://arxiv.org/html/2607.28826#bib.bib1),[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\. Existing work also demonstrated the potential of leveraging LLMs for cybersecurity\-related reasoning and decision\-making\[[6](https://arxiv.org/html/2607.28826#bib.bib8),[2](https://arxiv.org/html/2607.28826#bib.bib9),[9](https://arxiv.org/html/2607.28826#bib.bib10)\]\.
Collectively, these observations motivate investigating whether a large frontier cybersecurity model can be optimized for CybORG through prompt engineering and subsequently distilled into a lightweight RL agent capable of operating more efficiently within simulated enterprise networks\. Furthermore, the potential alignment limitations of teacher\-guided RL suggest that directly distilling expert\-guided defense strategies may provide a more stable alternative to jointly optimizing through environmental reward signals\. Finally, previous work in CybORG has largely focused on single simulated environments, creating an opportunity to evaluate transferability across varying network topologies to assess the robustness of the proposed methodology\.
## 3Methodology
This section outlines the phased approach we used in our work\. These phases include:
1. 1\.Distillation\. We distill the LLM’s knowledge into a lightweight RL agent while it actively interacts with the environment\. Action masking is employed to ensure the RL agent always performs adequately while a loss signal derived from the teacher’s feedback is used to optimize its underlying parameters\.
2. 2\.Prompt engineering\. We use an LLM pretrained on cybersecurity that has already been evaluated in CybORG as the frontier\-model teacher\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\. We then optimize the prompt using a constrained chain\-of\-thought template with reasoning rules and output constraints\.
3. 3\.Transferability\. We create 9 simulated scenarios ranging from 4 to 12 hosts to assess the transferability of our solution\.
### 3\.1Distillation
The distillation process we used employed action masking to ensure the RL agent always performed consistently with the teacher\[[20](https://arxiv.org/html/2607.28826#bib.bib1)\]\. In particular, we manually set the probability of selecting any action not recommended by the teacher to 0:
πmasked\(at\)=Mt\(at\)∑a′Mt\(a′\)\\pi\_\{\\mathrm\{masked\}\}\(a\_\{t\}\)=\\frac\{M\_\{t\}\(a\_\{t\}\)\}\{\\sum\_\{a^\{\\prime\}\}M\_\{t\}\(a^\{\\prime\}\)\}\(1\)
whereMt\(at\)M\_\{t\}\(a\_\{t\}\)is the masking matrix where every action not recommended by the teacher is 0 and the recommended action is set to 1\. We divide by the sum of all elements in the masking matrix∑a′Mt\(a′\)\\sum\_\{a^\{\\prime\}\}M\_\{t\}\(a^\{\\prime\}\)to ensure that a valid probability distribution is maintained\.
Masking ensures that there is an immediate performance improvement to our RL agent, but it does not update the RL agent’s underlying parameters\[[20](https://arxiv.org/html/2607.28826#bib.bib1)\]\. To distill the LLM’s policy into the RL agent itself, we incorporated a loss signal derived from the LLM’s recommendation\. In particular, we computed the loss signal as:
Lteacher\(θ\)=−log\(πθ\(atteacher\|st\)\)L^\{teacher\}\(\\theta\)=\-log\(\\pi\_\{\\theta\}\(a\_\{t\}^\{teacher\}\|s\_\{t\}\)\)\(2\)
whereLteacher\(θ\)L^\{teacher\}\(\\theta\)is the teacher\-guided loss function, andlog\(πθ\(atteacher\|st\)\)log\(\\pi\_\{\\theta\}\(a\_\{t\}^\{teacher\}\|s\_\{t\}\)\)is the log probability of selecting the teacher’s recommendation in the current policy\. If the agent is likely to select the teacher’s recommendation in its current policy, the loss is low, whereas if it is unlikely to select the teacher’s recommendation, loss will increase exponentially, encouraging the policy to align more closely with that of the teacher’s\.
We continued this distillation process using the teacher\-derived loss signal for 240 episodes before switching to RL\-agent only action selection\.
### 3\.2Prompt Engineering
We further refine the basic prompt engineering employed in previous work to better align the LLM with the operational constraints and decision\-making requirements of the modified CybORG environment\[[24](https://arxiv.org/html/2607.28826#bib.bib16),[22](https://arxiv.org/html/2607.28826#bib.bib6),[21](https://arxiv.org/html/2607.28826#bib.bib28)\]\.
For this, we employed a combination of role grounding, zero\-shot learning, rule semantics, and chain\-of\-thought scaffolding\[[12](https://arxiv.org/html/2607.28826#bib.bib17),[13](https://arxiv.org/html/2607.28826#bib.bib11),[7](https://arxiv.org/html/2607.28826#bib.bib18)\]\. We refer to this approach as chain\-of\-thought scaffolding instead of the well\-known chain\-of\-thought reasoning because we do not simply tell the LLM to explain its sequence, but instead provide a structured high\-level reasoning process intended to guide defensive decision\-making\. The complete prompt can be found in the accompanying GitHub repository\[[23](https://arxiv.org/html/2607.28826#bib.bib27)\]\.
To preserve operational realism and increase generalization, the information included in the prompt was intentionally constrained to information that would realistically be available to human defenders\. For example, privileged simulator state and adversarial tactics, techniques and procedures \(TTPs\) were not included in the prompt\.
The same procedure as done in Tholl et al\. was used to dynamically generate information for the prompt using CybORG’s state space as well as extracting the LLM’s recommended action using a combination of regex against the existing actions and falling back to semantic similarity if regex failed\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\. The high\-level process for creating the prompt and extracting an action from the LLM is shown in Fig\.[1](https://arxiv.org/html/2607.28826#S3.F1)\.
Figure 1:Overview of transforming CybORG’s state space into a coherent prompt and extracting an executable action from the LLM\.
### 3\.3Transferability
To assess the transferability of the proposed approach across different environments, we created nine additional scenarios ranging in complexity from 4 to 12 hosts\. A similar B\-line agent was used on the red side to infiltrate the networks, with minor modifications to make it function for the various simulated networks\[[3](https://arxiv.org/html/2607.28826#bib.bib19)\]\.
To keep the assessment fair, the same 8 billion parameter Cyber Risk Llama LLM was used with identical prompt structures\[[24](https://arxiv.org/html/2607.28826#bib.bib16)\]\. The only hyperparameter that changed throughout was the time at which we stopped distilling the LLM’s knowledge into the RL agent, as it was found that less time for knowledge transfer was required for smaller network topologies\.
## 4Evaluation
In this section, we present, evaluate, and interpret the results of our work with respect to their implications for autonomous cyber defense and RL\-guided defensive learning\. In particular, this section covers:
- •Prompt engineering\.Evaluating our optimized prompt against the baseline used in\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\.
- •Distillation\.Assessing the feasibility of distilling the teacher’s performance into an RL agent and comparing it against a baseline RL agent and the teacher\-guided approach used in\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\.
- •Learning stabilization\.Investigating whether teacher\-guided RL strategies can consistently surpass the optimized teacher policy during independent reward\-driven learning\.
- •Transferability\.We evaluate our solution’s performance across multiple simulated networks of varying complexity ranging from 4 to 12 hosts\.
### 4\.1Prompt Engineering
As discussed in Section[3](https://arxiv.org/html/2607.28826#S3), we further refined the prompt from what was used in Tholl et\. al’s work\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\. This was an iterative process of incrementally modifying the chain\-of\-thought scaffolding to guide the LLM’s decision\-making until superior performance was observed\. We present the performance of the optimized prompt versus the standard one in Figure[2](https://arxiv.org/html/2607.28826#S4.F2)\.
The metric used to evaluate the performance of the prompts is the reward obtained by the LLM with respect to CybORG’s reward signals\. We can see that using the optimized prompt yields an average reward of roughly 70, corresponding to an approximate performance increase of 35% with a notably lower standard error compared to the baseline prompt used previously\. The constrained chain\-of\-thought scaffolding likely improved policy consistency by reducing ambiguous action selection and explicitly structuring the LLM’s defensive prioritization process around operationally relevant features such as host criticality and suspicious activity\.
Figure 2:Evaluation of the Standard and Optimized Prompt across 10 independent runs for 50 episodes\. Per\-episode mean reward with a ±1 standard error\.
### 4\.2Distillation
After establishing an optimal prompt, we then proceeded to distill the 8 billion parameter LLM’s policy into the lightweight RL agent using the methodology discussed in Section[3](https://arxiv.org/html/2607.28826#S3)\. We compare our results against a baseline Proximal Policy Optimization \(PPO\) agent, the same on\-policy RL algorithm that has repeatedly shown success in previous ACO work\[[27](https://arxiv.org/html/2607.28826#bib.bib12),[11](https://arxiv.org/html/2607.28826#bib.bib13)\]\. We also compare the distilled agent’s performance against a teacher\-guided agent using the same teacher\-derived loss signal to guide initial training\. The difference is that rather than solely learning from the teacher, we decrease the influence of the teacher’s auxiliary loss signal and gradually transition to learning solely from the environment’s feedback\.
We can see in Figure[3](https://arxiv.org/html/2607.28826#S4.F3)that the distilled agent \(without post\-distillation learning\) outperforms both the baseline and teacher\-guided agent\. We completely remove the teacher’s guidance for our distilled agent by episode 240, while gradually decreasing it for the teacher\-guided agent\. While the teacher\-guided agent has comparable initial performance, it deteriorates noticeably and eventually converges with the baseline when learning from the environmental feedback, resulting in a policy misalignment between the LLM policy and the policy learned solely from CybORG’s signals\.
Our next question was whether the baseline PPO agent would ever surpass the performance of our distilled agent\. For this, we ran our baseline for 50,000 episodes \(1\.6 million timesteps\) as shown in Fig\.[4](https://arxiv.org/html/2607.28826#S4.F4)\.
From the Figure, we can see that while individual runs surpass our distilled agent’s performance, the mean performance never appears to stabilize beyond our agent\. There are points after 23,000 episodes where the mean performance of the baseline temporarily surpasses our distilled agent, but the performance does not remain stable throughout training, whereas our distilled agent remains consistently stable throughout the 50,000 episodes\.
Figure 3:Comparing distillation with teacher\-guided RL\. Per\-episode mean reward with a ±1 standard error for 2,000 episodes\.Figure 4:Comparing the distilled agent against the baseline PPO agent across 10 independent runs using a 10\-episode running average with a ±1 standard error for 50,000 episodes\.
### 4\.3Transferability
To assess the transferability of our solution, we created 9 additional CybORG environments ranging from 4\-12 hosts, ensuring that we modified the red agent for each respective environment to optimize its attack trajectory\.
For this implementation, we also incorporated action masking during the distillation phase to ensure the agent’s performance remained consistent with the teacher policy from the beginning\. The LLM’s prompt and every hyperparameter were kept constant with the exception of the transition point between distillation and independent RL, with environments containing more hosts requiring a longer distillation phase\. The results of our transferability evaluation can be found in Fig\.[7](https://arxiv.org/html/2607.28826#Pt0.A1.F7)under Appendix[0\.A](https://arxiv.org/html/2607.28826#Pt0.A1)\.
From Fig\.[7](https://arxiv.org/html/2607.28826#Pt0.A1.F7), we can see that our methodology for distilling an LLM’s policy into a lightweight RL agent transfers reasonably well across the evaluated CybORG environments\. The 7 host environment represents an exception, where the baseline RL agent starts to outperform the distilled agent by roughly episode 700 with respect to mean performance\. The 5 host, 6 host, 8 host, and 9 host scenarios all show that the baseline performs similarly to the LLM by roughly episode 1250; with the baseline showing similar performance by episode 500 for the 5 host environment with respect to mean\-performance\. It should be noted that the standard error for the distilled agent is much smaller than the baseline for each of the scenarios, demonstrating lower variability across runs, a property that is desirable in operational cybersecurity environments where stable defensive behavior is important\.
### 4\.4Improving Post\-Learning Performance
As shown in Fig\.[4](https://arxiv.org/html/2607.28826#S4.F4), when we gradually transition to environmental learning, the performance degrades from the baseline LLM\. As such, we attempted various techniques to rectify this drop in performance from teacher\-guided to independent RL\. Specifically, seven independent techniques were employed:
- •Adding the LLM’s recommendation to the critic loss\.PPO is an on\-policy RL algorithm which contains a critic network that computes a scalar reward signal for helping guide the policy\[[17](https://arxiv.org/html/2607.28826#bib.bib20)\]\. Here, we incorporate the LLM’s guidance not only as an auxiliary loss signal for the actor network, which computes the policy directly, but for the critic network as well\. Because the learning dynamic is entirely different for the critic network, we opted to integrate the teacher’s impact by making it train more strongly on states that were produced from teacher\-recommended actions\. In particular, we used the following: LLLM\(ϕ\)=𝔼\[1n∑i=1nMt\(ai\)∗\(reti−Vϕ\(si\)\)2\]L\_\{LLM\}\(\\phi\)=\\mathbb\{E\}\[\\frac\{1\}\{n\}\\sum^\{n\}\_\{i=1\}M\_\{t\}\(a\_\{i\}\)\*\(ret\_\{i\}\-V\{\\phi\}\(s\_\{i\}\)\)^\{2\}\]\(3\) wherennis the number of samples in the batch,Mt\(ai\)M\_\{t\}\(a\_\{i\}\)is the masking matrix applied, setting the action to 1 if it came from the LLM recommendation, and 0 otherwise\. The last two terms are standard in PPO critic learning:retiret\_\{i\}are the standard returns computed using the Generalized Advantage Estimate \(GAE\) andVϕ\(si\)V\{\\phi\}\(s\_\{i\}\)being the critic network’s output for the current state\[[17](https://arxiv.org/html/2607.28826#bib.bib20)\]\. We then combined this auxiliary loss term with the critic’s original loss to incorporate the LLM into the critic network’s learning process\. L\(ϕ\)=σLenv\(ϕ\)\+\(1−σ\)Lteacher\(ϕ\)L\(\\phi\)=\\sigma L^\{env\}\(\\phi\)\+\(1\-\\sigma\)L^\{teacher\}\(\\phi\)\(4\) whereσ\\sigmais gradually increased, reducing the teacher’s impact as the agent transitions to independent learning\.
- •Initializing the critic network with a pretrained one\. To validate whether the problem was the critic potentially lagging behind the actor network, we attempted to initialize the critic network with one that was trained for 10,000 episodes using the baseline PPO agent\.
- •Dynamically adjusting the actor and critic learning rates \(LRs\)\.We dynamically adjusted the critic and actor learning rates to facilitate a smoother transition from teacher\-guided to independent RL\. In particular, we increased the critic learning rate from1\.6e−31\.6e^\{\-3\}to3\.2e−33\.2e^\{\-3\}to estimate the returns of the LLM\-recommended actions, then gradually decreased it to0\.8e−30\.8e^\{\-3\}during the transition to independent RL\. The actor network’s learning rate was also decayed from1\.6e−31\.6e^\{\-3\}to0\.8e−30\.8e^\{\-3\}during the transition\.
- •Adding extra critic epochs\.We attempted to double the epochs for the critic network during the transition from teacher\-guided to independent RL to enable the PPO agent to quickly evaluate any new, unseen states\.
- •Incorporating the LLM’s guidance as a distribution\.Instead of computing the loss signal as the probability of selecting the LLM’s single recommendation in the RL agent’s policy, we mapped the LLM’s raw logits into a distribution across possible actions and used KL divergence to compute the loss\[[5](https://arxiv.org/html/2607.28826#bib.bib21)\]\.
- •Stopping critic learning\.We ceased the learning of the critic network only during the transition to independent RL, while keeping the actor network\.
- •Decay by a multiplicative factor\.Instead of linearly decreasing the impact the teacher has on the agent’s training, we decreased it by a multiplicative factor\. In particular, we used a multiplicative decay factor of 0\.99 per training interval instead of a linear subtraction\.
We present the results of the seven techniques discussed above in Figs\.[5](https://arxiv.org/html/2607.28826#S4.F5)and[6](https://arxiv.org/html/2607.28826#S4.F6)\.
Figure 5:First four attempts to increase performance beyond the LLM with the optimized prompt\. Techniques include: adding the LLM’s feedback to the critic’s loss, initializing the critic as a pretrained model over 10,000 episodes, dynamically changing the actor and critic LRs, and adding extra learning epochs to the critic during the transition from teacher\-guided to independent RL\. Plots show a mean reward after applying a 10\-episode running average with a ±1 standard error over 2,000 episodes\.Figure 6:Next three attempts to increase performance beyond the LLM with the optimized prompt\. Techniques include: incorporating the LLM’s guidance as a distribution \(instead of a single action\), stopping the critic learning after the transition to independent RL, and decaying the teacher’s impact by a multiplicative factor instead of a linear constant\. Plots show a mean reward after applying a 10\-episode running average with a ±1 standard error over 2,000 episodes\.From Fig\.[5](https://arxiv.org/html/2607.28826#S4.F5), we can see that adding the LLM to the critic loss, initializing the critic as a pretrained network, and adding extra training epochs to the critic network exhibit the same behavior, quickly converging to the teacher’s performance and then declining at roughly the same rate until they converge with the baseline by roughly episode 2,000\. Dynamically changing the learning rates for the critic and actor network slow down the decline from teacher performance, simply because we’re decreasing the extent to which the policy can deviate from the teacher during the transition to independent RL\.
In Fig\.[6](https://arxiv.org/html/2607.28826#S4.F6), stopping the critic learning after the transition to independent RL yields a very noticeable drop in performance from episodes 350 to 500, likely due to being unable to properly quantify any state not directly observed during the teacher\-guided phase\. Gradually decaying the LLM’s influence by a multiplicative factor instead of a linear constant shows similar performance to modifying the learning rate; the performance decreases during the transition to independent RL, just at a slower rate\. Finally, computing the LLM’s auxiliary loss signal by mapping its raw logits into a distribution over the action space yields poorer initial performance as the policy is now sampled from a distribution rather than directly from the LLM’s highest\-confidence recommendation; however, this approach shows increased performance after the transition to independent RL, but fails to converge to the LLM’s baseline performance\.
Overall, none of the seven solutions discussed above to stabilize the teacher\-guided RL techniques discussed in the work by Tholl et\. al consistently surpassed the LLM’s baseline performance\[[22](https://arxiv.org/html/2607.28826#bib.bib6)\]\.
### 4\.5Discussion
We have shown that by shifting focus to prompt engineering, a pretrained and generalized cybersecurity\-focused LLM can outperform RL agents specifically trained within the CybORG environment\. Furthermore, we demonstrated that the LLM’s superior defensive policy can be distilled into a lightweight RL agent in only 240 episodes\. These findings suggest that policy distillation may provide a practical pathway for compressing the defensive policies of large cybersecurity\-focused frontier models into lightweight agents capable of operating under significantly reduced computational constraints\. This is particularly valuable in operational cybersecurity settings, where continuous deployment of large models can be impractical due to increased latency, infrastructure demands, or resource limitations\.
We also investigated whether an RL agent guided by the LLM could eventually surpass the teacher policy through independent reward\-driven learning by implementing seven modifications to existing teacher\-guided RL techniques discussed in previous work; however, none of these approaches consistently surpassed the optimized teacher policy\[[20](https://arxiv.org/html/2607.28826#bib.bib1),[22](https://arxiv.org/html/2607.28826#bib.bib6),[25](https://arxiv.org/html/2607.28826#bib.bib25),[26](https://arxiv.org/html/2607.28826#bib.bib23),[4](https://arxiv.org/html/2607.28826#bib.bib24)\]\. Collectively, these findings suggest that environmental reward optimization within CybORG may encourage policies that diverge from expert\-guided defensive behavior, even when the resulting policies achieve competitive reward performance\.
## 5Conclusion
Reinforcement Learning \(RL\) has demonstrated significant promise for Autonomous Cyber Operations \(ACO\); however, conventional RL approaches remain limited by their dependence on environmental reward signals and their requirement to learn effective defense strategies through exploration from initially untrained policies\[[27](https://arxiv.org/html/2607.28826#bib.bib12),[11](https://arxiv.org/html/2607.28826#bib.bib13),[3](https://arxiv.org/html/2607.28826#bib.bib19)\]\. While teacher\-guided learning has mitigated some of these concerns, they remain susceptible to policy misalignment between the teacher and the policies reinforced through environmental feedback\.
In this work, we investigated an alternative paradigm, focused on directly optimizing the teacher’s policy, and distilling that policy into a lightweight RL agent\. Specifically, we demonstrated that a cybersecurity\-focused Large Language Model \(LLM\), optimized through prompt engineering rather than environment\-specific fine\-tuning, can outperform baseline RL agents within CybORG\. We further showed that this performance can be distilled into a compact RL agent several orders of magnitude smaller than the LLM with no noticeable drop in performance across multiple simulated environments\.
### 5\.1Contributions
This study has made several contributions to the fields of autonomous cyber defense, RL, and LLM\-guided cybersecurity:
- •Online LLM\-to\-RL policy distillation\.We proposed an online\-distillation framework capable of transferring the policy of an 8\-billion parameter LLM into a lightweight RL agent containing 64,910 parameters \(approximately 0\.0008% of the teacher model\)\.
- •Transferability evaluation across network topologies\.We evaluated the distilled policy across multiple CybORG scenarios ranging from 4 to 12 hosts, demonstrating that the proposed methodology transfers reasonably across the evaluated CybORG network topologies\.
- •Systematic evaluation of teacher\-guided stabilization techniques\.We implemented and evaluated multiple teacher\-guided RL stabilization strategies and showed that none consistently surpass the optimized teacher policy, highlighting potential policy\-alignment limitations between the environment’s reward signals and the teacher’s defense behavior\.
### 5\.2Limitations
Although this work has made meaningful contributions to autonomous cyber defense, several limitations should be considered:
- •Environment realism\.While CybORG provides a valuable benchmark for ACO research, it remains a simulated abstraction of operational enterprise environments\. Simplified attacker behavior, constrained action and observation spaces, reduced environmental noise, and limited operational complexity restrict the direct applicability of learned defense policies to real cybersecurity settings\.
- •Reward\-based evaluation metrics\.The primary metric used to evaluate the success of this work was CybORG’s scalar reward signal\. Although useful for standardized benchmarking, these rewards abstract the agent’s underlying behavior, and may not fully capture optimal defensive behavior\.
- •Dependence on teacher\.The proposed framework distills behavior from a static LLM policy, rather than discovering novel defensive strategies\. Consequently, any deficiencies inherent to the teacher, including hallucinations and sub\-optimal decision\-making, may propagate to the distilled agent\.
### 5\.3Future Work
The limitations identified in this work motivate several promising research directions:
- •Reward signal misalignmentFuture work should investigate whether current cybersecurity reward structures adequately capture desirable defense behavior and explore alternative evaluation paradigms to quantify the effectiveness of policies beyond scalar rewards\.
- •Multi\-teacher distillation\.The current framework distills knowledge from a single LLM policy\. Incorporating multiple teacher models may improve policy robustness and enhance generalization across different cybersecurity environments\.
- •Reasoning\-aware distillation\.The work distills action\-selection behavior\. Future approaches could incorporate intermediate reasoning representations or knowledge\-graph reasoning to improve policy transfer fidelity\.
- •Adversarial robustness\. While this work focused on autonomous cyber defense, it did not investigate attacks against either the teacher or student models themselves\. Future work should evaluate policy robustness against adversarial techniques such as observation poisoning and prompt injection\.
- •Operational Realism\.Future environments should incorporate operational realism through partial observability, increased benign activity, richer host artifacts, and more adaptive attacker behavior\.
Overall, this work suggests that teacher\-guided policy distillation may offer a viable pathway toward operationalizing the defensive capabilities of large frontier cybersecurity models through lightweight, resource\-efficient, and scalable autonomous cyber\-defense agents\.
## References
- \[1\]R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. Ramos, M\. Geist, and O\. Bachem\(2024\)ON\-POLICY DISTILLATION OF LANGUAGE MODELS: LEARNING FROM SELF\-GENERATED MISTAKES\.In12th International Conference on Learning Representations, ICLR 2024, May 7, 2024 \- May 11, 2024,12th International Conference on Learning Representations, ICLR 2024,Hybrid, Vienna, Austria,pp\. et al; Google Deepmind; Google Research; Meta; Microsoft\.Note:CompendexCited by:[§2\.3](https://arxiv.org/html/2607.28826#S2.SS3.p1.1)\.
- \[2\]T\. Ali and P\. Kostakos\(2023\-09\)HuntGPT: Integrating Machine Learning\-Based Anomaly Detection and Explainable AI with Large Language Models \(LLMs\)\.arXiv\(en\)\.Note:arXiv:2309\.16021 \[cs\]External Links:[Link](http://arxiv.org/abs/2309.16021),[Document](https://dx.doi.org/10.48550/arXiv.2309.16021)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p1.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1)\.
- \[3\]C\. Baillie, M\. Standen, J\. Schwartz, M\. Docking, D\. Bowman, and J\. Kim\(2020\-02\)CybORG: An Autonomous Cyber Operations Research Gym\.arXiv\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2002.10667)Cited by:[Appendix 0\.A](https://arxiv.org/html/2607.28826#Pt0.A1.p1.1),[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2607.28826#S3.SS3.p1.1),[§5](https://arxiv.org/html/2607.28826#S5.p1.1)\.
- \[4\]A\. Beikmohammadi and S\. Magnusson\(2023\-05\)TA\-Explore: Teacher\-Assisted Exploration for Facilitating Fast Reinforcement Learning\.InTA\-Explore: Teacher\-Assisted Exploration for Facilitating Fast Reinforcement Learning,London, United Kingdom\.Cited by:[§2\.2](https://arxiv.org/html/2607.28826#S2.SS2.p1.1),[§4\.5](https://arxiv.org/html/2607.28826#S4.SS5.p2.1)\.
- \[5\]J\. Cui, B\. Zhu, Q\. Xu, Z\. Tian, X\. Qi, B\. Yu, H\. Zhang, and R\. Hong\(2025\-03\)Generalized Kullback\-Leibler Divergence Loss\.arXiv\.Note:arXiv:2503\.08038 \[cs\]External Links:[Link](http://arxiv.org/abs/2503.08038),[Document](https://dx.doi.org/10.48550/arXiv.2503.08038)Cited by:[5th item](https://arxiv.org/html/2607.28826#S4.I2.i5.p1.1)\.
- \[6\]M\. Guastalla, Y\. Li, A\. Hekmati, and B\. Krishnamachari\(2024\)Application of Large Language Models to DDoS Attack Detection\.InSecurity and Privacy in Cyber\-Physical Systems and Smart Vehicles,Y\. Chen, C\. Lin, B\. Chen, and Q\. Zhu \(Eds\.\),Vol\.552,pp\. 83–99\(en\)\.Note:Series Title: Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications EngineeringExternal Links:ISBN 978\-3\-031\-51629\-0 978\-3\-031\-51630\-6,[Document](https://dx.doi.org/10.1007/978-3-031-51630-6%5F6)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p1.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1)\.
- \[7\]E\. Jawad\(2023\-09\)THE DEEP NEURAL NETWORK\-A REVIEW\.IJRDO \-JOURNAL OF MATHEMATICS9,pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.53555/m.v9i9.5842)Cited by:[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p2.1)\.
- \[8\]M\. Kiely, D\. Bowman, M\. Standen, and C\. Moir\(2023\)On autonomous agents in a cyber defence environment\.arXiv preprint arXiv:2309\.07388\.Cited by:[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p1.1)\.
- \[9\]J\. F\. Loevenich, E\. Adler, R\. Mercier, A\. Velazquez, and R\. R\. F\. Lopes\(2024\-04\)Design of an Autonomous Cyber Defence Agent using Hybrid AI models\.In2024 International Conference on Military Communication and Information Systems \(ICMCIS\),pp\. 1–10\.External Links:[Link](https://ieeexplore.ieee.org/document/10540988/?arnumber=10540988),[Document](https://dx.doi.org/10.1109/ICMCIS61231.2024.10540988)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p1.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1)\.
- \[10\]S\. Matthew and S\. Matthew\(2024\-09\)Prompt Engineering ChatGPT for Codenames\.\(en\-US\)\.External Links:[Link](https://ieeexplore-ieee-org.journal.rmc.ca/document/10645591)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p4.1)\.
- \[11\]G\. Mcdonald, L\. Li, and R\. A\. Mallah\(2024\)Finding the Optimal Security Policies for Autonomous Cyber Operations With Competitive Reinforcement Learning\.IEEE Access12,pp\. 120292–120305\.External Links:ISSN 2169\-3536,[Link](https://ieeexplore.ieee.org/document/10639381/),[Document](https://dx.doi.org/10.1109/ACCESS.2024.3446310)Cited by:[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p2.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1),[§4\.2](https://arxiv.org/html/2607.28826#S4.SS2.p1.1),[§5](https://arxiv.org/html/2607.28826#S5.p1.1)\.
- \[12\]N\. Nashid, M\. Sintaha, and A\. Mesbah\(2023\-05\)Retrieval\-Based Prompt Selection for Code\-Related Few\-Shot Learning\.In2023 IEEE/ACM 45th International Conference on Software Engineering \(ICSE\),Melbourne, Australia,pp\. 2450–2462\(en\)\.External Links:ISBN 978\-1\-66545\-701\-9,[Link](https://ieeexplore.ieee.org/document/10172590/),[Document](https://dx.doi.org/10.1109/ICSE48619.2023.00205)Cited by:[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p2.1)\.
- \[13\]G\. Palmer, C\. Parry, D\. J\. B\. Harrold, and C\. Willis\(2024\-09\)Deep Reinforcement Learning for Autonomous Cyber Operations: A Survey\.arXiv\(en\)\.Note:arXiv:2310\.07745 \[cs\]External Links:[Link](http://arxiv.org/abs/2310.07745),[Document](https://dx.doi.org/10.48550/arXiv.2310.07745)Cited by:[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p2.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1),[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p2.1)\.
- \[14\]M\. Pfeiffer, S\. Shukla, M\. Turchetta, C\. Cadena, A\. Krause, R\. Siegwart, and J\. Nieto\(2018\-10\)Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Mapless Navigation by Leveraging Prior Demonstrations\.IEEE Robotics and Automation Letters3\(4\),pp\. 4423–4430\.External Links:ISSN 2377\-3766, 2377\-3774,[Link](https://ieeexplore.ieee.org/document/8458422/),[Document](https://dx.doi.org/10.1109/LRA.2018.2869644)Cited by:[§2\.2](https://arxiv.org/html/2607.28826#S2.SS2.p1.1)\.
- \[15\]A\. Pozzi, A\. Incremona, D\. Tessera, and D\. Toti\(2025\-06\)Mitigating exposure bias in large language model distillation: an imitation learning approach\.Neural Computing and Applications37\(18\),pp\. 12013–12029\(en\)\.External Links:ISSN 1433\-3058,[Link](https://doi.org/10.1007/s00521-025-11162-0),[Document](https://dx.doi.org/10.1007/s00521-025-11162-0)Cited by:[§2\.3](https://arxiv.org/html/2607.28826#S2.SS3.p1.1)\.
- \[16\]F\. M\.P\. Santos, A\. Gonçalves, J\. M\.C\. Sousa, and S\. M\. Vieira\(2025\-07\)Distilling Knowledge from Deep Neural Networks to Neuro\-Fuzzy Inference Systems\.In2025 IEEE International Conference on Fuzzy Systems \(FUZZ\),pp\. 1–6\.Note:ISSN: 1558\-4739External Links:[Link](https://ieeexplore.ieee.org/document/11152050),[Document](https://dx.doi.org/10.1109/FUZZ62266.2025.11152050)Cited by:[§2\.3](https://arxiv.org/html/2607.28826#S2.SS3.p1.1)\.
- \[17\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\-08\)Proximal Policy Optimization Algorithms\.arXiv\(en\)\.Note:arXiv:1707\.06347 \[cs\]External Links:[Link](http://arxiv.org/abs/1707.06347)Cited by:[1st item](https://arxiv.org/html/2607.28826#S4.I2.i1.p1.1),[1st item](https://arxiv.org/html/2607.28826#S4.I2.i1.p3.4)\.
- \[18\]M\. Son and S\. Lee\(2025\-01\)Performance Analysis of Prompt\-Engineering Techniques for Large Language Model\.In2025 IEEE International Conference on Consumer Electronics \(ICCE\),pp\. 1–5\.Note:ISSN: 2158\-4001External Links:[Link](https://ieeexplore.ieee.org/document/10930066/),[Document](https://dx.doi.org/10.1109/ICCE63647.2025.10930066)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p4.1)\.
- \[19\]R\. S\. Sutton and A\. G\. Barto\(2014\)Reinforcement learning: an introduction\.2nd edition,MIT Press,Cambridge, MA\.Cited by:[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p1.1)\.
- \[20\]K\. Tholl, M\. El Mezouar, and R\. Al Mallah\(2025\-12\)A Comparative Evaluation of Teacher\-Guided Reinforcement Learning Techniques for Autonomous Cyber Operations\.In2025 IEEE Annual Congress on Artificial Intelligence of Things \(AIoT\),pp\. 845–849\.External Links:[Link](https://ieeexplore.ieee.org/document/11416350),[Document](https://dx.doi.org/10.1109/AIoT66900.2025.00136)Cited by:[§2\.2](https://arxiv.org/html/2607.28826#S2.SS2.p1.1),[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p3.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1),[§3\.1](https://arxiv.org/html/2607.28826#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2607.28826#S3.SS1.p4.1),[§4\.5](https://arxiv.org/html/2607.28826#S4.SS5.p2.1)\.
- \[21\]K\. Tholl, M\. E\. Mezouar, and R\. A\. Mallah\(2025\-08\)Towards Production\-Worthy Simulation for Autonomous Cyber Operations\.arXiv\.Note:arXiv:2508\.19278 \[cs\]External Links:[Link](http://arxiv.org/abs/2508.19278),[Document](https://dx.doi.org/10.48550/arXiv.2508.19278)Cited by:[§1](https://arxiv.org/html/2607.28826#S1.p4.1),[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p1.1)\.
- \[22\]K\. Tholl, F\. Rivest, M\. E\. Mezouar, A\. Taylor, and R\. A\. Mallah\(2026\-02\)Large Language Model Integration with Reinforcement Learning to Augment Decision\-Making in Autonomous Cyber Operations\.arXiv\.Note:arXiv:2509\.05311 \[cs\]External Links:[Link](http://arxiv.org/abs/2509.05311),[Document](https://dx.doi.org/10.48550/arXiv.2509.05311)Cited by:[§2\.4](https://arxiv.org/html/2607.28826#S2.SS4.p3.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1),[item 2](https://arxiv.org/html/2607.28826#S3.I1.i2.p1.1),[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p4.1),[1st item](https://arxiv.org/html/2607.28826#S4.I1.i1.p1.1),[2nd item](https://arxiv.org/html/2607.28826#S4.I1.i2.p1.1),[§4\.1](https://arxiv.org/html/2607.28826#S4.SS1.p1.1),[§4\.4](https://arxiv.org/html/2607.28826#S4.SS4.p5.1),[§4\.5](https://arxiv.org/html/2607.28826#S4.SS5.p2.1)\.
- \[23\]K\. Tholl\(2026\)LLMDistillationACO: distilling knowledge from large language models into lightweight reinforcement learning agents for autonomous cyber operations\.Note:[https://github\.com/Poly\-AIvsAI/LLMDistillationACO](https://github.com/Poly-AIvsAI/LLMDistillationACO)Cited by:[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p2.1)\.
- \[24\]Vanessasml\(2024\)Note:[https://huggingface\.co/Vanessasml/cyber\-risk\-llama\-3\-8b](https://huggingface.co/Vanessasml/cyber-risk-llama-3-8b)Cited by:[§3\.2](https://arxiv.org/html/2607.28826#S3.SS2.p1.1),[§3\.3](https://arxiv.org/html/2607.28826#S3.SS3.p2.1)\.
- \[25\]J\. Wang, T\. Wang, W\. Cai, L\. Xu, and C\. Sun\(2025\-01\)Boosting Efficient Reinforcement Learning for Vision\-and\-Language Navigation With Open\-Sourced LLM\.IEEE Robotics and Automation Letters10\(1\),pp\. 612–619\.Note:Conference Name: IEEE Robotics and Automation LettersExternal Links:ISSN 2377\-3766,[Link](https://ieeexplore.ieee.org/document/10777561),[Document](https://dx.doi.org/10.1109/LRA.2024.3511402)Cited by:[§4\.5](https://arxiv.org/html/2607.28826#S4.SS5.p2.1)\.
- \[26\]Z\. Wang, X\. Li, L\. Sun, H\. Zhang, H\. Liu, and J\. Wang\(2024\-02\)Learning State\-Specific Action Masks for Reinforcement Learning\.Algorithms17\(2\),pp\. 60\(en\)\.Note:Number: 2 Publisher: Multidisciplinary Digital Publishing InstituteExternal Links:ISSN 1999\-4893,[Link](https://www.mdpi.com/1999-4893/17/2/60),[Document](https://dx.doi.org/10.3390/a17020060)Cited by:[§2\.2](https://arxiv.org/html/2607.28826#S2.SS2.p1.1),[§4\.5](https://arxiv.org/html/2607.28826#S4.SS5.p2.1)\.
- \[27\]J\. Wiebe, R\. A\. Mallah, and L\. Li\(2023\-08\)Learning Cyber Defence Tactics from Scratch with Multi\-Agent Reinforcement Learning\.arXiv\(en\)\.Note:arXiv:2310\.05939 \[cs\]External Links:[Link](http://arxiv.org/abs/2310.05939),[Document](https://dx.doi.org/10.48550/arXiv.2310.05939)Cited by:[§2\.1](https://arxiv.org/html/2607.28826#S2.SS1.p2.1),[§2\.5](https://arxiv.org/html/2607.28826#S2.SS5.p1.1),[§4\.2](https://arxiv.org/html/2607.28826#S4.SS2.p1.1),[§5](https://arxiv.org/html/2607.28826#S5.p1.1)\.
- \[28\]Z\. Yu, S\. Li, and X\. Zhang\(2026\-01\)Language Model Distillation: A Temporal Difference Imitation Learning Perspective\.arXiv\.Note:arXiv:2505\.20335 \[cs\] version: 4External Links:[Link](http://arxiv.org/abs/2505.20335),[Document](https://dx.doi.org/10.48550/arXiv.2505.20335)Cited by:[§2\.3](https://arxiv.org/html/2607.28826#S2.SS3.p1.1)\.
## Appendix 0\.ATransferability
As discussed in Sections[3](https://arxiv.org/html/2607.28826#S3)and[4](https://arxiv.org/html/2607.28826#S4), we evaluated the transferability of the distilled agent by creating various scenarios in CybORG ranging from 4 to 12 hosts\[[3](https://arxiv.org/html/2607.28826#bib.bib19)\]\. The results of these experiments can be found in Figure[7](https://arxiv.org/html/2607.28826#Pt0.A1.F7)\.
Figure 7:Evaluation of the LLM\-distilled agent against the baseline agent across different CybORG scenarios ranging from 4 to 12 hosts\. The dotted line denotes the point at which the distilled agent has transitioned to acting independently of the LLM\. The results are across 10 independent runs, with a 10\-episode running average and a ± 1 standard error\.Similar Articles
Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models
This paper introduces LDM-v0, a large decision model trained offline on trajectories from thousands of diverse reinforcement learning environments, demonstrating that a single transformer policy can match the performance of task-specific policies across robotics, autonomous driving, inventory management, cybersecurity, trading, and video games.
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
This paper systematically investigates failure modes in reinforcement learning for small language models (70-500M parameters) using PPO, identifies silent LoRA freezing, numerical overflow, and catastrophic policy collapse, and proposes a robust system with merge-and-reinitialize adapters, float32 precision, and a safety mechanism. The approach converges stably and outperforms baselines with less data.
Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
This paper presents Connect the Dots (CoD), a framework for training LLMs via reinforcement learning to develop meta-capabilities for long-lifecycle agents, enabling continuous learning and cross-domain generalization.
Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
This arXiv paper presents a unified LLMOps architecture for real-time, enterprise-ready LLM deployments, integrating data ingestion, continual learning, RAG, and feedback loops. It introduces components like AIPO, STAR+FAR, and SAGE to address knowledge staleness, hallucination, and latency-cost trade-offs in regulated sectors.
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
This paper proposes a method for simulating large LLM-agent societies on a laptop by fitting low-parameter surrogate models from a few hundred queries, using a statistical-physics-based taxonomy to predict when this approximation holds. The approach is validated on EconAgent and several other simulations using DeepSeek-elicited agent behaviors.