robust-rl

Tag

Cards List
#robust-rl

ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

arXiv cs.LG · 2d ago Cached

ASGARD proposes a two-phase teacher–student pipeline using reinforcement learning to enhance UAV resilience against action-space attacks, ensuring mission completion through corrected commands and generalization to unseen attacks.

0 favorites 0 likes
#robust-rl

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

Hugging Face Daily Papers · 2026-07-27 Cached

This paper systematically investigates failure modes in reinforcement learning for small language models (70-500M parameters) using PPO, identifies silent LoRA freezing, numerical overflow, and catastrophic policy collapse, and proposes a robust system with merge-and-reinitialize adapters, float32 precision, and a safety mechanism. The approach converges stably and outperforms baselines with less data.

0 favorites 0 likes
← Back to home

Submit Feedback