Tag
Lev is an open-source decision engine that implements fast classifier models similar to Jev, using encoder models for quick probabilistic decisions with an escalation model for failure cases.
HuatuoGPT-3-27B is a medical language model built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), a reinforcement learning method for domain adaptation without supervised fine-tuning.
This paper proposes QRLQ, a quantum reinforcement learning framework that integrates parameterised quantum circuits with dueling double deep Q-networks to optimize cost and delay tradeoffs in quantum cloud orchestration, achieving lower costs and delays than heuristic baselines while using fewer parameters than classical DRL.
This paper proposes CC-OPD, a novel on-policy distillation method for multi-constraint instruction following that uses counterfactual ablations to enhance training signals, achieving superior performance where a 1.5B model surpasses its 7B teacher on benchmarks.
The paper introduces Wasserstein-Tilted Flow Maps (WTF), a simulation-free reinforcement learning algorithm for fine-tuning flow-based generative models to enhance reward alignment, achieving higher rewards with up to 280x less compute than baselines.
This paper identifies 'Preference Coverage Collapse' as a failure mode in hindsight relabeling for multi-objective reinforcement learning and introduces 'her_mix' to mitigate it, improving performance across various settings.
This paper shows that tool-result caching, even if marginally correct, can reverse the expected group-normalized policy updates in reinforcement learning, as demonstrated through mathematical analysis and experiments with a two-action model.
The paper introduces CoCA, an on-policy learning framework for dynamic capability allocation in long-horizon multimodal agents, addressing computational overhead and stage-dependent demands through conditional comparisons and reinforcement learning.
This paper advocates using category theory and environmental groupoids to structure reinforcement learning in partially observable environments, leveraging symmetries for improved sample efficiency and generalization.
SkillGym transforms human-written agent skills into executable training environments for LLMs, enabling supervised fine-tuning and reinforcement learning to enhance real-world problem-solving capabilities and performance on benchmarks.
This paper introduces Planned Test-Time Scaling (PTTS), a method that coordinates reasoning branches to enhance performance on challenging tasks, achieving significant gains over repeated sampling in mathematical reasoning benchmarks.
The paper introduces RECAP, a redundancy-aware learning method that improves the efficiency of large reasoning models by assigning credit to steps based on their structural role and efficacy, enhancing accuracy while reducing token usage.
This paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that decomposes trajectory reward into per-subtask advantages to improve credit assignment in reinforcement learning for language model agents, showing significant gains on high-heterogeneity agentic benchmarks.
This paper from Salesforce audits public RL environments for terminal agents, finding significant defects, and introduces RIVER, a training recipe that filters defective environments and penalizes repetitive behavior to enhance model performance.
Rufus-Air is an open and reproducible post-training recipe for LLMs, detailing an eight-stage pipeline to enhance capabilities from basic to advanced, with improvements over existing models like GLM-4.5-Air-Base.
This paper presents Qwen-Planner-Agent, a closed-loop AI-for-AI framework for scalable development of mobile planner agents, integrating data production, training, and deployment to improve performance on real-world tasks and benchmarks.
IterSynth introduces a role-decoupled iterative synthesis paradigm for deep search agents, using reinforcement learning to improve performance on long-horizon search tasks and surpassing prior methods on benchmarks.
The article argues that robotics needs post-training similar to language models to achieve high reliability, discussing challenges and potential approaches for universal post-training in robotic systems.
Skild AI trained a Unitree G1 robot to play soccer by simulating 140 years of practice against its own past versions, demonstrating significant progress in AI and robotics.
This paper proposes a proxy-guided hierarchical reinforcement learning framework to defend against diverse inference attacks on smart meter data by learning battery-based load-shaping policies that disrupt non-intrusive load monitoring patterns.