reinforcement-learning

Tag

Cards List
#reinforcement-learning

Introducing Lev

Lobsters Hottest ↗ · 8h ago Cached

Lev is an open-source decision engine that implements fast classifier models similar to Jev, using encoder models for quick probabilistic decisions with an escalation model for failure cases.

0 favorites 0 likes
#reinforcement-learning

FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

Reddit r/LocalLLaMA ↗ · 12h ago Cached

HuatuoGPT-3-27B is a medical language model built on Qwen3.8-27B using One-stage Policy Optimization (OnePO), a reinforcement learning method for domain adaptation without supervised fine-tuning.

0 favorites 0 likes
#reinforcement-learning

Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration

arXiv cs.LG ↗ · yesterday Cached

This paper proposes QRLQ, a quantum reinforcement learning framework that integrates parameterised quantum circuits with dueling double deep Q-networks to optimize cost and delay tradeoffs in quantum cloud orchestration, achieving lower costs and delays than heuristic baselines while using fewer parameters than classical DRL.

0 favorites 0 likes
#reinforcement-learning

Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following

arXiv cs.LG ↗ · yesterday Cached

This paper proposes CC-OPD, a novel on-policy distillation method for multi-constraint instruction following that uses counterfactual ablations to enhance training signals, achieving superior performance where a 1.5B model surpasses its 7B teacher on benchmarks.

0 favorites 0 likes
#reinforcement-learning

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

arXiv cs.LG ↗ · yesterday Cached

The paper introduces Wasserstein-Tilted Flow Maps (WTF), a simulation-free reinforcement learning algorithm for fine-tuning flow-based generative models to enhance reward alignment, achieving higher rewards with up to 280x less compute than baselines.

0 favorites 0 likes
#reinforcement-learning

On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

arXiv cs.LG ↗ · yesterday Cached

This paper identifies 'Preference Coverage Collapse' as a failure mode in hindsight relabeling for multi-objective reinforcement learning and introduces 'her_mix' to mitigate it, improving performance across various settings.

0 favorites 0 likes
#reinforcement-learning

Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates

arXiv cs.LG ↗ · yesterday Cached

This paper shows that tool-result caching, even if marginally correct, can reverse the expected group-normalized policy updates in reinforcement learning, as demonstrated through mathematical analysis and experiments with a two-action model.

0 favorites 0 likes
#reinforcement-learning

Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents

arXiv cs.AI ↗ · yesterday Cached

The paper introduces CoCA, an on-policy learning framework for dynamic capability allocation in long-horizon multimodal agents, addressing computational overhead and stage-dependent demands through conditional comparisons and reinforcement learning.

0 favorites 0 likes
#reinforcement-learning

Categorical Internalisation of Environmental Groupoids for Generalisable POMDP Solving

arXiv cs.AI ↗ · yesterday Cached

This paper advocates using category theory and environmental groupoids to structure reinforcement learning in partially observable environments, leveraging symmetries for improved sample efficiency and generalization.

0 favorites 0 likes
#reinforcement-learning

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

arXiv cs.CL ↗ · yesterday Cached

SkillGym transforms human-written agent skills into executable training environments for LLMs, enabling supervised fine-tuning and reinforcement learning to enhance real-world problem-solving capabilities and performance on benchmarks.

0 favorites 0 likes
#reinforcement-learning

Planned Test-Time Scaling with Coordinated Reasoning Paths

arXiv cs.CL ↗ · yesterday Cached

This paper introduces Planned Test-Time Scaling (PTTS), a method that coordinates reasoning branches to enhance performance on challenging tasks, achieving significant gains over repeated sampling in mathematical reasoning benchmarks.

0 favorites 0 likes
#reinforcement-learning

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

arXiv cs.CL ↗ · yesterday Cached

The paper introduces RECAP, a redundancy-aware learning method that improves the efficiency of large reasoning models by assigning credit to steps based on their structural role and efficacy, enhancing accuracy while reducing token usage.

0 favorites 0 likes
#reinforcement-learning

Reinforcement Learning with Decomposed Subtasks

arXiv cs.AI ↗ · yesterday Cached

This paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that decomposes trajectory reward into per-subtask advantages to improve credit assignment in reinforcement learning for language model agents, showing significant gains on high-heterogeneity agentic benchmarks.

0 favorites 0 likes
#reinforcement-learning

@dair_ai: Impressive paper from Salesforce. It discusses the importance of good verifiers for RL environments. Only 35.8% of the …

X AI KOLs Timeline ↗ · yesterday Cached

This paper from Salesforce audits public RL environments for terminal agents, finding significant defects, and introduces RIVER, a training recipe that filters defective environments and penalizes repetitive behavior to enhance model performance.

0 favorites 0 likes
#reinforcement-learning

Rufus-Air: An Open LLM Post-Training Recipe

Hugging Face Daily Papers ↗ · yesterday Cached

Rufus-Air is an open and reproducible post-training recipe for LLMs, detailing an eight-stage pipeline to enhance capabilities from basic to advanced, with improvements over existing models like GLM-4.5-Air-Base.

0 favorites 0 likes
#reinforcement-learning

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Hugging Face Daily Papers ↗ · yesterday Cached

This paper presents Qwen-Planner-Agent, a closed-loop AI-for-AI framework for scalable development of mobile planner agents, integrating data production, training, and deployment to improve performance on real-world tasks and benchmarks.

0 favorites 0 likes
#reinforcement-learning

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Hugging Face Daily Papers ↗ · yesterday Cached

IterSynth introduces a role-decoupled iterative synthesis paradigm for deep search agents, using reinforcement learning to improve performance on long-horizon search tasks and surpassing prior methods on benchmarks.

0 favorites 0 likes
#reinforcement-learning

Towards Universal Post-Training for Robotics (18 minute read)

TLDR AI ↗ · yesterday Cached

The article argues that robotics needs post-training similar to language models to achieve high reliability, discussing challenges and potential approaches for universal post-training in robotic systems.

0 favorites 0 likes
#reinforcement-learning

Skild AI trained a Unitree G1 robot to play soccer by simulating 140 years of practice against its own past versions

Reddit r/singularity ↗ · yesterday

Skild AI trained a Unitree G1 robot to play soccer by simulating 140 years of practice against its own past versions, demonstrating significant progress in AI and robotics.

0 favorites 0 likes
#reinforcement-learning

Learning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy

arXiv cs.LG ↗ · 2d ago Cached

This paper proposes a proxy-guided hierarchical reinforcement learning framework to defend against diverse inference attacks on smart meter data by learning battery-based load-shaping policies that disrupt non-intrusive load monitoring patterns.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback