deep-reinforcement-learning

Tag

Cards List
#deep-reinforcement-learning

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper analyzes the evaluation and design paradigms in deep reinforcement learning, revealing that performance rankings are not monotonic across data regimes and that common low-data regime benchmarks may lead to incorrect conclusions.

0 favorites 0 likes
#deep-reinforcement-learning

Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

arXiv cs.LG ↗ · 2026-07-09 Cached

This paper proposes a deep reinforcement learning framework (MORP-DRL) for multi-objective reliability-based portfolio optimization, jointly optimizing expected return and downside risk using CVaR and EVaR under practical constraints, and demonstrates performance on global equity indices across different market regimes.

0 favorites 0 likes
#deep-reinforcement-learning

Deep Reinforcement Learning for Dynamic Battery Management of Autonomous Order Pickers

arXiv cs.LG ↗ · 2026-07-08 Cached

This paper proposes a Proximal Policy Optimization (PPO)-based deep reinforcement learning framework for dynamic battery charging of autonomous mobile robots in warehouses, achieving up to 6% higher order-completion rates over baseline methods.

0 favorites 0 likes
#deep-reinforcement-learning

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

Hugging Face Daily Papers ↗ · 2026-07-08 Cached

This paper analyzes evaluation and design paradigms in deep reinforcement learning, demonstrating that canonical paradigms can lead to incorrect conclusions and providing insights into scaling, capacity, and complexity.

0 favorites 0 likes
#deep-reinforcement-learning

Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System

arXiv cs.LG ↗ · 2026-07-07 Cached

This work-in-progress paper proposes embedding a differentiable physics model into the PPO actor loss function to penalize anticipated safety violations in reinforcement learning, evaluated on a simulated 1-DoF helicopter system. The physics-informed soft regularizations reduce constraint violations while maintaining reliable target tracking.

0 favorites 0 likes
#deep-reinforcement-learning

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper presents PASE, a neuro-symbolic framework that uses LLMs to generate structured recovery plans for cloud systems and verifies them via a neural-symbolic world model, achieving over 40% reduction in recovery time.

0 favorites 0 likes
#deep-reinforcement-learning

A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management

arXiv cs.AI ↗ · 2026-07-01 Cached

A three-phase deep reinforcement learning system for personalized portfolio management that addresses ticker lock-in, monolithic objectives, and static user models, using a cross-asset encoder pretrained with self-supervised learning and the Chronos time series foundation model, fine-tuned with Mixture of Experts and PPO, and personalized via LoRA.

0 favorites 0 likes
#deep-reinforcement-learning

Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction

arXiv cs.LG ↗ · 2026-07-01 Cached

This paper proposes a novel offline reinforcement learning framework for active flow control that uses a sensor position-conditioned architecture with Point Attention layers to handle varying sensor configurations, enabling data-driven policy extraction without costly online interactions.

0 favorites 0 likes
#deep-reinforcement-learning

A3M: Adaptive, Adversarial and Multi-Objective Learning for Strategic Bidding in Repeated Auctions

arXiv cs.CL ↗ · 2026-06-30 Cached

Introduces A3M, a framework combining adaptive deep reinforcement learning, adversarial reasoning, and multi-objective reward design for strategic bidding in repeated auctions, achieving 30-40% regret reduction.

0 favorites 0 likes
#deep-reinforcement-learning

Continuous-time Optimal Stopping through Deep Reinforcement Learning

arXiv cs.LG ↗ · 2026-06-17 Cached

This paper introduces CARLOS, a deep reinforcement learning algorithm that learns continuous-time optimal stopping rules for American-style options using an aggregate deep neural network, effectively closing the Bermudan-American value gap with high computational efficiency.

0 favorites 0 likes
#deep-reinforcement-learning

A Deep Reinforcement Learning (DRL)-Based Transformer Method for Solving the Open Shop Scheduling Problem

arXiv cs.AI ↗ · 2026-06-15 Cached

Presents a Transformer-based scheduling policy trained with reinforcement learning for the open shop scheduling problem, showing that a model trained on small instances can generalize to much larger problems and compete with classical dispatching heuristics.

0 favorites 0 likes
#deep-reinforcement-learning

Performance Variation in Deep Reinforcement Learning

arXiv cs.LG ↗ · 2026-06-08 Cached

This paper identifies limitations of conventional uncertainty estimates for deep reinforcement learning and proposes percentile-based statistics and visualization to better assess run-to-run performance variation. Case studies demonstrate the method on PPO, SAC, TD-MPC, DQN, and Rainbow algorithms.

0 favorites 0 likes
#deep-reinforcement-learning

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

arXiv cs.LG ↗ · 2026-06-05 Cached

This paper argues that representation learning, not model-based planning, is the key to scalable multitask deep reinforcement learning. It introduces MR.Q, a simple model-free algorithm with auxiliary predictive objectives that outperforms prior world-model-based methods across diverse continuous control tasks.

0 favorites 0 likes
#deep-reinforcement-learning

A Unified Python Framework for Direct PPO-based Control of AHUs with Economizer Logic and CO2-Constrained Ventilation

arXiv cs.LG ↗ · 2026-05-26 Cached

A unified Python framework using PPO-based deep reinforcement learning for optimizing HVAC control with economizer logic and CO2-constrained ventilation is presented, showing improved energy efficiency and temperature stability over traditional PID controllers.

0 favorites 0 likes
#deep-reinforcement-learning

@tom_doerr: Hugging Face deep reinforcement learning course with practical exercises https://github.com/huggingface/deep-rl-class…

X AI KOLs Timeline ↗ · 2026-05-24 Cached

Hugging Face offers a deep reinforcement learning course with practical exercises, now in low-maintenance state but still a useful resource for learning theory and hands-on DRL.

0 favorites 0 likes
#deep-reinforcement-learning

OpenAI Fellows Fall 2018: Final projects

OpenAI Blog ↗ · 2019-05-17 Cached

OpenAI announces the completion of its Fall 2018 Fellows program and celebrates the fellows' research contributions. The organization also open-sourced part of the fellowship curriculum, including 'Spinning up in Deep RL,' an educational resource for learning reinforcement learning.

0 favorites 0 likes
#deep-reinforcement-learning

Spinning Up in Deep RL: Workshop review

OpenAI Blog ↗ · 2019-02-26 Cached

OpenAI held its first Spinning Up in Deep RL Workshop on February 2, engaging ~90 in-person participants and ~300 livestream viewers to provide education in deep RL, robotics, and AI safety through talks, mentorship, and hands-on projects.

0 favorites 0 likes
#deep-reinforcement-learning

Spinning Up in Deep RL

OpenAI Blog ↗ · 2018-11-08 Cached

OpenAI released 'Spinning Up in Deep RL,' an educational toolkit featuring introductory materials, curated paper lists, and clean standalone implementations of key RL algorithms (VPG, TRPO, PPO, DDPG, TD3, SAC) designed to help newcomers learn deep reinforcement learning from scratch.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback