Tag
The paper proposes PX-UAP, a method using probabilistic robustness and explainable AI to generate universal adversarial perturbations against deep reinforcement learning-based intrusion detection systems, demonstrating improved attack effectiveness in experiments.
This paper introduces a graph-based deep reinforcement learning framework for the one-dimensional bin packing problem, reducing optimality gaps compared to existing methods and enabling zero-shot generalization across instance sizes.
RefinePPO introduces iterative action refinement for continuous control policies, achieving performance comparable or better than standard PPO with faster convergence in benchmark tasks.
This paper proposes a composite-gradient learning method that integrates deep reinforcement learning and model predictive control for shared control authority in autonomous systems, with evaluations on traffic networks showing modest benefits under strong interaction.
The paper proposes a Decision Transformer-based approach for optimizing UAV-mounted RIS-assisted dynamic D2D communications, demonstrating cross-scenario generalization and efficient zero-shot transfer.
This paper conducts a systematic review of AI for data center energy optimization, identifies research gaps, and proposes the CLEAR-DC framework to close the optimizer-load loop by integrating energy, carbon, and water metrics.
This paper explores the impact of network topology and opponent information on the emergence of cooperation in multi-agent reinforcement learning systems, specifically in the Iterated Prisoner's Dilemma, finding that graph structure and information availability significantly influence cooperative strategies.
This research paper proposes a simulation framework using deep reinforcement learning for controlling connected and automated vehicle platoon joining in mixed traffic, showing that PPO achieves high success rates while highlighting trade-offs between safety and efficiency.
This paper presents a modified JAMPR deep reinforcement learning model to solve the Pickup and Delivery problem with Capacity and Time Window constraints (CPDPTW), offering fast optimal solutions for small to medium-sized instances and suboptimal solutions for larger ones.
This paper introduces the Travelling Thief Problem with Drone (TTP-D), which jointly optimizes ground routing, drone synchronization, and item selection using mixed-integer programming, metaheuristics, and attention-based deep reinforcement learning.
Berkeley's Deep RL class is now fully available online on YouTube for free, as announced by Sergey Levine.
The lectures for the Deep Reinforcement Learning course CS185/285 from UC Berkeley are now available online for public viewing.
This paper proposes a proxemics-based reward formulation for deep reinforcement learning social navigation, modeling human personal space as Gaussian-mixture fields to improve social compliance while maintaining navigation efficiency.
This paper presents a deep reinforcement learning approach for solving vehicle routing problems, demonstrated through three industrial truck planning case studies. The proposed method achieves over 10% cost reduction compared to baseline results and discusses generalization to more VRP variants.
This paper proposes PLAN, a lightweight parallel liquid-inspired approximation network for efficient representation learning in flexible job shop scheduling, achieving better makespan and lower inference latency with fewer parameters than state-of-the-art baselines.
This paper reports that deep reinforcement learning agents using frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, compressing task-relevant information through very few neurons without any sparsity-inducing objective.
This survey provides a unified view of progress reward modeling for robotic learning, organizing the field into three steps: interface, methods, and data/benchmarks.
This paper explores Relative Positional Encoding (RPE) as an additive bias in Transformer architectures to solve the Team Orienteering Problem, demonstrating consistent improvements in collected rewards and optimality gaps over vanilla Transformer architectures.
This paper systematically explores four deep reinforcement learning solutions (DQN, REINFORCE, PPO, and MuZero) for the asymmetric Nepali board game Baghchal, finding that MuZero achieves the best win rates due to model-based planning via Monte Carlo Tree Search.
Proposes a two-timescale multi-layer deep reinforcement learning framework with latent action space for joint service placement, computational delegation, and power control in hierarchical edge-cloud computing, achieving up to 20.8% latency reduction and 13% resource utilization improvement.