q-learning

Tag

Cards List
#q-learning

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

arXiv cs.LG ↗ · 2026-08-24 Cached

The paper develops model-free q-learning algorithms for reinforcement learning in continuous-time jump Markov decision processes, applied to network dynamic pricing, showing superior performance over benchmark methods.

0 favorites 0 likes
#q-learning

Q-Learning With World Models

arXiv cs.LG ↗ · 2026-08-19 Cached

This paper introduces QWM, a framework that integrates world models with Q-learning to enhance sample efficiency in reinforcement learning by using imagined trajectories for action selection without compromising training on real data. It demonstrates significant improvements over state-of-the-art methods on manipulation benchmarks.

0 favorites 0 likes
#q-learning

Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper revisits the overestimation bias in Q-learning under large discrete action spaces, proposing an action intersection strategy that enables semi-decoupling between two Q-functions to balance overestimation and underestimation. Experiments in tabular and deep RL settings show improved performance over several baselines.

0 favorites 0 likes
#q-learning

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.

0 favorites 0 likes
#q-learning

Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization

arXiv cs.AI ↗ · 2026-08-12 Cached

This paper presents RL2C, a Q-learning-based algorithm for optimizing laser cutting parameters (focal length, laser power) for optical films, reducing taper size and wastage. Experiments show it reduces optimization steps by up to 12.5% and processing time by up to 81.8% compared to existing RL methods.

0 favorites 0 likes
#q-learning

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper proposes an adaptive training controller for CVaR risk-aware Q-learning, improving finite-budget behavior, reducing Bellman residuals by ~85%, and yielding better risk-adjusted performance in daily Bitcoin trading.

0 favorites 0 likes
#q-learning

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

arXiv cs.LG ↗ · 2026-08-05 Cached

The paper proposes SADQ, a modification to Q-learning that uses one-step rollout predictions from a dynamics model to regularize TD target aggregation, reducing bootstrap-induced overestimation and improving training stability across benchmarks.

0 favorites 0 likes
#q-learning

Gated Q-learning: Add Off-Policy Bias to Taste

arXiv cs.LG ↗ · 2026-08-03 Cached

Introduces Gated Q-learning, a new Q(λ) framework that smoothly interpolates between Watkins' and Peng's Q-learning to trade off off-policy bias and multistep credit assignment. Provides theoretical guarantees and empirical validation in random-walk environments.

0 favorites 0 likes
#q-learning

Variance-Reduced Q-Learning over Static and Time-Varying Networks

arXiv cs.LG ↗ · 2026-07-27 Cached

Introduces VRDQ, a decentralized Q-learning algorithm for multi-agent reinforcement learning over static and time-varying networks, with finite-time convergence guarantees that achieve linear speedups in sample complexity with only Õ(1) communication.

0 favorites 0 likes
#q-learning

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper introduces DiPS, a Q-learning framework that dynamically selects persuasion strategies for high-stakes scenarios like wildfire evacuations, achieving higher success rates than zero-shot LLM and RAG baselines.

0 favorites 0 likes
#q-learning

Quantum Annealing Enhanced Reinforcement Learning for Accurate Remaining Useful Lifetime Prediction

arXiv cs.LG ↗ · 2026-06-18 Cached

This paper proposes a quantum annealing enhanced Q-learning framework for remaining useful life prediction, using the D-Wave system to solve QUBO formulations for action selection. It outperforms classical and quantum baselines on NASA C-MAPSS and predictive maintenance datasets.

0 favorites 0 likes
#q-learning

Reversal Q-Learning

arXiv cs.LG ↗ · 2026-06-17 Cached

This paper proposes Reversal Q-Learning (RQL), an offline reinforcement learning algorithm that trains a flow policy using an expanded Markov decision process framework and techniques to enable off-policy RL without backpropagation through time. It achieves state-of-the-art performance on challenging simulated robotic tasks.

0 favorites 0 likes
#q-learning

Moment Matching Q-Learning

arXiv cs.LG ↗ · 2026-05-29 Cached

Moment Matching Q-Learning (MoMa QL) uses maximum mean discrepancy to match all moment statistics for distribution-level convergence in offline RL, achieving computational efficiency and strong performance on D4RL tasks.

0 favorites 0 likes
#q-learning

FedQHD: Closed-Form Function-Space Federated Reinforcement Learning

arXiv cs.LG ↗ · 2026-05-29 Cached

This paper proposes FedQHD, a novel federated Q-learning method using hyperdimensional random-feature state encoders with linear readouts to enable closed-form function-space aggregation, addressing the federation gap due to heterogeneous client encoders.

0 favorites 0 likes
#q-learning

Sign-Separated Finite-Time Error Analysis of Q-Learning

arXiv cs.AI ↗ · 2026-05-18 Cached

This paper develops a sign-separated finite-time error analysis for constant step-size Q-learning, decomposing the error into negative and positive parts and providing bounds that reveal an asymmetry related to overestimation.

0 favorites 0 likes
#q-learning

A Switching System Theory of Q-Learning with Linear Function Approximation

arXiv cs.LG ↗ · 2026-05-13 Cached

This paper presents a switching-system theory for Q-learning with linear function approximation, using joint spectral radius to analyze convergence stability under deterministic, i.i.d., and Markovian observations.

0 favorites 0 likes
#q-learning

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

arXiv cs.AI ↗ · 2026-05-12 Cached

The paper introduces MemQ, a method that integrates Q-learning into self-evolving memory agents by using eligibility traces over provenance DAGs to solve credit assignment problems in episodic memory retrieval.

0 favorites 0 likes
#q-learning

Debiased Model-based Representations for Sample-efficient Continuous Control

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper introduces the DR.Q algorithm, which improves model-based representations for Q-learning by maximizing mutual information and using faded prioritized experience replay to reduce bias and overfitting in continuous control tasks.

0 favorites 0 likes
#q-learning

UCB exploration via Q-ensembles

OpenAI Blog ↗ · 2017-06-05 Cached

OpenAI presents a novel exploration strategy for deep reinforcement learning using ensembles of Q-functions with upper-confidence bounds (UCB), demonstrating significant performance improvements on the Atari benchmark.

0 favorites 0 likes
#q-learning

Equivalence between policy gradients and soft Q-learning

OpenAI Blog ↗ · 2017-04-21 Cached

OpenAI researchers demonstrate a precise mathematical equivalence between soft (entropy-regularized) Q-learning and policy gradient methods in reinforcement learning, providing theoretical insight into why Q-learning works despite inaccurate value estimates. They validate this equivalence empirically on the Atari benchmark and show a Q-learning method can closely match A3C's learning dynamics.

0 favorites 0 likes
← Back to home

Submit Feedback