markov-decision-processes

Tag

Cards List
#markov-decision-processes

Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes

arXiv cs.LG ↗ · 17h ago Cached

A HEC Montréal/MILA paper proposes fair policy optimization for major-minor weakly coupled MDPs, replacing the utilitarian objective with monotone concave fairness functions and introducing a count-proportion-based deep RL approach with a priority-based sampler, validated on machine replacement and NYC taxi pricing/relocation tasks.

0 favorites 0 likes
#markov-decision-processes

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

arXiv cs.LG ↗ · 2026-08-24 Cached

The paper develops model-free q-learning algorithms for reinforcement learning in continuous-time jump Markov decision processes, applied to network dynamic pricing, showing superior performance over benchmark methods.

0 favorites 0 likes
#markov-decision-processes

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.

0 favorites 0 likes
#markov-decision-processes

Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates

arXiv cs.LG ↗ · 2026-08-10 Cached

This paper presents a certificate-carrying sub-quadratic method for computing bisimulation metrics in Markov decision processes using approximate nearest neighbors, with coverage-augmented guarantees and two-sided bounds. Experiments show improved scaling and accurate metric recovery compared to baselines.

0 favorites 0 likes
#markov-decision-processes

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

arXiv cs.LG ↗ · 2026-08-10 Cached

This paper studies the sample complexity of robust average-reward Markov decision processes, deriving minimax-optimal learning rates via plug-in reductions under total-variation uncertainty sets.

0 favorites 0 likes
#markov-decision-processes

Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper formalises counterfactual policy optimisation for Markov Decision Processes under probabilistic nondeterministic causal models, which separate latent confounding from inherent stochasticity, and proposes a practical optimisation procedure for deriving robust counterfactual policies. The approach is validated on a sepsis treatment simulator with diabetes as an unobserved global confounder.

0 favorites 0 likes
#markov-decision-processes

Property-driven Causal Abstractions for Markov Decision Processes

arXiv cs.AI ↗ · 2026-07-31 Cached

This paper introduces a property-driven causal abstraction technique for factored Markov Decision Processes (MDPs), grouping states based on causal relations over state variable predicates to reduce model size while preserving property-relevant behavior. The approach is evaluated on standard benchmarks, yielding small abstractions that support near-optimal policy computation and often generalize to larger MDPs.

0 favorites 0 likes
#markov-decision-processes

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

arXiv cs.LG ↗ · 2026-07-28 Cached

This paper proposes efficient online learning algorithms for policy evaluation in MDPs with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation, introducing the UBSR-TD algorithm and demonstrating its convergence and practical effectiveness.

0 favorites 0 likes
#markov-decision-processes

@RitOnchain: Stanford computer science professor just revealed how to master Markov Decision Processes. 83-minutes. free. By Stanfor…

X AI KOLs Timeline ↗ · 2026-07-11 Cached

Stanford computer science professor offers a free 83-minute lecture on mastering Markov Decision Processes, covering policy evaluation, value iteration, and convergence limits.

0 favorites 0 likes
#markov-decision-processes

Performance-Driven Environment Abstraction with Multi-Timescale Learning

arXiv cs.LG ↗ · 2026-06-17 Cached

This paper proposes a performance-driven state abstraction method for reinforcement learning that directly optimizes decision quality, using a multi-timescale framework to jointly adapt the policy and a tree-structured abstraction. The algorithm refines or aggregates state space based on Q-value discrepancies, achieving better sample efficiency and faster replanning than baselines.

0 favorites 0 likes
#markov-decision-processes

Lyapunov-Based Sample Complexity Analysis for Weakly-Coupled MDPs

arXiv cs.LG ↗ · 2026-06-15 Cached

This paper studies the sample complexity of learning in average-reward weakly-coupled MDPs and restless bandits, establishing finite-sample PAC guarantees with polynomial complexity using a novel Lyapunov-based analysis framework.

0 favorites 0 likes
#markov-decision-processes

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets

arXiv cs.AI ↗ · 2026-06-10 Cached

This paper introduces Bellman-Taylor Score Decoding, a method to handle state-dependent feasible action sets in Markov decision processes, addressing a key challenge in applying deep reinforcement learning to operations research problems.

0 favorites 0 likes
#markov-decision-processes

Exact Unlearning in Reinforcement Learning

arXiv cs.LG ↗ · 2026-06-04 Cached

This paper formalizes exact unlearning in reinforcement learning, proposing a ρ-TV-stable RL algorithm for tabular MDPs that efficiently removes a user's data influence at a fraction of retraining cost, achieving near-minimax-optimal regret bounds. The work is accepted at ICML and establishes both upper and lower bounds for ρ-TV-stable RL algorithms.

0 favorites 0 likes
#markov-decision-processes

Answer-Set-Programming-based Abstractions for Reinforcement Learning

arXiv cs.AI ↗ · 2026-06-01 Cached

This paper presents an Answer Set Programming (ASP) based implementation of the CARCASS framework for constructing abstractions in reinforcement learning, demonstrating its effectiveness on Blocks World and Minigrid domains.

0 favorites 0 likes
#markov-decision-processes

Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs

arXiv cs.LG ↗ · 2026-05-26 Cached

This paper proposes a quantile Bayesian risk-aware MDP framework for online RL that adaptively balances robustness and exploration over time, providing theoretical regret bounds and demonstrating strong empirical performance.

0 favorites 0 likes
#markov-decision-processes

@RohOnChain: This 1 hour Stanford lecture on Markov Decision Processes will teach you more about the math behind systematic trading …

X AI KOLs Timeline ↗ · 2026-05-12 Cached

The article promotes a Stanford lecture on Markov Decision Processes as a valuable resource for understanding the mathematical foundations of systematic trading, claiming it offers more insight than a short-term internship at major financial firms.

0 favorites 0 likes
← Back to home

Submit Feedback