sequential-decision-making

Tag

Cards List
#sequential-decision-making

Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes

arXiv cs.LG ↗ · yesterday Cached

A HEC Montréal/MILA paper proposes fair policy optimization for major-minor weakly coupled MDPs, replacing the utilitarian objective with monotone concave fairness functions and introducing a count-proportion-based deep RL approach with a priority-based sampler, validated on machine replacement and NYC taxi pricing/relocation tasks.

0 favorites 0 likes
#sequential-decision-making

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

arXiv cs.CL ↗ · 2026-09-17 Cached

Introduces PrimeScientist, a method for strategic allocation of research effort in autonomous research agents using an adaptive MCTS-based policy to improve research quality and sample efficiency under resource constraints.

0 favorites 0 likes
#sequential-decision-making

Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary

arXiv cs.LG ↗ · 2026-09-15 Cached

This paper resolves an open problem by showing that a single algorithm achieves optimal switching regret for every S against an oblivious adversary in multi-armed bandits.

0 favorites 0 likes
#sequential-decision-making

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

arXiv cs.AI ↗ · 2026-09-11 Cached

This paper introduces Cros, a risk-constrained stopping layer for sequential clinical diagnosis agents that ensures finite-sample guarantees for diagnostic error and coverage, with evaluation on a benchmark derived from MIMIC.

0 favorites 0 likes
#sequential-decision-making

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

arXiv cs.AI ↗ · 2026-09-03 Cached

This paper proposes ProSE-Plan, a Bayes-adaptive planner that improves AI assistance by considering user evaluability under bounded rationality, using proposals as probes to learn preferences and enhance decision-making.

0 favorites 0 likes
#sequential-decision-making

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

arXiv cs.AI ↗ · 2026-08-19 Cached

The paper proposes RATTL, a framework that adjusts an agent's caution based on its Bayesian belief uncertainty using Wasserstein distance for safe sequential decision making, applicable to LLM-based systems.

0 favorites 0 likes
#sequential-decision-making

Let it Cook: Learning to Wait in Sequential Decision Making

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper introduces a reinforcement learning approach for training agents to wait strategically in sequential decision-making tasks, balancing task performance with resource conservation. Experiments show significant waiting behaviors across household and continuous-state environments.

0 favorites 0 likes
#sequential-decision-making

Rushes: A Human Preference Dataset for Pluralistic Alignment

arXiv cs.CL ↗ · 2026-07-24 Cached

Introduces Rushes, a large-scale dataset of human engagement preferences in AI-generated branching narratives, revealing that current LLMs like GPT-5 fail to outperform simple baselines in predicting user choices, highlighting the need for personalized alignment.

0 favorites 0 likes
#sequential-decision-making

Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

arXiv cs.CL ↗ · 2026-07-15 Cached

This paper investigates whether induced emotions can bias the sequential decision-making of LLMs using the Iowa Gambling Task as a testbed. The authors find that while emotional induction does not significantly affect average decision dynamics, anger can reduce penalty sensitivity and early-stage exploration.

0 favorites 0 likes
#sequential-decision-making

Stochastic Linear Bandits with Partially Observed Actions

arXiv cs.LG ↗ · 2026-07-13 Cached

This paper studies stochastic linear bandits where the agent only observes a random subset of action coordinates, proving that sublinear regret is possible when actions have low intrinsic dimension, and proposes the TOFU-POV algorithm with theoretical guarantees.

0 favorites 0 likes
#sequential-decision-making

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

arXiv cs.AI ↗ · 2026-06-29 Cached

Introduces a three-stage training paradigm to internalize world model planning in LLM agents, enabling future-aware decision-making. Outperforms baselines on search and mathematical reasoning tasks.

0 favorites 0 likes
#sequential-decision-making

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

arXiv cs.CL ↗ · 2026-06-25 Cached

This paper introduces Agent-Authored World Modeling (AAWM), a training procedure that constructs world-model supervision based on the policy's own decision needs rather than next-observation prediction, aligning the learning objective with the dynamics required for effective decision-making.

0 favorites 0 likes
#sequential-decision-making

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

arXiv cs.AI ↗ · 2026-05-11 Cached

This paper introduces Agentick, a unified benchmark for evaluating general sequential decision-making agents across RL, LLM, and VLM paradigms. It provides 37 procedurally generated tasks and reveals that no single approach currently dominates, highlighting significant room for improvement in agent autonomy.

0 favorites 0 likes
#sequential-decision-making

PRISM: Perception Reasoning Interleaved for Sequential Decision Making

arXiv cs.AI ↗ · 2026-05-08 Cached

This paper introduces PRISM, a framework that integrates Vision-Language Models and Large Language Models through a dynamic question-answering pipeline to improve sequential decision-making in embodied AI tasks.

0 favorites 0 likes
← Back to home

Submit Feedback