policy-evaluation

Tag

Cards List
#policy-evaluation

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

arXiv cs.LG · 2026-08-26 Cached

This paper proposes a robust gradient-based algorithm for learning behavior policies that reduce variance in online reinforcement learning policy evaluation, addressing uncertainties in transition functions with theoretical guarantees and numerical validation.

0 favorites 0 likes
#policy-evaluation

From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

arXiv cs.LG · 2026-07-30 Cached

This paper evaluates unsupervised subgrouping methods combined with causal discovery and policy evaluation for budget-constrained health interventions using observational data, finding no single method consistently outperforms others in held-out evaluation.

0 favorites 0 likes
#policy-evaluation

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

arXiv cs.LG · 2026-07-28 Cached

This paper proposes efficient online learning algorithms for policy evaluation in MDPs with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation, introducing the UBSR-TD algorithm and demonstrating its convergence and practical effectiveness.

0 favorites 0 likes
#policy-evaluation

Masked Visual Actions for Unified World Modeling

Hugging Face Daily Papers · 2026-07-21 Cached

Introduces Masked Visual Actions, a pixel-space control interface that expresses actions as partially revealed trajectories, enabling a single model to act as forward dynamics model, recover robot behavior, and support model-based planning and inverse modeling with only 15 hours of training data.

0 favorites 0 likes
#policy-evaluation

A Formally Grounded ODRL Evaluator: Implementation and Comparison

arXiv cs.AI · 2026-07-20 Cached

This paper presents a formally grounded ODRL evaluator with transparent semantics, supporting all rule types, and compares its performance with existing evaluators, addressing interoperability issues in policy modelling for data access and AI governance.

0 favorites 0 likes
#policy-evaluation

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

Hugging Face Daily Papers · 2026-07-02 Cached

This paper systematically studies world models for robotic policy evaluation, introduces the WMBench benchmark and GigaWorld-1 model, and shows that long-horizon rollout consistency is more critical than short-term visual realism.

0 favorites 0 likes
#policy-evaluation

WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

Hugging Face Daily Papers · 2026-06-11 Cached

WEAVER is a multi-view world model for robotic manipulation that achieves high fidelity, consistency, and efficiency using flow-matching loss, demonstrating superior performance in policy evaluation, improvement, and test-time planning with significant real-world improvements.

0 favorites 0 likes
#policy-evaluation

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

Hugging Face Daily Papers · 2026-05-29 Cached

StressDream enhances video world models by steering diffusion-based imaginations toward high-impact yet plausible outcomes through optimized noise initialization with semantic and plausibility objectives, enabling robust policy evaluation and improvement.

0 favorites 0 likes
#policy-evaluation

Robustness of Refugee-Matching Gains to Off-Policy Evaluation Choices

arXiv cs.LG · 2026-05-11 Cached

This paper demonstrates the robustness of refugee matching impact evaluations using off-policy methods like IPW and AIPW, confirming previous findings on algorithmic refugee assignment.

0 favorites 0 likes
#policy-evaluation

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

Hugging Face Daily Papers · 2026-04-14 Cached

RoboLab is a high-fidelity simulation benchmarking framework for evaluating task-generalist robotic policies, introducing the RoboLab-120 benchmark with 120 tasks across visual, procedural, and relational competency axes. It enables scalable, realistic task generation and systematic analysis of policy behavior under controlled perturbations to assess true generalization capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback