Tag
This paper introduces PolicyShiftBench, a benchmark for policy-adaptive image guardrails, and PolicyShiftGuard, a compact model trained with a two-stage method that improves performance under shifting safety policies.
OmniTacTune introduces a two-stage reinforcement learning pipeline for adapting tactile feedback to pretrained visual robot policies, achieving 85-100% success on contact-rich manipulation tasks within 40-80 minutes.
Warp RL replaces additive residual corrections in reinforcement learning with an invertible, state-conditioned transformation of the base policy's action distribution using monotonic rational-quadratic spline flows, enabling adaptation of distribution shape, scale, and geometry under dynamics shifts. It matches or outperforms residual correction in ManiSkill3 manipulation tasks and achieves 30% faster task completion in a real robot peg-insertion task.
SingGuard is a multimodal guardrail system from Ant Group that treats safety policy as an input, allowing dynamic adaptation via natural language. It is released under Apache 2.0 and covers text and image modalities.
Introduces WIZARD, a weight-space meta-learning framework that generates task-specific LoRA parameters for frozen VLA policies from language instructions and demonstration videos, enabling efficient task adaptation without fine-tuning.