Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities
Summary
RL4IL introduces a reinforcement learning-guided retrieval method that uses soft fusion over frozen demonstration libraries to handle missing sensor modalities in robotic imitation learning at inference time, achieving high success rates under complete camera dropout.
View Cached Full Text
Cached at: 06/18/26, 07:58 PM
Paper page - Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities
Source: https://huggingface.co/papers/2606.15514 RL4IL addresses a genuinely underexplored problem in imitation learning: what happens when sensors fail at deployment time? Most IL methods silently assume all modalities are always available, which is unrealistic for real robot deployments. Our key insight is that instead of retraining the policy for every possible dropout pattern, we can retrieve the right behaviour from a frozen demonstration library using a learned RL ranking policy. A few highlights that might interest the community:
The PPO policy operates over BFS-augmented candidate sets, giving it a richer and more label-diverse pool than plain kNN Soft cross-attention fusion over top-K ranked candidates consistently outperforms hard argmax selection, especially under noisy retrieval Zero-shot missing-modality handling at inference — no retraining needed when a camera fails On LIBERO benchmarks, RL4IL reaches up to 0.733 success rate under complete camera dropout, where the strongest prior method (DisDP) reaches only 0.295
Happy to discuss the retrieval design, the imputation pipeline, or the LIBERO experimental setup with anyone interested in robust robot learning.
Similar Articles
Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation
This paper introduces MIFS, a pipeline for synthesizing RL-ready multimodal data to enhance instruction following in MLLMs, achieving an 8.13% average improvement on benchmarks and faster training convergence.
MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning
MCite-RL is a citation-enhanced agentic reinforcement learning framework designed for reliable multimodal RAG, introducing iterative retrieval and reasoning for visual citations and a reward mechanism to jointly optimize answer accuracy and source traceability.
Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion
The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.
Region-Level Policy Optimization for Fine-grained MLLM Perception
Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.
LLMs help robots understand vague instructions and focus on key details
MIT CSAIL researchers developed Masked Inverse Reinforcement Learning (IRL), which uses large language models to clarify ambiguous instructions for robots and focus on key environmental details, reducing the need for extensive demonstration data.