Tag
LeRF introduces a method to enhance perspective-taking reasoning in vision-language models by learning reference coordinate frames, improving performance on benchmarks through supervised fine-tuning and reinforcement learning.