Tag
This paper proposes a test-time alignment approach for large vision-language models using trajectory-guided structured sampling and iterative MCMC refinement, improving visual reasoning accuracy without heavy post-training.
This paper analyzes test-time alignment methods that use a small aligned model as a proxy to guide a larger unaligned model's generation. The authors propose a novel rejection criterion based on conservative confidence betting and demonstrate improvements over existing approaches on multiple datasets.