Tag
Scal3R improves online 3D reconstruction for long videos by using multi-reference relative pose querying with lightweight tokens and pose-graph optimization, reducing drift and achieving state-of-the-art performance.
This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.
Introduces multi-reference image-grounded video captioning and proposes RefCaptioner, a two-stage post-training framework with mixed-data SFT and hierarchical coverage-discounted GRPO. The paper also presents MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding.
Oxygen-TryOn is a unified foundation model for any-item virtual try-on, achieving state-of-the-art consistency and realism across single and multi-item try-on tasks through a dedicated data engine and three-stage training pipeline.
This paper introduces MultiRef-Compass, a comprehensive benchmark for multi-reference-to-audio-video generation, comprising 350 curated samples and an evaluation protocol with four dimensions including Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.
Introduces PaliBench, a multi-reference benchmark for Pali-to-English translation using independent translations from multiple scholars, and a reusable methodology for creating similar benchmarks for classical languages.
Step-by-step tutorial shows how to use Higgsfield Cinema Studio’s preset-driven, multi-reference chaining to create a consistent 25-second cinematic chase scene from five reference images in under 15 minutes.