Tag
Introducing BIT, a bidirectional image-text diffusion bridge that enables direct transformation from text tokens to images and vice versa, allowing for caption tweaking and structural preservation across modalities.
BRAID is a framework that formulates interleaved text-image-text reasoning as a unified Markov decision process, enabling joint optimization of textual and visual generation via reinforcement learning with a VLM judge providing dense turn-level feedback.