Tag
SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, and expert optimization.
Kuaishou researchers propose UniGD, a unified generative-discriminative framework for industrial retrieval that integrates retrieval and relevance scoring into a single model, with techniques like CAGE and CAM to improve effectiveness and reduce latency. Online A/B tests show a 5.78% ad revenue increase and 33.1% inference latency reduction.
Hunyuan3D-Buffalo 1.0 is a unified multimodal model for 3D generation, understanding, and editing, trained on an 87M-scale 3D multimodal corpus. It combines Hunyuan3D-VLM and Hunyuan3D-DiT to achieve state-of-the-art performance on text-to-3D generation and 3D editing benchmarks.
This paper introduces SwanTale, a unified multi-speaker speech and audio generation model supporting both zero-shot and instruct tasks, along with SwanData-Caption for data annotation and SwanVAE for high-quality multi-audio-modality generation.
S1-Omni is a unified multimodal reasoning model for scientific tasks including understanding, prediction, and generation. It is trained on a corpus of 200 scientific tasks and outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks.
UniVR introduces VR-GRPO, a reinforcement learning paradigm for unified visual reasoning, learning complex reasoning and physical dynamics from pure visual demonstrations, achieving up to 25% improvement on the VR-X benchmark.
BRAID is a framework that formulates interleaved text-image-text reasoning as a unified Markov decision process, enabling joint optimization of textual and visual generation via reinforcement learning with a VLM judge providing dense turn-level feedback.
This paper presents SenseNova-Vision, a unified multimodal model that formulates computer vision tasks as generation problems, achieving performance comparable to specialized systems across diverse vision tasks. It introduces a large-scale instruction-response corpus and publicly releases the model and datasets.
This paper introduces Audex, a unified audio-text LLM from NVIDIA that achieves state-of-the-art performance across multiple audio and speech tasks while preserving strong text reasoning capabilities without regression.
Boogu-Image-0.1 is a new unified image generation and editing model family with 10B parameters, available under Apache 2.0 license. It features fast turbo inference in 4 steps, trained on 10x less data, and supports Chinese and English.
ILLUME-X is a unified multimodal model for free-form interleaved text-image generation, featuring improved data efficiency, stable training, and a comprehensive evaluation metric called ILScore. It outperforms previous models on tasks like style transfer, image decomposition, and storytelling.
Boogu has released a series of open-source unified image generation and editing models, including Base, Turbo, and Edit variants.
UniDDT proposes a decoupled diffusion transformer framework that unifies multimodal understanding and generation by leveraging a Noisy ViT encoder and LLM for semantic encoding, achieving strong performance on both tasks.
This paper introduces Audio-Interaction, a unified streaming audio model that combines offline task execution with real-time audio instruction following via an end-to-end framework. It proposes SoundFlow for the perceive-decide-respond loop and evaluates competitive performance across benchmarks.
Introduces Representation Forcing (RF), a technique that enables unified multimodal models to perform both perception and generation end-to-end without external VAE latent spaces, matching state-of-the-art VAE-based models in image generation while improving understanding.
Lumos-Nexus is a training-efficient video generation framework that uses a two-stage design with a lightweight generator for training and a high-capacity pretrained generator for inference, achieving enhanced visual fidelity through Unified Progressive Frequency Bridging.
Presents a unified neural scaling law that accurately models deep neural network scaling across multiple dimensions including parameters, dataset size, training steps, and compute, validated across diverse architectures and tasks.
UniT is a unified feed-forward model for geometry perception using a Group Autoregressive Transformer that integrates multiple paradigms (online/offline, multi-modal, long-horizon) while maintaining metric-scale accuracy via scale-adaptive loss and queue-style KV caching. It achieves state-of-the-art performance on ten benchmarks spanning seven tasks.
Uni-Edit proposes using intelligent image editing as a single general task to simultaneously improve unified multimodal models' understanding, generation, and editing capabilities, with an automated data synthesis pipeline creating complex editing instructions.
Lance is a unified multimodal model that leverages a dual-stream mixture-of-experts architecture and collaborative multi-task training to achieve strong performance in understanding, generation, and editing of both images and videos, outperforming existing open-source unified models.