SenseNova-U1.5: Towards Native Unified Visual Intelligence
Summary
SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, and expert optimization.
View Cached Full Text
Cached at: 09/11/26, 06:17 AM
Paper page - SenseNova-U1.5: Towards Native Unified Visual Intelligence
Source: https://huggingface.co/papers/2609.11929 Published on Sep 10
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, expert optimization, and on-policy distillation.
We launch SenseNova-U1.5, an8B-MoTnative unified multimodal modelthat understands, reasons about, and generates visual content within anencoder-freeandVAE-freearchitecture. We strengthen its visual interface throughspatially coherent patch reconstructionand scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts forvisual aesthetics,bilingual text rendering,infographic generation, and image editing, and consolidate their capabilities throughmulti-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, andinterleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning,reinforcement learning, and on-policy distillation.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.11929
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.11929 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.11929 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.11929 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
This paper introduces SenseNova-U1, a unified multimodal architecture that integrates understanding and generation tasks, releasing two variants (8B and 30B) that perform competitively in both perception and image synthesis.
sensenova/SenseNova-U1-8B-MoT
SenseNova U1 is a new series of native multimodal models that unify understanding and generation within a single architecture using the NEO-Unify framework, eliminating the need for separate visual encoders or VAEs.
sensenova/SenseNova-U1.5-8B-MoT
SenseNova-U1.5-8B-MoT is a native unified multimodal model for enhanced visual creation, featuring improvements in image generation quality, text rendering, and precise control.
Vision as Unified Multimodal Generation
This paper presents SenseNova-Vision, a unified multimodal model that formulates computer vision tasks as generation problems, achieving performance comparable to specialized systems across diverse vision tasks. It introduces a large-scale instruction-response corpus and publicly releases the model and datasets.
SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference
SenseNova U1.5-Lite is released, featuring expert training and OPD distillation to create a unified model for inference, with improvements in instruction following and image generation capabilities.