SenseNova-U1.5: Towards Native Unified Visual Intelligence

Hugging Face Daily Papers Papers

Summary

SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, and expert optimization.

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.
Original Article
View Cached Full Text

Cached at: 09/11/26, 06:17 AM

Paper page - SenseNova-U1.5: Towards Native Unified Visual Intelligence

Source: https://huggingface.co/papers/2609.11929 Published on Sep 10

#2 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs, achieving high fidelity and instruction following through patch reconstruction, curated data, expert optimization, and on-policy distillation.

We launch SenseNova-U1.5, an8B-MoTnative unified multimodal modelthat understands, reasons about, and generates visual content within anencoder-freeandVAE-freearchitecture. We strengthen its visual interface throughspatially coherent patch reconstructionand scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts forvisual aesthetics,bilingual text rendering,infographic generation, and image editing, and consolidate their capabilities throughmulti-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, andinterleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning,reinforcement learning, and on-policy distillation.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.11929

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.11929 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.11929 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.11929 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

sensenova/SenseNova-U1-8B-MoT

Hugging Face Models Trending

SenseNova U1 is a new series of native multimodal models that unify understanding and generation within a single architecture using the NEO-Unify framework, eliminating the need for separate visual encoders or VAEs.

sensenova/SenseNova-U1.5-8B-MoT

Hugging Face Models Trending

SenseNova-U1.5-8B-MoT is a native unified multimodal model for enhanced visual creation, featuring improvements in image generation quality, text rendering, and precise control.

Vision as Unified Multimodal Generation

Hugging Face Daily Papers

This paper presents SenseNova-Vision, a unified multimodal model that formulates computer vision tasks as generation problems, achieving performance comparable to specialized systems across diverse vision tasks. It introduces a large-scale instruction-response corpus and publicly releases the model and datasets.