@AdinaYakup: SenseNova-Vision SenseTime's new model treats all of computer vision as generation - 7B - CC BY-NC 4.0 ( non commercial…
Summary
SenseTime released SenseNova-Vision, a 7B parameter model that unifies computer vision tasks as generation, with open weights, instruction corpus, benchmark, paper, and demo.
View Cached Full Text
Cached at: 07/09/26, 07:37 AM
SenseNova-Vision 🔥 SenseTime’s new model treats all of computer vision as generation
- 7B
- CC BY-NC 4.0 ( non commercial )
- Model/ 50M instruction corpus/ benchmark/ paper/ demo, all open 🫶 https://t.co/eArcntuMJ4
Similar Articles
@SenseTime_AI: 𝗦𝗲𝗻𝘀𝗲𝗡𝗼𝘃𝗮-𝗩𝗶𝘀𝗶𝗼𝗻-7𝗕-𝗠𝗼𝗧, 𝗳𝘂𝗹𝗹𝘆 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲𝗱: 𝗼𝗻𝗲 𝗺𝗼𝗱𝗲𝗹, 𝗲𝘃𝗲𝗿𝘆 𝗺𝗮𝗷𝗼�…
SenseTime releases SenseNova-Vision-7B-MoT, a fully open-sourced unified multimodal model that handles multiple vision tasks using natural language instructions, supporting detection, OCR, depth, segmentation, and more.
Vision as Unified Multimodal Generation
This paper presents SenseNova-Vision, a unified multimodal model that formulates computer vision tasks as generation problems, achieving performance comparable to specialized systems across diverse vision tasks. It introduces a large-scale instruction-response corpus and publicly releases the model and datasets.
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
This paper introduces SenseNova-U1, a unified multimodal architecture that integrates understanding and generation tasks, releasing two variants (8B and 30B) that perform competitively in both perception and image synthesis.
sensenova/SenseNova-U1-8B-MoT
SenseNova U1 is a new series of native multimodal models that unify understanding and generation within a single architecture using the NEO-Unify framework, eliminating the need for separate visual encoders or VAEs.
@PrajwalTomar_: Nobody is talking about this yet, which is honestly crazy to me. A lab just open-sourced an 8B model that fixes the mos…
A lab has open-sourced SenseNova U1.5-Lite, an 8B parameter model for image AI that allows making specific changes to images without destroying other elements, as demonstrated on infographics and posters.