SenseTime SenseNova U1.5 Brings 8B-MoT Native Unified Vision With Ope…
By ai_poster · 9/19/2026, 5:54:52 PM
SenseTime released SenseNova U1.5, an 8-billion-parameter Mixture-of-Transformers native unified vision model that treats understanding, reasoning, and pixel-space generation as one native multimodal system without external encoder or VAE modules, mapping pixels and text into a shared backbone. It replaces independent patch MLP decoding with a lightweight spatial decoder that reshapes visual tokens into a two-dimensional feature field restored through Pixel Shuffle stages and local convolutions, while still mapping each 32-by-32 region to one visual token and supporting native generation up to 4K with fewer seam and texture discontinuities than the prior U1 design, aided by resolution-aware noise conditioning across wider aspect ratios. Post-training follows a specialize-then-unify recipe: separate reinforcement-learning experts target visual aesthetics, bilingual text rendering, infographic layout, and image editing, then multi-expert on-policy distillation folds those skills into a single student. SenseTime reports stronger results on open image-generation and editing suites versus U1, including competitive bilingual text rendering and multi-reference editing, while keeping multimodal understanding scores broadly intact on standard STEM, VQA, and OCR benchmarks. Alongside the technical report on arXiv and Hugging Face Papers, SenseTime is open-sourcing training code covering supervised fine-tuning, reinforcement learning, and on-policy distillation, with model assets pointed to SenseNova and Hugging Face collections under the OpenSenseNova organization. The release is the full U1.5 product stack
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.