Tag
Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.
CADENCE uses sparse autoencoders to decompose ECG foundation model representations into interpretable physiological concepts, significantly improving alignment with clinical phenotypes and waveform morphology.
Microsoft introduces Mage-VL, a codec-native streaming multimodal foundation model for image and video understanding that achieves up to 3.5x inference speedup by using a sparsity pattern inspired by video codecs, cutting visual tokens by over 75%.
Modus is a decoder-only model that predicts any modality from any combination of others, achieving strong performance across diverse benchmarks without modality-specific heads or losses.
Mage-VL is an efficient codec-native streaming multimodal foundation model that reduces visual token consumption by over 75% using a custom tokenizer, achieving up to 3.5x inference speedup while matching or outperforming existing models on static and video tasks.
Introduces N_0-VTLA, a vision-tactile-language-action foundation model for contact-rich manipulation, featuring large-scale tactile pretraining and advantage-conditioned offline policy improvement, with strong results on real-robot and simulation benchmarks.
Induction Labs introduces imagination models, a new foundation model architecture that learns from internet-scale video. Their first model, Photon-1, learns to use a computer by watching 18 years of screen recordings without action labels, achieving better results at 30× lower pretraining cost than Gemini 3.1 Flash.
Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly generates audio-visual content and, via collaboration with mimic robotics, enables video-action prediction for robot control, tested at Audi.
Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly learns from images, videos, and audio, enabling unified generation and understanding across modalities with early access now available.
MKB is a unified scientific multimodal foundation model that handles six scientific branches (DNA, RNA, proteins, small molecules, earth science, medical images) using a shared Transformer backbone with modality-tailored components, achieving competitive understanding and generation across domains.
Oxygen-TryOn is a unified foundation model for any-item virtual try-on, achieving state-of-the-art consistency and realism across single and multi-item try-on tasks through a dedicated data engine and three-stage training pipeline.
Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.
Delineate Anything v2 is a globally scalable foundation model for agricultural field boundary mapping, outperforming state-of-the-art by 103.3% relative gain in [email protected], using a 73-million-instance multi-resolution dataset spanning 61 countries and a manually curated 100-country evaluation benchmark.
Inertia-1 is a research project that systematically explores the full lifecycle of motion models—data, sensing, objectives, and scale—to produce a unified representation that transfers across body placements, devices, and tasks without retraining, leveraging self-supervised pretraining on 18 million hours of accelerometry data.
Soofi introduces Soofi S, a 30B parameter Mixture-of-Experts open source foundation model trained on 27 trillion tokens, targeting industrial AI applications in German and English. The model is part of a European sovereign AI initiative.
Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.
Xiaomi presents Robotics-1, a robot policy model trained via embodiment-free pre-training on 100,000 hours of data, showing clean scaling behavior and strong generalization to real-world tasks.
RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.
This article features an interview with Yang Zhilin, founder of Moonshot AI, discussing the challenges and vision of building foundation models and the AI assistant Kimi, reflecting on the past year of development.
Applied Computing, a London-based startup, has raised $20M for Orbital, a foundation AI model for oil, gas, and petrochemical plants that combines time series, physics, and language models to analyze sensor data and simulate operations, aiming to help operators use data more effectively.