foundation-model

Tag

Cards List
#foundation-model

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Hugging Face Daily Papers · 2026-07-30 Cached

Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.

0 favorites 0 likes
#foundation-model

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

arXiv cs.AI · 2026-07-29 Cached

CADENCE uses sparse autoencoders to decompose ECG foundation model representations into interpretable physiological concepts, significantly improving alignment with clinical phenotypes and waveform morphology.

0 favorites 0 likes
#foundation-model

microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model

Reddit r/LocalLLaMA · 2026-07-28 Cached

Microsoft introduces Mage-VL, a codec-native streaming multimodal foundation model for image and video understanding that achieves up to 3.5x inference speedup by using a sparsity pattern inspired by video codecs, cutting visual tokens by over 75%.

0 favorites 0 likes
#foundation-model

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Hugging Face Daily Papers · 2026-07-28 Cached

Modus is a decoder-only model that predicts any modality from any combination of others, achieving strong performance across diverse benchmarks without modality-specific heads or losses.

0 favorites 0 likes
#foundation-model

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Hugging Face Daily Papers · 2026-07-27 Cached

Mage-VL is an efficient codec-native streaming multimodal foundation model that reduces visual token consumption by over 75% using a custom tokenizer, achieving up to 3.5x inference speedup while matching or outperforming existing models on static and video tasks.

0 favorites 0 likes
#foundation-model

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

Hugging Face Daily Papers · 2026-07-26 Cached

Introduces N_0-VTLA, a vision-tactile-language-action foundation model for contact-rich manipulation, featuring large-scale tactile pretraining and advantage-conditioned offline policy improvement, with strong results on real-robot and simulation benchmarks.

0 favorites 0 likes
#foundation-model

@HaiyuWu1: Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to a…

X AI KOLs Following · 2026-07-24 Cached

Induction Labs introduces imagination models, a new foundation model architecture that learns from internet-scale video. Their first model, Photon-1, learns to use a computer by watching 18 years of screen recordings without action labels, achieving better results at 30× lower pretraining cost than Gemini 3.1 Flash.

0 favorites 0 likes
#foundation-model

Flux 3 X Mimic: The Next Generation of Video-Action Models

Hacker News Top · 2026-07-24 Cached

Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly generates audio-visual content and, via collaboration with mimic robotics, enables video-action prediction for robot control, tested at Audi.

0 favorites 0 likes
#foundation-model

Flux 3

Hacker News Top · 2026-07-24 Cached

Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly learns from images, videos, and audio, enabling unified generation and understanding across modalities with early access now available.

0 favorites 0 likes
#foundation-model

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

arXiv cs.LG · 2026-07-24 Cached

MKB is a unified scientific multimodal foundation model that handles six scientific branches (DNA, RNA, proteins, small molecules, earth science, medical images) using a shared Transformer backbone with modality-tailored components, achieving competitive understanding and generation across domains.

0 favorites 0 likes
#foundation-model

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

Hugging Face Daily Papers · 2026-07-23 Cached

Oxygen-TryOn is a unified foundation model for any-item virtual try-on, achieving state-of-the-art consistency and realism across single and multi-item try-on tasks through a dedicated data engine and three-stage training pipeline.

0 favorites 0 likes
#foundation-model

microsoft/Mage-Flow

Hugging Face Models Trending · 2026-07-21 Cached

Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.

0 favorites 0 likes
#foundation-model

Delineate Anything v2: A Global Foundation Model for Field Delineation

Hugging Face Daily Papers · 2026-07-21 Cached

Delineate Anything v2 is a globally scalable foundation model for agricultural field boundary mapping, outperforming state-of-the-art by 103.3% relative gain in [email protected], using a 73-million-instance multi-resolution dataset spanning 61 countries and a manually curated 100-country evaluation benchmark.

0 favorites 0 likes
#foundation-model

Inertia-1: An Open Exploration to a Unified Motion Foundation Model

Hacker News Top · 2026-07-20 Cached

Inertia-1 is a research project that systematically explores the full lifecycle of motion models—data, sensing, objectives, and scale—to produce a unified representation that transfers across body placements, devices, and tasks without retraining, leveraging self-supervised pretraining on 18 million hours of accelerometry data.

0 favorites 0 likes
#foundation-model

Soofi – Sovereign Open Source Foundation Models

Hacker News Top · 2026-07-20 Cached

Soofi introduces Soofi S, a 30B parameter Mixture-of-Experts open source foundation model trained on 27 trillion tokens, targeting industrial AI applications in German and English. The model is part of a European sovereign AI initiative.

0 favorites 0 likes
#foundation-model

@victormustar: Xiaomi-Robotics-1 just dropped on Hugging Face A robot foundation model trained on 100,000 hours of real-world manipula…

X AI KOLs Timeline · 2026-07-20 Cached

Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.

0 favorites 0 likes
#foundation-model

Xiaomi-Robotics-1

Hacker News Top · 2026-07-20 Cached

Xiaomi presents Robotics-1, a robot policy model trained via embodiment-free pre-training on 100,000 hours of data, showing clean scaling behavior and strong generalization to real-world tasks.

0 favorites 0 likes
#foundation-model

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Hugging Face Daily Papers · 2026-07-20 Cached

RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.

0 favorites 0 likes
#foundation-model

@zhang_benita: https://x.com/zhang_benita/status/2078716535548600458

X AI KOLs Timeline · 2026-07-19 Cached

This article features an interview with Yang Zhilin, founder of Moonshot AI, discussing the challenges and vision of building foundation models and the AI assistant Kimi, reflecting on the past year of development.

0 favorites 0 likes
#foundation-model

Applied Computing wants to give oil and gas operators an AI model for the entire plant

TechCrunch AI · 2026-07-16 Cached

Applied Computing, a London-based startup, has raised $20M for Orbital, a foundation AI model for oil, gas, and petrochemical plants that combines time series, physics, and language models to analyze sensor data and simulate operations, aiming to help operators use data more effectively.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback