Tag
The paper proposes Wiki Foundation Model (WFM), a novel foundation model for scalable, agent-native representation and retrieval using LLM Wiki, which couples dense text with link structure, achieving strong performance on benchmarks and 10.5x training acceleration.
This tweet thread claims that Brain-Computer Interface (BCI) technology will cause a 'cambrian explosion' in Human-Computer Interaction (HCI), detailing the development of a wearable BCI and the training of a brain foundation model.
This paper introduces a generalized multimodal foundation model capable of handling arbitrary modality combinations and prediction tasks, achieving competitive performance through training on large-scale synthetic datasets with diverse causal structures.
A research collaboration introduces ARLI, a method to enable RL fine-tuning of Vision-Language-Action models with asynchronous inference using real-time chunking, improving robot behaviors by allowing the RL policy to observe more recent images.
OneBid is a unified auto-bidding foundation model that addresses challenges in diverse oCPX advertising scenarios by using Mixture-of-Experts architecture and CROP optimization, with deployment at Kuaishou showing significant improvements like +13.1% in ROAS.
Cactus Needle 3 is a small, sliceable foundation model for automation tasks that runs on-device, achieving performance comparable to larger models on function calling and structured extraction.
Meta launches Muse as a standalone AI agent app, emphasizing isolation for security, but faces scrutiny over privacy and business model implications.
The paper proposes WFM, a Wiki Foundation Model for scalable, agent-native knowledge representation and retrieval, demonstrating improved performance in complex agentic reasoning tasks with a 10.5x training acceleration.
FoundAna introduces a foundation model for graph anomaly detection that integrates GNNs and transformers to achieve generalizable, cross-graph performance, demonstrating superior results on multiple benchmark datasets.
Odyssey-3 is a general-purpose world model that enables physical agents to control diverse systems like robots, vehicles, and drones with minimal task-specific data, demonstrating broad transfer learning capabilities.
NVIDIA released FoundationPose on Hugging Face, a unified foundation model for 6-DoF object pose estimation and tracking that works on novel objects without fine-tuning.
Odyssey-3 is a new foundation world model that can control robots, cars, drones, and virtual worlds by reusing world knowledge across machines with minimal experience.
LimiX-2 is a pretrained foundation model for structured data that uses contextual mechanism networks to achieve #1 on major tabular benchmarks, supporting multiple tasks without task-specific parameter updates.
IBM and NASA have open-sourced a lunar foundation model that integrates multi-sensor data to map the Moon, aiding in tasks like crater mapping, volcanic history investigation, and ice detection.
Atria Dawn Preview is a research agent from Shanghai that verifies its own work using external evidence to ensure verifiable and reproducible outcomes for long-horizon tasks.
Atria Dawn Preview is a foundation agentic language model designed for scientific research, achieving competitive benchmark results and demonstrating a shift toward human-AI project-level collaboration.
The paper presents StepAudio3Realtime, an audio-language foundation model for real-time spoken interaction, using a continuous listen-converse-think-act loop with Think-While-Speaking to achieve deep reasoning and low latency, with top-tier performance on benchmarks like MMSU and Full-Duplex Bench.
ZGCM-1 is a 7B open foundation model trained from scratch with extreme efficiency, combining internal reasoning and external tool use for math and agentic search tasks, achieving competitive performance with much larger models like Qwen3-235B-A22B and GLM-5.1.
Skild AI's S1 model enables robots to learn from video demonstrations without retraining, achieving $100M ARR in 10 months and deploying across multiple industries.
Skild AI has launched the S1 robot foundation model, which enables robots to learn new tasks from a single video demonstration using NVIDIA's Physical AI, improving adaptability in dynamic industrial environments.