Tag
This paper surveys smart glasses as unified first-person intelligence platforms, presenting a systematic framework for perception-state-interaction-action loops and evaluation protocols across constrained hardware and diverse applications.
This survey paper reviews the integration of large language models, knowledge bases, and reasoning capabilities for general embodied intelligence, proposing a unified framework and identifying key challenges for future development.
China's 29 embodied intelligence robot companies have made 125 external investments in their early financing stages, with AgiBot alone investing in 37 projects, using capital to lock in technology, teams, and market positions in advance.
SmartMage is a unified multimodal LLM that dynamically orchestrates visual and geometric modalities for query-dependent 3D scene understanding, achieving state-of-the-art results across five benchmarks.
CG-World is a large-scale world-state dataset and protocol derived from industrial computer graphics pipelines, explicitly recording multimodal world states, interventions, and counterfactual branches to support world model research. It demonstrates improvements in geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies.
Alibaba Cloud presents a visual walkthrough of the Qwen model family's evolution from LLMs to embodied intelligence, highlighting Model Studio's features and integrations.
This paper argues that the machine intelligence architectures proposed by Edmund C. Berkeley and David L. Heiserman remain relevant for embodied robotic cognition, emphasizing logical states, hardware-near control, and modular sensorimotor structure.
This paper introduces SIS-Bench, a benchmark for evaluating self-awareness and spatial cognition in UAV embodied intelligence using multimodal large language models, and explores motion-aware representations to improve performance.
LingBot-Video, a 30B parameter MoE-based video foundation model for embodied intelligence, has been released on Hugging Face with only 3B active parameters at inference, augmented with 70K hours of embodied data.
LingBot-Video, a 30B-parameter video model with sparse MoE, designed for embodied intelligence, is open-sourced. It outperforms existing models on RBench, trained on 70K+ hours of embodied data.
LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.
LingBot-Video presents a DiT-based video pretraining framework with Mixture-of-Experts architecture, specialized data augmentation, and multi-dimensional reward system for embodied intelligence applications.
A Peking University study published in Cell Reports finds that the human brain can adapt to virtual non-human limbs (wings) through VR training, showing dynamic plasticity in the occipitotemporal cortex.
Lin Junyang, former head of Alibaba's Qianwen team, closed his AI lab's first financing round at a $2B post-money valuation, with Gao Rong and Sequoia China each investing $100M and Tencent adding $20M. The lab will focus on world models and embodied intelligence rather than general LLMs.
This article delves into the concept of embodied intelligence, its intellectual origins (philosophy, cognitive science, AI robotics), and historical development (the failure of symbolism and Brooks' subsumption architecture). It analyzes its differences from pure software AI and the challenges it faces.
This paper proposes Embodied-BenchClaw, an autonomous multi-agent system that automatically constructs embodied spatial intelligence benchmarks from user intent through a five-stage pipeline with process quality control and an extensible Skill Library.
Shared an information gap report on the developer ecosystem, covering topics such as migrating openclaw-type projects to wearable devices like AI glasses and rings, open-source data and models for robotics and embodied AI, and niche open-source applications for AI API relay stations and routing.
GEM introduces a generative supervision method to improve embodied intelligence by leveraging generative models for training.
Qwen-VLA is a unified vision-language-action model for embodied decision-making, integrating manipulation, navigation, and trajectory prediction across different robot platforms. It uses a DiT-based action decoder and embodiment-aware prompt conditioning, achieving strong performance and out-of-distribution generalization.
Fei-Fei Li warns that AI is too focused on language models, emphasizing that the world is physical, visual, and spatial, and that most of the economy relies on embodied intelligence.