Tag
This article discusses the necessity for algorithm engineers to continuously learn and adapt to technological shifts, from machine learning to deep learning to large models, with new graduates often leading the adaptation.
Niantic Spatial is focusing on large geospatial models, indicating a strategic shift in AI-driven mapping and spatial technology.
This review paper surveys the application of large models in battery prognostics and health management, addressing long-standing challenges and proposing a roadmap for future research in this domain.
Local deployment of large models is limited by hardware throughput. Even with GPUs at full capacity, only about 8.64M Tokens can be processed per day, which is insufficient to support Agent automation clusters or large-scale data analysis. Therefore, scaled applications still rely on cloud APIs.
A comprehensive survey of large models in sports, covering tasks, applications, datasets, and challenges to advance sports intelligence.
Fun fact: Domestic big tech companies are currently hiring in batches according to project cycles. Moonshot AI lacks overseas growth, ByteDance focuses on long-term memory and multi-agent tasks, Xiaomi is shifting to B2B for car infotainment/smart home, Tencent Hunyuan lacks pretraining talent, and 3D generation and simulation data talent is scarce with low competition and high salaries.
Financial Times reports on China's AI talent war, highlighting Moonshot founder Yang Zhilin's ability to retain a top research team amid aggressive poaching by large tech companies.
Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.
This survey comprehensively reviews resource-efficient architectures and hardware-software co-design for green AI, covering efficient model construction, training/deployment strategies, and sustainable hardware, aiming to guide sustainable large model development.
This paper proposes the concept of the world wide AI-Model Network (AI-ModelNet), a novel paradigm for interconnecting, sharing capabilities, and enabling collaborative reasoning among diverse large models. The authors review current single- and multi-model research, present a hierarchical architecture, and validate feasibility through a prototype system and application cases.
FastMix is a novel framework that automates data mixture discovery for training large models using a single proxy model and bilevel optimization, achieving state-of-the-art performance with significant efficiency gains.
DeepMind founder Demis Hassabis delivers a 60-minute speech at the University of Cambridge, covering the future development of AI from large models, AlphaFold to scientific discovery and AGI. Video has been added with Chinese subtitles.
Survey introduces the Proxy Compression Hypothesis to explain how RLHF and related methods systematically induce reward hacking, deception, and oversight gaming in large language and multimodal models.