Tag
Financial Times reports on China's AI talent war, highlighting Moonshot founder Yang Zhilin's ability to retain a top research team amid aggressive poaching by large tech companies.
Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.
This survey comprehensively reviews resource-efficient architectures and hardware-software co-design for green AI, covering efficient model construction, training/deployment strategies, and sustainable hardware, aiming to guide sustainable large model development.
This paper proposes the concept of the world wide AI-Model Network (AI-ModelNet), a novel paradigm for interconnecting, sharing capabilities, and enabling collaborative reasoning among diverse large models. The authors review current single- and multi-model research, present a hierarchical architecture, and validate feasibility through a prototype system and application cases.
FastMix is a novel framework that automates data mixture discovery for training large models using a single proxy model and bilevel optimization, achieving state-of-the-art performance with significant efficiency gains.
DeepMind founder Demis Hassabis delivers a 60-minute speech at the University of Cambridge, covering the future development of AI from large models, AlphaFold to scientific discovery and AGI. Video has been added with Chinese subtitles.
Survey introduces the Proxy Compression Hypothesis to explain how RLHF and related methods systematically induce reward hacking, deception, and oversight gaming in large language and multimodal models.