Tag
The article introduces or discusses the concept of a scaling law in robotics, which may provide insights for advancing robotic systems and AI research.
The paper proposes a densing law for user representation learning that quantifies the relationship between data scale and tokenization capacity, and introduces an adaptive tokenization method ALGN to improve efficiency in billion-scale scenarios.
The tweet discusses simulation as a potential new scaling law in AI, highlighting insights from an interview with Simile AI's CEO on simulating human behavior for advanced AI research.
Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.
Chart Pathway's BDH-CQ, a 150M parameter reasoning model, achieves 29.5% on ARC-AGI-1 at a much lower cost per task compared to larger models like GPT-5.6 Luna, showcasing improved cost-accuracy trade-offs.
Recommending Stanford University's CS329A course, about self-improving AI agents, covering scaling laws, chain-of-thought, RLHF, reasoning models, and more.
The author summarized Professor Jianfeng Gao's presentation at the Berkeley AI Summit 2026, pointing out that the new AI modeling paradigm is based on Agentic Modeling and a new data flywheel, and emphasizing that Harness will become the key to application moats.
This paper presents a scaling law showing that the contextual influence of word order in human language decays approximately as 1/d with distance, as measured by the reduction in perplexity from large language models, across multiple languages and corpora.
DomainPilot introduces a domain-level loss-guided two-stage framework for data mixture optimization in LLM fine-tuning, achieving improvements on MMLU-Redux, AIME24, LiveCodeBench v5, and BFCL v3 without increasing data volume or training cost.
Researchers investigate how reinforcement learning (RL) interacts with pretraining quality, showing that stronger pretrained models benefit more from RL and that RL cannot fully compensate for weak pretraining. A joint scaling law is identified for pretraining and post-training compute allocation.
This paper investigates whether activation patterns in Polish Bielik LLMs can detect entity familiarity before generation, using unsupervised dispersion measures that achieve near-perfect separation of known vs. fabricated entities across model scales, while showing that factual reliability improves sharply with scale but is harder to predict from activations.
The ByteDance Seed team released the EdgeBench benchmark, which allows AI models to work continuously for 12-72 hours to evaluate their learning ability in long-horizon tasks. They discovered that the relationship between the time spent learning from the environment and performance follows a log-sigmoid curve, revealing a new scaling law.
EdgeBench reveals a new scaling law indicating that on-the-fly AI learning speed doubles every three months.
The paper argues that data-driven machine learning systems, including GPT-5, cannot achieve symbolic-level logical reasoning through scaling alone, due to inherent limitations in distinguishing logical structures from statistical regularities.
Sakana AI releases Fugu, a multi-agent orchestration system with only 0.6B parameters. By intelligently splitting tasks and coordinating multiple models, it achieves state-of-the-art performance while bypassing traditional parameter scaling. This marks the transition of multi-agent orchestration from a lab curiosity to a practical productivity tool.
A summary of the LatePost interview, reviewing Baidu US R&D's early AI布局, including investing in Cerebras, nearly investing in OpenAI and Anthropic, and the flow of talent from Baidu to these companies.
The article recounts Baidu Research US's investment in Cerebras, a wafer-scale chip company, a decade ago. It analyzes the shift in the AI chip market from training to inference and the importance of non-consensus investments.
This paper demonstrates that the weight norm causally controls the timescale of grokking in neural networks, reconciling conflicting accounts. Through interventions, it shows that grokking follows an exponential delay law and that norm magnitude dominates grokking time over learning rate across architectures.
Channel AI founder Luke Orthwine proposes a new software development methodology: shifting programming thinking from traditional chess-like single-threaded linear thinking to real-time strategy game (RTS) style high concurrency, macro scheduling, and saturation attack to achieve efficient development in the AI Agent era.
This article explores the deep connections between physics and deep learning, analyzes the isomorphism of phenomena such as Scaling Law and emergence with concepts like critical scaling laws and phase transitions in physics, and reviews the current status and prospects of applying physical methodologies in AI.