scaling-law

Tag

Cards List
#scaling-law

The Birth of the Robotics Scaling Law

Reddit r/singularity · 2h ago

The article introduces or discusses the concept of a scaling law in robotics, which may provide insights for advancing robotic systems and AI research.

0 favorites 0 likes
#scaling-law

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

Hugging Face Daily Papers · 2026-08-24 Cached

The paper proposes a densing law for user representation learning that quantifies the relationship between data scale and tokenization capacity, and introduces an adaptive tokenization method ALGN to improve efficiency in billion-scale scenarios.

0 favorites 0 likes
#scaling-law

@swyx: I think its easy to say "Simulation is a new scaling law" and treat it as marketing hyperbole, but midway along this in…

X AI KOLs Timeline · 2026-08-21 Cached

The tweet discusses simulation as a potential new scaling law in AI, highlighting insights from an interview with Simile AI's CEO on simulating human behavior for advanced AI research.

0 favorites 0 likes
#scaling-law

@seclink: 视频里是加速了的,实际上动作慢很多 ...

X AI KOLs Timeline · 2026-08-21 Cached

Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.

0 favorites 0 likes
#scaling-law

Intelligence per dollar is the new scaling law: A tiny reasoning model breaks the existing cost-accuracy Pareto frontier on Arc-AGI 1

Reddit r/artificial · 2026-08-17

Chart Pathway's BDH-CQ, a 150M parameter reasoning model, achieves 29.5% on ARC-AGI-1 at a much lower cost per task compared to larger models like GPT-5.6 Luna, showcasing improved cost-accuracy trade-offs.

0 favorites 0 likes
#scaling-law

@0xcryptowizard: Stanford's latest course, worth checking out. About self-evolving AI, covering scaling laws, chain-of-thought, RLHF, reasoning models, etc. YouTube link: https://youtu.be/6YnLB0XbTnI?si=Dg-aSXbDxymA4U…

X AI KOLs Timeline · 2026-08-04 Cached

Recommending Stanford University's CS329A course, about self-improving AI agents, covering scaling laws, chain-of-thought, RLHF, reasoning models, and more.

0 favorites 0 likes
#scaling-law

@arvin17x: The biggest takeaway today at Berkeley AI Summit 2026 came from Professor Jianfeng Gao's talk. A new paradigm for AI Modeling has emerged: - New training data: from harness to agentic modeli…

X AI KOLs Timeline · 2026-08-02 Cached

The author summarized Professor Jianfeng Gao's presentation at the Berkeley AI Summit 2026, pointing out that the new AI modeling paradigm is based on Agentic Modeling and a new data flywheel, and emphasizing that Harness will become the key to application moats.

0 favorites 0 likes
#scaling-law

A scaling law of contextual persistence in human language

arXiv cs.CL · 2026-07-29 Cached

This paper presents a scaling law showing that the contextual influence of word order in human language decays approximately as 1/d with distance, as measured by the reduction in perplexity from large language models, across multiple languages and corpora.

0 favorites 0 likes
#scaling-law

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

arXiv cs.LG · 2026-07-28 Cached

DomainPilot introduces a domain-level loss-guided two-stage framework for data mixture optimization in LLM fine-tuning, achieving improvements on MMLU-Redux, AIME24, LiveCodeBench v5, and BFCL v3 without increasing data volume or training cost.

0 favorites 0 likes
#scaling-law

@DeFiMinty: Can Reinforcement Learning compensate for weak pretraining? Researchers tested this with chess puzzles, where every pro…

X AI KOLs Timeline · 2026-07-21 Cached

Researchers investigate how reinforcement learning (RL) interacts with pretraining quality, showing that stronger pretrained models benefit more from RL and that RL cannot fully compensate for weak pretraining. A joint scaling law is identified for pretraining and post-training compute allocation.

0 favorites 0 likes
#scaling-law

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

arXiv cs.CL · 2026-07-09 Cached

This paper investigates whether activation patterns in Polish Bielik LLMs can detect entity familiarity before generation, using unsupervised dispersion measures that achieve near-perfect separation of known vs. fabricated entities across model scales, while showing that factual reliability improves sharply with scale but is harder to predict from activations.

0 favorites 0 likes
#scaling-law

@xiaogaifun: https://x.com/xiaogaifun/status/2073771786202939572

X AI KOLs Timeline · 2026-07-05 Cached

The ByteDance Seed team released the EdgeBench benchmark, which allows AI models to work continuously for 12-72 hours to evaluate their learning ability in long-horizon tasks. They discovered that the relationship between the time spent learning from the environment and performance follows a log-sigmoid curve, revealing a new scaling law.

0 favorites 0 likes
#scaling-law

EdgeBench Reveals the Next Scaling Law: On-the-Fly AI Learning Speed Doubles Every 3 Months

Reddit r/singularity · 2026-07-02

EdgeBench reveals a new scaling law indicating that on-the-fly AI learning speed doubles every three months.

0 favorites 0 likes
#scaling-law

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

arXiv cs.AI · 2026-06-26 Cached

The paper argues that data-driven machine learning systems, including GPT-5, cannot achieve symbolic-level logical reasoning through scaling alone, due to inherent limitations in distinguishing logical structures from statistical regularities.

0 favorites 0 likes
#scaling-law

@AYi_AInotes: Everyone is raving about Japan's Fugu beating GPT on benchmarks, but I bet 99% of people haven't understood what really makes it mind-blowing. First off, this isn't some giant monolithic model at all—it has only 0.6B parameters and essentially works as an AI project manager. It handles simple tasks on its own, automatically splits complex ones, and selects the most suitable models from a global pool of top-tier models...

X AI KOLs Timeline · 2026-06-23 Cached

Sakana AI releases Fugu, a multi-agent orchestration system with only 0.6B parameters. By intelligently splitting tasks and coordinating multiple models, it achieves state-of-the-art performance while bypassing traditional parameter scaling. This marks the transition of multi-agent orchestration from a lab curiosity to a practical productivity tool.

0 favorites 0 likes
#scaling-law

@paulwalker99318: This LatePost interview is packed with information about Baidu US R&D, Scaling Laws, OpenAI, Anthropic, and Cerebras. > "Dario joining Baidu was a very important step in his career. He was recruited by Greg Diamos. And before joining Baidu, Dario didn't have a computer science or AI background — he came from math, physics, and biology. Greg Diamos saw his intuition for AI and ability to train models."

X AI KOLs Timeline · 2026-06-22 Cached

A summary of the LatePost interview, reviewing Baidu US R&D's early AI布局, including investing in Cerebras, nearly investing in OpenAI and Anthropic, and the flow of talent from Baidu to these companies.

0 favorites 0 likes
#scaling-law

@SaitoWu: A group at Baidu Research US predicted ten years ago: Don't bet all AI compute on NVIDIA. So they actually invested in a 'wafer-scale' chip company — Cerebras. In 2016, Zhou Nan left investment banking for Baidu's US AI research institute. Andrew Ng was leading the team, budgets were ample, GPUs were bought freely. Dario (An…

X AI KOLs Timeline · 2026-06-17 Cached

The article recounts Baidu Research US's investment in Cerebras, a wafer-scale chip company, a decade ago. It analyzes the shift in the AI chip market from training to inference and the importance of non-consensus investments.

0 favorites 0 likes
#scaling-law

The Weight Norm Sets the Grokking Timescale: A Causal Delay Law

arXiv cs.LG · 2026-06-15 Cached

This paper demonstrates that the weight norm causally controls the timescale of grokking in neural networks, reconciling conflicting accounts. Through interventions, it shows that grokking follows an exponential delay law and that norm magnitude dominates grokking time over learning rate across architectures.

0 favorites 0 likes
#scaling-law

@gntalktalk: This is the best methodology in AI development recently: The scaling law of software engineering --- Channel AI founder Luke Orthwine proposes a new paradigm: ditch the single-threaded linear "chess thinking" and switch to a high-concurrency, macro-scheduling, saturation-attack "real-time strategy game…

X AI KOLs Timeline · 2026-06-14 Cached

Channel AI founder Luke Orthwine proposes a new software development methodology: shifting programming thinking from traditional chess-like single-threaded linear thinking to real-time strategy game (RTS) style high concurrency, macro scheduling, and saturation attack to achieve efficient development in the AI Agent era.

0 favorites 0 likes
#scaling-law

@snowboat84: https://x.com/snowboat84/status/2062686432335184321

X AI KOLs Timeline · 2026-06-05 Cached

This article explores the deep connections between physics and deep learning, analyzes the isomorphism of phenomena such as Scaling Law and emergence with concepts like critical scaling laws and phase transitions in physics, and reviews the current status and prospects of applying physical methodologies in AI.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback