@CalcCon: To truly reach AGI, we have to go beyond engineering hacks and establish the scientific principles behind why deep lear…

X AI KOLs Timeline Papers

Summary

A researcher claims to have established the scientific principles behind deep learning using Renormalization Group theory, moving beyond engineering hacks to potentially pave the way for AGI. This work builds on papers from 2021 in JMLR and Nature Communications.

To truly reach AGI, we have to go beyond engineering hacks and establish the scientific principles behind why deep learning works. One of the most amazing things we have discovered is that the best layers of the best deep neural networks all seem to converge with a universal signature This work dates back to papers we published in JMLR and Nature Communications back in 2021, for work actually done in 2018. Hard to beleive its been this long It had been clear to me from day one that these signatures resembled a univerisal critical exponent from Renormalization Group theory. And after several years of work, and numerous publications and talks, and a lot of sweat equity, I think I have finally cracked this and can provide a complete Wilsonian RG view of what's going on in learning (My only regret is that my PhD advisor passed away earl this year before I could share it with him, as he was one of the greats of RG theory.)
Original Article
View Cached Full Text

Cached at: 07/20/26, 01:27 PM

To truly reach AGI, we have to go beyond engineering hacks and establish the scientific principles behind why deep learning works.

One of the most amazing things we have discovered is that the best layers of the best deep neural networks all seem to converge with a universal signature

This work dates back to papers we published in JMLR and Nature Communications back in 2021, for work actually done in 2018. Hard to beleive its been this long

It had been clear to me from day one that these signatures resembled a univerisal critical exponent from Renormalization Group theory.

And after several years of work, and numerous publications and talks, and a lot of sweat equity, I think I have finally cracked this and can provide a complete Wilsonian RG view of what’s going on in learning

(My only regret is that my PhD advisor passed away earl this year before I could share it with him, as he was one of the greats of RG theory.)

To learn more about this, check out the open source weightwatcher project

https://weightwatcher.ai

and the research page on RG theory https://weightwatcher.ai/rg_theory_webpage/rg_theory.html…

Selected publications and presentations

KDD Workshop 2019 https://dl.acm.org/doi/abs/10.1145/3292500.3332294…

JMLR 2021 https://jmlr.org/papers/v22/20-410.html…

Nature Communications 2021 https://nature.com/articles/s41467-021-24025-8…

NeurIPS 2023 Spotlight paper https://neurips.cc/virtual/2023/poster/70435…

My 2023 NeurIPS Invited Talk https://nips.cc/virtual/2023/workshop/66530#wse-detail-83033…

SETOL 2023 archive paper (140 pages of stat mech!):
https://arxiv.org/abs/2507.17912

ICML 2025 Workshop paper (updated) https://arxiv.org/pdf/2605.12394

and please feel free to join us on the Weight lWatchers community channel to discuss

also, just add my late PhD advisor was not only one one of the best and nicest human beings you could choose to work with but also one of the best advisors you could hope for

There have been ben at most two dozen professors who have had two or more Nobel prize, winners as student students, and I was lucky enough to study in the same group as the two he produced

Yes, I do actually And in the most recent draft paper, you’ll see I make the case for this Both using arguments from statistical learning theory, and renormalization group theory

this is exactly the problem. I’m trying to solve. Thank you

most people are trying to solve this problem by analyzing the outputs of models

My approach is very different

I’m trying to understand if the model is trained properly or not given his data and the feedback loops it has been provided

Remarkably, if you look at the open source foundation scale models that are available many of them are actually suboptimal

Most notably, the open AI OSS models look horrible as if almost every layer is overfit to the training data

Other models are just not trained long enough and are giving away a huge amount of capacity

Of course, if you train a model with crazy data, it can give crazy results. I can’t get around that that’s a completely different problem and I’m glad other people are working on it

Similar Articles

Measuring progress toward AGI: A cognitive framework

Google DeepMind Blog

Google DeepMind released a paper proposing a cognitive framework to measure progress toward AGI, identifying ten key cognitive abilities and launching a Kaggle hackathon to build relevant evaluations.

The Next Paradigm (7 minute read)

TLDR AI

The article argues that training AI on millions of verifiable tasks across diverse RL environments could lead to AGI, and that scaling may overcome current limitations like sample inefficiency. It also examines why progress on computer use has been slower due to lack of grindable environments.

Taking a responsible path to AGI

Google DeepMind Blog

DeepMind publishes a comprehensive approach to AGI safety and security, outlining a systematic framework to address misuse, misalignment, accidents, and structural risks as artificial general intelligence approaches reality within the coming years.