@MindfulReturn: Today I saw an interview with Professor Huang Biwei (@huang_biwei) and learned about their new round of funding! After learning about the Aether AI solution and taking a closer look at their direction, let me share my thoughts: The next paradigm of AI is not bigger models, but causality. 1. Correlation Ceiling: Why the visuals are...

X AI KOLs Timeline News

Summary

This article offers an in-depth analysis of the Causal World Model (CWM) proposed by Aether AI (原识之智), arguing that the next AI paradigm will shift from correlation to causation. It discusses the theoretical foundations, technical architecture, and potential impact on video generation and embodied intelligence.

Today I saw an interview with Professor Huang Biwei (@huang_biwei) and learned about their new round of funding! After learning about the Aether AI solution and taking a closer look at their direction, let me share my thoughts: The next paradigm of AI is not bigger models, but causality. 1. Correlation Ceiling: Why the visuals are beautiful but the physics are fake Let's start with three sets of data. First set: Video generation. A paper from December 2025 tested the current strongest video generation models and found that the "gravity" they generate is only 1.81 m/s², equivalent to 18% of Earth's gravity, similar to the Moon. Two objects generated by the same model falling from the same height have different landing times. What Galileo wanted to prove at the Leaning Tower of Pisa, these models still haven't learned. Second set: Embodied intelligence. The Stanford AI Index 2026 report gave a set of numbers: humanoid robots have a success rate of 89.4% in simulation, but only 12% in the real world. A gap of 77 percentage points. The What-If World benchmark tested 9 SOTA world models, asking them to generate "videos after making a physical intervention to the scene": for example, changing object mass, friction coefficient, or lighting direction. As a result, no model exceeded 52% paired accuracy, with open-source models concentrated around 28%. Third set: Causal reasoning. HOCA-Bench divides AI video failures into two categories: ontological anomalies (objects disappearing, color flickering) and causal anomalies (wrong gravity direction, collision penetration, reversed buoyancy). Video understanding models score more than 20 percentage points lower on the latter category than the former. Models can recognize that "a cat using chopsticks" is strange, but cannot recognize that "a stone floating on water" violates physics. Because it learns statistical patterns, not physical rules. These three sets of data point to the same problem. In academia, it is called causal confusion. Today's AI learns correlation. Give it enough data, and it learns to predict the next word, the next image, the next video. But it doesn't know why. Ask it to generate "a glass cup falling onto a marble floor," it has seen similar images and generates similar images. But it hasn't understood gravity, hardness, or collision mechanics. So fragments fly upwards. In engineering, there is a very precise description: The model's output depends on the visual salience of the intervention, not on physical computability. In plain language: It doesn't check whether the physical laws are correct; it checks whether the scene looks like something seen in the training data. A robot trained to put a red block into the left box gets confused when the block is changed to blue. Raising the table height by 1 cm renders its learned grasping useless. This is not insufficient training data; it's that what was learned stays at the pixel level, without rising to the causal concept of "surface contact." Memorizing the answer bank and understanding the principle are two different things. 2. Four Jumps: The evolution of AI paradigms In her keynote at CVPR 2026, Huang Biwei broke down AI paradigms into four stages. First jump: small model × correlation — already past Second jump: small model × causality — academic reserve Third jump: large model × correlation — we are here now Fourth jump: large model × causality — next station We are currently at the third jump. GPT, Claude, Sora, Veo are essentially large models × correlation. They compress the texts, images, and videos of the internet into hundreds of billions of parameters, learning a super complex conditional probability distribution. The problem is, Ilya Sutskever himself announced at NeurIPS that the pretraining era is coming to an end. His exact words: "Data is the fossil fuel of AI. We have only one internet. Data will not grow anymore. We have reached peak data." It's not that Scaling Law has failed; it's that the object of scaling needs to change. This leads to the core statement in Huang Biwei's speech: "Compression is not complete intelligence. It should be structured compression that is intelligence." Brute-force data compression gives correlation, but extracting causal structure from the same data has a slope several times steeper. 3. What does the Causal World Model actually do differently? The Causal World Model (CWM) proposed by Aether AI has a fundamentally different underlying logic from today's video generation models and world models. Huang Biwei gave three hard criteria: First, learn causal feature representations. Today's models learn from raw data "what appears together with what." The causal world model must learn "what causes what": recovering interpretable latent factors from pixels. Object mass, surface friction, gravity direction, collision elasticity. These are not labels; they are causal variables that the model separates from the data by itself. Second, understand causal structure. It's not just knowing that "cup" and "shards" are related; it's knowing "cup drops -> hits the ground -> stress exceeds material strength -> shatters -> fragments fly following conservation of momentum." This is a causal graph, not a pixel map. Cross-level, from the macroscopic "cup breaks" to the microscopic "glass crack propagation," the model must be able to map. Third, capture causal dynamics. The world is not static. The same cup breaks when falling on marble but not on carpet. Today's models need to have seen both scenarios separately to generate them. The causal world model doesn't need that; it knows that the carpet absorbs the impact force, so the causal chain is cut at the "impact" step, and the subsequent "shattering" and "flying" won't happen. This is reasoning about the physical world. They designed a four-layer architecture to support this logic: System Layer — causality-driven agent system for decision-making and planning Foundation Model Layer — causal world model for core understanding and prediction Neural Architecture Layer — brain functional specialization-inspired modular design Infrastructure / Transformer Layer — modified Transformer that injects causal dependencies at the token level From underlying tokens to high-level decisions, causality is not an add-on; it runs throughout. 4. Who is doing this? Let's talk about the people. Huang Biwei, PhD from CMU (graduated 2022), supervised by Kun Zhang and Clark Glymour. Kun Zhang is a key figure in the field of causal discovery; Glymour, together with Spirtes, laid the foundation in 1989. This academic lineage deserves clarification. In 1989, Glymour and Spirtes published the foundational work on causal discovery algorithms. Over the next 37 years, the school's work has been very pure: letting machines discover causal relationships from observational data autonomously, instead of waiting for humans to feed them. The Causal-Learn library developed by Kun Zhang's team is the most mainstream open-source Python library in causal discovery, and Huang Biwei is a core contributor. She received the 2021 Apple Scholar award and was an assistant professor at UCSD's Halıcıoğlu Data Science Institute. She left academia to start a company in 2025, registered in Shanghai, called 「原识之智」 (Yuan Shi Zhi Zhi). The name is clear: the wisdom of recognizing origins. She is not an AI entrepreneur who got into the field halfway. She is someone who has accumulated ten years of academic work in the direction of causal discovery, and came out to industrialize with a complete theoretical framework and open-source tools. In July 2025, they first released Causal-Copilot: an autonomous analysis agent integrating more than 20 causal algorithms that outperformed GPT-4o on specific benchmarks. That was a warm-up. In June 2026 at CVPR, Huang Biwei formally released the Causal World Model framework. 5. Three Predictions Based on current information, I make three judgments. Prediction 1: Within 18 months, the causal world model will achieve a generational gap in the dimension of "physical plausibility" in video generation. Not higher quality, higher frame rate, or larger resolution. It means that in generated videos, a cup really breaks when it hits the ground, the direction of fragment splashes really follows conservation of momentum, and water really flows downward. There will no longer be "fragments flying upward." Once this happens, those video generation companies that rely on brute-forcing data quality will be in trouble. Because it's not about insufficient optimization — the underlying paradigm has been bypassed. Prediction 2: In embodied intelligence, the first breakthrough in the sim2real gap will come from causal representation learning, not from larger simulation data. In April 2026, DexWorldModel published a paper that achieved zero-shot sim-to-real transfer on a physical robot using a causal implicit world model, surpassing baselines fine-tuned on real data. This is not a coincidence. Causal representations capture cross-domain invariant causal variables like "surface contact" and "gravity," unaffected by domain-specific noise such as pixels, lighting, and texture. This path can work. Prediction 3: If Aether AI's direction succeeds in productization, its valuation logic will not be "another AI video company," but infrastructure layer. At the end of her CVPR keynote, Huang Biwei made a very precise statement: "The causal world model is the last piece of the puzzle for a general world model." This is not an exaggeration. Large language models solved reasoning in the symbolic world. The causal world model is meant to solve reasoning in the physical world. The two together make the whole. If this is validated, Aether AI's position will not be a player in a specific vertical; it will be a new foundational layer. 6. A Correction, and a Dividing Line Let me go back to the opening sentence. "Compression is not complete intelligence. It should be structured compression that is intelligence." When I first heard this correction, I felt a moment of awakening. "Compression is intelligence" is one of the most popular creeds in the AI community in recent years. It is beautiful, and it is also right. In the correlation paradigm, GPT compresses the internet into its parameters, and then you use prompts to decompress. But Huang Biwei's correction tells you: the method of compression matters more than the amount of compression. With the same data, brute-force correlation compression and extracting causal structure from it yield different kinds of "intelligence." The latter has a higher slope per bit of data. This reminds me of a more fundamental issue. Large language models teach AI to predict the next word. The causal world model is meant to teach AI to understand how the world works. These are two different things. Aether AI is betting on the second. References: Aether AI / 原识之智: http://aetherlabs.ai
Original Article
View Cached Full Text

Cached at: 06/18/26, 06:20 PM

Today I saw an interview with Professor Biwei Huang (@huang_biwei) and learned about their latest funding round! After looking into Aether AI’s approach and spending some time on their direction, here are my thoughts:

The next paradigm in AI: not bigger models, but causality.

I. The Ceiling of Correlation: Why the Picture Is Beautiful but the Physics Is Fake

Let’s start with three data points.

First: Video generation. A December 2025 paper tested the strongest video generation models and found that the “gravity” they produce is only 1.81 m/s² — 18% of Earth’s gravity, about the same as the Moon. When the same model drops two objects from the same height, they hit the ground at different times. What Galileo wanted to prove at the Leaning Tower of Pisa, these models still haven’t learned.

Second: Embodied intelligence. The Stanford AI Index 2026 report gives a number: humanoid robots achieve 89.4% success in simulation, but only 12% in the real world. A 77-percentage-point gap. The What-If World benchmark tested 9 state-of-the-art world models, asking them to generate videos of “a physical intervention on a scene” — changing object mass, friction coefficient, light direction. None scored above 52% on pairwise matching, and open-source models clustered around 28%.

Third: Causal reasoning. HOCA-Bench divides AI video failures into two categories: ontological anomalies (objects vanishing, color flickering) and causal anomalies (wrong gravity direction, collision clipping, floating buoyancy). Video understanding models score over 20 percentage points lower on the latter. A model can tell that “a cat using chopsticks” is strange, but can’t tell that “a rock floating on water” violates physics — because it learned statistical patterns, not physical rules.

These three data points point to the same problem. Academics call it causal confusion.

Today’s AI learns correlations. Feed it enough data, and it learns to predict the next word, next image, next video. But it doesn’t know why. Ask it to generate “a glass cup falling on a marble floor” — it’s seen similar images, so it generates similar images. But it hasn’t understood gravity, hardness, or collision mechanics. So the shards fly upward.

There’s a precise engineering description for this: the model’s output depends on the visual salience of the intervention, not on physical computability. In plain English: it doesn’t check whether the physics is correct; it checks whether the picture looks like something it saw in the training data.

A robot that puts a red cube into the left box is baffled when the cube is blue. Adjust the table height by 1 cm, and a previously learned grasp fails outright. This isn’t about insufficient training data — it’s about the learned representation staying at the pixel level, never rising to the causal concept of “surface contact.”

Memorizing the answer key is not the same as understanding the principle.

II. Four Jumps: The Evolution of AI Paradigms

In her keynote at CVPR 2026, Biwei Huang broke down AI paradigms into four stages:

Jump 1: Small model × Correlation — already past
Jump 2: Small model × Causality — academic reserve
Jump 3: Large model × Correlation — we are here now
Jump 4: Large model × Causality — the next station

We are in Jump 3. GPT, Claude, Sora, Veo — essentially, they are all large models × correlation. They compress text, images, and video from the internet into hundreds of billions of parameters, learning a super-complex conditional probability distribution.

The problem? Ilya Sutskever himself announced at NeurIPS that the pre-training era is ending. His exact words: “Data is the fossil fuel of AI. There’s only one internet. Data will not grow anymore. We have reached peak data.”

It’s not that Scaling Laws have broken down. It’s that the object of scale needs to change.

This brings us to the core line of Huang’s keynote:

“Compression is not intelligence. Structured compression is intelligence.”

Just brute-force cramming data to squeeze out correlations has a very different slope than extracting causal structure from the same data.

III. What the Causal World Model Does Differently

The Causal World Model (CWM) proposed by Aether AI operates on a fundamentally different logic from today’s video generation models and world models. Huang gave three hard criteria:

First: Learn causal feature representations.
Today’s models learn from raw data what co-occurs with what. A causal world model must learn what causes what: recovering interpretable latent factors from pixels — object mass, surface friction, gravity direction, collision elasticity. These aren’t labels; they are causal variables that the model separates from data on its own.

Second: Understand causal structure.
It’s not just knowing that “cup” and “shards” are associated, but knowing “cup falls → hits ground → stress exceeds material strength → shatters → shards fly following momentum conservation.” This is a causal graph, not a pixel map. The model must map across levels — from the macroscopic “cup breaks” to the microscopic “crack propagation in glass.”

Third: Capture causal dynamics.
The world is not static. The same cup falling on marble will break; falling on carpet will not. Today’s models need to have seen both scenarios separately to generate them. A causal world model doesn’t need that — it knows the carpet absorbs the impact, so the causal chain breaks at the “collision” step, and the subsequent “shatter” and “flying” never happen.

This is reasoning about the physical world.

They designed a four-layer architecture to support this logic:

  • System Layer — causality-driven agent system for decision-making and planning
  • Foundation Model Layer — causal world model for core understanding and prediction
  • Neural Architecture Layer — modular design inspired by brain functional specialization
  • Infrastructure / Transformer Layer — modified Transformer, injecting causal dependencies at the token level

From the underlying tokens to the top-level decisions, causality is not an add-on; it runs throughout.

IV. Who Is Doing This

Let’s talk about the people.

Biwei Huang — PhD from CMU (graduated 2022), advised by Kun Zhang and Clark Glymour. Kun Zhang is a central figure in causal discovery; Glymour, along with Spirtes, laid the foundational work in 1989.

This academic lineage is worth making clear. In 1989, Glymour and Spirtes published the seminal work on causal discovery algorithms. Over the following 37 years, this school’s mission has been pure: to let machines discover causal relationships from observational data on their own, rather than waiting to be fed them. Kun Zhang’s team developed Causal-Learn, the most mainstream open-source Python library in causal discovery, and Biwei Huang is a core contributor.

She received the 2021 Apple Scholar award and was an assistant professor at UCSD’s Halıcıoğlu Data Science Institute. In 2025 she founded the company, registered in Shanghai, under the name 「原识之智」(Yuán Shí Zhī Zhì — roughly “wisdom that recognizes origins”). The name is apt: wisdom that understands the root.

This is not a typical AI entrepreneur pivoting from something else. It’s someone with a decade of academic accumulation in causal discovery, bringing a complete theoretical framework and open-source tools to engineering.

In July 2025, they first released Causal-Copilot: an autonomous analysis agent integrating over 20 causal algorithms, which outperformed GPT-4o on specific benchmarks. That was a warm-up. Then at CVPR 2026, Biwei Huang officially released the Causal World Model framework.

V. Three Predictions

Based on current information, I make three judgments.

Prediction 1: Within 18 months, the causal world model will create a generational gap in the “physical plausibility” dimension of video generation.

Not higher resolution, faster frame rate, or bigger scale. But in the generated video, a cup dropped on the ground will actually break; the shards will fly in directions that obey momentum conservation; water truly flows downhill. No more “shards flying upward.”

Once this happens, video generation companies that rely solely on piling up data and competing on visual quality will be in trouble. Because it’s not about insufficient optimization — the underlying paradigm has been bypassed.

Prediction 2: In embodied intelligence, the first breakthrough in closing the sim-to-real gap will come from causal representation learning, not from larger simulation datasets.

In April 2026, DexWorldModel published a paper: using a causal implicit world model, they achieved zero-shot sim-to-real transfer on a physical robot, surpassing the baseline that was fine-tuned on real data. This is not accidental. Causal representations capture cross-domain invariant causal variables — like “surface contact” and “gravity” — while being unaffected by domain-specific noise like pixels, lighting, and texture. This path works.

Prediction 3: If Aether AI’s direction succeeds in productization, its valuation logic will not be “another AI video company” — it will be infrastructure.

Biwei Huang ended her CVPR keynote with a very precise statement: “The causal world model is the last piece of the puzzle for a general world model.” That’s not an exaggeration.

Large language models solved reasoning in the symbolic world. Causal world models aim to solve reasoning in the physical world. Only together are they complete. If this is validated, Aether AI’s position will not be that of a player in some vertical domain. It will be a new foundational layer.

VI. A Correction and a Dividing Line

Let me return to the opening sentence.

“Compression is not intelligence. Structured compression is intelligence.”

When I first heard this correction, it felt like a moment of clarity.

“Compression is intelligence” has been one of the most popular AI dogmas in recent years. It’s beautiful, and it’s right. Under the correlation paradigm, GPT compresses the internet into parameters, and you use prompts to decompress it.

But Huang’s correction tells you: how you compress matters more than how much you compress. From the same data, brute-force compression to extract correlations and structured extraction of causal structure yield fundamentally different “intelligence.” The latter has a much higher slope per bit of data.

This reminds me of a more fundamental question.

Large language models taught AI to predict the next word. Causal world models will teach AI to understand how the world works.

These are two different things.

Aether AI is betting on the second.


Sources: Aether AI / 原识之智: http://aetherlabs.ai


Aether AI — Causal World Models for Real-World Intelligence

Source: https://aetherlabs.ai/ Aether AI (https://aetherlabs.ai/index.html)About (https://aetherlabs.ai/index.html)Blog (https://aetherlabs.ai/blog.html)News (https://aetherlabs.ai/news.html)Careers (https://aetherlabs.ai/careers.html)Contact (https://aetherlabs.ai/contact.html)Manifesto · 2026Aether AI

Aether is building a new class of AI systems that understand mechanisms, reason under intervention, and operate reliably in real-world systems.

Real intelligence requires models of how the world works.

The next AI paradigm will not be built on pattern recognition alone. AI systems can now recognize, generate, imitate, and predict at extraordinary scale. But the most important systems in the world are not passive distributions. Physical environments, biological systems, and scientific experiments respond when we act, perturb, measure, and change them.

Real intelligence requires models of how the world works: what variables matter, how they interact, how interventions change future states, and why outcomes occur. We call these systems causal world models.

Causal world models move AI beyond passive prediction — toward reasoning about consequences, counterfactuals, and interventions.

They connect observation, latent state, mechanism, action, and outcome — so a system can understand not only what is likely to happen, but what can be changed.

§ 01.5Causal loop

Observation becomes intervention, then new evidence.

The system repeatedly infers structure, tests an action, observes the changed world, and updates the model.

Physical AI is our first proving ground.

Robotics makes the problem concrete. A robot cannot act reliably by recognizing objects alone. It must understand contact, force, friction, support, constraints, affordances — and the physical dynamics that determine how the world changes under action.

Much of today’s robotics AI still maps observations directly to actions. These systems can learn useful behaviors in familiar settings, but they become brittle when objects, environments, timing, or task structures change. In long-horizon tasks, small errors compound; without an internal model of why an action failed, recovery often requires more data, retraining, or manual engineering.

Aether is building the decision brain for Physical AI — the intelligence layer between perception and control, where scene understanding becomes physical reasoning, and physical reasoning becomes action.

The same principle extends to scientific discovery.

In biology, medicine, and longevity, progress depends on understanding mechanisms — not just detecting patterns. Aging, for example, is shaped by interacting processes across metabolism, inflammation, cellular senescence, mitochondrial function, epigenetic regulation, immune response, and environment.

A causal world model should help distinguish drivers from markers, predict how interventions propagate through downstream states, and suggest experiments that separate competing explanations.

Across domains, the challenge is the same: discover what changes what, understand why, and use that understanding to decide how to intervene.

The Aether approach.

Aether builds causal world models that connect state, action, mechanism, and outcome. These models discover stable causal structure, simulate possible futures, compare counterfactual alternatives, estimate uncertainty, and update from real-world feedback.

The approach is a loop: infer hidden state from observation; reason about interventions; test the model through action or experiment; and use the gap between expectation and outcome to update the representation.

In Physical AI, this becomes a decision brain for robots. In scientific discovery, it becomes a way to generate hypotheses, design experiments, and uncover mechanisms not visible from observation alone.

The next generation of AI will require both scale and structure. Scale provides capacity. Causal structure makes that capacity reliable, reusable, and grounded.

Aether is building AI that does not only predict outcomes, but learns the mechanisms that make reliable intervention possible.

Who We Are

Our founding team are leading experts in causal discovery, causal AI, causal foundation models, causal reinforcement learning, agentic systems, and foundation model training.

Biwei Huang (@huang_biwei): I’ve spent over a decade working on causal discovery and causal AI. A lot of late nights, a lot of papers, and a lot of open questions.

Today we’re putting something into the world. Aether AI has raised $20M to build causal world models that understand mechanisms. We believe the

Similar Articles

@gkxspace: LLM is likely just the first stop for AI large models. Professor Biwei Huang divides AI paradigms into four generations: First generation (1990s): Small models learn correlations. Second generation (2010s): Small models learn causation. Third generation (current LLMs): Large models learn correlations. Fourth generation (next step): Large models learn causation. Over 30 years, models have grown from small to large...

X AI KOLs Timeline

Professor Biwei Huang proposes a four-generation theory of AI paradigms, believing LLMs are just the first step, and the future lies in causal world models. Aether AI has completed a $20 million funding round, dedicated to building causal world models.

@dashen_wang: https://x.com/dashen_wang/status/2065053748746240161

X AI KOLs Timeline

The article delves into the naming philosophy behind Anthropic's release of the Fable and Mythos models, pointing out that the widespread application of AI is still dominated by 'reconstructing the known' (e.g., fixing bugs), while 'creating the unknown' is the truly scarce capability. It also discusses the trend of AI companies starting to hire philosophers, arguing that this marks the beginning of a mythological era of 'legislating for creation.'