@MindfulReturn: Today I saw an interview with Professor Huang Biwei (@huang_biwei) and learned about their new round of funding! After learning about the Aether AI solution and taking a closer look at their direction, let me share my thoughts: The next paradigm of AI is not bigger models, but causality. 1. Correlation Ceiling: Why the visuals are...
Summary
This article offers an in-depth analysis of the Causal World Model (CWM) proposed by Aether AI (原识之智), arguing that the next AI paradigm will shift from correlation to causation. It discusses the theoretical foundations, technical architecture, and potential impact on video generation and embodied intelligence.
View Cached Full Text
Cached at: 06/18/26, 06:20 PM
Today I saw an interview with Professor Biwei Huang (@huang_biwei) and learned about their latest funding round! After looking into Aether AI’s approach and spending some time on their direction, here are my thoughts:
The next paradigm in AI: not bigger models, but causality.
I. The Ceiling of Correlation: Why the Picture Is Beautiful but the Physics Is Fake
Let’s start with three data points.
First: Video generation. A December 2025 paper tested the strongest video generation models and found that the “gravity” they produce is only 1.81 m/s² — 18% of Earth’s gravity, about the same as the Moon. When the same model drops two objects from the same height, they hit the ground at different times. What Galileo wanted to prove at the Leaning Tower of Pisa, these models still haven’t learned.
Second: Embodied intelligence. The Stanford AI Index 2026 report gives a number: humanoid robots achieve 89.4% success in simulation, but only 12% in the real world. A 77-percentage-point gap. The What-If World benchmark tested 9 state-of-the-art world models, asking them to generate videos of “a physical intervention on a scene” — changing object mass, friction coefficient, light direction. None scored above 52% on pairwise matching, and open-source models clustered around 28%.
Third: Causal reasoning. HOCA-Bench divides AI video failures into two categories: ontological anomalies (objects vanishing, color flickering) and causal anomalies (wrong gravity direction, collision clipping, floating buoyancy). Video understanding models score over 20 percentage points lower on the latter. A model can tell that “a cat using chopsticks” is strange, but can’t tell that “a rock floating on water” violates physics — because it learned statistical patterns, not physical rules.
These three data points point to the same problem. Academics call it causal confusion.
Today’s AI learns correlations. Feed it enough data, and it learns to predict the next word, next image, next video. But it doesn’t know why. Ask it to generate “a glass cup falling on a marble floor” — it’s seen similar images, so it generates similar images. But it hasn’t understood gravity, hardness, or collision mechanics. So the shards fly upward.
There’s a precise engineering description for this: the model’s output depends on the visual salience of the intervention, not on physical computability. In plain English: it doesn’t check whether the physics is correct; it checks whether the picture looks like something it saw in the training data.
A robot that puts a red cube into the left box is baffled when the cube is blue. Adjust the table height by 1 cm, and a previously learned grasp fails outright. This isn’t about insufficient training data — it’s about the learned representation staying at the pixel level, never rising to the causal concept of “surface contact.”
Memorizing the answer key is not the same as understanding the principle.
II. Four Jumps: The Evolution of AI Paradigms
In her keynote at CVPR 2026, Biwei Huang broke down AI paradigms into four stages:
Jump 1: Small model × Correlation — already past
Jump 2: Small model × Causality — academic reserve
Jump 3: Large model × Correlation — we are here now
Jump 4: Large model × Causality — the next station
We are in Jump 3. GPT, Claude, Sora, Veo — essentially, they are all large models × correlation. They compress text, images, and video from the internet into hundreds of billions of parameters, learning a super-complex conditional probability distribution.
The problem? Ilya Sutskever himself announced at NeurIPS that the pre-training era is ending. His exact words: “Data is the fossil fuel of AI. There’s only one internet. Data will not grow anymore. We have reached peak data.”
It’s not that Scaling Laws have broken down. It’s that the object of scale needs to change.
This brings us to the core line of Huang’s keynote:
“Compression is not intelligence. Structured compression is intelligence.”
Just brute-force cramming data to squeeze out correlations has a very different slope than extracting causal structure from the same data.
III. What the Causal World Model Does Differently
The Causal World Model (CWM) proposed by Aether AI operates on a fundamentally different logic from today’s video generation models and world models. Huang gave three hard criteria:
First: Learn causal feature representations.
Today’s models learn from raw data what co-occurs with what. A causal world model must learn what causes what: recovering interpretable latent factors from pixels — object mass, surface friction, gravity direction, collision elasticity. These aren’t labels; they are causal variables that the model separates from data on its own.
Second: Understand causal structure.
It’s not just knowing that “cup” and “shards” are associated, but knowing “cup falls → hits ground → stress exceeds material strength → shatters → shards fly following momentum conservation.” This is a causal graph, not a pixel map. The model must map across levels — from the macroscopic “cup breaks” to the microscopic “crack propagation in glass.”
Third: Capture causal dynamics.
The world is not static. The same cup falling on marble will break; falling on carpet will not. Today’s models need to have seen both scenarios separately to generate them. A causal world model doesn’t need that — it knows the carpet absorbs the impact, so the causal chain breaks at the “collision” step, and the subsequent “shatter” and “flying” never happen.
This is reasoning about the physical world.
They designed a four-layer architecture to support this logic:
- System Layer — causality-driven agent system for decision-making and planning
- Foundation Model Layer — causal world model for core understanding and prediction
- Neural Architecture Layer — modular design inspired by brain functional specialization
- Infrastructure / Transformer Layer — modified Transformer, injecting causal dependencies at the token level
From the underlying tokens to the top-level decisions, causality is not an add-on; it runs throughout.
IV. Who Is Doing This
Let’s talk about the people.
Biwei Huang — PhD from CMU (graduated 2022), advised by Kun Zhang and Clark Glymour. Kun Zhang is a central figure in causal discovery; Glymour, along with Spirtes, laid the foundational work in 1989.
This academic lineage is worth making clear. In 1989, Glymour and Spirtes published the seminal work on causal discovery algorithms. Over the following 37 years, this school’s mission has been pure: to let machines discover causal relationships from observational data on their own, rather than waiting to be fed them. Kun Zhang’s team developed Causal-Learn, the most mainstream open-source Python library in causal discovery, and Biwei Huang is a core contributor.
She received the 2021 Apple Scholar award and was an assistant professor at UCSD’s Halıcıoğlu Data Science Institute. In 2025 she founded the company, registered in Shanghai, under the name 「原识之智」(Yuán Shí Zhī Zhì — roughly “wisdom that recognizes origins”). The name is apt: wisdom that understands the root.
This is not a typical AI entrepreneur pivoting from something else. It’s someone with a decade of academic accumulation in causal discovery, bringing a complete theoretical framework and open-source tools to engineering.
In July 2025, they first released Causal-Copilot: an autonomous analysis agent integrating over 20 causal algorithms, which outperformed GPT-4o on specific benchmarks. That was a warm-up. Then at CVPR 2026, Biwei Huang officially released the Causal World Model framework.
V. Three Predictions
Based on current information, I make three judgments.
Prediction 1: Within 18 months, the causal world model will create a generational gap in the “physical plausibility” dimension of video generation.
Not higher resolution, faster frame rate, or bigger scale. But in the generated video, a cup dropped on the ground will actually break; the shards will fly in directions that obey momentum conservation; water truly flows downhill. No more “shards flying upward.”
Once this happens, video generation companies that rely solely on piling up data and competing on visual quality will be in trouble. Because it’s not about insufficient optimization — the underlying paradigm has been bypassed.
Prediction 2: In embodied intelligence, the first breakthrough in closing the sim-to-real gap will come from causal representation learning, not from larger simulation datasets.
In April 2026, DexWorldModel published a paper: using a causal implicit world model, they achieved zero-shot sim-to-real transfer on a physical robot, surpassing the baseline that was fine-tuned on real data. This is not accidental. Causal representations capture cross-domain invariant causal variables — like “surface contact” and “gravity” — while being unaffected by domain-specific noise like pixels, lighting, and texture. This path works.
Prediction 3: If Aether AI’s direction succeeds in productization, its valuation logic will not be “another AI video company” — it will be infrastructure.
Biwei Huang ended her CVPR keynote with a very precise statement: “The causal world model is the last piece of the puzzle for a general world model.” That’s not an exaggeration.
Large language models solved reasoning in the symbolic world. Causal world models aim to solve reasoning in the physical world. Only together are they complete. If this is validated, Aether AI’s position will not be that of a player in some vertical domain. It will be a new foundational layer.
VI. A Correction and a Dividing Line
Let me return to the opening sentence.
“Compression is not intelligence. Structured compression is intelligence.”
When I first heard this correction, it felt like a moment of clarity.
“Compression is intelligence” has been one of the most popular AI dogmas in recent years. It’s beautiful, and it’s right. Under the correlation paradigm, GPT compresses the internet into parameters, and you use prompts to decompress it.
But Huang’s correction tells you: how you compress matters more than how much you compress. From the same data, brute-force compression to extract correlations and structured extraction of causal structure yield fundamentally different “intelligence.” The latter has a much higher slope per bit of data.
This reminds me of a more fundamental question.
Large language models taught AI to predict the next word. Causal world models will teach AI to understand how the world works.
These are two different things.
Aether AI is betting on the second.
Sources: Aether AI / 原识之智: http://aetherlabs.ai
Aether AI — Causal World Models for Real-World Intelligence
Source: https://aetherlabs.ai/ Aether AI (https://aetherlabs.ai/index.html)About (https://aetherlabs.ai/index.html)Blog (https://aetherlabs.ai/blog.html)News (https://aetherlabs.ai/news.html)Careers (https://aetherlabs.ai/careers.html)Contact (https://aetherlabs.ai/contact.html)Manifesto · 2026Aether AI
Aether is building a new class of AI systems that understand mechanisms, reason under intervention, and operate reliably in real-world systems.
Real intelligence requires models of how the world works.
The next AI paradigm will not be built on pattern recognition alone. AI systems can now recognize, generate, imitate, and predict at extraordinary scale. But the most important systems in the world are not passive distributions. Physical environments, biological systems, and scientific experiments respond when we act, perturb, measure, and change them.
Real intelligence requires models of how the world works: what variables matter, how they interact, how interventions change future states, and why outcomes occur. We call these systems causal world models.
Causal world models move AI beyond passive prediction — toward reasoning about consequences, counterfactuals, and interventions.
They connect observation, latent state, mechanism, action, and outcome — so a system can understand not only what is likely to happen, but what can be changed.
§ 01.5Causal loop
Observation becomes intervention, then new evidence.
The system repeatedly infers structure, tests an action, observes the changed world, and updates the model.
Physical AI is our first proving ground.
Robotics makes the problem concrete. A robot cannot act reliably by recognizing objects alone. It must understand contact, force, friction, support, constraints, affordances — and the physical dynamics that determine how the world changes under action.
Much of today’s robotics AI still maps observations directly to actions. These systems can learn useful behaviors in familiar settings, but they become brittle when objects, environments, timing, or task structures change. In long-horizon tasks, small errors compound; without an internal model of why an action failed, recovery often requires more data, retraining, or manual engineering.
Aether is building the decision brain for Physical AI — the intelligence layer between perception and control, where scene understanding becomes physical reasoning, and physical reasoning becomes action.
The same principle extends to scientific discovery.
In biology, medicine, and longevity, progress depends on understanding mechanisms — not just detecting patterns. Aging, for example, is shaped by interacting processes across metabolism, inflammation, cellular senescence, mitochondrial function, epigenetic regulation, immune response, and environment.
A causal world model should help distinguish drivers from markers, predict how interventions propagate through downstream states, and suggest experiments that separate competing explanations.
Across domains, the challenge is the same: discover what changes what, understand why, and use that understanding to decide how to intervene.
The Aether approach.
Aether builds causal world models that connect state, action, mechanism, and outcome. These models discover stable causal structure, simulate possible futures, compare counterfactual alternatives, estimate uncertainty, and update from real-world feedback.
The approach is a loop: infer hidden state from observation; reason about interventions; test the model through action or experiment; and use the gap between expectation and outcome to update the representation.
In Physical AI, this becomes a decision brain for robots. In scientific discovery, it becomes a way to generate hypotheses, design experiments, and uncover mechanisms not visible from observation alone.
The next generation of AI will require both scale and structure. Scale provides capacity. Causal structure makes that capacity reliable, reusable, and grounded.
Aether is building AI that does not only predict outcomes, but learns the mechanisms that make reliable intervention possible.
Who We Are
Our founding team are leading experts in causal discovery, causal AI, causal foundation models, causal reinforcement learning, agentic systems, and foundation model training.
Biwei Huang (@huang_biwei): I’ve spent over a decade working on causal discovery and causal AI. A lot of late nights, a lot of papers, and a lot of open questions.
Today we’re putting something into the world. Aether AI has raised $20M to build causal world models that understand mechanisms. We believe the
Similar Articles
@gkxspace: LLM is likely just the first stop for AI large models. Professor Biwei Huang divides AI paradigms into four generations: First generation (1990s): Small models learn correlations. Second generation (2010s): Small models learn causation. Third generation (current LLMs): Large models learn correlations. Fourth generation (next step): Large models learn causation. Over 30 years, models have grown from small to large...
Professor Biwei Huang proposes a four-generation theory of AI paradigms, believing LLMs are just the first step, and the future lies in causal world models. Aether AI has completed a $20 million funding round, dedicated to building causal world models.
@wanerfu: Top talents are quietly leaving ChatAI to take on Physical AI (the next OpenAI) · Fei-Fei Li → World Labs · LeCun → AMI Labs · DeepMind/Stanford/Berkeley → …
Top AI talent is shifting from language models to physical AI, such as Fei-Fei Li founding World Labs, LeCun joining AMI Labs, and Aether AI focusing on causal world models, aiming to build AI systems that understand mechanisms and causal relationships, applied to robotics and scientific discovery.
@cjziems: We're going live in 30 minutes, and we'd love to have you join Joined by @dorazhao9 and @Diyi_Yang, I'll be talking abo…
The article introduces the live discussion of the Augmented Mind podcast about the paper 'Reflections and New Directions for Human-Centered Large Language Models', emphasizing that AI development should shift from capability benchmarks to human flourishing and long-term well-being.
@dashen_wang: https://x.com/dashen_wang/status/2065053748746240161
The article delves into the naming philosophy behind Anthropic's release of the Fable and Mythos models, pointing out that the widespread application of AI is still dominated by 'reconstructing the known' (e.g., fixing bugs), while 'creating the unknown' is the truly scarce capability. It also discusses the trend of AI companies starting to hire philosophers, arguing that this marks the beginning of a mythological era of 'legislating for creation.'
@snowboat84: https://x.com/snowboat84/status/2070656715515932930
This article details the new paradigm of AI for Science (AI4S), from AI as an analysis tool to the transition to scientific agents, explaining autonomy levels, key cases, and future trends.