Jay from Wai Fei Si
Quantum Bit | Public Account QbitAI
0.2 seconds.
It was a moment when the robot paused urgently while closing the microwave oven door, due to someone suddenly stepping in.
There are even more interesting ones.
When a metal can is mixed with a cake plate and brought towards the microwave oven, the robotic arm will first pick out the can and place it aside. After ensuring that the plate is safe, it will then send the can into the oven.
To be honest, when watching these demos, I was quite stunned.
Friends who have been following embodied intelligence for a long time should have noticed that many robots now look quite cool, but most of the time they are actually doing something wrong. Once the lighting changes, something is kicked, or the instructions have some implicit intention, the robot instantly goes haywire.
But even like in the video, if you frantically shine a flashlight into the robot’s eyes, it still remains oblivious...
However, those scenes just now were not pre-recorded action replays, nor was it a coincidental correct statistical correlation.
In the face of continuous sudden disturbances, the system relies on its understanding of the causal laws in the physical world to infer this series of actions in real time.
The system that drives this robot is codenamed CRIS-0.
Behind it is the Causal Intelligence company founded by Huang Biwei, an assistant professor at UCSD and a scholar in the CMU Causal School.
Aether AI.
Three months ago, QuantumLeap reported on the causal intelligence path chosen by the company. At that time, many people thought that causal AI was still stuck in mathematical formulas and thought experiments.
Today, Aether AI has released its official Demo.
Say goodbye to mechanical "wasting time": Robots learn to understand cause and effect
In the implementation scenarios of embodied robots, what causes the most trouble for engineers is not the standard process, but the ubiquitous accidents.
Traditional robots, when faced with sudden disturbances, either mechanically repeat old coordinates leading to collisions, or simply stop and report an error, getting stuck. Meanwhile, end-to-end black box models are extremely prone to cumulative errors during long tasks of several dozen steps; a single misstep can cause everything to collapse.
The breakthrough solution provided by CRIS-0 is “causality”—
Let the robot move from “what it sees” to “understanding why things happen and what changes affect the outcome”.
First, look at the report card.
In the Coffee Preparation test where external sudden disturbances occur, the robot encountered artificial displacement and sudden changes in lighting when grasping the coffee machine. CRIS-0 identified the change in causal variables within 2 seconds on average and completed real-time replanning, achieving 9 effective recoveries out of 10 random disturbances.
And it isn’t just blindly starting from scratch for the entire process.
The coffee machine was pushed aside, and the system tracks the key causal variable of “the handle relative pose”. When the status of this variable shifts, it only corrects the approaching and grasping stages, without needing to restart from the original position.
Anti-interference is just basic skills; in long-distance autonomous missions, the advantages of causal driving are more pronounced.
In a continuous task that involves dozens of steps, such as “finding and organizing items in the living room,” the system breaks down the long-term goal into atomic, verifiable cause-and-effect stages, and checks preconditions and postconditions at each step.
What’s even more interesting is the common sense reasoning and physical perception abilities shown by CRIS-0:
When cleaning the desk, it organizes scattered ordinary books onto the coffee table, but puts bills related to personal privacy into the drawer;
Facing two drink cans that look similar, it first senses the weight of the can and the state of the liquid inside by grabbing and shaking them. Only after confirming that it is an empty can does it throw it into the trash can. For the undrunked drink, it returns it firmly to its original position.
Finally, when dealing with ambiguous needs, causal reasoning can significantly improve the generalization of prompt words, thereby enabling personalized understanding capabilities.
In the Personal Pick and Place test, when faced with 20 vague instructions that require implicit reasoning, such as “get me a bottle of drink suitable for morning running,” the system achieved 18 accurate grabs and placements.
It doesn’t rely on memorizing keywords to try its luck, but instead combines time, user status, and environmental context to extract “energy replenishment after exercise and avoidance of high sugar” as hidden causal variables, and automatically match the corresponding actions.
What is the actual technical architecture behind these Demos?
Disassemble CRIS-0: Transform the black box into a causal engine
Before answering this question, perhaps we should first clarify another issue—
Why did previous robots always seem “stupid” to people?
The physical world and the pure digital world are completely different. When writing code or chatting on a computer, you can undo mistakes, but in the real world, friction, gravity, collisions, and changes in light occur all the time every second due to various unexpected situations.
Traditional end-to-end models, such as the VLA scheme, learn quickly, but they have a fatal weakness when dealing with long-process tasks—error accumulation.
Step 1 has a slight deviation, step 2 has a larger deviation, and by step 10, the entire process may have completely changed. What’s worse is that when a single black box model encounters interference, it doesn’t know what’s happening, and often it can only keep pressing according to the original plan, or it may directly crash and stop working.
And another common Agent solution, although it breaks down the steps, its thinking logic basically remains at the level of pure text or probabilistic reasoning in multi-modal large models.
Every time the environment changes, it has to first convert the image into text. Then the model slowly figures out “what to do next”. By the time you finish reasoning, your hand might already be caught by the door...
How is it possible to achieve this anti-interference and coherence in CRIS-0?
Aether AI has designed a new set of logic:
Cause-and-effect native Agent architecture.
It can be understood from three dimensions—
1. Task is State: Extract Physical Causal Variables
This is what I consider to be the most important First Principle.
In the CRIS-0 system, there is a very low-level change called unified state and causal variables.
In the past, robots viewed the world through dense images of pixels. But in the CRIS-0, what matters is not how the world’s surface looks, but which step the task has reached. The task itself is a causal state progress bar.
Here’s a very practical example: for instance, a robot spreading tablecloths.
It doesn’t foolishly calculate the coordinates of every pixel on the tablecloth, but instead tracks several key causal variables in real time. For example, whether the tablecloth is actually captured by the grippers, how many centimeters away the corners of the tablecloth are from the target position, and whether there is any slipping when pulled.
Once the grab slips off, the system immediately knows that the causal variable “grab” is not satisfied. Then the task status directly returns to the stage of re-grabbing, rather than overturning the entire action and starting over, or blindly continuing to pull air further back.
In the officially published coffee-making interference test, facing the tormenting scenarios of the coffee machine shifting repeatedly and sudden changes in lighting, the system can identify the changes in key variables and complete replanning within 2 seconds on average, achieving effective recovery 9 out of 10 times during 10 external disturbances.
This is the power of causal variables: if something goes wrong, it clearly knows where the error occurred and which step to revert to.
2. Unified tool interface: Reject a single strategy dominating everything
How specifically does this system execute the Action?
Answer: Tool Call, with each tool performing its own function.
In its architecture, there is a Planner responsible for overall scheduling, which oversees a fully functional toolbox that integrates various modules through the Unified Tool Interface.
Rule-based functions: Combine visual positioning and inverse kinematics to efficiently handle large-scale spatial displacement and pre-positioning.
Skill Strategy Model (Policy models): Specifically designed for high-difficulty contact-based operations such as pouring and precise grabbing.
Navigation Module (SLAM): Responsible for global path guidance;
Verifier: Performs physical inspections at each stage node;
Causal World Model (CausalWM): Responsible for simulating the physical consequences of actions.
Among them, the most crucial trump card is precisely this causal world model.
Many friends when talking about the world model, the first reaction is to generate realistic videos for the next few seconds. However, in CausalWM in CRIS-0, more attention is given to the impact of actions on the environment.
Its logic is to first look at the current state, then predict which causal variables will change due to the candidate actions, and finally derive the future state.
In other words, before the robot reaches out, its brain is already doing calculations, for example: If I do this, will my hand be off-center? Will the cup tip over?
It is equivalent to adding an “rehearsal” step. The system first simulates in the world model, “What will the physical causal variables be like if I perform action B?” After confirming that the goal can be achieved, it calls the Action Head and outputs the underlying control.
3. Structured loops and native security built-in
And all of this, it repeats in a loop.
When facing complex tasks, the Planner breaks down large goals into smaller steps. First, it calls a rule function to quickly move the grippers near the target, and then it uses a strategy model to achieve precise contact and grasping. Complex problems that were originally difficult to solve end-to-end with a single model are broken down into several highly deterministic small steps.
Specifically, it is this dynamic evolving task graph: state identification → tool selection → execution → verification → adjustment.
After each step is completed, the verifier immediately checks whether the physical conditions are met. If not, the system has a clear three-level recovery mechanism: it first determines whether it’s possible to retry directly at this stage; if not, it re plans the path; and only if that fails does it indicate that manual intervention is required.
Once an unexpected adjustment succeeds, this successful recovery path will be recorded and stored in the task diagram. The next time a similar situation occurs, it will become more proficient.
This also explains why in that long-term organizing task in the living room, it can continue to execute dozens or even hundreds of steps without crashing.
Because every step it takes checks the causal state, the error is resolved on the spot as soon as it occurs, and there is no chance for it to spread like a snowball and cause a global collapse.
Similarly, this is also why it can achieve a safe emergency stop in 0.2 seconds.
In CRIS-0, the safety judgment is directly embedded in each causal stage.
At the closing stage, suddenly appearing hands are a dangerous variable that disrupts user safety conditions. The system can instantly detect this and trigger the highest-priority blocking action, stopping in 0.2 seconds.
But if the current task is to hand a glass of water to humans, then the extended hand is the target variable for interaction, and the system will not stop randomly.
In fact, the system-level capabilities demonstrated by CRIS-0 did not arise out of thin air; they are supported by a solid research foundation.
In the Agent capability assessment, previously the Aether AI causal agent framework RSIAgent achieved a Partial Score of 78.98% on OSWorld 2.0. Without updating model parameters, it was able to enhance the Agent capabilities of open-source models, surpassing closed-source models such as GPT-6 Astra.
Its first version of the causal world model CausalWM also topped the Leaderboard on the robot world model benchmark TriWorldBench, achieving Top-1.
And these two layers are the main contributors to this Demo.
Additionally, two more layers are also being advanced—the modular neural architecture: each module corresponds to different causal mechanisms, and can be combined or replaced; Causation Transformer: learning causal relationships rather than statistical dependencies.
The physical world doesn’t trust rote memorization
At this point, we could actually step away from the specific robot Demo and take a deeper look.
Why is the topic of causal intelligence being discussed more and more frequently these days?
To be honest, the recent surge in large language models was largely due to good luck. Language itself is a highly symbolic and abstract modality that humans have represented through texts. Many logic and rules have already been summarized by humans on the surface of texts. It’s possible to achieve miracles through statistical fitting based on probability correlations.
But the physical world doesn’t follow this set of rules.
Friction does not change just because you read a hundred thousand papers, and gravity also does not suddenly disappear due to probability statistics. In the real world filled with complex dynamics, simply relying on superficial correlations will eventually hit the Scaling Law limit.
In the case of severe physical data shortages, to create a system that can truly cope with complex physical environments, the model must learn to, like human scientists, extract underlying causal mechanisms from limited data, and understand how actions actually change the world.
This compression structured with the “causality” first-principle approach might be the embodied native Scaling Law.
Of course, this path is extremely difficult to traverse. The requirements for causal discovery and causal representation learning in mathematical and statistical theory are very high, and there are few teams worldwide that truly combine deep theoretical expertise with practical engineering capabilities.
Aether AI is willing to get involved, which is largely related to its unique academic framework too.
Huang Biwei has spent over a decade in this field. She studied under the founders of the CMU Causal School, Clark Glymour and Peter Spirtes, as well as second-generation scholars Bernhard Schölkopf and Kun Zhang. She is one of the key inheritors of the Causal Discovery School.
The theoretical framework of the Three Generations Causality School has ultimately become the underlying technical foundation of Aether AI today.
From the theoretical foundations of Carnegie Mellon University and Max Planck Institute, to the tangible implementation of CRIS-0 on real robotic arms today, causal intelligence has finally taken a crucial step towards the physical world.
When a system truly possesses the ability to extract causal variables, model world dynamics, and reason about consequences before taking action, its applications will far go beyond simply helping people turn on a microwave oven or pour a cup of coffee.
Looking further ahead, from trend predictions of complex dynamic systems to scientific discoveries in fields such as materials and life sciences, perhaps this logic can be applied across all of them.
Great times, friends.
We have been hoping that AI can truly step out of the pure numerical chat box and stand firmly in this real, complex, and unknown physical world.
And causality, perhaps, is that most important key.
— End —
QbitAI · Contracted headline account
Follow us to get up-to-date information on cutting-edge technology immediately
