By Tian Yanlin, from QbitAI
QbitAI | Official account QbitAI
Wow, in the World Action Model (WAM) track, a Chinese team has charged straight to global first?!
Moreover, a host of leading global embodied models, including GPT-6-Astra, Physical Intelligence's π0.5, and NVIDIA GR00T-N1.7, were all left behind this time.
The one making the move is a Chinese robotics company——Robot Era, the only directly held embodied intelligence company backed by Tsinghua。
Recently, Robot Era's self-developed world action model VPP2climbed to the top ofthe RoboDojo simulation leaderboard。
in one stroke, with a comprehensive average success rate of 32.26% and a comprehensive average score of 39.26 points, ranking first in both metrics.
The key is that this competition is really not easy to win.
RoboDojo is known asthe 'Mount Everest' of embodied intelligence, jointly built by the University of Hong Kong's MMLab and nearly 20 top academic institutions worldwide.
It deliberately sets questions in areas robots are bad at: can they still work in a different environment? Can they still grab things accurately when they're misplaced? Halfway through a task, do they still remember what happened before?
In short, muddling through by grinding proficiency isn't so easy.
And the result?
Robot Era's VPP2 not only took first place overall, but also topped the rankings in three dimensions:generalization ability, fine manipulation, and memory.
Reportedly, this time VPP2added no extra data, andused no enhancement methods like Agent RSI,, relying solely on the 'standard dataset'to outcompete a host of top players.
This means the VPP2 model itself is strong enough, with performance gains coming from pretraining and a base model general enough that it doesn't rely on enhancement plugins or 'score-boosting' through data expansion.
Wait, buddy, how exactly was such a rock-solid model trained???
We dug into the technical route, and you know what, VPP2's breakthrough this time really has something to it.
It targets a long-standing headache for world models: once the prediction is wrong, the actions that follow are wrong too.
The solution VPP2 offers is to first train its video prediction capability to be strong enough, and then let the robot actually move in unfamiliar environments.
From prediction to action, two kinds of generalization capability improve together.
Judging from the multiple test results disclosed so far, VPP2 has become one of the most competitive players in the world action model (WAM) track.
World action model WAM: a new champion emerges from the competition
The embodied AI industry has a rather awkward phenomenon.
Robot demos are getting flashier and flashier, and model leaderboard scores keep climbing, but if you really ask: whose robot is smarter and more capable of getting work done?
It's genuinely hard to answer.
For example, have a robot grab a cup. On a table it saw during training, it might hit the mark every time. But change the cup or shift its position, and it immediately starts questioning its existence.
Not to mention tasks that require executing multiple consecutive steps.
This is why the RoboDojo evaluation suite is worth paying attention to.
Its goal is to establish a unified, reproducible evaluation standard for embodied intelligence. The simulation testsinclude 42 dual-arm manipulation taskscovering five dimensions: generalization, precise manipulation, long-horizon tasks, memory, and open-vocabulary instruction understanding.
Put simply, it means pulling robots out of their comfort zone.
Traditional robot models may score near full marks on familiar tasks, but once the environment, object positions, or task combinations change, their success rates can drop noticeably.
RoboDojo specifically examines these kinds of complex situations, measuring how much generalization ability a robot has when facing changes.
And the scorecard VPP2 has delivered this time is quite impressive.
An average success rate of 32.26% and an average score of 39.26 points — both indicators rank first.
For comparison, in the post-training evaluation of the RoboDojo simulation benchmark,GPT-6-Astraachieved an average success rate of 22.48% and an average score of 28.97 points.
whileVPP2leads by 9.78 percentage points and 10.29 points respectively.
Screenshot from the paper; baseline model data taken from the official leaderboard
The two indicators each have a focus: the success rate measures whether the robot can complete a task in full; the average score measures how far the task progressed, reflecting stage-by-stage performance even when the task wasn't fully completed.
Both indicators combine test results across five major capability dimensions.
In other words,VPP2 not only completes tasks at a higher rate, but also demonstrates stronger stage-by-stage completion ability on tasks it couldn't finish.
Moreover, this lead is not limited to overall scores.
Breaking down the five capability dimensions, VPP2 also took first place in three of them: generalization, precise manipulation, and memory.
These three capabilities correspond to several hurdles robots must clear in the real world: can it still work in a different environment, can it operate with precision, and can it remember previous steps when executing tasks in sequence.
Leading in overall scores while also being strong in key capabilities.
However, no matter how hard RoboDojo is, the testing ground is still a "simulation environment." In the physical world, object friction, deformation, and positional deviations introduce new variables.
Can the generalization ability VPP2 demonstrated in simulationstill hold on a real robot?
Robot Era also ran experiments.
This time, the team directly deployed VPP2on a real ALOHA dual-arm robotand tested 10 categories of zero-shot manipulation tasks, including grasping, placing, stacking, folding, and pouring.
That is, without any additional fine-tuning for these test tasks, the robot had to dive right in.
As a result, VPP2 once again delivered a leading performance:an average success rate of 58.5%, higher than π0.5's 40%.
Among the 10 task categories, VPP2 achieved the best results in 9 of them.
Now this is getting interesting.
Ranking first on the leaderboard proves its capability under standard evaluation systems; zero-shot operation on real hardware further tests whether those capabilities can transfer to the real physical environment.
With the two report cards side by side, the technical advantages of VPP2 are even more worth digging into.
You should know that VPP2 takes theWorld Action Model (WAM) route。
The basic idea of this kind of model is to use video prediction to understand how the physical world will change next, and then convert that prediction into robot actions.
But this route has long had a problem: beautiful video predictions don't mean the robot can actually get work done.
And this time, VPP2 is aimed squarely at that problem.
From prediction generalization to action generalization, how does VPP2 do it?
To understand VPP2's breakthrough, first consider a simple task: getting a robot to put the cup on the table into a box.
For humans, reaching out, grasping, lifting, and placing it in requires almost no thought.
But the robot has to determine the cup's position, the trajectory of the robotic arm, when the gripper should close, and also predict what changes will occur in the physical world throughout the entire operation. A deviation at any step could cause a failure.
Existing world action models (WAM) are precisely prone to failing here.
On the one hand, ordinary video models are good at generating visuals, but visuals that look plausible and visuals that conform to real physical laws are two different things.
For example, the model may be asked to grab the cup on the left but predicts grabbing the one on the right instead. The video can still be generated, yet when the robot executes the action it will go off track.
On the other hand, directly adding action learning into video models may also damage the original generalization capability.
The robot learned the action, but when placed in a different environment it could no longer perform it.
VPP2 puts forward a very direct judgment:The quality of video prediction determines the upper limit of the action.
Following this approach, Xingdong Jiyuan chose to train video prediction and action learning in separate stages; the focus is not on "stewing video and action together," but on keeping the two"Decoupling", again"Sort"。
Xingdong Jiyuan designed a set for this purposeThree-stage training strategy:
- In the first stage, event-level video continues pretraining, teaching the model to predict complete operation processes.
- In the second stage, fixed-duration video post-training and distillation keep the prediction capability in step with the speed at which robots execute in real time.
- In the third stage, action expert training converts video prediction into concrete actions while trying to preserve the original generalization ability.
These three stages build on each other step by step, ultimately pointing towardtwo core breakthroughs: prediction that generalizes, and action that generalizes as well.
The first breakthrough: giving video prediction generalization capability
What VPP2 must first solve is whether a robot can accurately predict the physical changes in an unfamiliar task.
Galbot builds on Wan2.1-I2V-14B, an open-source model from Alibaba, integrating diverse data including robot manipulation, human activities, and general videos, covering different robot embodiments and manipulation methods.
What is truly interesting, however, is its handling of the data.
The team did not simply feed videos to the model in bulk; instead, they first segmented the manipulation processes into semantically complete clips, then paired them with detailed descriptions.
For example, you cannot just tell the model "put the cup into the box"; you must also specify which robotic arm, which cup, which box, and the entire manipulation process.
Because a single vague instruction may correspond to countless action trajectories.
If the model has not even figured out the objects of the operation, it naturally cannot accurately predict the future.
VPP2 simply describes tasks in finer detail, forming a more stable correspondence between language instructions and physical operations.
A more crucial step is event-level video prediction training.
Traditional short-horizon prediction focuses on what happens in the next few frames, whereas VPP2 directly lets the model learn the changes of one complete operation from start to finish.
From the robotic arm approaching the cup, to grasping the cup, to placing it into the box, what the model must understand is how an entire event unfolds.
This way, when facing new objects, positions, or even manipulation tasks, it has a better chance of predicting reasonable physical changes.
In the instruction-following test on robot manipulation videos,the 14B-parameterVPP2 achieved90%a success rate, while the 64B-parameter Cosmos3 model achieved 78%.
14B beats 64B.
This also shows that in physical manipulation prediction, parameter scale is not the only decisive factor—whether the model truly understands the manipulation process matters just as much.
However, no matter how accurate the predictions are, they are useless if the robot cannot actually move.
The second breakthrough: turning prediction generalization into action generalization
There are two hurdles here: first,speed, and second,how to preserve the video model's original generalization capability while learning actions.
The first hurdle is easy to understand. If a robot has to spend several seconds generating a video before each action, even the strongest predictive ability is hard to put to use.
Xingdong Jiyuan (Robot Era) chose to first speed up the prediction model.
The team further adjusted the event-level prediction model to predict fixed 8-second video segments, and then used consistency distillation to compress multi-step computation into single-step generation.
In the end, predicting visual changes over the next 8 seconds takes about 0.12 seconds of computation.
The second hurdle is a bit trickier.
After finally getting the video model to learn to predict unfamiliar tasks, once action training was added, the model instead only executed familiar operations.
Action capability grows, but generalization is lost? What VPP2 aims to do is clear both hurdles at once.
VPP2 introduces a0.9B-parameterdiffusion Transformer (Action DiT),using a MoT architecture, to learn how to generate robot actions based on the predicted future states.
However, Xingdong Jiyuan did not directly train the video prediction model (Video DiT) and the action expert together from scratch.
In the early stage of action training, the base parameters of the video model were frozen and adapted only via LoRA, to prevent action training from destroying the existing predictive generalization ability.
This is equivalent to first letting the robot understand how the physical world changes, and then training it to convert this "understanding" into "actions".
The two capabilities grow in stages, while complementing each other.
In the end, video prediction takes about 0.12 seconds, the action expert takes about 0.1 seconds,and the overall action-segment generation latency is about 0.22 seconds.
Of course, no matter how elegant the architecture is, the final results are what matter.
The previously mentionedALOHA real-robot zero-shot testhas already demonstrated VPP2's potential to turn predictive ability into unfamiliar-task operation.
Two other sets of generalization tests further validated this route.
In LIBERO-Pro, which examines operation ability after changes in object positions and task requirements, VPP2 achieved an overall success rate of 45.0%, while the best baseline models reached only 11.0%.
LIBERO-OODexamines compositional generalization, recombining familiar objects, layouts and task goals into new tasks never seen before.
VPP2 achieved an overall success rate of 63.9%.
Putting these results together, VPP2's technical approach becomes clear.
First, through multi-source data, detailed task descriptions and event-level prediction training, the model learns to predict physical changes more accurately.
Then, through distillation acceleration and staged action learning, this predictive ability is converted into actual operation, while preserving the original generalization ability as much as possible.
The performance ceiling of embodied intelligence clearly has not yet reached the point where scale is the only thing that matters.
VPP2 once again shows thatoptimizing data pipelines and adjusting training paradigmsthese seemingly plain engineering approaches,still have the opportunity to bring considerable improvements in model performance.
From GPT to WAM, physical AI is forming a new paradigm
Prediction that generalizes, and action that also generalizes—this is the most noteworthy core breakthrough of VPP2 as a World Action Model (WAM).
But real tasks are often far more complex. A robot not only needs to know how to do something, it must also judge what to do first and what to do later.
This requires the involvement of higher-level cognition and planning capabilities.
Here, Xingdong Jiyuan (Robot Era) contributed an idea:GPT does the thinking, VPP2 does the doing.
The former is responsible for understanding, reasoning and planning, breaking complex tasks down into explicit steps; the latter is responsible for predicting how the physical world will change, and then converting the plan into concrete actions.
This idea already has preliminary experimental support in the VPP2 paper.
The team actually employed a VLM high-level planner, which is responsible for semantic understanding, memory and task decomposition, and then hands explicit subtasks to VPP2 for execution.
The results are quite intuitive. In selected long-horizon task tests on RoboDojo,after introducing VLM subtask planning,the average success rate improved from 27.6% to57.6%。
This means that for robots to complete complex tasks, strong motion capabilities alone are not enough. Whether high-level planning and low-level execution can coordinate also affects task completion.
Following this line of thought,GPT+VPP2is expected to form anew paradigm of physical AI technology: general cognition + general physical execution.
General models like GPT are responsible for understanding intent, reasoning, and planning, while a WAM like VPP2 is responsible for predicting physical changes and generating actions, then continuously adjusting through execution feedback.
From cognition, planning, and prediction, to execution and feedback, a complete physical AI closed loop is gradually becoming clear.
And this gives the value of WAM greater room for imagination.
Model capabilities must interact with the physical world through the robot's body, and real-world use will expose failures and generate feedback, driving continued improvement of models and hardware.
Xingdong Jiyuan develops the brain, the body, and dexterous hands simultaneously, and its "deep full-stack" strategy can be explained by this logic:What the company is trying to master includes not only the training stage, but also the stages of capability deployment and feedback flow.
u1s1, no matter how elegant the technical approach is, it ultimately has to do real work in the real world.
In this regard, Xingdong Jiyuan has already started cooperation with companies such as China Post and SF Express, conductingroutine operations in more than 10 logistics centers across 5 provinces and cities nationwide。
In logistics scenarios, cargo dimensions, placement positions, and workflows may all change. With the same set of operations, a robot might have to readapt when switching to another warehouse.
This also places higher demands on the robot's generalization capability.
Whether a model can transfer learned skills to unfamiliar tasks directly affects the efficiency and cost of future cross-scenario deployment.
Of course, the commercialization progress of existing logistics business does not mean VPP2 has achieved large-scale deployment. Whether the model can run stably over the long term in more real-world scenarios still needs further verification.
But VPP2's results this time have already released a signal worth watching.
The competition among world action models is moving from predicting the future toward converting prediction into general action.
Looking back at RoboDojo topping the leaderboard, what deserves attention is no longer just a first place.
From simulation to real robots, from predictive generalization to action generalization, VPP2's series of experimental results show:that staged training of video prediction and action learning can indeed improve a robot's manipulation capability on unfamiliar tasks.
And as general cognitive models like GPT are further combined with WAMs, physical AI is also expected to form a more complete capability system.
At the end of the day, the next round of embodied intelligence competition is about whether robots can carry what they have learned into more unfamiliar scenarios.
Accurate prediction is not enough—you also have to be able to do it. This is the key for world action models to truly enter the physical world.
PS:VPP2 is now open-sourced,Interested readers can try it out and reproduce the paper's results.
If you really run into any surprises, remember to come back and share them with us~
1. Open-source code
: https://github.com/roboterax/video-prediction-policy-2
2. Project homepage:
https://robert-gyj.github.io/video-prediction-policy-2
