Today for AI

量子位官网 · 10/9/2026, 11:52:21

Tsinghua-backed Robot Era Open-Sources VPP2: Decoupling Video Prediction and Action Learning Tops RoboDojo Benchmark

By 田, 晏林Original title: 清华具身模型登顶全球第一!突围GPT-6、英伟达,不靠外挂和额外数据
78AI Score
Executive Summary

Tsinghua-backed robotics firm Robot Era released its world action model VPP2, which topped the global RoboDojo benchmark with a composite score of 39.26. By decoupling video prediction from action learning—training strong predictive capabilities before action control—the model achieved superior generalization, fine manipulation, and memory without relying on extra data or augmentation strategies; the code is now open-source.

SOURCE COVERAGEOriginal coverage

Contents5 sections

Xingdong Era chose to train video prediction and action learning in separate stages. The focus is not on "throwing everything into one pot" (jointly training video and actions), but rather on "decoupling" the two and reordering their training sequence.

Tian Yanlin reporting from Ao Fei Si QbitAI | WeChat Official Account: QbitAI

Wow, has a Chinese team just surged to global number one in the World Action Model (WAM) track?!

Moreover, top-tier global embodied models such as GPT-6-Astra, Physical Intelligence’s π0.5, and NVIDIA’s GR00T-N1.7 were all left behind this time.

The company making this move is a Chinese robotics firm—Xingdong Era, the only direct subsidiary of Tsinghua University in the embodied AI space.

Recently, Xingdong Era’s self-developed World Action Model VPP2 took the top spot on the RoboDojo simulation leaderboard.

With an overall average success rate of 32.26% and an overall average score of 39.26, it ranked first in both metrics.

Crucially, this competition was far from easy.

RoboDojo is hailed as the "Mount Everest" of embodied intelligence, led by MMLab at the University of Hong Kong and co-built by nearly 20 top academic institutions worldwide.

It specifically tests areas where robots struggle: Can they still perform tasks when the environment changes? Can they grasp objects accurately if they are misaligned? Do they remember what happened earlier when a task is half-finished?

In short, trying to pass by merely memorizing proficiency levels is no longer viable.

So, what was the result?

Xingdong Era’s VPP2 not only secured the overall first place but also topped the charts in three key dimensions: generalization capability, fine-grained manipulation, and memory capability.

Reportedly, for this run, VPP2 did not use additional data, nor did it employ enhancement methods like Agent RSI. It relied solely on "standard datasets" to outperform a host of strong competitors.

This indicates that the VPP2 model itself is sufficiently powerful. Its performance gains stem from pre-training, with a base model robust enough for generalization. Therefore, it does not rely on external enhancement strategies or data expansion to "boost scores."

Wait, how exactly was such a formidable model trained???

We dug into the technical roadmap, and indeed, VPP2’s breakthrough this time has some real substance.

It targets a longstanding challenge in World Action Models: if the prediction is wrong, the subsequent action follows suit and fails.

VPP2’s solution is to first train video prediction capabilities to be sufficiently strong, and then enable the robot to actually act in unfamiliar environments.

From prediction to action, both types of generalization capabilities improve simultaneously.

Based on multiple disclosed test results, VPP2 has become one of the most competitive contenders in the World Action Model (WAM) track.

A New Champion Emerges in the WAM Track

There is a somewhat awkward phenomenon in the embodied AI industry.

Robot demos are becoming increasingly flashy, and leaderboard scores are rising, but if you ask which robot is smarter and more capable of doing actual work, it’s hard to answer.

For example, consider having a robot grab a cup. On a table surface seen during training, it might achieve perfect accuracy. But change the cup or shift its position slightly, and it immediately starts doubting its existence.

Let alone tasks requiring the continuous execution of multiple steps.

This is why the RoboDojo evaluation suite deserves attention.

Its goal is to establish unified, reproducible evaluation standards for embodied intelligence. The simulation tests include 42 dual-arm manipulation tasks, covering five dimensions: generalization, precise manipulation, long-horizon tasks, memory, and open-vocabulary instruction understanding.

Simply put, it pulls robots out of their comfort zones.

Traditional robot models may approach perfect scores on familiar tasks, but once the environment, object positions, or task combinations change, success rates can drop significantly.

RoboDojo specifically examines these complex scenarios to assess how much generalization capability a robot possesses when facing changes.

And VPP2’s report card this time is quite impressive.

An average success rate of 32.26% and an average score of 39.26, ranking first in both metrics.

For comparison, in the post-training evaluation of the RoboDojo simulation benchmark, GPT-6-Astra achieved an average success rate of 22.48% and an average score of 28.97.

VPP2 leads by 9.78 percentage points and 10.29 points, respectively.

Screenshot from the paper; baseline model data taken from the official leaderboard.

These two metrics have different focuses: the success rate measures whether the robot can complete the task entirely; the average score measures how far the task progressed. Even if the task wasn’t fully completed, it reflects the robot’s partial performance.

Both metrics aggregate results across the five major capability dimensions.

In other words, VPP2 not only completes tasks at a higher rate but also demonstrates stronger partial completion capabilities when tasks aren't fully finished.

Moreover, this lead is not limited to overall scores.

Breaking down the five capability dimensions, VPP2 also took first place in generalization, precise manipulation, and memory.

These three capabilities correspond to several hurdles robots face when entering the real world: Can they still work in a new environment? Can they manipulate objects precisely? And can they remember previous steps during continuous task execution?

Leading in overall scores while also excelling in key capabilities.

However, no matter how difficult RoboDojo is, the testing ground remains a "simulation environment." In the physical world, factors like friction, deformation, and positional deviations introduce sudden variables.

Does the generalization capability VPP2 demonstrated in simulation hold up on real hardware?

Xingdong Era conducted experiments to find out.

This time, the team directly deployed VPP2 onto a real ALOHA dual-arm robot, testing 10 categories of zero-shot manipulation tasks, including grasping, placing, stacking, folding, and pouring.

That means the robot had to perform directly without any additional fine-tuning for these specific test tasks.

The result: VPP2 again delivered leading performance: an average success rate of 58.5%, surpassing π0.5’s 40%.

Out of the 10 task categories, VPP2 achieved the best performance in 9 of them.

Now things get interesting.

Ranking first on the leaderboard demonstrates its capabilities within standard evaluation frameworks; zero-shot operation on real robots further tests whether these capabilities can transfer to real-world physical environments.

Placing these two sets of results side by side makes VPP2’s technical advantages even more worthy of deep exploration.

It is important to note that VPP2 follows the World Action Model (WAM) approach.

The fundamental idea behind this type of model is to use video prediction to understand how the physical world will change next, and then translate those predictions into robot actions.

However, this approach has long faced a critical issue: beautiful video predictions do not necessarily mean the robot can actually perform tasks effectively.

VPP2 specifically targets this problem.

From Predictive Generalization to Action Generalization: How Does VPP2 Achieve This?

To understand VPP2’s breakthrough, consider a simple task: instructing a robot to place a cup from a table into a box.

For humans, reaching out, grasping, lifting, and placing the cup requires almost no conscious thought.

But for a robot, it must determine the cup’s position, calculate the robotic arm’s trajectory, decide when to close the gripper, and predict how the physical world will change throughout the entire manipulation process. A deviation in any step could lead to failure.

Existing World Action Models (WAMs) are particularly prone to issues at this stage.

On one hand, standard video models excel at generating visuals, but what looks plausible in a video is not necessarily consistent with real-world physical laws.

For example, if instructed to grab the left cup, the model might predict grabbing the right one instead. The video generation may still look coherent, but the robot will execute the wrong action.

On the other hand, directly integrating action learning into video models can compromise their original generalization capabilities.

The result is that while the robot learns specific actions, it fails to adapt when placed in new environments.

VPP2 proposes a straightforward premise: The quality of video prediction determines the upper limit of action performance.

Based on this insight, Robot Era chose to train video prediction and action learning in separate stages. The focus was not on "mixing video and action training together," but rather on decoupling the two and reordering the training process.

Robot Era designed a three-stage training strategy for this purpose:

  • Stage 1: Event-level video continued pre-training, enabling the model to learn to predict complete manipulation processes.
  • Stage 2: Fixed-duration video post-training and distillation, ensuring prediction speed matches the robot’s real-time execution requirements.
  • Stage 3: Action expert training, translating video predictions into concrete actions while preserving existing generalization capabilities as much as possible.

These three progressive stages ultimately lead to two core breakthroughs: predictive generalization and actionable generalization.

First Breakthrough: Enabling Video Prediction to Possess Generalization Capabilities

The primary challenge VPP2 addresses is whether the robot can accurately predict physical changes in unfamiliar tasks.

Robot Era built upon Alibaba’s open-source Wan2.1-I2V-14B model, integrating diverse data sources including robot manipulations, human activities, and general videos, covering various robot embodiments and manipulation methods.

However, the truly interesting aspect lies in how they processed the data.

Instead of simply feeding raw videos into the model, the team segmented manipulation processes into semantically complete clips and paired them with detailed descriptions.

For instance, rather than just telling the model "put the cup in the box," the description specifies which robotic arm, which cup, which box, and details the entire manipulation sequence.

This is because a single vague instruction can correspond to countless possible motion trajectories.

If the model cannot clearly identify the objects involved in the operation, it naturally struggles to make accurate future predictions.

VPP2 opts for finer-grained task descriptions, establishing a more stable correspondence between language instructions and physical operations.

A more critical step is event-level video prediction training.

Traditional short-term prediction focuses on what happens in the next few frames, whereas VPP2 trains the model to learn the changes occurring throughout an entire manipulation process, from start to finish.

From the robotic arm approaching the cup, to grasping it, to placing it in the box, the model needs to understand how the whole event unfolds.

This enables the model to better predict reasonable physical changes when facing new objects, positions, or even entirely new manipulation tasks.

In instruction-following tests using robot manipulation videos, the 14B parameter VPP2 achieved a 90% success rate, compared to 78% for the 64B parameter Cosmos3 model.

14B outperforms 64B.

This indicates that in physical manipulation prediction, parameter scale is not the sole determinant of success; truly understanding the manipulation process is equally important.

However, no matter how accurate the predictions are, if the robot cannot act, it is useless.

Second Breakthrough: Transforming Predictive Generalization into Actionable Generalization

Here lie two hurdles: speed, and how to preserve the video model’s original generalization capabilities while learning actions.

The first hurdle is easy to understand. If generating a video takes several seconds before each robot action, even powerful prediction capabilities become impractical.

Robot Era chose to accelerate the prediction model first.

The team adjusted the event-level prediction model to generate fixed 8-second video segments, then used consistency distillation to compress multi-step calculations into single-step generation.

Ultimately, predicting visual changes over the next 8 seconds takes approximately 0.12 seconds.

The second hurdle is slightly more tricky.

After painstakingly teaching the video model to predict unfamiliar tasks, adding action training often causes the model to revert to only performing familiar operations.

Did the robot gain action capabilities at the cost of losing generalization? VPP2 aims to overcome both hurdles simultaneously.

VPP2 introduces a 0.9B parameter Diffusion Transformer (Action DiT), utilizing a MoT architecture, to learn how to generate robot actions based on predicted future states.

However, Robot Era did not train the video prediction model (Video DiT) and the action expert from scratch together.

During the initial phase of action training, the base parameters of the video model were frozen, adapting only via LoRA. This minimizes the risk of action training disrupting the existing predictive generalization capabilities.

Essentially, the robot first learns to understand how the physical world changes, and then learns how to translate that "understanding" into "action."

Two capabilities grow in stages yet complement each other.

Finally, video prediction takes about 0.12 seconds, and the action expert takes about 0.1 seconds, resulting in an overall action segment generation latency of approximately 0.22 seconds.

Of course, no matter how elegant the architecture, final performance is what counts.

The previously mentioned ALOHA real-robot zero-shot tests have already demonstrated VPP2’s potential to translate predictive capabilities into operations on unfamiliar tasks.

Two additional generalization tests further validate this approach.

LIBERO-Pro evaluates manipulation capabilities under changes in object positions and task requirements. VPP2 achieved an overall success rate of 45.0%, while the highest-performing baseline models reached only 11.0%.

LIBERO-OOD assesses compositional generalization by recombining familiar objects, layouts, and task goals to form new tasks not seen during training.

VPP2 achieved an overall success rate of 63.9%.

Viewing these results together clarifies VPP2’s technical strategy.

First, it uses multi-source data, detailed task descriptions, and event-level prediction training to enable the model to predict physical changes more accurately.

Second, through distillation for acceleration and staged action learning, it translates this predictive capability into actual operations while preserving original generalization abilities as much as possible.

Clearly, the performance ceiling for embodied intelligence has not yet reached a point where scaling alone is the only solution.

VPP2 once again demonstrates that optimizing data pipelines and adjusting training paradigms—seemingly straightforward engineering methods—still offer significant opportunities for improving model performance.

From GPT to WAM: A New Paradigm for Physical AI

The ability to generalize predictions and actions is VPP2’s most notable core breakthrough as a World Action Model (WAM).

However, real-world tasks are often more complex. Robots need not only to know how to act but also to determine what to do first and what to do next.

This requires higher-level cognitive and planning capabilities.

Xingdong Era (RobotEra) contributes a key insight here: GPT handles thinking, while VPP2 handles doing.

The former is responsible for understanding, reasoning, and planning, breaking down complex tasks into clear steps; the latter predicts how the physical world will change and translates plans into specific actions.

This approach already has preliminary experimental support in the VPP2 paper.

The team actually employs a VLM-based high-level planner, which handles semantic understanding, memory, and task decomposition, then passes explicit sub-tasks to VPP2 for execution.

The results are quite intuitive. In long-horizon task tests selected from RoboDojo, introducing VLM sub-task planning increased the average success rate from 27.6% to 57.6%.

This means that for robots to complete complex tasks, strong motor skills alone are insufficient. The coordination between high-level planning and low-level execution significantly impacts task completion.

Following this line of thought, GPT + VPP2 is poised to form a new technical paradigm for Physical AI: General Cognition + General Physical Execution.

General models like GPT handle intent understanding, reasoning, and planning, while WAMs like VPP2 predict physical changes and generate actions, continuously adjusting based on execution feedback.

From cognition and planning to prediction, execution, and feedback, a complete Physical AI closed loop is gradually taking shape.

This also expands the imaginative potential of WAMs’ value.

Model capabilities must interact with the physical world via the robot body. Real-world usage exposes failures and generates feedback, driving further improvements in both models and hardware.

Xingdong Era develops the "brain," the robot body, and dexterous hands simultaneously. Their "Deep Full-Stack" strategy can be explained by this logic: The company aims to control not just the training phase, but also the deployment and feedback loops.

To be honest, no matter how beautiful the technical roadmap, it ultimately needs to work in the real world.

In this regard, Xingdong Era has already collaborated with companies like China Post and SF Express, operating routinely in over 10 logistics centers across 5 provinces and cities nationwide.

In logistics scenarios, cargo sizes, placement positions, and operational workflows may all vary. The same set of operations might require the robot to adapt anew when moving to a different warehouse.

This places higher demands on the robot’s generalization capabilities.

Whether the model can transfer learned skills to unfamiliar tasks directly affects the efficiency and cost of future cross-scenario deployments.

Of course, existing commercial progress in logistics does not mean VPP2 has achieved large-scale deployment. Whether the model can run stably over the long term in more real-world scenarios still requires further validation.

But VPP2’s current achievements send a signal worth noting.

Competition among World Action Models is shifting from predicting the future to translating predictions into generalized actions.

Looking back at topping the RoboDojo leaderboard, what matters is not just the number one spot.

From simulation to real robots, from prediction generalization to action generalization, VPP2’s series of experimental results demonstrate: Staged training of video prediction and action learning indeed improves robot manipulation capabilities on unfamiliar tasks.

As general cognitive models like GPT combine further with WAMs, Physical AI is expected to form a more complete capability system.

Ultimately, the next round of competition in embodied intelligence hinges on whether robots can bring their learned skills into more unfamiliar scenarios.

Accurate prediction must be matched by effective execution. This is the key for World Action Models to truly enter the physical world.

PS: VPP2 is now open source. Interested readers are encouraged to try it out and reproduce the paper’s results.

If you achieve any surprising results, remember to share them.