机器之心 (微信公众号) · 10/8/2026, 20:19:00
Stanford Releases OpenWAM: A Modular Framework Decoupling Video Prediction and Action Control for Robots
Stanford's Fei-Fei Li team released OpenWAM, an open-source unified framework for World-Action Models (WAMs) designed to address attribution challenges caused by coupled variables in existing research. By decoupling video backbones, training data, and inference flows via a shared MoT architecture, the framework enables independent iteration of 'imagination' and 'execution' modules. Experiments show that swapping base models significantly boosts robot task success rates, providing a standardized testbed for building interpretable, modular robotic brains.
SOURCE COVERAGEOriginal coverage
Contents2 sections

Open Source OpenWAM

Edited by | Panda
On September 22, Black Forest Labs, renowned for its FLUX image models, crossed over into robotics with the open-source release of FLUX 3 Action, a 7B-parameter World-Action Model (WAM). According to their official blog, this model surpassed Cosmos 3 Nano—the previously strongest open-source model—on NVIDIA’s RoboLab-120 benchmark, despite having less than half the parameter count. The fact that an image generation company is now competing in robot control suggests that the "physical intuition" accumulated within video models is increasingly being adopted as the foundation for robotic brains by more and more teams.
This approach is known as the World-Action Model (WAM). Since NVIDIA explicitly introduced this term in the DreamZero paper in February of this year, related papers have emerged rapidly. However, behind the hype lies an awkward problem: these systems almost simultaneously change the video backbone, training data, interaction mechanisms between video and action, and inference pipelines. Consequently, it remains unclear which specific design choice actually drives performance improvements.
This week, a team from Stanford University, including Fei-Fei Li, Jiajun Wu, Ehsan Adeli, and others, published a paper introducing OpenWAM, an open framework for world-action modeling, aiming to provide a systematic answer to this question.

Overview of the OpenWAM Framework: Three-stage training, shared MoT architecture, four configurable interaction modes, and a locally reusable dynamics model that can be frozen.

- Paper Title: OpenWAM: An Open Framework for Composable World-Action Models
- Paper Link: https://arxiv.org/abs/2610.07922
- Code Repository: https://github.com/OpenWAM/OpenWAM
Why We Need a "Unified Exam"
Traditional Vision-Language-Action (VLA) models are like drivers relying on conditioned reflexes: they see the scene, hear the command, and directly output actions.
WAMs add a step of "mental rehearsal": simultaneously predicting how the visual scene will evolve over the next few seconds and determining how to act accordingly. Videos naturally record how objects fall, collide, and are pushed; models pre-trained on massive video datasets therefore possess inherent physical priors.
The issue is that there are many ways to couple "predicting visuals" and "generating actions": imagine first then act, act first then imagine consequences, generate both simultaneously, or compute them independently. Different teams make different choices regarding backbones and data. Comparing them is akin to evaluating chefs who simultaneously changed ingredients, stoves, and cooking sequences—it becomes difficult to determine which specific step contributed to the dish's quality.
A second, more subtle problem exists. Robot training data mostly consists of successful demonstrations, telling the model only "how to take the correct path," while rarely containing information about "what happens if the arm extends two centimeters too far." This is similar to a student who has only watched perfect slalom videos from a coach: they memorize the route but do not understand how much steering input causes the car to deviate.
OpenWAM addresses the former issue with a shared backbone and switchable interaction modes, and the latter with large-scale "counterfactual" data.
Same Skeleton, Four Modes of "Think First or Act First"
OpenWAM starts with Alibaba’s open-source video model Wan2.2-5B. It undergoes continued pre-training on approximately 3.34 million robot-human interaction videos, totaling roughly 14,600 hours of nominal duration. No action labels are used during this phase, and training was conducted for 14 days on 32 B200 GPUs.
The key modification is "causalization": each video block can only attend to current and previous content, without peeking at future frames. The original model acted like someone repeatedly editing lines with the entire script in hand; now, it must perform sequentially in real-time, just like live broadcasting. This is precisely what is required for robots acting step-by-step.
Subsequently, the team uses a Mixture-of-Transformers (MoT) architecture to combine a 5B video expert and a 2B action expert. Each retains its own normalization, projection, and feed-forward layers, exchanging information only through a single joint attention layer.

Shared MoT Architecture: Video and action experts perform joint attention on packed tokens.
On this skeleton, the team defines four "interaction programs":
- VTA: Imagine visuals first, then generate actions.
- ATV: Reverse order (act first, then imagine).
- Joint: Generate both simultaneously with mutual reference.
- Decoupled: Ensure the two modalities cannot see each other’s future parts.
This is like the same band playing the same piece, where only the starting sequence and whether they can hear each other changes. All four share the same backbone, tokenization, and flow matching objectives. The only variables are the generation order and cross-modal attention, giving the comparison a true "controlled variable" nature.
Attention masks for the four interaction programs (A–D) and two local dynamics programs (E–F).
Extracting the "Action Translator" for Independent Training
A more innovative aspect of the paper is decomposing the Inverse Dynamics Model (IDM) and Forward Dynamics Model (FDM) into independently trainable components.
The IDM acts like a translator: taking the current visual frame, robot proprioceptive state, and an imagined future video as input, it outputs specific actions. The FDM works in reverse, predicting how the visual scene will change based on given actions.
The key lies in the "local" scope: neither component receives task instructions or earlier history, focusing solely on "what action corresponds to this small segment of visual change." Consequently, they are not bound to specific tasks and can be paired with any compatible video predictor. The video model determines "what should happen," while the IDM handles "how to achieve it."
However, a translator trained only on successful demonstrations has limited exposure. To address this, the team constructed the LIBERO-Long-CF dataset: by restoring simulator states from demonstrations and executing modified actions—such as stopping, reversing, adding noise, or altering gripper opening/closing timing—they generated 32,000 segments comprising approximately 4.1 million control records. This is 29.7 times the volume of the original demonstrations, with 72.3% of segments resulting in changed object placements. Many branches fail to complete the task, but that is precisely the point. Returning to the driving school analogy, this is akin to an instructor asking students to deliberately turn the steering wheel half a circle too far; what the student learns is no longer just the route, but the car's inherent behavior.
Experiments: Strong Foundation, Portable Translator
On the four LIBERO subsets, all four interaction programs achieved high success rates. VTA averaged 98.6%, slightly exceeding previously reported results for LingBot-VA (98.5%) and π0.5 (96.9%). However, since LIBERO is approaching saturation, a more noteworthy detail is observed on LIBERO-Long: the Decoupled approach (97.0%), where future parts are mutually invisible, performs comparably to the Joint approach (96.6%), which features bidirectional interaction.

Closed-loop success rates (%) on the four LIBERO subsets; baselines are previously reported results.
On real hardware, the team tested two Franka FR3 robotic arms on tasks including toasting bread, solving the last layer of a 2×2 Rubik's Cube, and sorting cups by color. VTA and Joint achieved average success rates of 92.1% and 91.9%, respectively.
Left: Three real-world dual-arm tasks. Right: Four LIBERO-90 tasks used for component transfer.
Ablation studies highlighted the importance of the foundation: keeping other conditions constant, initializing with the original Wan2.2 resulted in a VTA success rate of only 68.4% on LIBERO-Long. Switching to a causal foundation pre-trained on robot videos raised this to 97.8%, an improvement of 29.4 percentage points. The authors emphasize that this gain stems from the combined effect of data and causal architecture modifications.
Impact of video backbone initialization methods on LIBERO-Long success rates.
The most critical experiment was component transfer. The team froze the IDM trained on LIBERO-Long and paired it with a video predictor fine-tuned for four new LIBERO-90 tasks.
Results showed that the local IDM trained on demonstrations plus counterfactual data achieved an average success rate of 84.0%. In contrast, an IDM receiving full context reached only 47.0%, and a local IDM trained solely on demonstrations dropped to 21.5%.
"Looking only locally" and "experiencing sufficiently diverse consequences" are both indispensable; their combination yields an action translator capable of portability. It is important to note that the video predictor for target tasks was fine-tuned using demonstrations containing action labels. This is not zero-shot transfer, a limitation explicitly acknowledged by the authors.

Success rates of the frozen IDM transferred to four LIBERO-90 tasks.
Conclusions for the FDM were consistent. Starting from the same state and executing 16 different actions, the model was asked to determine which predicted frame corresponded to the actual outcome. Random guessing yields approximately 6%; an FDM trained only on demonstrations achieved 21.1%, while adding counterfactual data raised performance to 71.3%. The model finally learned to distinguish between "pushing like this" and "pushing like that."
Prediction performance of the local FDM on counterfactual transitions.
Conclusion
The authors candidly acknowledge limitations: component reuse was primarily validated in simulation, whereas the cost of "deliberately making mistakes" is significantly higher on real robots; additionally, the FDM covers only short-horizon predictions. However, the value of OpenWAM does not lie in squeezing out another fraction of a percentage point. With WAM papers emerging weekly, a public testbed featuring a fixed foundation and single-variable changes helps shift the field's focus from "who scores higher" to "why they score higher." Furthermore, reusable action translators suggest that future robotic brains need not be monolithic black boxes trained end-to-end; modules responsible for imagination and execution can iterate independently and be assembled as needed. Those trajectories that failed to complete tasks are precisely the nourishment that allows models to understand the physical world.
Currently, the team has open-sourced the training, evaluation, and composition code, along with pre-trained video model weights.

© THE END
Please contact this official account for reprint authorization.
For submissions or media inquiries: [email protected]