Unified state–action interface
A unified 34-dimensional state–action interface maps egocentric human manipulation, UMI, real-robot, and simulation data to fixed semantic slots.

Learning how the world changes — and how to act in it.
Explore the project ↓
Learning how the world changes — and how to act in it.

A unified 34-dimensional state–action interface maps egocentric human manipulation, UMI, real-robot, and simulation data to fixed semantic slots.
Current State combines VLM context with Current 3D Geometry. Transition uses 3D Motion, and Future State uses Future Semantics.
Layer-aligned joint attention couples the 3D stream, semantic stream, and action expert throughout continuous action generation.
Structured World Transition organizes world-state changes as Current State–Transition–Future State. Current State combines VLM context with Current 3D Geometry. 3D Motion describes action-induced three-dimensional state transitions, while Future Semantics represents task-relevant future outcomes.

Task context and present physical structure.
Action-induced three-dimensional state transitions.
Task-relevant future outcomes.
Figure 2 still uses the earlier “Future 3D Motion” label in its artwork; the current manuscript text calls this representation “3D Motion.”
The VLM encodes multi-view observations, a language instruction, and a proprioceptive state into task context. Current 3D Geometry and 3D Motion share the hidden states of the 3D stream; a separate semantic stream models Future Semantics. Frozen Track4World and DINOv3 provide latent supervision only during training. At inference, the internal 3D and semantic streams generate the world representations.

The 3D stream, semantic stream, and action expert exchange information at layers 4, 8, 12, 16, 20, and 24 through joint attention. Action hypotheses inform future motion and semantic predictions, realizing action-conditioned world transition. Updated world hidden states inform action velocity-field predictions, realizing world-informed action generation. This bidirectional interaction continues at every flow integration step.

The manipulation corpus contains approximately 2.014 million valid episodes from nine datasets across four acquisition domains: egocentric human manipulation, UMI, real robots, and simulation. The vision–language corpus contains approximately 2.61 million samples. Manipulation and vision–language batches are mixed at a 9:1 ratio.

The unified 34-dimensional state–action interface allocates seven joint dimensions, one gripper opening dimension, three end-effector position dimensions, and six rotation dimensions to each side. Unavailable dimensions are zero-padded and excluded from losses by validity masks. Egocentric human data provide virtual end-effector poses and gripper openings; UMI trajectories are converted into target-robot supervision through kinematic retargeting.
Action chunks use H = 50. Future observations are selected using each source’s action sampling multiplier, aligning 3D Motion and Future Semantics supervision with the action chunk’s temporal coverage. Joint pre-training combines flow matching, three world representation objectives, FAST discrete action supervision, and autoregressive vision–language modeling.
RoboDojo-Sim comprises 42 simulation tasks and summarizes policy capabilities along five dimensions: Generalization, Precision, Long-Horizon, Memory, and Open. Each dimension reports a process Score followed by success rate (SR, %); Generalization averages the Gen-Std and Gen-Rand splits.
27 methods · Agent, VLA, and WAM · Score / success rate (%)
| Method | Type | Open-source | Generalization | Precision | Long-Horizon | Memory | Open | Average |
|---|---|---|---|---|---|---|---|---|
| GPT-6-AstraAgent · Open-source: No | Agent | No | Generalization33.36/30.50 | Precision12.65/4.00 | Long-Horizon21.45/8.25 | Memory43.04/38.67 | Open34.36/31.00 | Average28.97/22.48 |
| PhysicalRSIAgent + VLA · Open-source: No | Agent + VLA | No | Generalization21.30/15.62 | Precision38.05/32.50 | Long-Horizon46.37/37.33 | Memory46.74/46.56 | Open28.88/24.92 | Average36.27/31.38 |
| VLActVLA · Open-source: Yes | VLA | Yes | Generalization9.54/6.28 | Precision20.57/15.17 | Long-Horizon20.12/13.67 | Memory0.66/0.56 | Open2.37/2.25 | Average10.65/7.58 |
| StarVLA-PI_v3VLA · Open-source: Yes | VLA | Yes | Generalization11.22/8.05 | Precision17.77/12.50 | Long-Horizon18.46/11.00 | Memory4.59/4.00 | Open2.03/2.00 | Average10.81/7.51 |
| InternVLA-A1.5VLA · Open-source: Yes | VLA | Yes | Generalization10.35/6.83 | Precision15.23/10.17 | Long-Horizon23.80/13.75 | Memory4.93/3.56 | Open1.43/1.42 | Average11.15/7.14 |
| π₀.₅VLA · Open-source: Yes | VLA | Yes | Generalization13.38/8.17 | Precision12.40/5.50 | Long-Horizon23.54/14.67 | Memory5.89/4.67 | Open1.98/1.67 | Average11.44/6.93 |
| Spatial ForcingVLA · Open-source: Yes | VLA | Yes | Generalization14.12/9.34 | Precision17.32/10.58 | Long-Horizon23.26/14.58 | Memory5.43/4.11 | Open1.78/1.58 | Average12.38/8.04 |
| SimpleMemVLAVLA · Open-source: Yes | VLA | Yes | Generalization6.36/3.95 | Precision7.42/2.92 | Long-Horizon14.58/5.50 | Memory33.71/33.22 | Open0.85/0.75 | Average12.58/9.27 |
| KinRTVLA · Open-source: Yes | VLA | Yes | Generalization14.02/8.61 | Precision15.65/9.92 | Long-Horizon26.40/18.08 | Memory4.82/3.56 | Open4.23/3.83 | Average13.02/8.80 |
| Hy-Embodied-0.5-VLAVLA · Open-source: Yes | VLA | Yes | Generalization11.78/8.39 | Precision13.81/8.00 | Long-Horizon25.74/14.92 | Memory13.37/12.11 | Open0.65/0.58 | Average13.07/8.80 |
| Meituan-Robotics-0VLA · Open-source: Yes | VLA | Yes | Generalization13.75/8.17 | Precision16.77/7.75 | Long-Horizon29.61/18.58 | Memory10.06/8.89 | Open4.54/4.25 | Average14.95/9.53 |
| Xiaomi-Robotics-1VLA · Open-source: Yes | VLA | Yes | Generalization23.54/17.00 | Precision26.69/18.83 | Long-Horizon38.39/23.67 | Memory7.81/6.56 | Open3.94/3.58 | Average20.07/13.93 |
| GalaxeaVLA (G0.5)VLA · Open-source: Yes | VLA | Yes | Generalization18.46/12.83 | Precision28.25/20.42 | Long-Horizon44.12/32.25 | Memory8.61/7.33 | Open1.73/1.58 | Average20.23/14.88 |
| DM0.5VLA · Open-source: Yes | VLA | Yes | Generalization15.77/10.95 | Precision24.82/16.75 | Long-Horizon33.70/19.50 | Memory47.74/47.44 | Open2.43/2.08 | Average24.90/19.34 |
| Simate-betaVLA · Open-source: No | VLA | No | Generalization35.09/27.95 | Precision34.35/26.92 | Long-Horizon57.84/43.42 | Memory33.33/33.00 | Open9.12/8.50 | Average33.95/27.96 |
| Liber-0 PreviewVLM + WAM · Open-source: No | VLM + WAM | No | Generalization24.99/18.61 | Precision38.28/33.17 | Long-Horizon45.98/33.17 | Memory37.77/37.33 | Open6.68/5.33 | Average30.74/25.52 |
| Fast-WAMWAM · Open-source: Yes | WAM | Yes | Generalization2.34/1.11 | Precision1.96/0.00 | Long-Horizon9.14/5.17 | Memory3.55/3.44 | Open0.42/0.42 | Average3.48/2.03 |
| AHA-WAMWAM · Open-source: Yes | WAM | Yes | Generalization5.79/3.28 | Precision5.86/2.42 | Long-Horizon8.61/2.67 | Memory2.97/2.78 | Open0.88/0.83 | Average4.82/2.39 |
| GigaWorld-Policy-0WAM · Open-source: No | WAM | No | Generalization5.35/2.89 | Precision6.15/1.83 | Long-Horizon15.51/8.92 | Memory3.46/2.22 | Open0.54/0.50 | Average6.20/3.27 |
| X-WAMWAM · Open-source: Yes | WAM | Yes | Generalization7.39/3.33 | Precision6.72/1.83 | Long-Horizon17.47/9.08 | Memory6.32/4.67 | Open0.57/0.25 | Average7.69/3.83 |
| OpenWAM-αWAM · Open-source: Yes | WAM | Yes | Generalization20.71/14.83 | Precision18.45/9.25 | Long-Horizon34.93/25.33 | Memory10.41/9.11 | Open1.41/1.08 | Average17.18/11.92 |
| ME-U0WAM · Open-source: No | WAM | No | Generalization17.53/10.22 | Precision23.95/15.42 | Long-Horizon36.98/22.33 | Memory8.42/7.00 | Open1.41/0.92 | Average17.66/11.18 |
| Liber-0 LiteWAM · Open-source: No | WAM | No | Generalization25.36/19.11 | Precision36.60/31.08 | Long-Horizon45.50/32.92 | Memory35.68/35.44 | Open3.06/2.58 | Average29.24/24.23 |
| InternW0-ΔWAM · Open-source: Yes | WAM | Yes | Generalization30.09/22.78 | Precision31.98/23.25 | Long-Horizon46.29/29.33 | Memory34.67/34.00 | Open10.84/10.17 | Average30.77/23.91 |
| VPP2-PreviewWAM · Open-source: No | WAM | No | Generalization31.34/24.62 | Precision33.75/25.58 | Long-Horizon35.31/22.17 | Memory50.78/50.33 | Open5.83/5.42 | Average31.40/25.62 |
| Awomo-0.5WAM · Open-source: No | WAM | No | Generalization30.31/23.84 | Precision39.65/33.50 | Long-Horizon40.36/26.75 | Memory42.27/41.22 | Open24.09/22.92 | Average35.34/29.64 |
| Magic-W0WAM · Open-source: Yes | WAM | Yes | Generalization37.48/29.17 | Precision37.14/30.73 | Long-Horizon51.80/36.00 | Memory52.68/51.67 | Open4.65/4.25 | Average36.75/30.36 |
Magic-W0 achieves the highest Average Score on RoboDojo-Sim, ranking No. 1 across all evaluated Agent+VLA, VLA+WAM, VLA, and WAM methods. It reaches an Average Score of 36.75 with an average success rate of 30.36%. Compared with the strongest Agent+VLA method, PhysicalRSI, Magic-W0 improves the Average Score by 0.48 points; compared with the strongest WAM method, Awomo-0.5, by 1.41 points; and compared with the strongest VLA method, Simate-beta, by 2.80 points. Magic-W0 also achieves the best overall performance in Generalization, with a Score of 37.48, and in Memory, with 52.68 Score / 51.67% Success Rate, both ranking first among all compared methods. The video below shows a single inference rollout, while the table summarizes the full evaluation results.
Score / SR (%). Generalization averages Gen-Std and Gen-Rand. Open-source means both code and model weights are publicly available.
27 supplied head-camera episodes across five benchmark categories. Choose a category, then select a task to play.
After fine-tuning from the pretrained checkpoint, Magic-W0 has the highest average success rate among methods compared in the manuscript, with the highest Spatial and Goal results.
Magic-W0 and π₀.₅ use identical demonstrations, training budgets, visual and proprioceptive observations, action spaces, and initial evaluation states. The five tasks are clothes folding, bottle uprighting, pen storage, object storage, and kitchen storage, with 100 independent trials per task.
Mean success after downstream fine-tuning, versus 91.8% for π₀.₅ under matched training and evaluation conditions.
100 trials per task · five tasks
Magic-W0 achieves an average success rate of 94.6%, exceeding π₀.₅’s 91.8% by 2.8 percentage points. It improves on four tasks and matches 100% on bottle uprighting. Object storage improves from 90% to 95%, and kitchen storage from 91% to 95%.
In the lowest-noise interval, cross-sample action replacement increases Future Semantics loss by 73.5%. Masking the 3D-to-semantic connection increases that loss by 60.8%, averaged over all flow intervals. These interventions show that world predictions respond to action conditions and that future semantic prediction depends on shared 3D representations.


After supervised fine-tuning in the target domain, the 3D stream, semantic stream, and action expert are jointly updated at every flow integration step to generate continuous action chunks. The recordings below show real-robot task executions; the five-task quantitative comparison is reported above.
7 tasks · 14 recordings. All recordings are shown above; use the player controls to pause or view full screen.