One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

University of California, Berkeley
(*Equal Contribution, Contact: yichen_xie@berkeley.edu)
University of California, Berkeley

RoboActualizer selects the task-conditioned future from the multiple plausible ones contained by the representation space. With as few as 60 M parameters, RoboActualizer outperforms baselines with 100x more parameters on simulation benchmarks.

Abstract

RoboActualizer selects a successful future from a latent world model and achieves strong performance with fewer trainable parameters.

A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures.

In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100 times fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin~2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.

Methodology

What the world model leaves undetermined.
The world model captures all plausible futures that may follow the current scene, forming a mixture over different task-conditioned outcomes. However, without knowing the task instruction, it cannot determine which future will actually be realized.

\[ \begin{aligned} p(\mathbf{z}^+_{t:t+H}\mid z_t) &=\sum_{\ell\in\mathcal{L}} p(\ell\mid z_t)\, p(\mathbf{z}^+_{t:t+H}\mid z_t,\ell). \end{aligned} \tag{1} \]

What the actualizer remains to learn.
Given a task instruction, the actualizer selects the corresponding future from the world-model prior and predicts the actions needed to realize it. In this view, the frozen world model provides what can happen, while the actualizer learns which future to pursue and how to achieve it.

\[ \begin{aligned} p(\mathbf z^+_{t:t+H},\mathbf a_{t:t+H}\mid z_t,\ell) &=\underbrace{p(\mathbf z^+_{t:t+H}\mid z_t)}_{\mathbf{prior}\text{: given by }g_\phi} \underbrace{\frac{p(\ell\mid \mathbf z^+_{t:t+H},z_t)}{p(\ell\mid z_t)}}_{\mathbf{selection}\text{: learned}} \underbrace{p(\mathbf a_{t:t+H}\mid \mathbf z^+_{t:t+H},z_t, \ell)}_{\mathbf{realization}\text{: learned}}. \end{aligned} \tag{2} \]

RoboActualizer overview: a frozen latent world model predicts future latents, and a lightweight actualizer with latent and action DiT branches turns them into robot actions.

We first use a frozen latent world model to predict how the scene may evolve in the future. Our actualizer then jointly predicts task-conditioned future latents and robot actions with two interacting DiT branches. The world-model future predictions supervise the latent branch, helping the policy generate actions that are consistent with plausible scene dynamics.

Performance on simulation and real world

Our method achieves better performance across the LIBERO, LIBERO-Plus, and RoboTwin 2.0 benchmarks, covering both single-arm and bimanual simulation. These results show that our method delivers strong performance with substantially fewer trainable parameters.
Success rates (%) on the LIBERO benchmark. “Embodied PT” denotes additional embodied-data pretraining before downstream task training.
Method Trainable Params. Embodied PT Spatial Object Goal Long Overall Latency (ms)
OpenVLA279MYes84.788.479.253.776.5145.5
π03.3BYes96.898.895.885.294.1120.4
π0.53.3BYes98.898.298.092.496.9128.5
UniVLA123MYes96.596.895.692.095.2157.3
Motus5.9BYes96.899.896.697.697.72230.0
WorldVLA7.0BNo87.696.283.460.081.8397.5
Fast-WAM6.0BNo98.2100.097.095.297.6111.7
RoboActualizer (ours)60MNo98.4100.096.696.898.039.9
Success rates (%) on the LIBERO-Plus benchmark. * No open-source code is available for measuring latency.
Method Trainable Params. Camera Robot Lang. Light Bg. Noise Layout Overall Latency (ms)
OpenVLA279M0.83.523.08.134.815.228.515.6145.5
π03.3B13.86.058.885.081.479.068.953.6120.4
π0-Fast3.3B65.121.661.073.273.274.468.861.668.5
UniVLA123M1.846.269.669.081.021.231.942.9157.3
WorldVLA7.04B0.127.941.643.717.110.938.025.0397.5
Fast-WAM6.02B16.444.568.978.253.737.760.751.5111.7
DC-WAM6.00B23.951.783.491.761.354.269.860.9—*
RoboActualizer (ours)60M39.560.570.787.256.959.067.863.139.9
Success rates (%) on RoboTwin 2.0 (50 demos per task).
Method Trainable Params. Success Rates (%)
π0.53.3B31.4
LingBot-VA5.3B17.2
Fast-WAM6.0B41.9
RoboActualizer (ours)60M58.8

We also evaluate five real-world tasks across two platforms: a stationary G1 humanoid with a Unitree Dex3-1 hand and head-mounted RealSense camera for bimanual manipulation, and a FANUC CRX-i10 arm with two third-person cameras and one in-hand camera for single-arm manipulation. Our method consistently outperforms the corresponding baselines across all five tasks, demonstrating effective learning on both single-arm and bimanual platforms.

Further Analysis

How does a lightweight actualizer help?
Joint latent-and-action prediction is essential: removing the latent target reduces success from 60.0% to 17.9%, showing that an action-only policy on frozen V-JEPA features generalizes poorly. By explicitly actualizing future latents, the lightweight model unlocks the world model's predictive structure, while PCA visualizations show that it faithfully predicts both task-conditioned and instruction-dependent diverse futures.

PCA visualization of task-conditioned future latent predictions on LIBERO and RoboTwin tasks.

How is RoboActualizer different from latent world action models?
Unlike latent world action models that use large networks to learn future dynamics in semantic feature spaces, RoboActualizer reuses predictive structure already present in a frozen video world model and only learns to select and realize a task-conditioned future. The representation ablation confirms that temporal predictivity matters more than representation strength alone, enabling a much smaller actualizer.

Ablation study on latent representation. Future predictivity within the latent space is more important than representation ability.
Representation Params. Input Range Latent Predictivity Success Rates (%)
DINOv3300M1 frame✕39.7
WAN VAE127M4 frames✕25.9
V-JEPA 2.1image300M1 frame✕46.4
V-JEPA 2.1image300M4 frames✕44.8
V-JEPA 2.1video300M4 frames✓60.0

How do we allocate the computation?
Scaling the frozen world model improves success more than scaling the trainable actualizer. This indicates that predictive capacity is best placed in the pretrained prior, while a compact actualizer is sufficient for future selection and action realization.

Computation allocation between world model and actualizer.
World Model Actualizer Success Rates (%)
V-JEPA 2.1 (80M)DiT-S (60M)50.7
V-JEPA 2.1 (80M)DiT-B (245M)53.2
V-JEPA 2.1 (300M)DiT-S (60M)60.0
V-JEPA 2.1 (300M)DiT-B (245M)61.0

BibTeX

@article{du2026roboactualizer,
  title={One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions},
  author={Bang Du and Yichen Xie and Shuqi Zhao and Yuxin Chen and Menglin Wu and Masayoshi Tomizuka},
  journal={arXiv preprint arXiv:2609.36413},
  year={2026}
}