World action models · Adaptive inference

AnyStep-WAM

Budget-Aligned Distillation and Adaptive Inference
for World Action Models

Rui Wang1,3,*Xiangyu Wang1,*Donglin Yang2Yibo Li1Canyang Chen1Zhongrui Wang1,†Xiaojuan Qi2,3,†
1 Southern University of Science and Technology2 The University of Hong Kong3 Shenzhen Loop Area Institute

* Equal contribution   ·   † Corresponding authors

See the difference · Real-world comparison

Fewer steps.
More responsive robot actions.

Base-Motus and AnyStep-Motus perform the same task side by side. AnyStep adapts its denoising budget to the current scene.

See how AnyStep works
Put Block / Unitree G1DBase-Motus AnyStep-Motus2× speed
1–10Selectable denoising steps
3 backbonesMotus · FastWAM · LingBotVA
50 + 6Simulation + real-world tasks
6.14×Peak reported real-world inference speedup

Overview

Adaptive computation for world action models.

World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed number of denoising steps. However, manipulation tasks contain action chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. We introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation.

Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student–teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements.

Learn effective generation across budgets, then adapt the budget to each action chunk. The reductions illustrated here refer to Motus.

Method

Learn across budgets.
Choose the steps at inference.

01 / DISTILL

Budget-aligned learning

Frozen-teacher trajectories provide explicit targets for finite-time transitions. Shared LoRA adapters train interval-conditioned flow maps for one-step prediction through multi-step refinement.

02 / SELECT

Risk–benefit scheduling

A single preview predicts teacher-curvature-based difficulty and student–teacher fidelity for each candidate budget. The scheduler selects the smallest budget that meets the risk-adaptive fidelity requirements.

03 / RECYCLE

Reuse the preview

For a one-step budget, return the preview. For larger budgets, recycle its update into the first scheduled state and perform the remaining K − 1 updates. Scheduling needs no online teacher or candidate rollouts.

The selected budget applies to each active denoising branch. Motus jointly predicts video and action; FastWAM uses the action branch at inference; LingBotVA follows sequential video-to-action generation.
What does the scheduler predict?

Risk measures teacher-trajectory difficulty, rather than a calibrated probability of task failure. Benefit estimates budget-specific student–teacher fidelity. Feasibility requires the fidelity threshold to be met for every active modality; if no candidate is feasible, the scheduler falls back to the budget with the largest sum of predicted modality fidelities. See the paper for the complete supervision and selection rules.

Real-world demonstrations

Three backbones. Six tasks.

Unitree G1D dual-arm robot.
Task recordings play at their original 1× speed.

6 task recordings

Bowl Pouring

01 / 06

Clean Table

02 / 06

Put Block

03 / 06

Put Cup

04 / 06

Stack Blocks

05 / 06

Stack Bowls

06 / 06

Showing six AnyStep-FastWAM demonstrations.

Experiments

Less denoising.
Comparable manipulation success.

Evaluated on 50 RoboTwin 2.0 tasks and six real-world tasks. Adaptive inference preserves comparable or higher average success rates while reducing the number of denoising steps and inference latency.

RoboTwin 2.0

Clean + randomized settings · Average results
RoboTwin 2.0 average success rates, steps, and latency, from Table 1 and Section 5.1 of the paper
BackboneSuccess rate ↑Base → AnyStepDenoising steps ↓Base → AnyStepStep reduction ↓Latency ↓Base → AnyStep
Motus87.84 → 87.96%10 → 3.9860.2%1.932 → 0.859 s
FastWAM91.83 → 91.59%10 → 5.0349.8%‡0.496 → 0.295 s
LingBotVA92.24 → 92.22%25 / 50 → 3.68 / 3.6885.28% / 92.64%9.732 → 3.247 s

Latency: seconds per inference call on one NVIDIA A100. LingBotVA values separated by “/” denote video / action. ‡ FastWAM's reduction is reported as 49.8% in the paper; its displayed mean steps are rounded to 5.03.

Cross-budget capability

Stronger one-step prediction.

AnyStep training improves average one-step success rates on RoboTwin 2.0 across all three backbones.

Base, 1 stepAnyStep, 1 step
Motus+7.07 pp
75.05%
82.12%
FastWAM+12.08 pp
77.20%
89.28%
LingBotVA+8.94 pp
73.77%
82.71%

Success rate (%) · Bars share a 0–100% scale · pp = percentage points

Real-world manipulation

Six tasks · Unitree G1D
Real-world average results, from Table 2 and Section 5.2 of the paper
BackboneSuccess rate ↑Base → AnyStepDenoising steps ↓Base → AnyStepLatency ↓Base → AnyStepSpeedup ↑
Motus65.83 → 69.17%10 → 4.061.868 → 1.117 s1.67×
FastWAM70.00 → 71.67%10 → 3.610.280 → 0.122 s2.30×
LingBotVA68.33 → 70.00%25 / 50 → 3.54 / 3.544.654 → 0.758 s6.14×

Latency: seconds per inference call on one NVIDIA RTX 5090. Each base model and its AnyStep variant use 50 demonstrations per task, 8,000 optimization steps, four NVIDIA H100 GPUs, and a global batch size of 32. All results above are reported in arXiv v1.

Citation

BibTeX

@article{wang2026anystepwam,
  title   = {AnyStep-WAM: Budget-Aligned Distillation and Adaptive
             Inference for World Action Models},
  author  = {Wang, Rui and Wang, Xiangyu and Yang, Donglin and Li, Yibo
             and Chen, Canyang and Wang, Zhongrui and Qi, Xiaojuan},
  journal = {arXiv preprint arXiv:2609.33748},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.33748}
}

Figure