Budget-aligned learning
Frozen-teacher trajectories provide explicit targets for finite-time transitions. Shared LoRA adapters train interval-conditioned flow maps for one-step prediction through multi-step refinement.
Budget-Aligned Distillation and Adaptive Inference
for World Action Models
See the difference · Real-world comparison
Base-Motus and AnyStep-Motus perform the same task side by side. AnyStep adapts its denoising budget to the current scene.
See how AnyStep worksOverview
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed number of denoising steps. However, manipulation tasks contain action chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. We introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation.
Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student–teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements.
Method
Frozen-teacher trajectories provide explicit targets for finite-time transitions. Shared LoRA adapters train interval-conditioned flow maps for one-step prediction through multi-step refinement.
A single preview predicts teacher-curvature-based difficulty and student–teacher fidelity for each candidate budget. The scheduler selects the smallest budget that meets the risk-adaptive fidelity requirements.
For a one-step budget, return the preview. For larger budgets, recycle its update into the first scheduled state and perform the remaining K − 1 updates. Scheduling needs no online teacher or candidate rollouts.
Risk measures teacher-trajectory difficulty, rather than a calibrated probability of task failure. Benefit estimates budget-specific student–teacher fidelity. Feasibility requires the fidelity threshold to be met for every active modality; if no candidate is feasible, the scheduler falls back to the budget with the largest sum of predicted modality fidelities. See the paper for the complete supervision and selection rules.
Real-world demonstrations
Unitree G1D dual-arm robot.
Task recordings play at their original 1× speed.
Showing six AnyStep-FastWAM demonstrations.
Experiments
Evaluated on 50 RoboTwin 2.0 tasks and six real-world tasks. Adaptive inference preserves comparable or higher average success rates while reducing the number of denoising steps and inference latency.
| Backbone | Success rate ↑Base → AnyStep | Denoising steps ↓Base → AnyStep | Step reduction ↓ | Latency ↓Base → AnyStep |
|---|---|---|---|---|
| Motus | 87.84 → 87.96% | 10 → 3.98 | 60.2% | 1.932 → 0.859 s |
| FastWAM | 91.83 → 91.59% | 10 → 5.03 | 49.8%‡ | 0.496 → 0.295 s |
| LingBotVA | 92.24 → 92.22% | 25 / 50 → 3.68 / 3.68 | 85.28% / 92.64% | 9.732 → 3.247 s |
Latency: seconds per inference call on one NVIDIA A100. LingBotVA values separated by “/” denote video / action. ‡ FastWAM's reduction is reported as 49.8% in the paper; its displayed mean steps are rounded to 5.03.
Cross-budget capability
AnyStep training improves average one-step success rates on RoboTwin 2.0 across all three backbones.
| Backbone | Success rate ↑Base → AnyStep | Denoising steps ↓Base → AnyStep | Latency ↓Base → AnyStep | Speedup ↑ |
|---|---|---|---|---|
| Motus | 65.83 → 69.17% | 10 → 4.06 | 1.868 → 1.117 s | 1.67× |
| FastWAM | 70.00 → 71.67% | 10 → 3.61 | 0.280 → 0.122 s | 2.30× |
| LingBotVA | 68.33 → 70.00% | 25 / 50 → 3.54 / 3.54 | 4.654 → 0.758 s | 6.14× |
Latency: seconds per inference call on one NVIDIA RTX 5090. Each base model and its AnyStep variant use 50 demonstrations per task, 8,000 optimization steps, four NVIDIA H100 GPUs, and a global batch size of 32. All results above are reported in arXiv v1.
Citation
@article{wang2026anystepwam,
title = {AnyStep-WAM: Budget-Aligned Distillation and Adaptive
Inference for World Action Models},
author = {Wang, Rui and Wang, Xiangyu and Yang, Donglin and Li, Yibo
and Chen, Canyang and Wang, Zhongrui and Qi, Xiaojuan},
journal = {arXiv preprint arXiv:2609.33748},
year = {2026},
url = {https://arxiv.org/abs/2609.33748}
}