Learn across time scales
Mixed-span goals and variable-length chunks jointly train latent prediction and action generation.
Goals can be far away. Action chunks need not have a fixed length.
FlexiWorld jointly learns latent dynamics and goal-conditioned action generation through mixed-span goal supervision and variable-length action chunks. Built on JEPA, it predicts future states in a learned latent space. The same trained model supports different planning chunk lengths without retraining.
Mixed-span goals and variable-length chunks jointly train latent prediction and action generation.
An autoregressive actor enables search-free Direct control. ARCEM optionally refines actions through residual search.
Longer chunks reduce latent prediction steps, without retraining or changing the primitive-action execution rate.
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately 1.3× on average while maintaining comparable average success.
Watch FlexiWorld and INTACT pursue the same visual goal. Every frame below comes from the recorded benchmark evaluations.
A causal action encoder handles variable-length chunks. An autoregressive actor generates primitive actions, and Student Forcing exposes it to generated prefixes during training.
ARCEM generalizes POPLIN's action-residual replanning to autoregressive chunks. Each perturbed action becomes part of the prefix for subsequent actions; latent predictions carry its effects across chunks.
Visual goal-reaching on PushT, Cube, Reacher, and TwoRoom, evaluated over 25, 50, 75, and 100 primitive steps with up to one replan.
Main-paper results. FlexiWorld uses three training seeds; baselines use one checkpoint per task. All use three evaluation seeds and 100 episodes per distance and evaluation seed. Error values are sample SD: across training seeds for FlexiWorld, across evaluation seeds for baselines. INTACT's main-table configuration is Guarded-A. Task-specific outcomes vary.
| Method | PushT | Cube | Reacher | TwoRoom | Average |
|---|
Longer chunks make fewer latent predictions. At D = 50, ARCEM planning is approximately 1.3× faster, with comparable four-task mean success in the matched chunk-length study.
Matched timing inputs, including encoding and planning but excluding environment execution. Each input uses the median of 20 synchronized solves after warm-up. These latency measurements are separate from the selected videos and main-table success evaluations.
A closer look at action prediction and learned features.
On PushT, FlexiWorld improves all four mean action-prediction metrics at D = 50, 75, and 100. INTACT predicts better at D = 25, while Direct success remains comparable.
Frozen-feature probes predict demonstrated time gaps better with FlexiWorld. Action-probe R2 matches or exceeds INTACT at every tested distance.
These probe gains do not translate into a clear actor-free control advantage: both models perform similarly with CEM alone.
@misc{ren2026flexiworld,
title = {FlexiWorld: Learning and Planning via Flexible Action Chunks
Across Multiple Time Scales},
author = {Shidu Ren and Qilin Gu and Zhenghao Ni and Junhan Sun
and Jiaqi Wang and Damien Scieur and Yunze Liu},
year = {2026},
eprint = {2609.35138},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2609.35138}
}
Evaluated in simulation; real-world deployment remains untested.