FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

1 University of Toronto2 Zhejiang University3 Tencent Jarvis Lab4 Mila & Université de Montréal5 Samsung SAIL6 Tsinghua University

* Equal Contribution † Corresponding Author

University of Toronto 清华大学 Mila 浙江大学 Tencent Jarvis Lab (Tencent logo) Samsung SAIL (Samsung logo)
89.29%Mean success with ARCEM
+5.31ppOver the strongest baseline
1.3×Faster planning with longer action chunks
01 / THE IDEA

A world model.
More than one stride.

Goals can be far away. Action chunks need not have a fixed length.

FlexiWorld jointly learns latent dynamics and goal-conditioned action generation through mixed-span goal supervision and variable-length action chunks. Built on JEPA, it predicts future states in a learned latent space. The same trained model supports different planning chunk lengths without retraining.

FIGURE 1 / OVERVIEW Training and planning interfaces above; selected PushT configurations and four-benchmark mean success below.
01

Learn across time scales

Mixed-span goals and variable-length chunks jointly train latent prediction and action generation.

02

Generate, then refine

An autoregressive actor enables search-free Direct control. ARCEM optionally refines actions through residual search.

03

Change the granularity

Longer chunks reduce latent prediction steps, without retraining or changing the primitive-action execution rate.

02 / ABSTRACT

Abstract

Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately 1.3× on average while maintaining comparable average success.

03 / RECORDED EXECUTION

Same goal.
Different outcomes.

Watch FlexiWorld and INTACT pursue the same visual goal. Every frame below comes from the recorded benchmark evaluations.

INITIAL OBSERVATIONFIRST PLAN0 / 150
04 / HOW IT WORKS

Flexible by training.
Autoregressive by design.

A causal action encoder handles variable-length chunks. An autoregressive actor generates primitive actions, and Student Forcing exposes it to generated prefixes during training.

Mixed-span learning

Mixed goal spans and variable-length chunks jointly train the world model and actor.

Autoregressive actor

Both methods predict states between chunks; FlexiWorld also feeds generated actions back within each chunk.

ARCEM

Action-residual search with within-chunk feedback and latent prediction at chunk boundaries.
FIGURE 2 / JOINT TRAINING Variable-length transitions couple the action encoder, latent predictor, and goal-conditioned actor.
ACTOR-RESIDUAL CROSS-ENTROPY METHOD

Refine actions.
Let later actions respond.

ARCEM generalizes POPLIN's action-residual replanning to autoregressive chunks. Each perturbed action becomes part of the prefix for subsequent actions; latent predictions carry its effects across chunks.

01 Sample residuals02 Generate candidates03 Score & refit
FIGURE 3 / TEST-TIME SEARCH Within-chunk action feedback. Chunk-level latent prediction. No residual-policy training.
05 / THE EVIDENCE

Go further.
Stay goal-directed.

Visual goal-reaching on PushT, Cube, Reacher, and TwoRoom, evaluated over 25, 50, 75, and 100 primitive steps with up to one replan.

Main-paper results. FlexiWorld uses three training seeds; baselines use one checkpoint per task. All use three evaluation seeds and 100 episodes per distance and evaluation seed. Error values are sample SD: across training seeds for FlexiWorld, across evaluation seeds for baselines. INTACT's main-table configuration is Guarded-A. Task-specific outcomes vary.

Full benchmark table
Success rate (%), mean ± sample SD
Method PushT Cube Reacher TwoRoom Average
WITHOUT RETRAINING

Trade granularity
for planning speed.

Longer chunks make fewer latent predictions. At D = 50, ARCEM planning is approximately 1.3× faster, with comparable four-task mean success in the matched chunk-length study.

Matched timing inputs, including encoding and planning but excluding environment execution. Each input uses the median of 20 synchronized solves after warm-up. These latency measurements are separate from the selected videos and main-table success evaluations.

06 / DIAGNOSTICS

Where do the
gains come from?

A closer look at action prediction and learned features.

ACTION PREDICTION

Better action predictions at longer goal distances.

On PushT, FlexiWorld improves all four mean action-prediction metrics at D = 50, 75, and 100. INTACT predicts better at D = 25, while Direct success remains comparable.

APPENDIX D.1 Action-prediction quality across goal distances; hollow points show individual checkpoints.
FROZEN FEATURES

Useful temporal and action information, without the actor.

Frozen-feature probes predict demonstrated time gaps better with FlexiWorld. Action-probe R2 matches or exceeds INTACT at every tested distance.

APPENDIX D.2 Temporal-gap and action probes on frozen representations.

These probe gains do not translate into a clear actor-free control advantage: both models perform similarly with CEM alone.

Citation

@misc{ren2026flexiworld,
  title = {FlexiWorld: Learning and Planning via Flexible Action Chunks
           Across Multiple Time Scales},
  author = {Shidu Ren and Qilin Gu and Zhenghao Ni and Junhan Sun
            and Jiaqi Wang and Damien Scieur and Yunze Liu},
  year = {2026},
  eprint = {2609.35138},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url = {https://arxiv.org/abs/2609.35138}
}

Evaluated in simulation; real-world deployment remains untested.