Cosmos3-Super SO-101 Forward Dynamics β€” chunk_length=16 (iter 4000)

Action-conditioned world model for the SO-101 robot arm, post-trained from nvidia/Cosmos3-Super (64B). Feed one real camera frame and a recorded/candidate action sequence; the model generates the video of the robot executing it. Intended use is robot policy evaluation and failure analysis without running the physical robot.

This is the chunk_length=16 checkpoint: the model predicts 16 action steps (17 frames at 30 fps, ~0.53 s) per chunk. Longer rollouts are produced by autoregressively chaining chunks. See the chunk_length=32 variant for the horizon-targeted follow-up.

Model details

Base model nvidia/Cosmos3-Super (64B, MoT)
Adaptation LoRA rank 16 / alpha 32 on q/k/v/o_proj_moe_gen + unfrozen action pathway (21.1M params: action_modality_embed, action2llm.*, llm2action.*)
Mode forward_dynamics, action-conditioned video generation
Action space 10-D Cartesian EE (dx, dy, dz, 6-D rotation, gripper), quantile-normalized
Chunk length 16 action steps (17 frames @ 30 fps)
Camera single top view, ego_view, 480p
Training iteration 4000 (of a 30000-iter schedule, lr 2e-4 LambdaCosine)
Dataset geonmin-kim/so101_merged_v2 β€” 1444 episodes / 486k frames / 33 tasks
Parallelism FSDP 4-way shard, bf16 (fp32 master), 4Γ— A100 80GB
Effective batch grad_accum 4, max 24000 tokens after packing

Checkpoint format

PyTorch Distributed Checkpoint (DCP), saved from a 4-rank FSDP run:

model/      __0_0.distcp ... __3_0.distcp   (~120 GB total, full model + LoRA)
optim/      optimizer state (LoRA + action pathway only, ~250 MB)
scheduler/  LR scheduler state
trainer/    trainer bookkeeping (iteration counter, RNG)

This is not a safetensors/HF-format checkpoint. Load it through cosmos-framework (pinned to commit 5e67049) with the SO-101 overlay from nota-github/xpu-cosmos3-simulator (branch feat/so101-a100-port-and-action-pathway).

Usage

export COSMOS3_SIM_ROOT=/path/to/working/root

# 1. environment (see the simulator repo README for full setup)
git clone -b feat/so101-a100-port-and-action-pathway \
  https://github.com/nota-github/xpu-cosmos3-simulator.git
cd xpu-cosmos3-simulator
git clone https://github.com/NVIDIA/cosmos-framework.git "$COSMOS3_SIM_ROOT/packages/cosmos-framework"
git -C "$COSMOS3_SIM_ROOT/packages/cosmos-framework" checkout 5e67049
./overlay/apply_overlay.sh "$COSMOS3_SIM_ROOT/packages/cosmos-framework"
./scripts/setup_venv313.sh

# 2. this checkpoint
hf download geonmin-kim/cosmos3-super-so101-fd-chunk16 \
  --local-dir "$COSMOS3_SIM_ROOT/ckpt/chunk16_iter4000"

# 3. action normalization stats (REQUIRED β€” see warning below)
export SO101_ACTION_STATS="$COSMOS3_SIM_ROOT/ckpt/chunk16_iter4000/so101_stats_stride1_v2.json"

# 4. build conditioning inputs from a LeRobot episode, then roll out
source ./env.sh
python scripts/make_inputs.py --episodes 0 63 119
./scripts/run_rollout.sh 0,1 "$COSMOS3_SIM_ROOT/ckpt/chunk16_iter4000" \
    "$COSMOS3_SIM_ROOT/out/rollouts" \
    --input-dirs "$COSMOS3_SIM_ROOT"/out/inputs/ep*_droid_lerobot_s1_* \
    --modes autoregressive teacher_forced

# 5. score against the recorded episode
python scripts/evaluate.py --rollout-dir "$COSMOS3_SIM_ROOT/out/rollouts"
python scripts/make_comparison.py --rollout-dir "$COSMOS3_SIM_ROOT/out/rollouts"

⚠️ Normalization stats are part of the model contract. This checkpoint was trained with so101_stats_stride1_v2.json (bundled in this repo). Using the pre-megamix v1 stats silently mis-scales actions by up to 19%. Do not mix.

Sampling configuration used at eval time: num_steps=30, guidance=1.0, shift=10.0, sigma_max=80.0, resolution 480, 16:9, fps 30.

Evaluation (at iter 4000)

Replay evaluation on held-out episodes (teacher-forced = ground-truth frame re-injected each chunk; autoregressive = model consumes its own output):

Metric Value
PSNR / SSIM (autoregressive) 18.8 dB / 0.844
Motion ratio (1.0 = matches reality) 0.87
Motion correlation 0.50
Axis separation (90Β° β‰ˆ ceiling) 86Β° (base model: 32Β°)
Usable horizon ~0.5 s
Real-time factor 0.067

Axis-probe verdict (synthetic constant-direction probes, 16-frame horizon): x-axis sign respected and trustworthy (cos βˆ’0.95); z-axis sign respected but low arm visibility; y-axis inconclusive. Known limitation: the model holds up across chunks with real conditioning frames (teacher-forced PSNR improves 31.1 β†’ 34.4) but degrades when consuming its own output (31.1 β†’ 18.0) β€” the motivation for the chunk_length=32 variant.

Training provenance

  • Launched via train/run_train.sh super_action with CLI overrides model.config.activation_checkpointing.mode=selective model.config.compile.enabled=True
  • SO101_TRAIN_ACTION_PATHWAY=1 (required β€” without it the LoRA injector freezes the action pathway and training silently degenerates into first-frame video prediction)
  • Base checkpoint: nvidia/Cosmos3-Super converted to DCP via cosmos_framework.scripts.convert_model_to_dcp

License

Base model and framework: OpenMDW-1.1 (NVIDIA Cosmos3). Fine-tuned weights released under the same terms. Training data: see geonmin-kim/so101_merged_v2 (Apache-2.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for geonmin-kim/cosmos3-super-so101-fd-chunk16

Adapter
(2)
this model

Dataset used to train geonmin-kim/cosmos3-super-so101-fd-chunk16