StarVLA QwenPI_v3 Qwen3-VL-4B for RoboDojo (100k)

This directory contains one StarVLA QwenPI_v3 checkpoint initialized from Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture. The VLM, VLM interface, and action model were trained end to end; this is not a LoRA or adapter-only checkpoint.

Model details

Item Value
Framework StarVLA QwenPI_v3
Base VLM Qwen3-VL-4B-Instruct
Action model 36-layer LayerwiseFM
Action representation 14D absolute joint position (abs_qpos)
Action horizon 50
State dimension 14
Camera input Head, left wrist, right wrist; resized to 224 x 224
Inference flow steps 4
Checkpoint step 100,000
Checkpoint format Complete StarVLA framework state dict (.pt)

The policy predicts a normalized 50 x 14 action chunk. RoboDojo evaluation must use the saved arx_x5 normalization statistics and execute 16 actions before requesting the next chunk.

Files

README.md
config.yaml
config.full.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt

Only the requested 100k checkpoint is included. Keep the configuration and dataset statistics beside the checkpoints/ directory; StarVLA uses them to reconstruct the framework and unnormalize actions.

Training details

Setting Value
Dataset mixture robodojo_v21_all_h50_q99
Training tasks 35
Per-GPU batch size 16
Gradient accumulation 1
Frozen modules None
Optimizer AdamW, betas (0.9, 0.95), epsilon 1e-8
VLM learning rate 1e-5
VLM-interface learning rate 1e-5
Action-model learning rate 1e-4
Schedule Cosine, 5,000 warmup steps, minimum LR 5e-7
Gradient checkpointing Enabled
Random seed 42

Official RoboDojo evaluation

All policies below use the official complete 42-task protocol: 50 episodes per task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This directory's policy is bolded. Higher is better for both SR and Score.

Group summary

Policy Average Generalization Precision Long-Horizon Memory Open
QwenOFT 4.86 / 8.01 4.33 / 6.42 11.75 / 17.54 5.50 / 12.95 1.67 / 1.77 0.50 / 0.60
QwenGR00T 3.81 / 7.35 3.50 / 6.52 5.75 / 10.09 6.50 / 15.46 3.33 / 4.37 0.00 / 0.00
QwenPI_v3 6.19 / 9.60 4.17 / 7.28 14.00 / 19.06 10.00 / 17.84 2.00 / 2.32 0.75 / 0.88

Task details

Task values are SR (%) / Score; each task uses 50 episodes. The bolded column is the policy in this directory.

Evaluation group / task QwenOFT QwenGR00T QwenPI_v3
Generalization 4.33 / 6.42 3.50 / 6.52 4.17 / 7.28
stack_bowls 18.00 / 21.00 10.00 / 14.80 14.00 / 16.70
push_T 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
pack_objects_into_box 0.00 / 3.10 0.00 / 7.80 0.00 / 6.80
fold_clothes 10.00 / 12.80 8.00 / 12.40 2.00 / 9.60
hang_mugs 0.00 / 3.60 0.00 / 3.00 0.00 / 3.50
sweep_blocks 0.00 / 0.00 0.00 / 0.00 2.00 / 2.00
pour_liquid_into_cup 14.00 / 14.00 14.00 / 14.00 12.00 / 12.00
make_toast 0.00 / 1.00 2.00 / 5.00 2.00 / 5.00
arrange_largest_number 0.00 / 1.90 2.00 / 4.10 2.00 / 5.70
sort_nesting_dolls_by_size 0.00 / 0.00 4.00 / 4.00 6.00 / 6.00
store_laptop_and_headphones 4.00 / 11.20 0.00 / 7.20 2.00 / 8.40
stack_blocks 6.00 / 8.40 2.00 / 5.90 8.00 / 11.60
Precision 11.75 / 17.54 5.75 / 10.09 14.00 / 19.06
fasten_screws 4.00 / 8.00 0.00 / 2.00 0.00 / 6.00
plug_in_charger 6.00 / 6.00 2.00 / 2.00 4.00 / 4.00
insert_tubes 40.00 / 51.60 28.00 / 40.40 44.00 / 56.80
pour_balls_into_vase 8.00 / 8.00 0.00 / 0.00 2.00 / 2.00
play_Xylophone 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
deposit_coin 0.00 / 3.20 2.00 / 5.60 6.00 / 7.60
insert_key 0.00 / 12.90 0.00 / 9.90 0.00 / 11.10
build_tower 36.00 / 50.60 14.00 / 20.80 56.00 / 65.00
Long-Horizon 5.50 / 12.95 6.50 / 15.46 10.00 / 17.84
put_bottles_into_dustbin 22.00 / 40.90 26.00 / 44.40 64.00 / 73.60
fill_pen_holder 4.00 / 11.70 6.00 / 14.40 8.00 / 23.00
classify_objects 2.00 / 5.50 6.00 / 11.50 0.00 / 7.50
play_tic_tac_toe 0.00 / 12.40 2.00 / 16.40 0.00 / 6.80
fill_egg_holder 0.00 / 0.60 0.00 / 0.00 0.00 / 0.80
organize_table 0.00 / 16.50 0.00 / 25.00 4.00 / 27.00
make_kong 16.00 / 16.00 12.00 / 12.00 4.00 / 4.00
play_stacking_toy 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
Memory 1.67 / 1.77 3.33 / 4.37 2.00 / 2.32
cover_blocks 0.00 / 0.60 6.00 / 12.10 0.00 / 1.50
match_and_pick_from_conveyor 10.00 / 10.00 14.00 / 14.00 12.00 / 12.00
swap_blocks 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
swap_T 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
press_by_number 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
imitate_sorting_sequence 0.00 / 0.00 0.00 / 0.10 0.00 / 0.40
Open 0.50 / 0.60 0.00 / 0.00 0.75 / 0.88
align_blocks 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
general_pickup 4.00 / 4.00 0.00 / 0.00 6.00 / 6.00
stack_blocks_by_language 0.00 / 0.80 0.00 / 0.00 0.00 / 0.80
solve_equation 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
classify_objects_by_language 0.00 / 0.00 0.00 / 0.00 0.00 / 0.20
pick_from_conveyor_by_image 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
store_tools_in_toolbox 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00
pour_by_language 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00

Evaluation

This is a StarVLA checkpoint, not a Hugging Face from_pretrained() directory. Use the StarVLA model server and the XPolicyLab StarVLA adapter.

Start the model server from the StarVLA repository:

export CKPT=/path/to/MODEL_DIR/checkpoints/steps_100000_pytorch_model.pt
python deployment/model_server/server_policy.py \
  --ckpt_path "$CKPT" \
  --port 57700 \
  --use_bf16

Run a RoboDojo task from XPolicyLab/policy/starVLA:

STARVLA_CKPT_PATH="$CKPT" \
STARVLA_INCLUDE_STATE=True \
STARVLA_UNNORM_KEY=arx_x5 \
STARVLA_EXECUTE_HORIZON=16 \
bash eval.sh \
  RoboDojo build_tower qwenpi_v3_steps_100000 \
  arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>

The final arguments are the seed, policy GPU, simulator GPU, policy environment, and RoboDojo environment. Use the RoboDojo task registry's native episode counts when producing an official aggregate.

Intended use

This checkpoint is intended for RoboDojo simulation research with the ARX X5 dual-arm embodiment. Performance with different camera calibration, state/action ordering, normalization, robots, or real-world hardware has not been established.

Downloads last month
23
Video Preview
loading

Model tree for StarVLA/StarVLA-Qwen3vl4b-PIv3-RoboDojo

Finetuned
(377)
this model