StarVLA QwenPI_v3 Qwen3-VL-4B for RoboDojo (100k)
This directory contains one StarVLA QwenPI_v3 checkpoint initialized from
Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture.
The VLM, VLM interface, and action model were trained end to end; this is not a
LoRA or adapter-only checkpoint.
Model details
| Item | Value |
|---|---|
| Framework | StarVLA QwenPI_v3 |
| Base VLM | Qwen3-VL-4B-Instruct |
| Action model | 36-layer LayerwiseFM |
| Action representation | 14D absolute joint position (abs_qpos) |
| Action horizon | 50 |
| State dimension | 14 |
| Camera input | Head, left wrist, right wrist; resized to 224 x 224 |
| Inference flow steps | 4 |
| Checkpoint step | 100,000 |
| Checkpoint format | Complete StarVLA framework state dict (.pt) |
The policy predicts a normalized 50 x 14 action chunk. RoboDojo evaluation
must use the saved arx_x5 normalization statistics and execute 16 actions
before requesting the next chunk.
Files
README.md
config.yaml
config.full.yaml
dataset_statistics.json
summary.jsonl
checkpoints/
└── steps_100000_pytorch_model.pt
Only the requested 100k checkpoint is included. Keep the configuration and
dataset statistics beside the checkpoints/ directory; StarVLA uses them to
reconstruct the framework and unnormalize actions.
Training details
| Setting | Value |
|---|---|
| Dataset mixture | robodojo_v21_all_h50_q99 |
| Training tasks | 35 |
| Per-GPU batch size | 16 |
| Gradient accumulation | 1 |
| Frozen modules | None |
| Optimizer | AdamW, betas (0.9, 0.95), epsilon 1e-8 |
| VLM learning rate | 1e-5 |
| VLM-interface learning rate | 1e-5 |
| Action-model learning rate | 1e-4 |
| Schedule | Cosine, 5,000 warmup steps, minimum LR 5e-7 |
| Gradient checkpointing | Enabled |
| Random seed | 42 |
Official RoboDojo evaluation
All policies below use the official complete 42-task protocol: 50 episodes per
task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This
directory's policy is bolded. Higher is better for both SR and Score.
Group summary
| Policy | Average | Generalization | Precision | Long-Horizon | Memory | Open |
|---|---|---|---|---|---|---|
| QwenOFT | 4.86 / 8.01 | 4.33 / 6.42 | 11.75 / 17.54 | 5.50 / 12.95 | 1.67 / 1.77 | 0.50 / 0.60 |
| QwenGR00T | 3.81 / 7.35 | 3.50 / 6.52 | 5.75 / 10.09 | 6.50 / 15.46 | 3.33 / 4.37 | 0.00 / 0.00 |
| QwenPI_v3 | 6.19 / 9.60 | 4.17 / 7.28 | 14.00 / 19.06 | 10.00 / 17.84 | 2.00 / 2.32 | 0.75 / 0.88 |
Task details
Task values are SR (%) / Score; each task uses 50 episodes. The bolded
column is the policy in this directory.
| Evaluation group / task | QwenOFT | QwenGR00T | QwenPI_v3 |
|---|---|---|---|
| Generalization | 4.33 / 6.42 | 3.50 / 6.52 | 4.17 / 7.28 |
| stack_bowls | 18.00 / 21.00 | 10.00 / 14.80 | 14.00 / 16.70 |
| push_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pack_objects_into_box | 0.00 / 3.10 | 0.00 / 7.80 | 0.00 / 6.80 |
| fold_clothes | 10.00 / 12.80 | 8.00 / 12.40 | 2.00 / 9.60 |
| hang_mugs | 0.00 / 3.60 | 0.00 / 3.00 | 0.00 / 3.50 |
| sweep_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| pour_liquid_into_cup | 14.00 / 14.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| make_toast | 0.00 / 1.00 | 2.00 / 5.00 | 2.00 / 5.00 |
| arrange_largest_number | 0.00 / 1.90 | 2.00 / 4.10 | 2.00 / 5.70 |
| sort_nesting_dolls_by_size | 0.00 / 0.00 | 4.00 / 4.00 | 6.00 / 6.00 |
| store_laptop_and_headphones | 4.00 / 11.20 | 0.00 / 7.20 | 2.00 / 8.40 |
| stack_blocks | 6.00 / 8.40 | 2.00 / 5.90 | 8.00 / 11.60 |
| Precision | 11.75 / 17.54 | 5.75 / 10.09 | 14.00 / 19.06 |
| fasten_screws | 4.00 / 8.00 | 0.00 / 2.00 | 0.00 / 6.00 |
| plug_in_charger | 6.00 / 6.00 | 2.00 / 2.00 | 4.00 / 4.00 |
| insert_tubes | 40.00 / 51.60 | 28.00 / 40.40 | 44.00 / 56.80 |
| pour_balls_into_vase | 8.00 / 8.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| play_Xylophone | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| deposit_coin | 0.00 / 3.20 | 2.00 / 5.60 | 6.00 / 7.60 |
| insert_key | 0.00 / 12.90 | 0.00 / 9.90 | 0.00 / 11.10 |
| build_tower | 36.00 / 50.60 | 14.00 / 20.80 | 56.00 / 65.00 |
| Long-Horizon | 5.50 / 12.95 | 6.50 / 15.46 | 10.00 / 17.84 |
| put_bottles_into_dustbin | 22.00 / 40.90 | 26.00 / 44.40 | 64.00 / 73.60 |
| fill_pen_holder | 4.00 / 11.70 | 6.00 / 14.40 | 8.00 / 23.00 |
| classify_objects | 2.00 / 5.50 | 6.00 / 11.50 | 0.00 / 7.50 |
| play_tic_tac_toe | 0.00 / 12.40 | 2.00 / 16.40 | 0.00 / 6.80 |
| fill_egg_holder | 0.00 / 0.60 | 0.00 / 0.00 | 0.00 / 0.80 |
| organize_table | 0.00 / 16.50 | 0.00 / 25.00 | 4.00 / 27.00 |
| make_kong | 16.00 / 16.00 | 12.00 / 12.00 | 4.00 / 4.00 |
| play_stacking_toy | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| Memory | 1.67 / 1.77 | 3.33 / 4.37 | 2.00 / 2.32 |
| cover_blocks | 0.00 / 0.60 | 6.00 / 12.10 | 0.00 / 1.50 |
| match_and_pick_from_conveyor | 10.00 / 10.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| swap_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| swap_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| press_by_number | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| imitate_sorting_sequence | 0.00 / 0.00 | 0.00 / 0.10 | 0.00 / 0.40 |
| Open | 0.50 / 0.60 | 0.00 / 0.00 | 0.75 / 0.88 |
| align_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| general_pickup | 4.00 / 4.00 | 0.00 / 0.00 | 6.00 / 6.00 |
| stack_blocks_by_language | 0.00 / 0.80 | 0.00 / 0.00 | 0.00 / 0.80 |
| solve_equation | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| classify_objects_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.20 |
| pick_from_conveyor_by_image | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| store_tools_in_toolbox | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pour_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
Evaluation
This is a StarVLA checkpoint, not a Hugging Face from_pretrained() directory.
Use the StarVLA model server and the XPolicyLab StarVLA adapter.
Start the model server from the StarVLA repository:
export CKPT=/path/to/MODEL_DIR/checkpoints/steps_100000_pytorch_model.pt
python deployment/model_server/server_policy.py \
--ckpt_path "$CKPT" \
--port 57700 \
--use_bf16
Run a RoboDojo task from XPolicyLab/policy/starVLA:
STARVLA_CKPT_PATH="$CKPT" \
STARVLA_INCLUDE_STATE=True \
STARVLA_UNNORM_KEY=arx_x5 \
STARVLA_EXECUTE_HORIZON=16 \
bash eval.sh \
RoboDojo build_tower qwenpi_v3_steps_100000 \
arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>
The final arguments are the seed, policy GPU, simulator GPU, policy environment, and RoboDojo environment. Use the RoboDojo task registry's native episode counts when producing an official aggregate.
Intended use
This checkpoint is intended for RoboDojo simulation research with the ARX X5 dual-arm embodiment. Performance with different camera calibration, state/action ordering, normalization, robots, or real-world hardware has not been established.
- Downloads last month
- 23
Model tree for StarVLA/StarVLA-Qwen3vl4b-PIv3-RoboDojo
Base model
Qwen/Qwen3-VL-4B-Instruct