VLADrop-pi05-LIBERO-flops-drop12-gate

Checkpoint for Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?.

DTR (Drop-Then-Recovery) removes transformer blocks from a pretrained VLA model and recovery-fine-tunes the smaller dense model. Code: https://github.com/s1ghhh/VLADrop

This checkpoint

Paper row Table 4 'Drop-12' (FLOPs-matched; best average)
Dropped blocks LLM blocks [3,4,7,9,10,11,12,13,14,15,16,17] (GateProbe; keep [0,1,2,5,6,8])
Recovery training batch size 64, 34235 steps (FLOPs-matched), lr 5e-5
LIBERO success rate Spatial 96.8 / Object 98.4 / Goal 94.4 / Long 85.0 / Avg 93.7

Usage

This is an openpi-format pi0.5 checkpoint (PyTorch). Use with the VLADrop fork: https://github.com/s1ghhh/VLADrop

python scripts/serve_policy_batch_drop.py \
    --config pi05_libero_dropped \
    --dir <this_repo_local_path> \
    --port 8000

Important: the drop lists are NOT stored inside the checkpoint. Pass the exact llm_drop_attn_list / llm_drop_mlp_list shown above (via config or CLI) when serving, otherwise layers will be mismatched. assets/ contains the LIBERO norm stats. The optimizer state (train_state/) is not included.

Citation

@article{sun2026vladrop,
  title={Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?},
  author={Sun, Guoheng and Feng, Kaixi and He, Shwai and Gong, Xiaochuan and He, Yexiao and Wang, Ziyao and Shen, Zheyu and Ye, Wanghao and Kompella, Ramana Rao and Liu, Gaowen and Li, Ang},
  journal={arXiv preprint arXiv:2606.27755},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Collection including s1ghhh/VLADrop_pi05_LIBERO_Drop12Block_GateProbe_FLOPsMatched

Paper for s1ghhh/VLADrop_pi05_LIBERO_Drop12Block_GateProbe_FLOPsMatched