Kaminoikari commited on
Commit
3ac08be
verified
1 Parent(s): f71ad8c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +52 -52
README.md CHANGED
@@ -1,52 +1,52 @@
1
- ---
2
- license: apache-2.0
3
- library_name: lerobot
4
- base_model: lerobot/smolvla_base
5
- datasets:
6
- - lerobot/svla_so101_pickplace
7
- tags:
8
- - robotics
9
- - lerobot
10
- - smolvla
11
- - vision-language-action
12
- - so-101
13
- - imitation-learning
14
- pipeline_tag: robotics
15
- ---
16
-
17
- # SmolVLA 路 SO-101 Pick-and-Place (fine-tuned)
18
-
19
- Fine-tune of [`lerobot/smolvla_base`](https://huggingface.co/lerobot/smolvla_base) (450M VLA)
20
- on the real-robot SO-101 dataset
21
- [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
22
- (50 teleoperated episodes, 2 cameras `up`/`side`, 6-DoF state).
23
- [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
24
- [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
25
- [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
26
- (50 teleoperated episodes, 2 cameras `up`/`side`, 6-DoF state).
27
-
28
- ## Training
29
-
30
- | | |
31
- |---|---|
32
- | Base | `lerobot/smolvla_base` |
33
- | Dataset | `lerobot/svla_so101_pickplace` (50 ep / 11,939 frames) |
34
- | Steps | 2,000 (batch size 8) |
35
- | GPU | single T4 (~78 min) |
36
- | Camera mapping | `up to camera1`, `side to camera2` via `--rename_map` |
37
- | Loss | 0.410 to 0.141 (monotonic) |
38
-
39
- ## Intended use & honest limitations
40
-
41
- This is a **pipeline-validation / learning run**, not a production policy.
42
-
43
- - Demonstrates the full real-robot imitation-learning loop: load a real teleoperation dataset, fine-tune a pretrained VLA, converge, ship a checkpoint.
44
- - Only 2,000 steps (~1.3 epochs). The SmolVLA paper uses ~20k steps; expect this checkpoint to under-perform a fully trained one.
45
- - No closed-loop success rate. Evaluation on the physical SO-101 arm (`lerobot-record`) was not run (no hardware). Reported signal is training-loss
46
- convergence only, which proves the model is learning, not real-world task success.
47
-
48
- ## Load
49
-
50
- ```python
51
- from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy
52
- policy = SmolVLAPolicy.from_pretrained("Kaminoikari/smolvla-so101-pickplace-ft")
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: lerobot
4
+ base_model: lerobot/smolvla_base
5
+ datasets:
6
+ - lerobot/svla_so101_pickplace
7
+ tags:
8
+ - robotics
9
+ - lerobot
10
+ - smolvla
11
+ - vision-language-action
12
+ - so-101
13
+ - imitation-learning
14
+ pipeline_tag: robotics
15
+ ---
16
+
17
+ # SmolVLA 路 SO-101 Pick-and-Place (fine-tuned)
18
+
19
+ Fine-tune of [`lerobot/smolvla_base`](https://huggingface.co/lerobot/smolvla_base) (450M VLA)
20
+ on the real-robot SO-101 dataset
21
+ [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
22
+ (50 teleoperated episodes, 2 cameras `up`/`side`, 6-DoF state).
23
+ [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
24
+ [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
25
+ [`lerobot/svla_so101_pickplace`](https://huggingface.co/datasets/lerobot/svla_so101_pickplace)
26
+ (50 teleoperated episodes, 2 cameras `up`/`side`, 6-DoF state).
27
+
28
+ ## Training
29
+
30
+ | | |
31
+ |---|---|
32
+ | Base | `lerobot/smolvla_base` |
33
+ | Dataset | `lerobot/svla_so101_pickplace` (50 ep / 11,939 frames) |
34
+ | Steps | 2,000 (batch size 8) |
35
+ | GPU | single T4 (~78 min) |
36
+ | Camera mapping | `up to camera1`, `side to camera2` via `--rename_map` |
37
+ | Loss | 0.410 to 0.141 (monotonic) |
38
+
39
+ ## Intended use & honest limitations
40
+
41
+ This is a **pipeline-validation / learning run**, not a production policy.
42
+
43
+ - Demonstrates the full real-robot imitation-learning loop: load a real teleoperation dataset, fine-tune a pretrained VLA, converge, ship a checkpoint.
44
+ - Only 2,000 steps (~1.3 epochs). The SmolVLA paper uses ~20k steps; expect this checkpoint to under-perform a fully trained one.
45
+ - No closed-loop success rate. Evaluation on the physical SO-101 arm (`lerobot-record`) was not run (no hardware). Reported signal is training-loss
46
+ convergence only, which proves the model is learning, not real-world task success.
47
+
48
+ ## Load
49
+
50
+ ```python
51
+ from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy
52
+ policy = SmolVLAPolicy.from_pretrained("Kaminoikari/smolvla-so101-pickplace-ft")