zhudanburujiandan's picture
Initial upload: Qwen2-VL-7B LoRA finetune on DriveLM-nuScenes v1.1
616c60f verified
|
Raw
History Blame Contribute Delete
7.63 kB
---
base_model: Qwen/Qwen2-VL-7B-Instruct
library_name: peft
pipeline_tag: image-text-to-text
tags:
- lora
- peft
- vision-language
- autonomous-driving
- driving
- nuscenes
- drivelm
- qwen2-vl
datasets:
- OpenDriveLab/DriveLM
language:
- en
license: apache-2.0
---
# Qwen2-VL-7B-DriveLM-LoRA
A LoRA adapter that fine-tunes **Qwen/Qwen2-VL-7B-Instruct** on the **DriveLM-nuScenes v1.1** graph visual question-answering dataset. Turns the base Qwen2-VL into a driving perception + prediction + planning assistant that reads a front-view dashcam image and answers structured driving questions in the DriveLM convention (Perception / Prediction / Planning / Behavior).
## TL;DR
| Metric on 100 held-out val QAs (scene-based 90/10 split, seed=42) | Base Qwen2-VL-7B (zero-shot) | **This adapter** | Δ |
|---|---|---|---|
| Exact match | 9 / 100 | **62 / 100** | **+53** |
| Fuzzy substring match | 11 / 100 | 65 / 100 | +54 |
| Verbose (pred > 3× GT length) | 14 / 100 | **0 / 100** | -14 |
Per DriveLM level (25 QAs each):
| Level | Zero-shot exact | Adapter exact | Δ |
|---|---|---|---|
| Perception | 5 | 16 | +11 |
| Prediction | 4 | 22 | +18 |
| Planning | 0 | 13 | +13 |
| Behavior | 0 | 11 | +11 |
## Intended use
- Front-view dashcam / camera image → driving-focused Q&A in the DriveLM style.
- Zero-shot on driving datasets that share the DriveLM schema.
- Research + educational; **not** a production driving policy.
## Non-intended use
- Multi-camera fusion (this LoRA was trained on a single CAM_FRONT view; sideways/rear questions filtered out during training).
- Speed estimation from a single frame (a documented ceiling — see "Known limitations" below).
- Any safety-critical driving decisions.
## How to use
```python
from transformers import Qwen2VLForConditionalGeneration, AutoProcessor
from peft import PeftModel
from qwen_vl_utils import process_vision_info
import torch
BASE = "Qwen/Qwen2-VL-7B-Instruct"
ADAPTER = "zhudanburujiandan/Qwen2-VL-7B-DriveLM-LoRA"
model = Qwen2VLForConditionalGeneration.from_pretrained(
BASE, dtype=torch.bfloat16, device_map="cuda",
attn_implementation="sdpa",
)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()
processor = AutoProcessor.from_pretrained(BASE,
min_pixels=100*28*28, max_pixels=512*28*28)
messages = [
{"role": "system",
"content": "You are a driving perception assistant. You see the front "
"camera view of the ego vehicle. Answer the question concisely "
"based only on what is visible."},
{"role": "user", "content": [
{"type": "image", "image": "path/to/front_cam.jpg"},
{"type": "text", "text": "Predict the behavior of the ego vehicle."},
]},
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, _ = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs,
padding=True, return_tensors="pt").to("cuda")
with torch.inference_mode():
out = model.generate(**inputs, max_new_tokens=192, do_sample=False)
print(processor.batch_decode(out[:, inputs.input_ids.shape[1]:],
skip_special_tokens=True)[0])
```
Expected style of output on driving images:
- Perception: `"There are two cars and one truck to the front of the ego car."`
- Prediction: `"Yes."` / `"The ego vehicle."`
- Planning: `"Keep going at the same speed."`
- Behavior: `"The ego vehicle is going straight. The ego vehicle is driving slowly."`
## Training details
### Base model
- **Qwen/Qwen2-VL-7B-Instruct** (8.3B total: 675M DFN-ViT + 7.6B Qwen2 LLM).
- Vision encoder + 2×2 vision-merger MLP kept **frozen**.
### Dataset
- **OpenDriveLab/DriveLM-nuScenes v1.1** train split (696 scenes, 4,072 keyframes, 377,956 QAs).
- Scene-based 90/10 split (626 train scenes, 70 val scenes) with `seed=42`.
- Filtered to QAs answerable from CAM_FRONT alone (drops references to non-front cameras and back-zone questions). **154,363 train QAs after filter.**
### LoRA config
```python
LoraConfig(
r=32, lora_alpha=64, lora_dropout=0.05, bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj","k_proj","v_proj","o_proj",
"gate_proj","up_proj","down_proj"],
)
```
Trainable parameters: **80,740,352 / 8,372,115,968 ≈ 0.96 %**.
### Optimization
- 1 epoch, effective batch size **8** (`per_device_batch_size=1`, `gradient_accumulation_steps=8`).
- `optim="adamw_8bit"` (bitsandbytes) + `bf16=True` + `gradient_checkpointing=True`.
- Cosine schedule, `lr=2e-5`, `warmup_ratio=0.03`, `weight_decay=0`.
- Vision preprocessing: `max_pixels=512*28*28` (~480 vision tokens per image), `max_length=1024`.
- `attn_implementation="sdpa"`.
### Compute
- 1× **NVIDIA L40S 48 GB** (AWS EC2, CUDA 13.2, driver 595.64).
- **19 h 33 min** wall time. ~17 GB VRAM peak, `3.6 s/step`.
- Final `train_loss ≈ 0.30 (epoch avg)`, `eval_loss ≈ 0.18` (plateau by 52 % of epoch).
## Data scaling curve
| Train QAs | Wall time | Val exact | Val fuzzy |
|---|---|---|---|
| 500 (smoke, 3.2 ep) | 12 min | 40 | 40 |
| 20,000 (1 ep) | 2 h 33 min | 56 | 59 |
| **154,363 (1 ep, this release)** | **19 h 33 min** | **62** | **65** |
Marginal cost/benefit degrades sharply past 20 k: **7.7× more compute buys +6 exact points**. If you re-train on new driving data, **20 k QAs is likely the sweet spot** for the first pass.
## Known limitations
1. **Speed judgment**: single-frame training has no temporal signal, so `"driving fast / slowly / normally"` predictions in the Behavior category rely on static visual cues (traffic density, road curvature). Behavior exact match tops out around 44 % (11 / 25 in the eval). Real temporal input (multi-frame from nuScenes) is expected to lift this further.
2. **Side / rear scenes**: this adapter was intentionally trained on the front camera only; queries about `"to the back"` / `"back-left/right"` / `"behind"` will hallucinate. A 3-camera panorama variant was attempted and **did not help** (dropped exact from 62 → 52 due to per-camera resolution loss).
3. **DriveLM QA schema is narrow**: the adapter learns DriveLM's specific short-answer conventions (`"Yes."` / `"The ego vehicle."` / `"Please proceed."`). It may under-generate on other datasets that expect longer, unstructured answers.
4. **English-only**: DriveLM QAs are English; Chinese/other-language driving QAs are out of distribution (though the Qwen2 base has multilingual capabilities the adapter did not train them on driving).
## Reproducibility
Training pipeline: [collator with token-search label masking](https://huggingface.co/zhudanburujiandan/Qwen2-VL-7B-DriveLM-LoRA/blob/main/README.md) (`<|im_start|>assistant\n` boundary is found in tokenized `input_ids` — more robust than the naive `apply_chat_template` prompt-length approach which suffers from BPE boundary drift).
Eval harness: 100 QAs sampled balanced across 4 DriveLM levels (25 each) from the 70 val scenes, `seed=42`. Metrics computed on lowercase-stripped strings.
## Citation
If you use this adapter, cite the underlying works:
```bibtex
@article{qwen2vl,
title={{Qwen2-VL}: Enhancing Vision-Language Model's Perception of the World at Any Resolution},
author={Wang, Peng and Bai, Shuai and others},
journal={arXiv preprint arXiv:2409.12191},
year={2024},
}
@inproceedings{drivelm,
title={{DriveLM}: Driving with Graph Visual Question Answering},
author={Sima, Chonghao and Renz, Katrin and others},
booktitle={ECCV},
year={2024},
}
```
## License
Apache-2.0 (inherits from Qwen2-VL-7B-Instruct and DriveLM-nuScenes v1.1).