File size: 5,504 Bytes
e81a5b4
6ab2916
 
 
 
 
 
 
 
 
 
 
 
 
 
e81a5b4
 
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
 
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
 
 
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
 
 
 
 
 
 
 
 
 
 
e81a5b4
6ab2916
e81a5b4
 
 
6ab2916
e81a5b4
6ab2916
 
195c26a
 
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
 
 
 
e81a5b4
6ab2916
 
e81a5b4
6ab2916
 
 
e81a5b4
6ab2916
 
 
e81a5b4
6ab2916
 
 
 
 
 
e81a5b4
6ab2916
 
e81a5b4
6ab2916
e81a5b4
6ab2916
 
e81a5b4
6ab2916
 
 
 
 
 
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
 
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
e81a5b4
6ab2916
 
 
 
 
 
 
e81a5b4
6ab2916
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
---
license: apache-2.0
library_name: peft
base_model: WinKawaks/vit-tiny-patch16-224
tags:
  - lora
  - peft
  - image-classification
  - vit
  - food101
datasets:
  - food101
pipeline_tag: image-classification
metrics:
  - accuracy
---

# ViT-tiny LoRA adapter on Food-101

A LoRA adapter that teaches [`WinKawaks/vit-tiny-patch16-224`](https://huggingface.co/WinKawaks/vit-tiny-patch16-224) to classify images from the [Food-101](https://huggingface.co/datasets/food101) dataset (101 food categories) while leaving the original pretrained weights mathematically untouched.

- **Base model:** `WinKawaks/vit-tiny-patch16-224` (~5.7M params)
- **Dataset:** [Food-101](https://huggingface.co/datasets/food101) (75,750 train / 25,250 test, 101 classes)
- **Method:** LoRA on attention `query` + `value` projections + a fresh 101-way classification head
- **Demo Space:** [`turhancan97/vit-tiny-imagenet-demo`](https://huggingface.co/spaces/turhancan97/vit-tiny-imagenet-demo)

## How it works

The backbone is never fine-tuned. Instead a low-rank update $\Delta W = BA$ (with rank $r = 8$) is added to each attention projection, and a separate 101-class linear head is trained on top of the pooled CLS features. The full artifact is tiny (~1–2 MB) and additive — disabling the adapter at inference time recovers the exact original ImageNet-1k model.

```text
adapter_config.json           # PEFT LoRA config
adapter_model.safetensors     # LoRA weights (B, A matrices)
classifier.pt                 # 101-way Linear head (state_dict)
labels.json                   # {"0": "apple_pie", "1": "baby_back_ribs", ...}
preprocessor_config.json      # image processor (224x224, standard ImageNet norm)
```

## Training

Trained with the script at [`turhancan97/vit-tiny-imagenet-demo/train_lora.py`](https://huggingface.co/spaces/turhancan97/vit-tiny-imagenet-demo/blob/main/train_lora.py):

```bash
python train_lora.py \
  --rank 8 --alpha 16 --dropout 0.1 \
  --target-modules query value \
  --epochs 5 --batch-size 64 --lr 5e-4 \
  --warmup-ratio 0.03 --weight-decay 0.0 \
  --push-to-hub turhancan97/vit-tiny-lora-food101
```

**Hyperparameters**

| Setting | Value |
|---|---|
| LoRA rank | 8 |
| LoRA alpha | 16 |
| LoRA dropout | 0.1 |
| Target modules | `query`, `value` |
| Optimizer | AdamW (HF `Trainer` default) |
| Learning rate | 5e-4 |
| Batch size | 64 |
| Epochs | 5 |
| Warmup ratio | 0.03 |
| Weight decay | 0.0 |
| Precision | FP16 |
| Augmentation | RandomResizedCrop(0.8–1.0), RandomHorizontalFlip |

**Trainable parameters:** ~93k of ~5.6M total (**~1.7%**).

## Evaluation

Evaluated on the Food-101 test split (25,250 images).

| Metric | Value |
|---|---|
| Top-1 accuracy | 85 % |
| Top-5 accuracy | 90 % |

## Usage

The adapter uses the standard PEFT format plus a sidecar `classifier.pt` and `labels.json`. Minimal loader:

```python
import json
import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from torch import nn
from transformers import AutoImageProcessor, AutoModelForImageClassification

BASE = "WinKawaks/vit-tiny-patch16-224"
ADAPTER = "turhancan97/vit-tiny-lora-food101"

processor = AutoImageProcessor.from_pretrained(BASE, use_fast=True)
base = AutoModelForImageClassification.from_pretrained(BASE)
model = PeftModel.from_pretrained(base, ADAPTER)

id2label = {int(k): v for k, v in json.loads(
    open(hf_hub_download(ADAPTER, "labels.json")).read()
).items()}

head_state = torch.load(
    hf_hub_download(ADAPTER, "classifier.pt"), map_location="cpu", weights_only=True
)
head = nn.Linear(base.config.hidden_size, len(id2label))
head.load_state_dict(head_state)
model.base_model.model.classifier = head

model.eval()
```

**Inference:**

```python
from PIL import Image

image = Image.open("my_food.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.inference_mode():
    logits = model(**inputs).logits[0]
topk = logits.softmax(-1).topk(5)
for score, idx in zip(topk.values, topk.indices):
    print(f"{id2label[idx.item()]:30s} {score.item():.3f}")
```

**Switching back to the base model** (ImageNet-1k, 1000 classes) without unloading:

```python
with model.disable_adapter():
    logits = base(**inputs).logits  # uses the pristine pretrained weights
```

## Intended use

- Educational / demo use for showing how LoRA adds new capabilities to a frozen backbone.
- Classifying photos of prepared food into the Food-101 taxonomy.

## Limitations

- Only 101 food categories; anything outside the taxonomy will be misclassified.
- Trained on Food-101 which is mostly western/restaurant-style dishes, with label noise in the original data.
- ViT-tiny is a low-capacity backbone; a larger base model would likely get higher accuracy with the same adapter recipe.

## License

Apache-2.0, matching the base model and the [Food-101 dataset license](https://data.vision.ee.ethz.ch/cvl/datasets_extra/food-101/).

## Citation

If you use this adapter, please cite the underlying works:

```bibtex
@inproceedings{hu2022lora,
  title={{LoRA}: Low-Rank Adaptation of Large Language Models},
  author={Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu},
  booktitle={ICLR},
  year={2022}
}

@inproceedings{bossard2014food101,
  title={Food-101 -- Mining Discriminative Components with Random Forests},
  author={Bossard, Lukas and Guillaumin, Matthieu and Van Gool, Luc},
  booktitle={ECCV},
  year={2014}
}
```