File size: 3,994 Bytes
89353f5
 
86c9baf
 
89353f5
86c9baf
89353f5
 
 
 
86c9baf
 
 
 
 
 
 
 
 
89353f5
86c9baf
89353f5
 
86c9baf
 
 
89353f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
babd32c
89353f5
babd32c
 
0f9b52d
89353f5
 
babd32c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0f9b52d
 
 
 
 
 
89353f5
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
---
license: apple-amlr
library_name: mlx
pipeline_tag: image-text-to-text
base_model: apple/FastVLM-7B
base_model_relation: quantized
tags:
- fastvlm
- mlx
- mlx-vlm
- safetensors
- transformers
- llava_qwen2
- image-text-to-text
- text-generation
- conversational
- custom_code
- vision-language
- multimodal
- apple-silicon
- quantized
- 7B
- q4 / 4-bit quantized
- 4-bit
- 4-bit precision
- q4
---

# fastvlm-7b-mlx-q4

This repository contains an MLX-converted FastVLM checkpoint.

## Model

- Base model: `apple/FastVLM-7B`
- Parameters: `7B`
- Precision: `q4 / 4-bit quantized`
- Approx. folder size: `4.9G`

The checkpoint was converted from Apple FastVLM using the official FastVLM model export workflow and patched `mlx-vlm`.

## Files

This repository should include:

- `config.json`
- MLX model weights
- tokenizer files
- `fastvithd.mlpackage` vision tower

## Example Usage

```bash
hf download dbaek111/fastvlm-7b-mlx-q4 --local-dir ./fastvlm-7b-mlx-q4

python -m mlx_vlm.generate \
  --model ./fastvlm-7b-mlx-q4 \
  --image /path/to/your/image.jpg \
  --prompt "Explain the image." \
  --max-tokens 64 \
  --temp 0.0
```

## Benchmark

Benchmarked on an Apple Silicon Mac with the patched FastVLM `mlx-vlm` workflow.

- Task: pedestrian wayfinding captioning
- Images: three local test images resized to 512px and 1024px long edge
- Prompt: `Describe what is visible for pedestrian wayfinding in one short sentence. Do not list categories. Do not mention anything you cannot see. Keep under 30 words.`
- Max tokens: `64`
- Temperature: `0.0`
- Timing: model loaded once per image set, then three images processed sequentially

### Model Selection

| Model | Size | Precision | Avg 512px | Avg 1024px | Load | Recommended use |
|---|---:|---|---:|---:|---:|---|
| fastvlm-0.5b-mlx-q4 | 819M | q4 | 0.338s | 0.370s | 2.82s | Smallest and fastest; rough real-time captions |
| fastvlm-0.5b-mlx-q8 | 1.1G | q8 | 0.419s | 0.414s | 2.57s | Fast, with richer captions than 0.5B q4 |
| fastvlm-0.5b-mlx-fp16 | 1.6G | fp16 | 0.435s | 0.421s | 2.79s | Small FP16 baseline |
| fastvlm-1.5b-mlx-q4 | 1.4G | q4 | 0.447s | 0.464s | 2.65s | Best real-time balance for pedestrian wayfinding |
| fastvlm-1.5b-mlx-q8 | 2.2G | q8 | 0.552s | 0.541s | 2.68s | More detail while staying sub-second |
| fastvlm-1.5b-mlx-fp16 | 3.8G | fp16 | 0.636s | 0.557s | 3.08s | 1.5B FP16 reference variant |
| **fastvlm-7b-mlx-q4** | **4.9G** | **q4** | **1.263s** | **1.241s** | **3.04s** | **Best quality/latency tradeoff among 7B variants** |
| fastvlm-7b-mlx-q8 | 8.0G | q8 | 1.497s | 1.495s | 3.85s | Higher precision 7B, slower than q4 |
| fastvlm-7b-mlx-fp16 | 15G | fp16 | 1.834s | 1.874s | 45.48s | Full precision reference; expensive to load |

### Per-Image Timing

Each cell is `img1 / img2 / img3` inference time in seconds.

| Model | 512px images | 1024px images |
|---|---:|---:|
| fastvlm-0.5b-mlx-q4 | 0.318 / 0.353 / 0.343 | 0.374 / 0.370 / 0.366 |
| fastvlm-0.5b-mlx-q8 | 0.393 / 0.465 / 0.400 | 0.424 / 0.387 / 0.431 |
| fastvlm-0.5b-mlx-fp16 | 0.394 / 0.490 / 0.421 | 0.430 / 0.395 / 0.437 |
| fastvlm-1.5b-mlx-q4 | 0.454 / 0.458 / 0.430 | 0.465 / 0.472 / 0.456 |
| fastvlm-1.5b-mlx-q8 | 0.591 / 0.611 / 0.453 | 0.594 / 0.552 / 0.477 |
| fastvlm-1.5b-mlx-fp16 | 0.701 / 0.709 / 0.496 | 0.520 / 0.635 / 0.516 |
| **fastvlm-7b-mlx-q4** | **1.161 / 1.330 / 1.297** | **1.081 / 1.343 / 1.298** |
| fastvlm-7b-mlx-q8 | 1.364 / 1.639 / 1.487 | 1.341 / 1.712 / 1.432 |
| fastvlm-7b-mlx-fp16 | 1.561 / 2.041 / 1.902 | 1.595 / 2.277 / 1.749 |

## Compatibility

This is an MLX export of FastVLM for Apple Silicon Macs. It includes the CoreML FastViTHD vision tower as `fastvithd.mlpackage`.

This repository is not a standard PyTorch Transformers checkpoint and is not intended for vLLM, SGLang, or Linux GPU inference.

## Notes

This is a converted and quantized derivative of Apple FastVLM.

Please refer to the original Apple FastVLM repository and model card for license and usage conditions.