fastvlm-7b-mlx-q4 / README.md
dbaek111's picture
Update benchmark tables with 512px and 1024px comparison
babd32c verified
|
Raw
History Blame
3.99 kB
---
license: apple-amlr
library_name: mlx
pipeline_tag: image-text-to-text
base_model: apple/FastVLM-7B
base_model_relation: quantized
tags:
- fastvlm
- mlx
- mlx-vlm
- safetensors
- transformers
- llava_qwen2
- image-text-to-text
- text-generation
- conversational
- custom_code
- vision-language
- multimodal
- apple-silicon
- quantized
- 7B
- q4 / 4-bit quantized
- 4-bit
- 4-bit precision
- q4
---
# fastvlm-7b-mlx-q4
This repository contains an MLX-converted FastVLM checkpoint.
## Model
- Base model: `apple/FastVLM-7B`
- Parameters: `7B`
- Precision: `q4 / 4-bit quantized`
- Approx. folder size: `4.9G`
The checkpoint was converted from Apple FastVLM using the official FastVLM model export workflow and patched `mlx-vlm`.
## Files
This repository should include:
- `config.json`
- MLX model weights
- tokenizer files
- `fastvithd.mlpackage` vision tower
## Example Usage
```bash
hf download dbaek111/fastvlm-7b-mlx-q4 --local-dir ./fastvlm-7b-mlx-q4
python -m mlx_vlm.generate \
--model ./fastvlm-7b-mlx-q4 \
--image /path/to/your/image.jpg \
--prompt "Explain the image." \
--max-tokens 64 \
--temp 0.0
```
## Benchmark
Benchmarked on an Apple Silicon Mac with the patched FastVLM `mlx-vlm` workflow.
- Task: pedestrian wayfinding captioning
- Images: three local test images resized to 512px and 1024px long edge
- Prompt: `Describe what is visible for pedestrian wayfinding in one short sentence. Do not list categories. Do not mention anything you cannot see. Keep under 30 words.`
- Max tokens: `64`
- Temperature: `0.0`
- Timing: model loaded once per image set, then three images processed sequentially
### Model Selection
| Model | Size | Precision | Avg 512px | Avg 1024px | Load | Recommended use |
|---|---:|---|---:|---:|---:|---|
| fastvlm-0.5b-mlx-q4 | 819M | q4 | 0.338s | 0.370s | 2.82s | Smallest and fastest; rough real-time captions |
| fastvlm-0.5b-mlx-q8 | 1.1G | q8 | 0.419s | 0.414s | 2.57s | Fast, with richer captions than 0.5B q4 |
| fastvlm-0.5b-mlx-fp16 | 1.6G | fp16 | 0.435s | 0.421s | 2.79s | Small FP16 baseline |
| fastvlm-1.5b-mlx-q4 | 1.4G | q4 | 0.447s | 0.464s | 2.65s | Best real-time balance for pedestrian wayfinding |
| fastvlm-1.5b-mlx-q8 | 2.2G | q8 | 0.552s | 0.541s | 2.68s | More detail while staying sub-second |
| fastvlm-1.5b-mlx-fp16 | 3.8G | fp16 | 0.636s | 0.557s | 3.08s | 1.5B FP16 reference variant |
| **fastvlm-7b-mlx-q4** | **4.9G** | **q4** | **1.263s** | **1.241s** | **3.04s** | **Best quality/latency tradeoff among 7B variants** |
| fastvlm-7b-mlx-q8 | 8.0G | q8 | 1.497s | 1.495s | 3.85s | Higher precision 7B, slower than q4 |
| fastvlm-7b-mlx-fp16 | 15G | fp16 | 1.834s | 1.874s | 45.48s | Full precision reference; expensive to load |
### Per-Image Timing
Each cell is `img1 / img2 / img3` inference time in seconds.
| Model | 512px images | 1024px images |
|---|---:|---:|
| fastvlm-0.5b-mlx-q4 | 0.318 / 0.353 / 0.343 | 0.374 / 0.370 / 0.366 |
| fastvlm-0.5b-mlx-q8 | 0.393 / 0.465 / 0.400 | 0.424 / 0.387 / 0.431 |
| fastvlm-0.5b-mlx-fp16 | 0.394 / 0.490 / 0.421 | 0.430 / 0.395 / 0.437 |
| fastvlm-1.5b-mlx-q4 | 0.454 / 0.458 / 0.430 | 0.465 / 0.472 / 0.456 |
| fastvlm-1.5b-mlx-q8 | 0.591 / 0.611 / 0.453 | 0.594 / 0.552 / 0.477 |
| fastvlm-1.5b-mlx-fp16 | 0.701 / 0.709 / 0.496 | 0.520 / 0.635 / 0.516 |
| **fastvlm-7b-mlx-q4** | **1.161 / 1.330 / 1.297** | **1.081 / 1.343 / 1.298** |
| fastvlm-7b-mlx-q8 | 1.364 / 1.639 / 1.487 | 1.341 / 1.712 / 1.432 |
| fastvlm-7b-mlx-fp16 | 1.561 / 2.041 / 1.902 | 1.595 / 2.277 / 1.749 |
## Compatibility
This is an MLX export of FastVLM for Apple Silicon Macs. It includes the CoreML FastViTHD vision tower as `fastvithd.mlpackage`.
This repository is not a standard PyTorch Transformers checkpoint and is not intended for vLLM, SGLang, or Linux GPU inference.
## Notes
This is a converted and quantized derivative of Apple FastVLM.
Please refer to the original Apple FastVLM repository and model card for license and usage conditions.