Fourier-VLM: Compressing Vision Tokens in the Frequency Domain for Large Vision-Language Models
Paper โข 2508.06038 โข Published
Official checkpoints for Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models.
| Model | Base Model | Visual Tokens | Compression | Weights |
|---|---|---|---|---|
| Fourier-LLaVA-v1.5-7B-256 | LLaVA-v1.5-7B | 256 | 55.6% | ๐ค HF |
| Fourier-LLaVA-v1.5-7B-144 | LLaVA-v1.5-7B | 144 | 75.0% | ๐ค HF |
| Fourier-LLaVA-v1.5-7B-64 | LLaVA-v1.5-7B | 64 | 88.9% | ๐ค HF |
| Fourier-LLaVA-v1.5-7B-36 | LLaVA-v1.5-7B | 36 | 93.8% | ๐ค HF |
| Fourier-LLaVA-v1.5-13B-144 | LLaVA-v1.5-13B | 144 | 75.0% | ๐ค HF |
| Fourier-Qwen2-VL-2B-0.67 | Qwen2-VL-2B-Instruct | Dynamic | 55.6% | ๐ค HF |
| Fourier-Qwen2.5-VL-3B-0.67 | Qwen2.5-VL-3B-Instruct | Dynamic | 55.6% | ๐ค HF |
Base model
Qwen/Qwen2.5-VL-3B-Instruct