--- title: OpenCaption-4B-VL-SFT emoji: 📝 colorFrom: indigo colorTo: pink sdk: gradio sdk_version: 6.24.0 app_file: app.py short_description: Dense fine-grained image captioning with Qwen3-VL-4B python_version: "3.12" startup_duration_timeout: 30m license: apache-2.0 --- # OpenCaption-4B-VL-SFT-v1.0 Dense, fine-grained image captioning with **OpenCaption-4B-VL-SFT-v1.0** — a vision-language model fine-tuned from Qwen3-VL-4B-Instruct for rich, structured image descriptions. Upload an image and the model produces a detailed, structured caption with thematic sections covering subjects, backgrounds, lighting, atmosphere, and fine visual details. ## Model - **Base**: [Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) - **Fine-tune**: [prithivMLmods/OpenCaption-4B-VL-SFT-v1.0](https://huggingface.co/prithivMLmods/OpenCaption-4B-VL-SFT-v1.0) - **License**: Apache-2.0