Video-Text-to-Text
Transformers
Safetensors
English
qwen3_5
image-text-to-text
video
multimodal
video-captioning
temporal-grounding
qwen
text-generation
VLM
gptq
4-bit precision
quantized
custom_code
Instructions to use prasannaJagadesh/marlin-2B-GPTQ-4BITS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prasannaJagadesh/marlin-2B-GPTQ-4BITS with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("prasannaJagadesh/marlin-2B-GPTQ-4BITS", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("prasannaJagadesh/marlin-2B-GPTQ-4BITS", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| Marlin 2B | |
| Copyright (c) 2026 NemoStation | |
| This product includes weights derived from Qwen3.5-2B | |
| (https://huggingface.co/Qwen/Qwen3.5-2B), Copyright (c) 2025 Alibaba Cloud, | |
| used under the Apache License, Version 2.0. | |
| Modifications by NemoStation include: integration of a video-capable | |
| visual tower, custom training data curation (~400K clip-level annotations | |
| with Gemini-3-Flash teacher distillation), two-stage SFT + SimPO | |
| post-training, and custom modeling code (modeling_marlin.py) exposing | |
| the .caption() and .find() inference modes. | |
| Marlin 2B is distributed under the Apache License, Version 2.0 | |
| (see LICENSE). | |