RoadTones-VL-CoT

This model accompanies the paper RoadTones: Tone Controllable Text Generation from Road Event Videos

Model Summary

RoadTones-VL-CoT is an open-source large multimodal model with superior tone-controlled road video captioning capabilities. Built on the foundation of Qwen3-VL-8B-Instruct, it has been finetuned on RoadTones-51k dataset with the Chain-of-Thought (CoT) intermediate drafts as well for better interpretability. Evaluated on the RoadTones-Eval Metrics, Its performance is on par with the best benchmarked model (Gemini-2.5-pro) while displaying superior Tone Adherance, thereby demonstrating the RoadTones-51K dataset's capability in improving the tone-controllability for text generation in VideoLLMs.

For further details, please refer to the following resources:

Citation

@misc{parikh2026roadtonestonecontrollabletext,
      title={RoadTones: Tone Controllable Text Generation from Road Event Videos}, 
      author={Chirag Parikh and Siddhi Pravin Lipare and Ravi Kiran Sarvadevabhatla},
      year={2026},
      eprint={2605.21411},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2605.21411}, 
}
Downloads last month
3
Safetensors
Model size
770k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for siddhi-lipare/RoadTones-VL-CoT

Finetuned
(2)
this model

Dataset used to train siddhi-lipare/RoadTones-VL-CoT

Paper for siddhi-lipare/RoadTones-VL-CoT