AudioRubrics

The model from Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning: Qwen2.5-Omni-7B post-trained with GRPO using self-evolving, audio-grounded rubric rewards and an overthinking penalty.

This is the full merged checkpoint (thinker merged back into the complete Omni model) and can be served directly with vLLM:

vllm serve umd-zhou-lab/AudioRubrics --served-model-name omni --trust-remote-code \
  --max-model-len 8192 --limit-mm-per-prompt '{"audio":1}'

See the GitHub repository for training and evaluation instructions.

Downloads last month
-
Safetensors
Model size
12B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for umd-zhou-lab/AudioRubrics

Finetuned
(59)
this model

Collection including umd-zhou-lab/AudioRubrics