pavlichenko's picture
Update model card and add evaluation results (#2)
d16d8d8
Raw
History Blame Contribute Delete
658 Bytes
# Mellum 2 Thinking — evaluation results
# Self-reported by JetBrains. Post-training evaluations from Mellum 2 Technical Report
# (Table 10, RL column).
# Only entries for benchmarks confirmed to be registered as HF Hub Benchmarks are listed
# (a Benchmark is a dataset repo with `eval.yaml` at its root).
- dataset:
id: Idavidrein/gpqa
task_id: diamond
value: 57.6
date: "2026-05-27"
notes: "post-training eval (SFT + RLVR), with thinking, no-tools"
- dataset:
id: gorilla-llm/Berkeley-Function-Calling-Leaderboard
task_id: bfclv3
value: 69.4
date: "2026-05-27"
notes: "post-training eval (SFT + RLVR), with thinking, tools"