# Mellum 2 Thinking — evaluation results # Self-reported by JetBrains. Post-training evaluations from Mellum 2 Technical Report # (Table 10, RL column). # Only entries for benchmarks confirmed to be registered as HF Hub Benchmarks are listed # (a Benchmark is a dataset repo with `eval.yaml` at its root). - dataset: id: Idavidrein/gpqa task_id: diamond value: 57.6 date: "2026-05-27" notes: "post-training eval (SFT + RLVR), with thinking, no-tools" - dataset: id: gorilla-llm/Berkeley-Function-Calling-Leaderboard task_id: bfclv3 value: 69.4 date: "2026-05-27" notes: "post-training eval (SFT + RLVR), with thinking, tools"