Buckets:
| # Mellum 2 Thinking — evaluation results | |
| # Self-reported by JetBrains. Post-training evaluations from Mellum 2 Technical Report | |
| # (Table 10, RL column). | |
| # Only entries for benchmarks confirmed to be registered as HF Hub Benchmarks are listed | |
| # (a Benchmark is a dataset repo with `eval.yaml` at its root). | |
| - dataset: | |
| id: Idavidrein/gpqa | |
| task_id: diamond | |
| value: 57.6 | |
| date: "2026-05-27" | |
| notes: "post-training eval (SFT + RLVR), with thinking, no-tools" | |
| - dataset: | |
| id: gorilla-llm/Berkeley-Function-Calling-Leaderboard | |
| task_id: bfclv3 | |
| value: 69.4 | |
| date: "2026-05-27" | |
| notes: "post-training eval (SFT + RLVR), with thinking, tools" | |
Xet Storage Details
- Size:
- 658 Bytes
- Xet hash:
- 186e6d2c2edfe21a19ece433352ecb753c69b501a98f28ef6b7e047bf85b3b8d
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.