--- sdk: static tags: - icml2026-repro - paper-STcIzNrUBB orid: STcIzNrUBB title: "Calibrated Preference Learning: The Case of Label Ranking" emoji: 🔬 current_claim_count: 5 colorFrom: indigo colorTo: purple --- # Reproducibility logbook This logbook evaluates the five current anchored claims for the paper above. The evidence pages give a decisive verdict for every claim and state the exact mathematical, empirical, and dataset scope used. ## Official released inputs - Paper: [arXiv:2605.30447](https://arxiv.org/abs/2605.30447) - PDF SHA-256: `3bf53f877c75908058377c9c7ae5fb840c625e63f4112056c6375a8ccf7add48` - Source archive SHA-256: `8cba0290408c2108e45f5c5ba563305769e04c35bf93dd0332f221c8842f7f23` - Abstract HTML SHA-256: `0fc7355568c11097dd5fe522d2dbad666b6f22a73e852063484cc28c646ba4c7` - Official implementation commit: [`dc0b53324cfc98625517f6e7cdeb5c65a77b89a6`](https://github.com/Advueu963/Calibrated_Preference_Learning/tree/dc0b53324cfc98625517f6e7cdeb5c65a77b89a6) - Official RewardBench2 score-artifact listing: [allenai/reward-bench-2-results](https://huggingface.co/datasets/allenai/reward-bench-2-results/tree/main/eval-set-scores) The empirical pages use released CSV result artifacts from the official repository and released raw score JSON artifacts from the official RewardBench2 result dataset. No checkpoint or training artifact is needed for these checks. ## Authorship and evidence boundary The evidence was authored from the paper, its released source, and the official released artifacts listed above. No peer logbook, peer Space, peer repository, or peer sentence was used as evidence.