Add evaluation results

#12
by SaylorTwift HF Staff - opened

PR Description: Add Evaluation Results for zai-org/GLM-5.3-Flash

Summary

This PR adds evaluation results extracted from the model card's benchmark graph for zai-org/GLM-5.3-Flash to the .eval_results/ directory, following the Hugging Face Hub evaluation-results specification.

Benchmarks Added

Benchmarks Skipped (Not Registered on Hub)

The following benchmarks were present in the model card but could not be added because they do not have a registered eval.yaml on the Hugging Face Hub:

  • Agent's Last Exam: 26.3 — no registered eval.yaml found on the Hub.
  • AutomationBench v1.0.6: 48.8 — no registered eval.yaml found on the Hub.
  • GDPval-AA v2: 1773 — no registered eval.yaml found on the Hub.

These can be added once the benchmark authors register their eval.yaml on the Hub.

Source

  • Model card: https://huggingface.co/zai-org/GLM-5.3-Flash
  • Note: the model card cites arxiv:2602.15763 as "the GLM-5 Technical report", but that paper (published Feb 2026) is the original GLM-5 report and predates this model — it does not contain GLM-5.3-Flash-specific numbers, so it wasn't used as a source.

Files Added

  • .eval_results/GLM-5.3-Flash.yaml

Verification

These results were extracted from the model card's published benchmark chart (bench_53.png, an image embedded in the README, read visually — the announcement blog is a client-rendered SPA and wasn't fetchable). No verified token is provided as these were not run via HF Jobs with inspect-ai.


To upload this file to the Hub, run:

hf upload zai-org/GLM-5.3-Flash --type model --include .eval_results/*.yaml --commit-message "Add evaluation results"
Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment