# Evaluation Results Scores mirror the Macaron-V1-Venti model card benchmark table. Higher is better. | Benchmark | Macaron V1 | GLM 5.2 | GPT 5.5 | Claude Opus 4.8 | Gemini 3.1 Pro | Qwen 3.7 Max | Minimax M3 | |---|---:|---:|---:|---:|---:|---:|---:| | ChatBench | 58.3 | 54.5 | 55.5 | 52.8 | 52.0 | 52.5 | 49.1 | | LivingBench | 64.0 | 60.5 | 61.9 | 63.8 | 52.1 | 56.1 | 57.1 | | VitaBench | 60.0 | 55.8 | 55.8 | 56.5 | 55.2 | 61.2 | 56.8 | | PinchBench | 94.0 | 88.1 | 89.0 | 91.8 | 82.9 | 93.4 | 86.1 | | ClawGym | 77.7 | 74.6 | 82.5 | 80.5 | 77.5 | 75.7 | 76.2 | | SWE Verified | 85.6 | 80.4 | 82.9 | 88.6 | 80.6 | 80.4 | 80.5 | | TerminalBench 2.1 | 87.6 | 82.7 | 83.4 | 78.9 | 70.7 | 73.5 | 66.0 | | DeepSWE | 58.4 | 54.9 | 70.0 | 58.0 | 10.0 | 18.0 | 20.0 | | SWE Atlas QnA | 49.5 | 48.9 | 45.4 | 57.3 | 13.5 | 22.6 | 37.9 | | UI4ABench | 87.8 | 67.1 | 72.1 | 75.9 | 60.3 | 62.5 | 63.0 | Files: - `swe_bench_verified.yaml`: Hugging Face ingestible SWE-bench Verified result. The full model-card benchmark table is available in machine-readable form at [`../evaluation/benchmark_summary.yaml`](../evaluation/benchmark_summary.yaml).