# MA-bench seed-0 sweep — results snapshot 2026-08-09 20:07 UTC Worker `gpt-5.6-luna` throughout. Proposer arms: `gpt-5.6-sol` vs `claude-opus-5`. MID-RUN: these are each method's own selection metric, NOT the sealed test. ## MetaHarness — dev (3-rep) | bench | tier | arm | iters | start | best | delta | |---|---|---|---|---|---|---| | charxiv | medium | opus5 | 81 | 0.7121 | 0.7879 | **+0.0758** | | charxiv | medium | sol | 17 | 0.7172 | 0.7626 | **+0.0454** | | charxiv | minimal | opus5 | 81 | 0.7222 | 0.7929 | **+0.0707** | | charxiv | minimal | sol | 23 | 0.7374 | 0.7626 | **+0.0252** | | charxiv | strong | opus5 | 17 | 0.7475 | 0.8333 | **+0.0858** | | charxiv | strong | sol | 9 | 0.7424 | 0.7525 | **+0.0101** | | gpqa | medium | opus5 | 34 | 0.8209 | 0.8955 | **+0.0746** | | gpqa | medium | sol | 27 | 0.8408 | 0.9055 | **+0.0647** | | gpqa | minimal | opus5 | 15 | 0.8259 | 0.8905 | **+0.0646** | | gpqa | minimal | sol | 25 | 0.8358 | 0.8905 | **+0.0547** | | gpqa | strong | opus5 | 35 | 0.8806 | 0.9154 | **+0.0348** | | gpqa | strong | sol | 23 | 0.8856 | 0.9204 | **+0.0348** | | tau2 | medium | opus5 | 21 | 0.6970 | 0.6970 | **+0.0000** | | tau2 | medium | sol | 20 | 0.5576 | 0.6909 | **+0.1333** | | tau2 | minimal | opus5 | 23 | 0.5818 | 0.6848 | **+0.1030** | | tau2 | minimal | sol | 20 | 0.5273 | 0.6667 | **+0.1394** | | tau2 | strong | opus5 | 23 | 0.5576 | 0.6667 | **+0.1091** | | tau2 | strong | sol | 19 | 0.4667 | 0.6727 | **+0.2060** | ## GEPA — valset dev | bench | tier | arm | iters | start | best | delta | |---|---|---|---|---|---|---| | charxiv | medium | opus5 | 125 | 0.7172 | 0.7778 | **+0.0606** | | charxiv | medium | sol | 157 | 0.7222 | 0.7626 | **+0.0404** | | charxiv | minimal | opus5 | 177 | 0.7323 | 0.7879 | **+0.0556** | | charxiv | minimal | sol | 218 | 0.7172 | 0.7475 | **+0.0303** | | charxiv | strong | opus5 | 149 | 0.7525 | 0.7929 | **+0.0404** | | charxiv | strong | sol | 185 | 0.7626 | 0.7778 | **+0.0152** | | gpqa | medium | opus5 | 86 | 0.8408 | 0.8905 | **+0.0498** | | gpqa | medium | sol | 88 | 0.8756 | 0.8905 | **+0.0149** | | gpqa | minimal | opus5 | 185 | 0.8358 | 0.8806 | **+0.0448** | | gpqa | minimal | sol | 209 | 0.8408 | 0.8657 | **+0.0249** | | gpqa | strong | opus5 | 58 | 0.8756 | 0.9154 | **+0.0398** | | gpqa | strong | sol | 53 | 0.9005 | 0.9055 | **+0.0050** | | tau2 | medium | opus5 | 28 | 0.6788 | 0.6788 | **+0.0000** | | tau2 | medium | sol | 28 | 0.6364 | 0.6848 | **+0.0485** | | tau2 | minimal | opus5 | 28 | 0.5879 | 0.6606 | **+0.0727** | | tau2 | minimal | sol | 27 | 0.5939 | 0.6788 | **+0.0848** | | tau2 | strong | opus5 | 29 | 0.5636 | 0.6788 | **+0.1152** | | tau2 | strong | sol | 26 | 0.6364 | 0.6545 | **+0.0182** | ## AdaEvolve — train fitness | bench | tier | arm | iters | start | best | delta | |---|---|---|---|---|---|---| | charxiv | minimal | opus5 | 250 | 0.8088 | 0.8824 | **+0.0735** | | charxiv | minimal | sol | 250 | 0.7941 | 0.8529 | **+0.0588** | | gpqa | minimal | opus5 | 300 | 0.8923 | 0.9538 | **+0.0615** | | gpqa | minimal | sol | 300 | 0.9231 | 0.9538 | **+0.0308** | | tau2 | minimal | opus5 | 74 | 0.0000 | 0.6481 | **+0.6481** | | tau2 | minimal | sol | 82 | 0.0000 | 0.6667 | **+0.6667** | --- ## Read before comparing 1. **Mid-run, and these are each method's OWN metric.** MetaHarness = dev (3-rep); GEPA = its valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed test. Sealed test scores are in each cell's `summary.json` once it finishes. 2. **Improvement is measured against each cell's own starting point**, not against the ROSTER baselines — those were measured on Bedrock Haiku, and the worker here is `gpt-5.6-luna`. 3. **tau2 deviates from the paper backbone**: its user-simulator and nl-assertion judge are `gpt-5.4-mini`, not `gpt-4.1`, which the Copilot forwarder serves on an exhausted quota (every call 429s, verified at concurrency 1). Uniform across all tau2 cells. 4. **AdaEvolve's tau2 rows start at 0.0000** — the initial program scored nothing, so its +0.65 "delta" is a cold start, not an improvement of that size. 5. **`mh × sol` restarted 2026-08-07** after a codex bug (its sandbox could not execute any command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed reporting the SEED's score). Those cells have less elapsed search than the rest. 6. **Single seed.** Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12. These rank the arms suggestively; they do not separate them. ## Layout - `RESULTS.md` — this file - `cells/---/` — per cell: `summary.json` (config, provenance, sealed test when finished, spend, infra_failures), `scores.jsonl` (MH), `candidates.json` (GEPA), `best/` (AdaEvolve best program + info) Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them.