MA-bench seed-0 sweep — results snapshot 2026-08-09 20:07 UTC
Worker gpt-5.6-luna throughout. Proposer arms: gpt-5.6-sol vs claude-opus-5.
MID-RUN: these are each method's own selection metric, NOT the sealed test.
MetaHarness — dev (3-rep)
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 81 | 0.7121 | 0.7879 | +0.0758 |
| charxiv | medium | sol | 17 | 0.7172 | 0.7626 | +0.0454 |
| charxiv | minimal | opus5 | 81 | 0.7222 | 0.7929 | +0.0707 |
| charxiv | minimal | sol | 23 | 0.7374 | 0.7626 | +0.0252 |
| charxiv | strong | opus5 | 17 | 0.7475 | 0.8333 | +0.0858 |
| charxiv | strong | sol | 9 | 0.7424 | 0.7525 | +0.0101 |
| gpqa | medium | opus5 | 34 | 0.8209 | 0.8955 | +0.0746 |
| gpqa | medium | sol | 27 | 0.8408 | 0.9055 | +0.0647 |
| gpqa | minimal | opus5 | 15 | 0.8259 | 0.8905 | +0.0646 |
| gpqa | minimal | sol | 25 | 0.8358 | 0.8905 | +0.0547 |
| gpqa | strong | opus5 | 35 | 0.8806 | 0.9154 | +0.0348 |
| gpqa | strong | sol | 23 | 0.8856 | 0.9204 | +0.0348 |
| tau2 | medium | opus5 | 21 | 0.6970 | 0.6970 | +0.0000 |
| tau2 | medium | sol | 20 | 0.5576 | 0.6909 | +0.1333 |
| tau2 | minimal | opus5 | 23 | 0.5818 | 0.6848 | +0.1030 |
| tau2 | minimal | sol | 20 | 0.5273 | 0.6667 | +0.1394 |
| tau2 | strong | opus5 | 23 | 0.5576 | 0.6667 | +0.1091 |
| tau2 | strong | sol | 19 | 0.4667 | 0.6727 | +0.2060 |
GEPA — valset dev
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 125 | 0.7172 | 0.7778 | +0.0606 |
| charxiv | medium | sol | 157 | 0.7222 | 0.7626 | +0.0404 |
| charxiv | minimal | opus5 | 177 | 0.7323 | 0.7879 | +0.0556 |
| charxiv | minimal | sol | 218 | 0.7172 | 0.7475 | +0.0303 |
| charxiv | strong | opus5 | 149 | 0.7525 | 0.7929 | +0.0404 |
| charxiv | strong | sol | 185 | 0.7626 | 0.7778 | +0.0152 |
| gpqa | medium | opus5 | 86 | 0.8408 | 0.8905 | +0.0498 |
| gpqa | medium | sol | 88 | 0.8756 | 0.8905 | +0.0149 |
| gpqa | minimal | opus5 | 185 | 0.8358 | 0.8806 | +0.0448 |
| gpqa | minimal | sol | 209 | 0.8408 | 0.8657 | +0.0249 |
| gpqa | strong | opus5 | 58 | 0.8756 | 0.9154 | +0.0398 |
| gpqa | strong | sol | 53 | 0.9005 | 0.9055 | +0.0050 |
| tau2 | medium | opus5 | 28 | 0.6788 | 0.6788 | +0.0000 |
| tau2 | medium | sol | 28 | 0.6364 | 0.6848 | +0.0485 |
| tau2 | minimal | opus5 | 28 | 0.5879 | 0.6606 | +0.0727 |
| tau2 | minimal | sol | 27 | 0.5939 | 0.6788 | +0.0848 |
| tau2 | strong | opus5 | 29 | 0.5636 | 0.6788 | +0.1152 |
| tau2 | strong | sol | 26 | 0.6364 | 0.6545 | +0.0182 |
AdaEvolve — train fitness
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | minimal | opus5 | 250 | 0.8088 | 0.8824 | +0.0735 |
| charxiv | minimal | sol | 250 | 0.7941 | 0.8529 | +0.0588 |
| gpqa | minimal | opus5 | 300 | 0.8923 | 0.9538 | +0.0615 |
| gpqa | minimal | sol | 300 | 0.9231 | 0.9538 | +0.0308 |
| tau2 | minimal | opus5 | 74 | 0.0000 | 0.6481 | +0.6481 |
| tau2 | minimal | sol | 82 | 0.0000 | 0.6667 | +0.6667 |
Read before comparing
- Mid-run, and these are each method's OWN metric. MetaHarness = dev (3-rep); GEPA = its
valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed
test. Sealed test scores are in each cell's
summary.jsononce it finishes. - Improvement is measured against each cell's own starting point, not against the ROSTER
baselines — those were measured on Bedrock Haiku, and the worker here is
gpt-5.6-luna. - tau2 deviates from the paper backbone: its user-simulator and nl-assertion judge are
gpt-5.4-mini, notgpt-4.1, which the Copilot forwarder serves on an exhausted quota (every call 429s, verified at concurrency 1). Uniform across all tau2 cells. - AdaEvolve's tau2 rows start at 0.0000 — the initial program scored nothing, so its +0.65 "delta" is a cold start, not an improvement of that size.
mh × solrestarted 2026-08-07 after a codex bug (its sandbox could not execute any command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed reporting the SEED's score). Those cells have less elapsed search than the rest.- Single seed. Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12. These rank the arms suggestively; they do not separate them.
Layout
RESULTS.md— this filecells/<arm>-<bench>-<tier>-<method>/— per cell:summary.json(config, provenance, sealed test when finished, spend, infra_failures),scores.jsonl(MH),candidates.json(GEPA),best/(AdaEvolve best program + info)
Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them.