| # MA-bench seed-0 sweep β results snapshot 2026-08-09 20:07 UTC |
|
|
| Worker `gpt-5.6-luna` throughout. Proposer arms: `gpt-5.6-sol` vs `claude-opus-5`. |
| MID-RUN: these are each method's own selection metric, NOT the sealed test. |
|
|
|
|
| ## MetaHarness β dev (3-rep) |
|
|
| | bench | tier | arm | iters | start | best | delta | |
| |---|---|---|---|---|---|---| |
| | charxiv | medium | opus5 | 81 | 0.7121 | 0.7879 | **+0.0758** | |
| | charxiv | medium | sol | 17 | 0.7172 | 0.7626 | **+0.0454** | |
| | charxiv | minimal | opus5 | 81 | 0.7222 | 0.7929 | **+0.0707** | |
| | charxiv | minimal | sol | 23 | 0.7374 | 0.7626 | **+0.0252** | |
| | charxiv | strong | opus5 | 17 | 0.7475 | 0.8333 | **+0.0858** | |
| | charxiv | strong | sol | 9 | 0.7424 | 0.7525 | **+0.0101** | |
| | gpqa | medium | opus5 | 34 | 0.8209 | 0.8955 | **+0.0746** | |
| | gpqa | medium | sol | 27 | 0.8408 | 0.9055 | **+0.0647** | |
| | gpqa | minimal | opus5 | 15 | 0.8259 | 0.8905 | **+0.0646** | |
| | gpqa | minimal | sol | 25 | 0.8358 | 0.8905 | **+0.0547** | |
| | gpqa | strong | opus5 | 35 | 0.8806 | 0.9154 | **+0.0348** | |
| | gpqa | strong | sol | 23 | 0.8856 | 0.9204 | **+0.0348** | |
| | tau2 | medium | opus5 | 21 | 0.6970 | 0.6970 | **+0.0000** | |
| | tau2 | medium | sol | 20 | 0.5576 | 0.6909 | **+0.1333** | |
| | tau2 | minimal | opus5 | 23 | 0.5818 | 0.6848 | **+0.1030** | |
| | tau2 | minimal | sol | 20 | 0.5273 | 0.6667 | **+0.1394** | |
| | tau2 | strong | opus5 | 23 | 0.5576 | 0.6667 | **+0.1091** | |
| | tau2 | strong | sol | 19 | 0.4667 | 0.6727 | **+0.2060** | |
|
|
| ## GEPA β valset dev |
|
|
| | bench | tier | arm | iters | start | best | delta | |
| |---|---|---|---|---|---|---| |
| | charxiv | medium | opus5 | 125 | 0.7172 | 0.7778 | **+0.0606** | |
| | charxiv | medium | sol | 157 | 0.7222 | 0.7626 | **+0.0404** | |
| | charxiv | minimal | opus5 | 177 | 0.7323 | 0.7879 | **+0.0556** | |
| | charxiv | minimal | sol | 218 | 0.7172 | 0.7475 | **+0.0303** | |
| | charxiv | strong | opus5 | 149 | 0.7525 | 0.7929 | **+0.0404** | |
| | charxiv | strong | sol | 185 | 0.7626 | 0.7778 | **+0.0152** | |
| | gpqa | medium | opus5 | 86 | 0.8408 | 0.8905 | **+0.0498** | |
| | gpqa | medium | sol | 88 | 0.8756 | 0.8905 | **+0.0149** | |
| | gpqa | minimal | opus5 | 185 | 0.8358 | 0.8806 | **+0.0448** | |
| | gpqa | minimal | sol | 209 | 0.8408 | 0.8657 | **+0.0249** | |
| | gpqa | strong | opus5 | 58 | 0.8756 | 0.9154 | **+0.0398** | |
| | gpqa | strong | sol | 53 | 0.9005 | 0.9055 | **+0.0050** | |
| | tau2 | medium | opus5 | 28 | 0.6788 | 0.6788 | **+0.0000** | |
| | tau2 | medium | sol | 28 | 0.6364 | 0.6848 | **+0.0485** | |
| | tau2 | minimal | opus5 | 28 | 0.5879 | 0.6606 | **+0.0727** | |
| | tau2 | minimal | sol | 27 | 0.5939 | 0.6788 | **+0.0848** | |
| | tau2 | strong | opus5 | 29 | 0.5636 | 0.6788 | **+0.1152** | |
| | tau2 | strong | sol | 26 | 0.6364 | 0.6545 | **+0.0182** | |
|
|
| ## AdaEvolve β train fitness |
|
|
| | bench | tier | arm | iters | start | best | delta | |
| |---|---|---|---|---|---|---| |
| | charxiv | minimal | opus5 | 250 | 0.8088 | 0.8824 | **+0.0735** | |
| | charxiv | minimal | sol | 250 | 0.7941 | 0.8529 | **+0.0588** | |
| | gpqa | minimal | opus5 | 300 | 0.8923 | 0.9538 | **+0.0615** | |
| | gpqa | minimal | sol | 300 | 0.9231 | 0.9538 | **+0.0308** | |
| | tau2 | minimal | opus5 | 74 | 0.0000 | 0.6481 | **+0.6481** | |
| | tau2 | minimal | sol | 82 | 0.0000 | 0.6667 | **+0.6667** | |
|
|
| --- |
|
|
| ## Read before comparing |
|
|
| 1. **Mid-run, and these are each method's OWN metric.** MetaHarness = dev (3-rep); GEPA = its |
| valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed |
| test. Sealed test scores are in each cell's `summary.json` once it finishes. |
| 2. **Improvement is measured against each cell's own starting point**, not against the ROSTER |
| baselines β those were measured on Bedrock Haiku, and the worker here is `gpt-5.6-luna`. |
| 3. **tau2 deviates from the paper backbone**: its user-simulator and nl-assertion judge are |
| `gpt-5.4-mini`, not `gpt-4.1`, which the Copilot forwarder serves on an exhausted quota |
| (every call 429s, verified at concurrency 1). Uniform across all tau2 cells. |
| 4. **AdaEvolve's tau2 rows start at 0.0000** β the initial program scored nothing, so its |
| +0.65 "delta" is a cold start, not an improvement of that size. |
| 5. **`mh Γ sol` restarted 2026-08-07** after a codex bug (its sandbox could not execute any |
| command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed |
| reporting the SEED's score). Those cells have less elapsed search than the rest. |
| 6. **Single seed.** Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12. |
| These rank the arms suggestively; they do not separate them. |
|
|
| ## Layout |
|
|
| - `RESULTS.md` β this file |
| - `cells/<arm>-<bench>-<tier>-<method>/` β per cell: `summary.json` (config, provenance, sealed |
| test when finished, spend, infra_failures), `scores.jsonl` (MH), `candidates.json` (GEPA), |
| `best/` (AdaEvolve best program + info) |
| |
| Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them. |
| |