temp_file_1 / RESULTS-20260809-2007.md
simonycl's picture
seed-0 sweep results snapshot (mid-run)
3d6d975 verified
|
Raw
History Blame Contribute Delete
4.95 kB
# MA-bench seed-0 sweep β€” results snapshot 2026-08-09 20:07 UTC
Worker `gpt-5.6-luna` throughout. Proposer arms: `gpt-5.6-sol` vs `claude-opus-5`.
MID-RUN: these are each method's own selection metric, NOT the sealed test.
## MetaHarness β€” dev (3-rep)
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 81 | 0.7121 | 0.7879 | **+0.0758** |
| charxiv | medium | sol | 17 | 0.7172 | 0.7626 | **+0.0454** |
| charxiv | minimal | opus5 | 81 | 0.7222 | 0.7929 | **+0.0707** |
| charxiv | minimal | sol | 23 | 0.7374 | 0.7626 | **+0.0252** |
| charxiv | strong | opus5 | 17 | 0.7475 | 0.8333 | **+0.0858** |
| charxiv | strong | sol | 9 | 0.7424 | 0.7525 | **+0.0101** |
| gpqa | medium | opus5 | 34 | 0.8209 | 0.8955 | **+0.0746** |
| gpqa | medium | sol | 27 | 0.8408 | 0.9055 | **+0.0647** |
| gpqa | minimal | opus5 | 15 | 0.8259 | 0.8905 | **+0.0646** |
| gpqa | minimal | sol | 25 | 0.8358 | 0.8905 | **+0.0547** |
| gpqa | strong | opus5 | 35 | 0.8806 | 0.9154 | **+0.0348** |
| gpqa | strong | sol | 23 | 0.8856 | 0.9204 | **+0.0348** |
| tau2 | medium | opus5 | 21 | 0.6970 | 0.6970 | **+0.0000** |
| tau2 | medium | sol | 20 | 0.5576 | 0.6909 | **+0.1333** |
| tau2 | minimal | opus5 | 23 | 0.5818 | 0.6848 | **+0.1030** |
| tau2 | minimal | sol | 20 | 0.5273 | 0.6667 | **+0.1394** |
| tau2 | strong | opus5 | 23 | 0.5576 | 0.6667 | **+0.1091** |
| tau2 | strong | sol | 19 | 0.4667 | 0.6727 | **+0.2060** |
## GEPA β€” valset dev
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 125 | 0.7172 | 0.7778 | **+0.0606** |
| charxiv | medium | sol | 157 | 0.7222 | 0.7626 | **+0.0404** |
| charxiv | minimal | opus5 | 177 | 0.7323 | 0.7879 | **+0.0556** |
| charxiv | minimal | sol | 218 | 0.7172 | 0.7475 | **+0.0303** |
| charxiv | strong | opus5 | 149 | 0.7525 | 0.7929 | **+0.0404** |
| charxiv | strong | sol | 185 | 0.7626 | 0.7778 | **+0.0152** |
| gpqa | medium | opus5 | 86 | 0.8408 | 0.8905 | **+0.0498** |
| gpqa | medium | sol | 88 | 0.8756 | 0.8905 | **+0.0149** |
| gpqa | minimal | opus5 | 185 | 0.8358 | 0.8806 | **+0.0448** |
| gpqa | minimal | sol | 209 | 0.8408 | 0.8657 | **+0.0249** |
| gpqa | strong | opus5 | 58 | 0.8756 | 0.9154 | **+0.0398** |
| gpqa | strong | sol | 53 | 0.9005 | 0.9055 | **+0.0050** |
| tau2 | medium | opus5 | 28 | 0.6788 | 0.6788 | **+0.0000** |
| tau2 | medium | sol | 28 | 0.6364 | 0.6848 | **+0.0485** |
| tau2 | minimal | opus5 | 28 | 0.5879 | 0.6606 | **+0.0727** |
| tau2 | minimal | sol | 27 | 0.5939 | 0.6788 | **+0.0848** |
| tau2 | strong | opus5 | 29 | 0.5636 | 0.6788 | **+0.1152** |
| tau2 | strong | sol | 26 | 0.6364 | 0.6545 | **+0.0182** |
## AdaEvolve β€” train fitness
| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | minimal | opus5 | 250 | 0.8088 | 0.8824 | **+0.0735** |
| charxiv | minimal | sol | 250 | 0.7941 | 0.8529 | **+0.0588** |
| gpqa | minimal | opus5 | 300 | 0.8923 | 0.9538 | **+0.0615** |
| gpqa | minimal | sol | 300 | 0.9231 | 0.9538 | **+0.0308** |
| tau2 | minimal | opus5 | 74 | 0.0000 | 0.6481 | **+0.6481** |
| tau2 | minimal | sol | 82 | 0.0000 | 0.6667 | **+0.6667** |
---
## Read before comparing
1. **Mid-run, and these are each method's OWN metric.** MetaHarness = dev (3-rep); GEPA = its
valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed
test. Sealed test scores are in each cell's `summary.json` once it finishes.
2. **Improvement is measured against each cell's own starting point**, not against the ROSTER
baselines β€” those were measured on Bedrock Haiku, and the worker here is `gpt-5.6-luna`.
3. **tau2 deviates from the paper backbone**: its user-simulator and nl-assertion judge are
`gpt-5.4-mini`, not `gpt-4.1`, which the Copilot forwarder serves on an exhausted quota
(every call 429s, verified at concurrency 1). Uniform across all tau2 cells.
4. **AdaEvolve's tau2 rows start at 0.0000** β€” the initial program scored nothing, so its
+0.65 "delta" is a cold start, not an improvement of that size.
5. **`mh Γ— sol` restarted 2026-08-07** after a codex bug (its sandbox could not execute any
command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed
reporting the SEED's score). Those cells have less elapsed search than the rest.
6. **Single seed.** Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12.
These rank the arms suggestively; they do not separate them.
## Layout
- `RESULTS.md` β€” this file
- `cells/<arm>-<bench>-<tier>-<method>/` β€” per cell: `summary.json` (config, provenance, sealed
test when finished, spend, infra_failures), `scores.jsonl` (MH), `candidates.json` (GEPA),
`best/` (AdaEvolve best program + info)
Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them.