File size: 4,954 Bytes
3d6d975
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
# MA-bench seed-0 sweep — results snapshot 2026-08-09 20:07 UTC

Worker `gpt-5.6-luna` throughout. Proposer arms: `gpt-5.6-sol` vs `claude-opus-5`.
MID-RUN: these are each method's own selection metric, NOT the sealed test.


## MetaHarness — dev (3-rep)

| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 81 | 0.7121 | 0.7879 | **+0.0758** |
| charxiv | medium | sol | 17 | 0.7172 | 0.7626 | **+0.0454** |
| charxiv | minimal | opus5 | 81 | 0.7222 | 0.7929 | **+0.0707** |
| charxiv | minimal | sol | 23 | 0.7374 | 0.7626 | **+0.0252** |
| charxiv | strong | opus5 | 17 | 0.7475 | 0.8333 | **+0.0858** |
| charxiv | strong | sol | 9 | 0.7424 | 0.7525 | **+0.0101** |
| gpqa | medium | opus5 | 34 | 0.8209 | 0.8955 | **+0.0746** |
| gpqa | medium | sol | 27 | 0.8408 | 0.9055 | **+0.0647** |
| gpqa | minimal | opus5 | 15 | 0.8259 | 0.8905 | **+0.0646** |
| gpqa | minimal | sol | 25 | 0.8358 | 0.8905 | **+0.0547** |
| gpqa | strong | opus5 | 35 | 0.8806 | 0.9154 | **+0.0348** |
| gpqa | strong | sol | 23 | 0.8856 | 0.9204 | **+0.0348** |
| tau2 | medium | opus5 | 21 | 0.6970 | 0.6970 | **+0.0000** |
| tau2 | medium | sol | 20 | 0.5576 | 0.6909 | **+0.1333** |
| tau2 | minimal | opus5 | 23 | 0.5818 | 0.6848 | **+0.1030** |
| tau2 | minimal | sol | 20 | 0.5273 | 0.6667 | **+0.1394** |
| tau2 | strong | opus5 | 23 | 0.5576 | 0.6667 | **+0.1091** |
| tau2 | strong | sol | 19 | 0.4667 | 0.6727 | **+0.2060** |

## GEPA — valset dev

| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | medium | opus5 | 125 | 0.7172 | 0.7778 | **+0.0606** |
| charxiv | medium | sol | 157 | 0.7222 | 0.7626 | **+0.0404** |
| charxiv | minimal | opus5 | 177 | 0.7323 | 0.7879 | **+0.0556** |
| charxiv | minimal | sol | 218 | 0.7172 | 0.7475 | **+0.0303** |
| charxiv | strong | opus5 | 149 | 0.7525 | 0.7929 | **+0.0404** |
| charxiv | strong | sol | 185 | 0.7626 | 0.7778 | **+0.0152** |
| gpqa | medium | opus5 | 86 | 0.8408 | 0.8905 | **+0.0498** |
| gpqa | medium | sol | 88 | 0.8756 | 0.8905 | **+0.0149** |
| gpqa | minimal | opus5 | 185 | 0.8358 | 0.8806 | **+0.0448** |
| gpqa | minimal | sol | 209 | 0.8408 | 0.8657 | **+0.0249** |
| gpqa | strong | opus5 | 58 | 0.8756 | 0.9154 | **+0.0398** |
| gpqa | strong | sol | 53 | 0.9005 | 0.9055 | **+0.0050** |
| tau2 | medium | opus5 | 28 | 0.6788 | 0.6788 | **+0.0000** |
| tau2 | medium | sol | 28 | 0.6364 | 0.6848 | **+0.0485** |
| tau2 | minimal | opus5 | 28 | 0.5879 | 0.6606 | **+0.0727** |
| tau2 | minimal | sol | 27 | 0.5939 | 0.6788 | **+0.0848** |
| tau2 | strong | opus5 | 29 | 0.5636 | 0.6788 | **+0.1152** |
| tau2 | strong | sol | 26 | 0.6364 | 0.6545 | **+0.0182** |

## AdaEvolve — train fitness

| bench | tier | arm | iters | start | best | delta |
|---|---|---|---|---|---|---|
| charxiv | minimal | opus5 | 250 | 0.8088 | 0.8824 | **+0.0735** |
| charxiv | minimal | sol | 250 | 0.7941 | 0.8529 | **+0.0588** |
| gpqa | minimal | opus5 | 300 | 0.8923 | 0.9538 | **+0.0615** |
| gpqa | minimal | sol | 300 | 0.9231 | 0.9538 | **+0.0308** |
| tau2 | minimal | opus5 | 74 | 0.0000 | 0.6481 | **+0.6481** |
| tau2 | minimal | sol | 82 | 0.0000 | 0.6667 | **+0.6667** |

---

## Read before comparing

1. **Mid-run, and these are each method's OWN metric.** MetaHarness = dev (3-rep); GEPA = its
   valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed
   test. Sealed test scores are in each cell's `summary.json` once it finishes.
2. **Improvement is measured against each cell's own starting point**, not against the ROSTER
   baselines — those were measured on Bedrock Haiku, and the worker here is `gpt-5.6-luna`.
3. **tau2 deviates from the paper backbone**: its user-simulator and nl-assertion judge are
   `gpt-5.4-mini`, not `gpt-4.1`, which the Copilot forwarder serves on an exhausted quota
   (every call 429s, verified at concurrency 1). Uniform across all tau2 cells.
4. **AdaEvolve's tau2 rows start at 0.0000** — the initial program scored nothing, so its
   +0.65 "delta" is a cold start, not an improvement of that size.
5. **`mh × sol` restarted 2026-08-07** after a codex bug (its sandbox could not execute any
   command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed
   reporting the SEED's score). Those cells have less elapsed search than the rest.
6. **Single seed.** Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12.
   These rank the arms suggestively; they do not separate them.

## Layout

- `RESULTS.md` — this file
- `cells/<arm>-<bench>-<tier>-<method>/` — per cell: `summary.json` (config, provenance, sealed
  test when finished, spend, infra_failures), `scores.jsonl` (MH), `candidates.json` (GEPA),
  `best/` (AdaEvolve best program + info)

Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them.