temp_file_1 / RESULTS-20260809-2007.md
simonycl's picture
seed-0 sweep results snapshot (mid-run)
3d6d975 verified
|
Raw
History Blame Contribute Delete
4.95 kB

MA-bench seed-0 sweep — results snapshot 2026-08-09 20:07 UTC

Worker gpt-5.6-luna throughout. Proposer arms: gpt-5.6-sol vs claude-opus-5. MID-RUN: these are each method's own selection metric, NOT the sealed test.

MetaHarness — dev (3-rep)

bench tier arm iters start best delta
charxiv medium opus5 81 0.7121 0.7879 +0.0758
charxiv medium sol 17 0.7172 0.7626 +0.0454
charxiv minimal opus5 81 0.7222 0.7929 +0.0707
charxiv minimal sol 23 0.7374 0.7626 +0.0252
charxiv strong opus5 17 0.7475 0.8333 +0.0858
charxiv strong sol 9 0.7424 0.7525 +0.0101
gpqa medium opus5 34 0.8209 0.8955 +0.0746
gpqa medium sol 27 0.8408 0.9055 +0.0647
gpqa minimal opus5 15 0.8259 0.8905 +0.0646
gpqa minimal sol 25 0.8358 0.8905 +0.0547
gpqa strong opus5 35 0.8806 0.9154 +0.0348
gpqa strong sol 23 0.8856 0.9204 +0.0348
tau2 medium opus5 21 0.6970 0.6970 +0.0000
tau2 medium sol 20 0.5576 0.6909 +0.1333
tau2 minimal opus5 23 0.5818 0.6848 +0.1030
tau2 minimal sol 20 0.5273 0.6667 +0.1394
tau2 strong opus5 23 0.5576 0.6667 +0.1091
tau2 strong sol 19 0.4667 0.6727 +0.2060

GEPA — valset dev

bench tier arm iters start best delta
charxiv medium opus5 125 0.7172 0.7778 +0.0606
charxiv medium sol 157 0.7222 0.7626 +0.0404
charxiv minimal opus5 177 0.7323 0.7879 +0.0556
charxiv minimal sol 218 0.7172 0.7475 +0.0303
charxiv strong opus5 149 0.7525 0.7929 +0.0404
charxiv strong sol 185 0.7626 0.7778 +0.0152
gpqa medium opus5 86 0.8408 0.8905 +0.0498
gpqa medium sol 88 0.8756 0.8905 +0.0149
gpqa minimal opus5 185 0.8358 0.8806 +0.0448
gpqa minimal sol 209 0.8408 0.8657 +0.0249
gpqa strong opus5 58 0.8756 0.9154 +0.0398
gpqa strong sol 53 0.9005 0.9055 +0.0050
tau2 medium opus5 28 0.6788 0.6788 +0.0000
tau2 medium sol 28 0.6364 0.6848 +0.0485
tau2 minimal opus5 28 0.5879 0.6606 +0.0727
tau2 minimal sol 27 0.5939 0.6788 +0.0848
tau2 strong opus5 29 0.5636 0.6788 +0.1152
tau2 strong sol 26 0.6364 0.6545 +0.0182

AdaEvolve — train fitness

bench tier arm iters start best delta
charxiv minimal opus5 250 0.8088 0.8824 +0.0735
charxiv minimal sol 250 0.7941 0.8529 +0.0588
gpqa minimal opus5 300 0.8923 0.9538 +0.0615
gpqa minimal sol 300 0.9231 0.9538 +0.0308
tau2 minimal opus5 74 0.0000 0.6481 +0.6481
tau2 minimal sol 82 0.0000 0.6667 +0.6667

Read before comparing

  1. Mid-run, and these are each method's OWN metric. MetaHarness = dev (3-rep); GEPA = its valset; AdaEvolve = train fitness. They are NOT mutually comparable, and none is the sealed test. Sealed test scores are in each cell's summary.json once it finishes.
  2. Improvement is measured against each cell's own starting point, not against the ROSTER baselines — those were measured on Bedrock Haiku, and the worker here is gpt-5.6-luna.
  3. tau2 deviates from the paper backbone: its user-simulator and nl-assertion judge are gpt-5.4-mini, not gpt-4.1, which the Copilot forwarder serves on an exhausted quota (every call 429s, verified at concurrency 1). Uniform across all tau2 cells.
  4. AdaEvolve's tau2 rows start at 0.0000 — the initial program scored nothing, so its +0.65 "delta" is a cold start, not an improvement of that size.
  5. mh × sol restarted 2026-08-07 after a codex bug (its sandbox could not execute any command in this pod, so the proposer read nothing and proposed nothing; eight cells had sealed reporting the SEED's score). Those cells have less elapsed search than the rest.
  6. Single seed. Finite-test SE is ~0.06 on these benchmarks; effects here are +0.02..+0.12. These rank the arms suggestively; they do not separate them.

Layout

  • RESULTS.md — this file
  • cells/<arm>-<bench>-<tier>-<method>/ — per cell: summary.json (config, provenance, sealed test when finished, spend, infra_failures), scores.jsonl (MH), candidates.json (GEPA), best/ (AdaEvolve best program + info)

Full trajectories and raw traces are NOT in this bundle (~51 GB); ask if you want them.