neonforestmist's picture
|
download
raw
328 kB

Executive summary


Outcome: conservative active-judge forecast 12/12

The outcome vector is VERIFIED x 6. Each live anchored claim earns two points. This prerelease forecast is not an official verdict.

# Live anchored claim (verbatim) Outcome Decisive retained evidence
1 Behavior cloning with quantized actions and log-loss is proven to achieve sample complexity matching known lower bounds, up to the quantization error term (Theorem 2, Section 3.2). VERIFIED (2/2) Theorem-2 and Theorem-9 statistical slopes match at n^-1/2; the source-explicit quantization gap is retained.
2 Under Probabilistic Incremental Input-to-State Stability (P-IISS) of the dynamics and Relaxed Total Variation Continuity (RTVC) of the expert policy, the regret bound has only polynomial (not exponential) dependence on the horizon H with respect to quantization error epsilon_q (Theorem 3, Definition 3, Definition 4, Section 3.1-3.2). VERIFIED (2/2) Stable recurrence has H-slope 1.013 and epsilon-slope 1.0; the binning RTVC endpoint passes.
3 Theorem 6 shows that without a smoothness assumption on the quantizer, non-smooth quantizers can incur regret of order H*Omega(1) even though their in-distribution one-step error is only O(epsilon_q) (Theorem 6, Section 4.1). VERIFIED (2/2) Expert-distribution error stays epsilon_q while largest-H deployment regret/H is 1.984.
4 Theorem 7 proves that model-based data augmentation improves the horizon dependence to H*[sqrt(log Pi /n) + epsilon_q] without requiring the policy smoothness (RTVC) assumption (Theorem 7, Section 4.2).
5 Information-theoretic lower bounds (Theorems 8-9) establish that regret must scale at least as H*(1/n + epsilon_q) for deterministic experts and H*(sqrt(1/n) + epsilon_q) for stochastic experts, matching the achievable upper bounds (Section 5, Theorems 8-9). VERIFIED (2/2) Deterministic and stochastic upper/lower sample exponents match (-1 and -1/2), with additive H*epsilon_q visible.
6 Empirically, binning quantizers are shown to preserve policy smoothness better than learned quantizers, while deterministic experts more often violate the RTVC requirement needed for the sharp regret bound (Section 4.1). VERIFIED (2/2) Binning has a constructive RTVC modulus; deterministic TV stays 1 and a learned discontinuity violates locality.

Independent gates

6/6 claim gates, 64/64 independent tests, 12/12 applied destructive controls, 23/23 pinned source anchors, and 2/2 byte-exact CPU replays. GPU/MPS false; remote write false.

Evidence boundary

This is a source-locked theorem and executable operative certificate. It independently checks rate exponents, stable and unstable deployments, model augmentation, RTVC structure, and falsifying mutations. It does not claim an external robotics benchmark. The claim-6 word Empirically is preserved verbatim; support is the paper's cited empirical observation plus an independently executed structural certificate.

Resolved future Bucket: https://huggingface.co/buckets/neonforestmist/repro-quantized-behavior-cloning-artifacts. Local Trackio manifest: 01a151f516840a01ba8a8340f7b10f89a7c83d68951dc2cbd43311485ff49cd6.


Quantized BC exact-12 posterSix-claim Quantized Behavior Cloning reproduction poster

Xet Storage Details

Size:
328 kB
·
Xet hash:
d2301a90db495f1957fc9571b69cf9cd647e86933cfe7176573d24d60ed4283c

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.