# Executive summary --- Claims 2 and 3 are **verified at mechanism level**: the official DSAC-AID implementation disentangles epistemic vs aleatoric variance and applies LCB/UCB bounds exactly as described. Claim 1 is **partially verified at scaled scope**: DSAC-AID trains stably on Gym Ant-v4 and reaches ~980 return at 5% of paper training length on a single T4 GPU (~18 min). Full SOTA across all MuJoCo/DMC tasks at 1.5M steps was not attempted due to compute cost. ## Scope & cost | | This reproduction | Full replication | |---|---|---| | Scope | Ant-v4; DSAC-AID mechanism + scaled training; DSACT baseline in progress | All Gym-MuJoCo + DMC tasks, 1.5M iters, multi-seed | | Hardware | 1× T4 (HF Job) + local CPU | GPU cluster per paper | | Compute time | ~18 min GPU + ~15 min local CPU | days–weeks | | Cost | ~$0.20 GPU | hundreds of USD | | Outcome | mechanism reproduced; Ant performance plausible at scale | SOTA tables not fully replicated | --- Figure source: `repro_aid/poster/poster_embed.html` ````html
Epistemic/aleatoric disentanglement verified; LCB suppresses impulse, UCB drives exploration.
Scaled DSAC-AID vs DSACT on Gym-MuJoCo Ant (HF GPU job + local CPU).