Text Generation
PyTorch
English
FockMultiXiPARFLM
research
semsimula
conservative-language-model
scalar-potential
lagrangian-mechanics
energy-based-model
physics-informed
fock-space
parflm
parf
multi-channel-xi
structured-vtheta
analytical-gradient
interpretable
gaussian-mixture-potential
virtual-registers
controlled-conservativity
sparse-routing
tinystories
non-transformer
attention-free
constant-memory-inference
riemannian-geometry
riemannian-geodesics
Eval Results (legacy)
File size: 5,927 Bytes
7b6330e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | # Fock-PARFLM v2.1 (SQ3 structured V_theta) — Causal Leak Fix vs. Leaky Baseline (`A2`, seed 0)
Comparison of two runs of the same arm, `A2` (SQ3: mixture of `K_mix=8` diagonal
quadratic wells, `tau=1.0`, `K_xi=4`, `lambda_V=0.01`), differing only in
whether the cross-layer Fock register state carries the causal leak identified
in [`Fock-PARFLM_Causal_Leak_Audit_Results.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companion_notes/Fock-PARFLM_Causal_Leak_Audit_Results.md):
- **leaky** — legacy register lifecycle (`prefix_causal_registers=False`
behavior; the field didn't exist yet in this run). Registers are one bank
per sequence, evolved across the full window, so token *t*'s force in
later layers can be influenced by tokens after *t*.
- **fixed** (`checkpoint/ckpt_best.pt` in this repo) — `prefix_causal_registers=True`.
Registers are one bank per *position*; the bank consulted at position *t*
is built exclusively from tokens 1…t.
Same architecture in both: `d=256, L=8, V_theta=SQ3(K_mix=8, tau=1.0), K_xi=4,
lambda_V=0.01, M=16, d_k=64, top_k=8`, same seed (0), same 16,000 steps,
TinyStories (5M train tokens). This is the structured-V_theta counterpart of
the MLP-V_theta [Fock-PARFLM v2.1 leak fix](https://huggingface.co/dimitarpg13/semsimula-fock-parflm/blob/main/causal_leak_fix_verification.md);
the base architecture, register mechanism, and fix are identical — only the
scalar potential parameterisation differs.
## Headline: the fix costs a modest, honest tax on standard PPL
| | leaky | fixed | Δ |
|---|---|---|---|
| Best val PPL (step 14,400) | **10.36** | **10.90** | +0.54 (+5.2%) |
| Final val PPL (step 16,000) | 10.68 | 11.36 | +0.68 |
The leaky number (10.36) is the one originally published on this model card.
The fixed run is consistently worse across training — the same qualitative
signature as the MLP variant's leak-closure tax (+0.40 PPL there), just
somewhat larger here (+0.54 PPL). Consistent with that pattern, the SQ3
mixture potential appears to lean on the leaked future-token signal slightly
more than the MLP potential does at this scale, though both effects are small
compared to the ~33× (+3.51 nats) inflation measured on the much larger
OpenWebText-scale variant.
## The proof the fix actually worked, at the trained scale
**Future-perturbation test, run on the live trained model** at steps 8,000 and
16,000: perturbing tokens after the position of interest changed past logits
by exactly **0.0** — bit-exact, not "small":
```json
"probe_max_dlogit_past": 0.0,
"probe_mean_dnll_past_nats": 0.0,
"probe_gate_zero_control": 0.0
```
**Honest PPL vs. standard PPL**, same live model:
| step | standard PPL (mid-window) | honest PPL (last-position, leak-free) | paired diff | SE |
|---|---|---|---|---|
| 8,000 | 16.06 | 15.51 | −0.0346 nats | 0.0252 |
| 16,000 | 13.44 | 14.00 | +0.0408 nats | 0.0210 |
Both differences are within ~1.4–2.0 standard errors of zero — not
statistically distinguishable from no leak, the same qualitative result as
the MLP variant (where the two checks landed at <1 SE). The architectural
causal probe (`event: causal_probe`) also reports the deterministic,
no-noise check `eval_max_delta = 0.0` bit-exact at all four checkpoints
(steps 4k/8k/12k/16k) — the `passed: False` flag in the log is a known
overly-strict artifact of the noisy `.train()`-mode branch (Gumbel-softmax
routing jitter, `train_max_delta` ~1e-5–4e-5), not a residual leak; see the
MLP variant's verification note for the full explanation of why this flag
can be ignored in favor of `eval_max_delta`.
## A parameter-count correction found along the way
This model card previously reported **17,407,980 total parameters** (copied
from the MLP-V_theta sibling model's card). Reconstructing the exact SQ3
architecture and loading `ckpt_best.pt` with **zero missing/unexpected keys**
gives the true count:
| | Reported (was wrong) | Actual |
|---|---|---|
| Total parameters | 17,407,980 | **18,194,420** |
| \\(V_\theta\\) parameters | not previously broken out | **4,207,625** |
The SQ3 mixture's per-position hypernetwork (`mu_proj`, `a_proj`, `pi_proj`,
`b_proj`, each mapping the `K_xi·d = 1024`-dim xi-context to `K_mix·d = 2048`-dim
outputs) is larger than the 3-layer MLP it replaces, despite producing a much
flatter landscape (see below) — this model has *more* parameters devoted to
\\(V_\theta\\) than the MLP baseline (4.2M vs 3.4M) even though the resulting
potential is nearly flat.
## V_theta landscape statistics, recomputed on the fixed checkpoint
| Metric | leaky (published) | fixed (this checkpoint) |
|---|---|---|
| Mean | 0.008 | 0.017 |
| Std | 0.26 | 0.42 |
| Min | −5.49 | −1.23 |
| Max | 13.61 | 14.78 |
| Range | 19.10 | 16.01 |
Both checkpoints show the same qualitative picture — a near-flat \\(V_\theta\\)
landscape (mean close to 0, most of the force budget carried by \\(V_\phi\\)
and the Fock registers) — the fixed run is an independently-trained model
with different weights, so exact values shift, but the "flattest landscape
among the three architectures" conclusion is unaffected.
## Data sources
- leaky rerun: `semsimula_fock_multixi_structured_vtheta_buggy2/A2/seed0/`
- fixed (this repo's checkpoint): `semsimula_fock_multixi_structured_vtheta/A2/seed0/`
- Landscape stats recomputed via a fresh forward pass of the reconstructed
SQ3 architecture over TinyStories validation batches, loading
`ckpt_best.pt` with `strict=True` semantics (0 missing / 0 unexpected keys).
- Root-cause audit: [`Fock-PARFLM_Causal_Leak_Audit_Results.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companion_notes/Fock-PARFLM_Causal_Leak_Audit_Results.md)
- MLP-V_theta sibling verification: [`semsimula-fock-parflm/causal_leak_fix_verification.md`](https://huggingface.co/dimitarpg13/semsimula-fock-parflm/blob/main/causal_leak_fix_verification.md)
|