File size: 8,213 Bytes
e740e03 c5d516b e740e03 c5d516b e740e03 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 | ---
license: apache-2.0
language:
- en
library_name: t3-reference
tags:
- transformer
- interpretability
- geometric-algebra
- clifford-algebra
- adaptive-computation
pipeline_tag: text-generation
base_model: gpt2
datasets:
- HuggingFaceFW/fineweb-edu
- mlfoundations/dclm-baseline-1.0
- HuggingFaceTB/smollm-corpus
model-index:
- name: t3-124m-v36
results:
- task: { type: text-generation, name: WikiText-103 perplexity }
dataset: { type: wikitext, name: WikiText-103 }
metrics: [{ type: perplexity, value: 27.76 }]
- task: { type: multiple-choice, name: BoolQ }
dataset: { type: boolq, name: BoolQ }
metrics: [{ type: accuracy, value: 0.6046 }]
- task: { type: multiple-choice, name: ARC-Easy }
dataset: { type: arc, name: ARC-Easy }
metrics: [{ type: accuracy, value: 0.4331 }]
- task: { type: multiple-choice, name: ARC-Challenge }
dataset: { type: arc, name: ARC-Challenge }
metrics: [{ type: accuracy, value: 0.2176 }]
- task: { type: multiple-choice, name: PIQA }
dataset: { type: piqa, name: PIQA }
metrics: [{ type: accuracy, value: 0.6050 }]
- task: { type: multiple-choice, name: HellaSwag }
dataset: { type: hellaswag, name: HellaSwag }
metrics: [{ type: accuracy, value: 0.3040 }]
- task: { type: multiple-choice, name: WinoGrande }
dataset: { type: winogrande, name: WinoGrande }
metrics: [{ type: accuracy, value: 0.5043 }]
- task: { type: multiple-choice, name: COPA }
dataset: { type: copa, name: COPA }
metrics: [{ type: accuracy, value: 0.6000 }]
- task: { type: multiple-choice, name: RTE }
dataset: { type: rte, name: RTE }
metrics: [{ type: accuracy, value: 0.5235 }]
---
# T³ 124M v3.6 (run-3 release)
Inference-ready checkpoint for **T³**, a Clifford-algebra-augmented
transformer architecture. 124M parameters, GPT-2 Small substrate, 5B
training tokens.
This is the canonical reference checkpoint for the v3.6 lineage. Companion
artifacts:
- **Code:** [`mirrorethic/t3-reference`](https://github.com/mirrorethic/t3-reference) (Apache-2.0)
- **Trace library + benchmarks:** <https://t3atlas.dev>
- **Sibling checkpoint:** [`mirrorethic/t3-124m-v36-pcloss`](https://huggingface.co/mirrorethic/t3-124m-v36-pcloss) — same
architecture, trained with the inter-stage predictive-coding loss un-detached.
Slightly worse PPL (28.53 vs 27.76), neutral on reasoning. Use the pair for
the controlled inter-stage-PC ablation.
## Quick start
```bash
pip install t3-reference
```
```python
from huggingface_hub import hf_hub_download
from t3 import T3Model
ckpt = hf_hub_download("mirrorethic/t3-124m-v36", "pytorch_model.bin")
model = T3Model.from_checkpoint(ckpt)
model.eval()
import torch
input_ids = torch.randint(0, 50257, (1, 16))
with torch.no_grad():
logits, *_ = model(input_ids)
```
To generate a schema-v1 ecology trace:
```python
from t3.tracing import generate_trace
generate_trace(model, "The capital of France is",
prompt_id="factual", n_tokens=32,
out_path="trace.jsonl")
```
To re-run the published lm-eval-harness benchmarks:
```python
from t3.benchmarks import run_benchmark_suite
results = run_benchmark_suite("path/to/pytorch_model.bin")
```
## Architecture
T³ extends standard multi-head attention with a per-head **ecology** of
six conjugate primitives `(E, I, F, V, C, K)` coupled through bivector
composition in `Cl(3,3)` geometric algebra. Heads interact through a
learned blockade-and-cosurvival graph and ponder adaptively per stage via
output-entropy halt. Full technical specification:
[`docs/ARCHITECTURE.md`](https://github.com/mirrorethic/t3-reference/blob/main/docs/ARCHITECTURE.md).
| Field | Value |
|---|---|
| Parameters | 124,500,000 |
| Stages | 3, with `layers_per_stage = [4, 3, 5]` (12 transformer blocks total) |
| `d_model` | 768 |
| `n_heads` | 12 |
| `d_ff` | 3072 |
| `vocab_size` | 50257 (GPT-2 tokenizer) |
| `max_seq_len` | 1024 |
| Substrate | GPT-2 Small initialization |
| Training data | 5B tokens (FineWeb-Edu 40%, DCLM 20%, StackEdu 10%, FineMath 10%, Cosmopedia 10%, Wikipedia 10%) |
| Cumulative training step | 138,000 (135.5K substrate + 2,500 v3.6 increment) |
| Hamiltonian coupling ω | 0.02 |
| Trivectors | off (the trivectors-on variant is a planned v3.7 follow-up release) |
| Inter-stage predictive coding | on (`weight = 0.05`) |
| Scratchpad heads | on (`scratchpad_inject_entropy = (0.0, 0.0, 0.03)` — S2-only) |
| ACT | output-entropy halt + per-stage 4-step cap |
## Evaluation
All numbers are full lm-eval-harness 0.4.x runs (no subset). Reproduce with
`examples/run_benchmarks.py` from the reference repo.
| Task | Metric | Value | stderr |
|---|---|---:|---:|
| WikiText-103 (val) | perplexity | **27.76** | — |
| BoolQ | acc | **0.6046** | 0.0086 |
| ARC-Easy | acc | 0.4331 | 0.0102 |
| ARC-Challenge | acc | 0.2176 | 0.0121 |
| PIQA | acc | 0.6050 | 0.0114 |
| HellaSwag | acc | 0.3040 | 0.0046 |
| WinoGrande | acc | 0.5043 | 0.0141 |
| COPA | acc | 0.6000 | 0.0492 |
| RTE | acc | 0.5235 | 0.0301 |
For comparison panels (parameter-efficiency vs vanilla GPT-2 same-data,
compute frontier), see <https://t3atlas.dev/benchmarks/>.
### vs vanilla GPT-2 (same 5B-token training data)
T³-124M-v36 vs `gpt2/124M_vanilla_5b_same_data` baseline. The interesting
delta is on multi-step reasoning — T³ pondering allocates extra forward
compute per token (S1 averages 2.7–3.7 ponder loops on these tasks).
## Intended use
- **Research and interpretability.** The checkpoint is designed to be
inspected via the trace library, not deployed for production text
generation. The model is small (124M), English-only, and not
instruction-tuned.
- **Architectural comparison.** A reference point for novel-attention work
(Mamba, RWKV, xLSTM, etc.) that's matched to vanilla GPT-2 on training
data and parameter count.
- **Ecology / dynamics analysis.** The trace JSONL records per-head, per-stage
ecology state across forward passes — useful for studying how
Clifford-algebra-coupled state evolves during inference.
## Limitations
- 124M parameters: too small to be a useful generative chat model.
- English only.
- No instruction tuning, no RLHF, no safety tuning.
- Trained without grade-3 trivector terms (the static bivector Ω is the
full Cl(3,3) rotation in this checkpoint). The state-dependent variant
is a planned v3.7 Medium-scale follow-up.
- The [`v36-pcloss` sibling](https://huggingface.co/mirrorethic/t3-124m-v36-pcloss)
is the same architecture trained with the inter-stage predictive-coding
loss un-detached. Slightly worse PPL (28.53), neutral on reasoning — the
K-predictor learning a real cross-stage map (r=0.59) doesn't translate
into downstream gains at this scale.
## Capabilities probe
The checkpoint declares the following dynamics in its `config` and
state dict (consumed by `t3atlas` viewer for trace-rendering):
```json
{
"has_coupling": true,
"has_trivectors": false,
"has_dyn_omega": false,
"has_inter_stage_pc": true,
"has_scratchpad": true,
"n_primitives": 6,
"null_cone_strength": 0.02,
"hamiltonian_coupling": 0.02,
"sigma_hidden": 16,
"scratchpad_inject_entropy": [0.0, 0.0, 0.03]
}
```
## Citation
```bibtex
@misc{sutherland2026t3,
author = {Sutherland, Garret},
title = {T³: A Clifford-Algebra-Augmented Transformer Architecture},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/mirrorethic/t3-124m-v36}
}
```
## License
Apache-2.0. Both code (`mirrorethic/t3-reference`) and weights (this
repository).
## Contact
Garret Sutherland (MirrorEthic LLC) — `gsutherland@mirrorethic.com`.
---
*Released 2026-05-03. The `pytorch_model.bin` here is a stripped
inference-ready copy (498 MB) of the canonical `best.pt` from the v3.6
training campaign (run-3, step 2500, val PPL 27.76 on WikiText-103). The
optimizer state and data-loader state were dropped; everything T3Model
needs at inference is preserved (model_state, ecology_state, config, and
provenance metadata).*
|