--- license: apache-2.0 language: - en base_model: GSAI-ML/LLaDA-8B-Base library_name: transformers pipeline_tag: text-generation tags: - diffusion-language-model - register-tokens - chunked-reasoning - dllm-registers - chunked-grpo --- # llada-8b-dllm-registers-longarith-rl-r4 Chunked GRPO on LongArithmetic starting from the mix60k R4 SFT main-triad checkpoint. Released at step 400, the step reported in tab:rl_results. This checkpoint is part of the *dLLM Registers* project — register tokens as a bounded, trained, continuous channel for carrying decoding state across denoising windows in diffusion language models. - **Paper section**: tab:rl_results (LongArithmetic, registers arm) - **Carry channel**: continuous register tokens (`channel_mode=registers`, `num_registers=4`, `tail_length=0`) - **Base model**: [GSAI-ML/LLaDA-8B-Base](https://huggingface.co/GSAI-ML/LLaDA-8B-Base) - **Training data**: LongArithmetic (5-8 operations, add/sub20, on-policy GRPO rollouts) - **Training config**: [`diffu-grpo/amlt/mix60k_r4_longarith_expr5678_state_bonete54.yaml`](https://github.com/lbertge/d1-registers/blob/main/diffu-grpo/amlt/mix60k_r4_longarith_expr5678_state_bonete54.yaml) - **Chunk size**: C = 128 tokens - **Prompt dropout rate** (Bernoulli per-trace CSG mask): 0.3 - **RL task**: LongArithmetic - **GRPO updates**: 400 - **Initialized from**: [`albertge/llada-8b-dllm-registers-mix60k-r4`](https://huggingface.co/albertge/llada-8b-dllm-registers-mix60k-r4) - **Date uploaded**: 2026-06-13 ## How to load ```python from transformers import AutoModel, AutoTokenizer model = AutoModel.from_pretrained("albertge/llada-8b-dllm-registers-longarith-rl-r4", trust_remote_code=True) tok = AutoTokenizer.from_pretrained("albertge/llada-8b-dllm-registers-longarith-rl-r4", trust_remote_code=True) ``` To use the carry channel correctly at inference, see the evaluator at [`eval/eval.py`](https://github.com/lbertge/d1-registers/blob/main/eval/eval.py) and the wrapper [`eval/run_mix60k_full_eval.sh`](https://github.com/lbertge/d1-registers/blob/main/eval/run_mix60k_full_eval.sh). Key flags for this checkpoint: `--num_registers 4 --channel_mode registers --tail_length 0` --use_mask_token_for_registers ## Repository Training and eval code: [https://github.com/lbertge/d1-registers](https://github.com/lbertge/d1-registers) ## Citation If you use this checkpoint, please cite the dLLM Registers paper: ```bibtex @misc{dllm-registers-2026, title = {Register Tokens for Unbounded Reasoning in Diffusion Language Models}, author = {Albert Ge and collaborators}, year = {2026}, note = {Preprint}, url = {https://github.com/lbertge/d1-registers} } ```