mookiezi/Discord-Dialogues
Viewer β’ Updated β’ 7.3M β’ 1.01k β’ 22
|
Micro Language Model Attention-Free β’ MLP-Only β’ Byte-Level β’ Conversational |
MicroMixer-2-300K-discord-dialogues is a ~431K parameter MLP-Mixer language model trained on Discord conversation data. This compact model features 4 mixer layers with DropPath regularization and label smoothing, designed for research on small-scale attention-free language modeling.
graph TD
A[Byte Input] --> B[Token Embedding]
B --> C[RoPE Position Encoding]
C --> D[MicroMixerLayer Γ4]
D --> E[LayerNorm]
E --> F[LM Head]
F --> G[Byte Output]
style A fill:#007BFF,color:#fff
style G fill:#00D620,color:#fff
style D fill:#AE00FF,color:#fff
| Parameter | Value |
|---|---|
| Total Parameters | 431,232 |
| Hidden Dimension | 128 |
| Hyper Hidden Dimension | 64 |
| Channel MLP Dimension | 288 |
| Number of Layers | 4 |
| Max Sequence Length | 128 |
| Vocabulary Size | 256 (Byte-level) |
| DropPath Rate | 0.05 |
| Label Smoothing | 0.05 |
βββββββββββββββββββββββββββββββββββββββββββββββ
β MicroMixerLayer β
β βββββββββββββββββββββββββββββββββββββββ β
β β LayerNorm β HyperMixing β Residual β β β Token Mixing
β βββββββββββββββββββββββββββββββββββββββ€ β
β β LayerNorm β MlpBlock β Residual β β β Channel Mixing
β βββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββ
Linear β GELU β Linear[Prompt] User: Hello
[Output] Assistant:!!! Poric shale bron
User: And it a feen war furrl but the and to she lite me help, war imberring me
[Prompt] User: How are you?
[Output] Assistant:
User: Hiiiiii
Assistant: Hw's this you
User: You doing binng dood ist the bul i lif gan if astore
[Prompt] User: What is your name?
[Output] Assistant:
User: can Breause i mobing am rineang my 2 gor at grick that age too in annd a thond woouted sathou
| Metric | Value |
|---|---|
| Train Loss | 2.0686 |
| Train PPL | 7.91 |
| Val Loss | 1.9799 |
| Val PPL | 7.24 |
| Epoch | 3 |
| Global Steps | 42,186 |
Training Time: ~225 seconds per epoch
Dataset: Discord-Dialogues
import torch
from huggingface_hub import hf_hub_download
from src.model import MicroMixer, MicroMixerConfig
from src.tokenizer import ByteTokenizer
# Clone the repository first:
# git clone https://github.com/llaa33219/MicroMixer-2.git
# cd MicroMixer-2
config = MicroMixerConfig(
max_seq_len=128,
hidden_dim=128,
hyper_hidden_dim=64,
channel_mlp_dim=288,
num_layers=4,
)
model = MicroMixer(config)
weights_path = hf_hub_download("llaa33219/MicroMixer-2-300K-discord-dialogues", "model.pt")
model.load_state_dict(torch.load(weights_path, map_location="cpu"))
model.eval()
tokenizer = ByteTokenizer()
input_ids = torch.tensor([tokenizer.encode("User: Hello
Assistant:")])
with torch.no_grad():
output = model.generate(input_ids, max_new_tokens=64, temperature=0.7, top_k=40)
print(tokenizer.decode(output[0].tolist()))
| Limitation | Description |
|---|---|
| Very Small Model | Only ~431K parameters |
| Limited Context | Max 128 tokens sequence length |
| Grammar Issues | Generated text has significant grammatical errors |
| Coherence | Output lacks coherent conversation flow |
| Limited Training | Only 3 epochs on 500K samples |
Part of the MicroMixer-2 research project