Instructions to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged") model = AutoModelForCausalLM.from_pretrained("NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
Use Docker
docker model run hf.co/NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with Ollama:
ollama run hf.co/NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
- Unsloth Studio
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged to start chatting
- Docker Model Runner
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with Docker Model Runner:
docker model run hf.co/NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
- Lemonade
How to use NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Q4_K_M
Run and chat with the model
lemonade run user.ilo-toki-1.3-MiLMMT-46-1b-merged-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:# Run inference directly in the terminal:
llama cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:# Run inference directly in the terminal:
./llama-cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:# Run inference directly in the terminal:
./build/bin/llama-cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:Use Docker
docker model run hf.co/NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:ilo toki 1.3 — MiLMMT-46 1B
A translator between Toki Pona and English, Russian and Vietnamese. Small enough to run on a phone: it powers ilo toki, which does all of its translation on device.
This repository holds both the merged weights and GGUF builds, so there is one place to look rather than a repository per format.
Version 1.3 replaces
ilo-toki-1.1-MiLMMT-46-1b-merged.
See what changed and, before relying on it,
known limitations — several of them are inherited rather than
new.
Prompt format
The model keeps the prompt format of its base, and there is no chat template — do not wrap the input in one.
Translate this from Toki Pona to English:
Toki Pona: jan li moku e kili
English:
The translation follows the final <target language>: line and ends at the model's
end-of-generation token. Either side can be the source:
Translate this from Russian to Toki Pona:
Russian: Я тебя люблю.
Toki Pona:
Language names are written out in full — Toki Pona, English, Russian,
Vietnamese. Getting the format wrong does not fail loudly: the model keeps
producing fluent text while silently ignoring the requested target language.
Toki Pona is written in lower case; capitalization in the input is not something the model expects. Terminal punctuation is optional and barely changes the answer.
Which file to use
| File | Size | Notes |
|---|---|---|
ilo-toki-1.3-MiLMMT-46-1b-Q4_K_M.gguf |
0.94 GB | Smallest. |
ilo-toki-1.3-MiLMMT-46-1b-Q5_K_M.gguf |
1.00 GB | |
ilo-toki-1.3-MiLMMT-46-1b-Q6_K.gguf |
1.24 GB | |
ilo-toki-1.3-MiLMMT-46-1b-Q8_0.gguf |
1.29 GB | What the app ships — see below. |
model.safetensors |
2.48 GB | Merged weights, bf16, for transformers. |
The quantizations sit unusually close together because the 262k-token embedding matrix is about a third of the model and quantizes the same way in all of them. Q8_0 therefore costs only 0.05 GB more than Q6_K and 0.35 GB more than Q4_K_M, which is why the app ships it: on a phone the difference between these files is small, while the difference between fitting in RAM and not is enormous.
Running it
With llama.cpp:
llama-completion -m ilo-toki-1.3-MiLMMT-46-1b-Q8_0.gguf --temp 0 --top-k 1 \
-p "Translate this from Toki Pona to English:
Toki Pona: jan li moku e kili
English:"
Greedy decoding is what this model is meant to be run with. There is one right answer per input, and sampling only ever walks away from it.
With transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "Translate this from Toki Pona to English:\nToki Pona: jan li moku e kili\nEnglish:"
inputs = tokenizer(prompt, return_tensors="pt")
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=64, do_sample=False)[0]))
How it was built
A LoRA adapter (NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b)
trained with TRL SFT — rank 64, targeting the attention and MLP projections, plus
15 761 individual embedding rows through PEFT's trainable_token_indices — merged
into MiLMMT-46-1B-v0.1
and quantized with llama.cpp. The checkpoint is the one at 15 000 steps, taken
before the validation loss turns rather than at the end of training.
The base is a 46-language translation model from Xiaomi Research, so the fine-tune starts from a model that already translates rather than from a general purpose one.
Training data
| Dataset | What it contributes |
|---|---|
tokipona-mined-pairs |
Mined parallel sentences. |
tokipona-proper-names-mt |
Proper names, which Toki Pona transliterates rather than borrows. |
tokipona-wiki-titles-mt |
Wikipedia titles. |
tokipona-wiki-parallel-mt |
Parallel Wikipedia text. |
lipu-sewi |
lipu sewi. |
tatoeba-tokipona |
Tatoeba sentence pairs. |
Alongside tok↔x pairs the mix includes x↔y pairs between the natural
languages, at half the rate of 1.1, meant to keep their generation fluent without
crowding out the pairs where Toki Pona is one side.
What changed in 1.3
Measured against 1.0 and 1.1 over 103 prompts in both directions across the three languages, Q8_0 against Q8_0, greedy throughout.
toki ponais no longer answered about as another language. 1.1 turnedmi sona e toki ponainto «I know Russian» andsina sona ala sona e toki ponainto «Do you know Russian?». Both are right again.- Invented specifics are mostly gone. 1.1 rendered
jan li moku e kiliinto Vietnamese as «people in the state of Oregon eat delicious food», and put a sleeping animal «in the kitten room». 1.3 says what the sentence says. - Terminal punctuation moves the answer far less. Of fourteen sentences tried bare and with their final mark, 1.1 changed its answer on five and 1.3 on one.
- Only
.is dropped from the source during training now. Dropping?and!had made a declarative source map to an interrogative target, which is label noise rather than augmentation.
Known limitations
alais sometimes reversed. Two of ten negation probes come back meaning the opposite:jan li lape alagives «someone is sleeping»,mi pilin ike la mi moku alagives «when I feel bad then I eat too much». 1.1 gets the same two wrong, so this is inherited rather than new — but a negation that reads fluently and means the opposite is the worst thing here, and it is the first target of the next round.- «Что ты делаешь» without a question mark comes back as
o tawa. With the mark, and in English either way, it is right. lais read as a conditional where the relation is causal or temporal:ilo mi li pakala la mi ken ala toki tawa sinagives «if my computer breaks down then…» rather than «because». 1.0 handled this better; every version since has not.- Unmarked features get a default rather than a reading. Toki Pona marks
neither number nor tense, and bare
miusually comes back as «we», unmarked verbs as past. Both readings are valid —mi muteis optional — but the model does not choose by context, it just picks.
Licence
Gemma Terms of Use, inherited through the base model.
- Downloads last month
- 14
Model tree for NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged
Base model
google/gemma-3-1b-pt
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged:# Run inference directly in the terminal: llama cli -hf NetherQuartz/ilo-toki-1.3-MiLMMT-46-1b-merged: