Instructions to use mitsutani/mahjonglm-100m-q4-k-m-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mitsutani/mahjonglm-100m-q4-k-m-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
Use Docker
docker model run hf.co/mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use mitsutani/mahjonglm-100m-q4-k-m-gguf with Ollama:
ollama run hf.co/mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
- Unsloth Studio
How to use mitsutani/mahjonglm-100m-q4-k-m-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for mitsutani/mahjonglm-100m-q4-k-m-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for mitsutani/mahjonglm-100m-q4-k-m-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for mitsutani/mahjonglm-100m-q4-k-m-gguf to start chatting
- Docker Model Runner
How to use mitsutani/mahjonglm-100m-q4-k-m-gguf with Docker Model Runner:
docker model run hf.co/mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
- Lemonade
How to use mitsutani/mahjonglm-100m-q4-k-m-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mitsutani/mahjonglm-100m-q4-k-m-gguf:Q4_K_M
Run and chat with the model
lemonade run user.mahjonglm-100m-q4-k-m-gguf-Q4_K_M
List all available models
lemonade list
- Atomic Chat
MahjongLM 100M Q4_K_M GGUF
This repository contains the Q4_K_M GGUF export of mitsutani/mahjonglm-100m.
GGUF file:
mahjonglm-100m-Q4_K_M.gguf
Training
The source model was trained on mitsutani/mahjonglm-dataset, using complete, imperfect, and omniscient MahjongLM views over Tenhou logs from 2011 through 2024.
Prompt Format
<bos> rule_player_4 rule_length_hanchan view_complete game_start
For view_omniscient, use:
<bos> rule_player_4 rule_length_hanchan view_omniscient game_start round_start wall
and then provide or generate the 136 wall tile tokens.
llama.cpp WordLevel Tokenizer Patch
MahjongLM uses a fixed WordLevel tokenizer over Mahjong log tokens rather than a byte-pair or sentencepiece tokenizer. The GGUF file contains the vocabulary metadata, but upstream llama.cpp builds may not know how to interpret this tokenizer type.
This repository includes the minimal llama.cpp patch in patches/:
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
git apply /path/to/patches/0001-wordlevel-tokenizer-support.patch
cmake -B build
cmake --build build --config Release
After building the patched binary, run the model normally with llama-cli or compatible llama.cpp tools. The patch is tokenizer-only; the model architecture is plain Qwen3.
- Downloads last month
- 6
4-bit