h4bbo commited on
Commit
87e6a8e
·
verified ·
1 Parent(s): df99ae9

README: GGUFs no longer shipped; document local conversion to llama.cpp

Browse files
Files changed (1) hide show
  1. README.md +9 -5
README.md CHANGED
@@ -69,11 +69,17 @@ print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
69
 
70
  ### llama.cpp (completion mode)
71
 
72
- The repo includes GGUF files (`FuseLLM-112M.fp16.gguf`, `FuseLLM-112M.Q4_K_M.gguf`)
73
- verified to load and generate in `llama.cpp`.
74
 
75
  ```bash
76
- # Completion mode pass the raw code seed, do NOT use chat/conversation mode.
 
 
 
 
 
 
77
  llama-cli -m FuseLLM-112M.Q4_K_M.gguf -cnv -st --no-jinja \
78
  -f seed.txt -n 64 --temp 0.0 --repeat-penalty 1.1 --no-display-prompt < /dev/null
79
  ```
@@ -85,8 +91,6 @@ isn't chat-tuned, so conversation mode is not meaningful for this model).
85
 
86
  - `model.safetensors`, `config.json`, `generation_config.json` — HF model
87
  - `tokenizer.json`, `tokenizer_config.json`, `chat_template.jinja` — tokenizer + ChatML template
88
- - `FuseLLM-112M.fp16.gguf` — lossless fp16 GGUF (~220 MB)
89
- - `FuseLLM-112M.Q4_K_M.gguf` — 4-bit quantized GGUF (~88 MB), the practical llama.cpp file
90
 
91
  ## Notes
92
 
 
69
 
70
  ### llama.cpp (completion mode)
71
 
72
+ No GGUF is shipped in this repo. The HF model is **verified** to convert and run in
73
+ `llama.cpp`; generate the GGUF locally:
74
 
75
  ```bash
76
+ # 1) convert HF -> lossless fp16 GGUF
77
+ python convert_hf_to_gguf.py h4bbo/FuseLLM-112M --outtype f16 \
78
+ --model-name FuseLLM-112M --outfile FuseLLM-112M.fp16.gguf
79
+ # (optional) 4-bit quantize
80
+ llama-quantize FuseLLM-112M.fp16.gguf FuseLLM-112M.Q4_K_M.gguf Q4_K_M
81
+
82
+ # 2) completion mode — pass the raw code seed, do NOT use chat/conversation mode.
83
  llama-cli -m FuseLLM-112M.Q4_K_M.gguf -cnv -st --no-jinja \
84
  -f seed.txt -n 64 --temp 0.0 --repeat-penalty 1.1 --no-display-prompt < /dev/null
85
  ```
 
91
 
92
  - `model.safetensors`, `config.json`, `generation_config.json` — HF model
93
  - `tokenizer.json`, `tokenizer_config.json`, `chat_template.jinja` — tokenizer + ChatML template
 
 
94
 
95
  ## Notes
96