Nick Emb commited on
Commit
baac2ef
·
verified ·
1 Parent(s): 883875d

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +105 -0
README.md ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - gguf
8
+ - context-1
9
+ - chroma
10
+ - moe
11
+ - qwen
12
+ - llama.cpp
13
+ - quantized
14
+ base_model: chromadb/context-1
15
+ ---
16
+
17
+ # Context-1 GGUF Quantizations
18
+
19
+ GGUF quantized versions of [chromadb/context-1](https://huggingface.co/chromadb/context-1), converted for inference with [llama.cpp](https://github.com/ggml-org/llama.cpp), [LM Studio](https://lmstudio.ai/), and other GGUF-compatible engines.
20
+
21
+ ## About Context-1
22
+
23
+ **Context-1** is a 20.9B parameter Mixture-of-Experts (MoE) causal language model developed by [Chroma](https://www.trychroma.com/). It uses the `GptOssForCausalLM` architecture with 32 experts and 4 active per token, providing strong performance with efficient inference.
24
+
25
+ | Detail | Value |
26
+ |--------|-------|
27
+ | **Architecture** | GptOssForCausalLM (MoE) |
28
+ | **Total Parameters** | ~20.9B |
29
+ | **Active Parameters** | ~3B per token (4 of 32 experts) |
30
+ | **Hidden Size** | 2880 |
31
+ | **License** | Apache-2.0 |
32
+
33
+ ## Quantization Details
34
+
35
+ Quantized from the F16 safetensors in the original repository using [llama.cpp](https://github.com/ggml-org/llama.cpp)'s `llama-quantize` tool, running on NVIDIA H100 via [Modal](https://modal.com/).
36
+
37
+ | File | Format | Size |
38
+ |------|--------|------|
39
+ | `context-1-Q4_K_M.gguf` | Q4_K_M | 15.81 GB |
40
+
41
+ More quantization levels (Q3, Q5, Q8, etc.) may be added in the future.
42
+
43
+ ## Usage
44
+
45
+ ### llama.cpp
46
+
47
+ ```bash
48
+ # Download
49
+ huggingface-cli download nicoism/context-1-GGUF context-1-Q4_K_M.gguf --local-dir .
50
+
51
+ # Run
52
+ ./llama-cli -m context-1-Q4_K_M.gguf -p "Your prompt here" -ngl 99
53
+ ```
54
+
55
+ ### LM Studio
56
+
57
+ Search for `nicoism/context-1-GGUF` in LM Studio's model browser and download the desired quantization.
58
+
59
+ ### Python (llama-cpp-python)
60
+
61
+ ```python
62
+ from llama_cpp import Llama
63
+
64
+ llm = Llama.from_pretrained(
65
+ repo_id="nicoism/context-1-GGUF",
66
+ filename="context-1-Q4_K_M.gguf",
67
+ n_gpu_layers=-1,
68
+ )
69
+
70
+ response = llm.create_chat_completion(
71
+ messages=[{"role": "user", "content": "Hello!"}]
72
+ )
73
+ print(response)
74
+ ```
75
+
76
+ ## Chat Template
77
+
78
+ This model uses a custom chat template based on the OpenAI/Oss architecture with support for multi-channel output (analysis, commentary, final), tool calling, and built-in browser/python tools. The template is embedded in the GGUF files.
79
+
80
+ Format overview:
81
+
82
+ ```
83
+ <|start|>system<|message|>...<|end|>
84
+ <|start|>developer<|message|>...<|end|>
85
+ <|start|>user<|message|>...<|end|>
86
+ <|start|>assistant<|channel|>final<|message|>...<|end|>
87
+ ```
88
+
89
+ For the full template, see [`chat_template.jinja`](https://huggingface.co/chromadb/context-1/blob/main/chat_template.jinja) in the original repository.
90
+
91
+ ## Limitations
92
+
93
+ - GGUF quantization introduces minor quality degradation compared to the original F16 weights. Q4_K_M provides a good balance of quality and size.
94
+ - This model inherits any biases and limitations from the base model.
95
+ - Requires a GPU or sufficient RAM for inference (16 GB+ for Q4_K_M).
96
+
97
+ ## License
98
+
99
+ Apache-2.0 — same as the [original model](https://huggingface.co/chromadb/context-1).
100
+
101
+ ## Acknowledgements
102
+
103
+ - [Chroma](https://www.trychroma.com/) for the original model
104
+ - [llama.cpp](https://github.com/ggml-org/llama.cpp) for the quantization tooling
105
+ - [Modal](https://modal.com/) for the GPU compute infrastructure