HERMES
Safetensors
English
qwen3_5_moe
qwen3.6
Mixture of Experts
agentic
tool-calling
qlora
unsloth
carnice
Instructions to use samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- HERMES
How to use samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B with HERMES:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B", max_seq_length=2048, )
Update README with FP8 variant and formats table
Browse files
README.md
CHANGED
|
@@ -29,7 +29,20 @@ This is the successor to [Carnice-MoE-35B-A3B](https://huggingface.co/samuelcard
|
|
| 29 |
|
| 30 |
Training methodology adapted from **[kai-os/Carnice-9b](https://huggingface.co/kai-os/Carnice-9b)** — same two-stage approach and datasets, applied to the larger MoE architecture. Key inspiration: training on actual Hermes Agent execution traces for native agentic behavior.
|
| 31 |
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
## Model Details
|
| 35 |
|
|
|
|
| 29 |
|
| 30 |
Training methodology adapted from **[kai-os/Carnice-9b](https://huggingface.co/kai-os/Carnice-9b)** — same two-stage approach and datasets, applied to the larger MoE architecture. Key inspiration: training on actual Hermes Agent execution traces for native agentic behavior.
|
| 31 |
|
| 32 |
+
## Available Formats
|
| 33 |
+
|
| 34 |
+
| Format | Size | Location | Use Case |
|
| 35 |
+
|---|---|---|---|
|
| 36 |
+
| **BF16 SafeTensors** | 67 GB | Root | Full precision, Transformers / vLLM |
|
| 37 |
+
| **FP8 Dynamic** | 34 GB | `fp8/` | vLLM optimized, ~2x faster inference |
|
| 38 |
+
| **GGUF** | 19-65 GB | [GGUF repo](https://huggingface.co/samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B-GGUF) | llama.cpp, Ollama, LM Studio |
|
| 39 |
+
|
| 40 |
+
### FP8 Usage (vLLM)
|
| 41 |
+
|
| 42 |
+
```bash
|
| 43 |
+
# Clone the repo and point vLLM to the fp8/ subfolder
|
| 44 |
+
vllm serve samuelcardillo/Carnice-Qwen3.6-MoE-35B-A3B --quantization fp8 --dtype auto
|
| 45 |
+
```
|
| 46 |
|
| 47 |
## Model Details
|
| 48 |
|