How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf tflsxyy/Qwen3-235B-A22B-IQ2_S:IQ2_S
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default tflsxyy/Qwen3-235B-A22B-IQ2_S:IQ2_S
Run Hermes
hermes
Quick Links

Attn: Q4_K

Experts: IQ2_S

Please refer to unsloth for running this model.

./llama.cpp/llama-quantize --imatrix /work/yanzhi_group/models/unsloth/Qwen3-235B-A22B-GGUF/imatrix_unsloth.dat --keep-split /work/yanzhi_group/models/unsloth/Qwen3-235B-A22B-GGUF/BF16/Qwen3-235B-A22B-BF16-00001-of-00010.gguf /scratch/xie.yany/Qwen/Qwen3-235B-A22B-IQ2_S/Qwen3-235B-A22B-IQ2_S.gguf IQ2_S
Downloads last month
25
GGUF
Model size
235B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tflsxyy/Qwen3-235B-A22B-IQ2_S

Quantized
(55)
this model