Instructions to use Alittlehammmer/Ornith-1.0-397B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Alittlehammmer/Ornith-1.0-397B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alittlehammmer/Ornith-1.0-397B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
- Ollama
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Ollama:
ollama run hf.co/Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
- Unsloth Studio
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Alittlehammmer/Ornith-1.0-397B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Alittlehammmer/Ornith-1.0-397B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Alittlehammmer/Ornith-1.0-397B-GGUF to start chatting
- Pi
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Docker Model Runner:
docker model run hf.co/Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
- Lemonade
How to use Alittlehammmer/Ornith-1.0-397B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Alittlehammmer/Ornith-1.0-397B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ornith-1.0-397B-GGUF-Q4_K_M
List all available models
lemonade list
Are other quants on the horizon?
Hello,
Thank you for being part of the few who are making quants of this model. Will you be creating smaller quants as well? Specifically within the range of 2-3 bit? The base model for this release is very resilient to quantization even down to the 2 bit range.
Thank you.
Bartowski indicated he might do some as well:
https://www.reddit.com/r/LocalLLaMA/comments/1ufykja/comment/otzkomw/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
Bartowski indicated he might do some as well:
https://www.reddit.com/r/LocalLLaMA/comments/1ufykja/comment/otzkomw/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
oh that's good to see, thanks for the heads up!
Yeah, admittedly I was only planning on uploading a Q4_K_M, Q6_K, and the BF16 GGUFs, under the guise that I'm sure other groups are doing more thorough releases. This was more intended as a quick compile since I know myself and others have been clamoring to at least try the 397B variant ha.
I also didn't use any imatrix calibration sets as I am new to this, so if I tried to pump out a lower quant it would for sure be a degradation in comparison to Bartowski's. While I could just use his imatrix dataset, that feels unfair to the amount of work he's put into that to get his own quants up and running.
All in all, I was treating the quant as a bit of a learning exercise for me with the added benefit of some real, usable files as the output. Actually makes me want to experiment with imatrix's some, see if I can't try to target generalized agentic use as a calibration set, and compare benchmark runs since the activations will be different. Might totally be a waste of time too lol, but that's a bit of the route I want to take.
Yeah, admittedly I was only planning on uploading a
Q4_K_M,Q6_K, and theBF16GGUFs, under the guise that I'm sure other groups are doing more thorough releases. This was more intended as a quick compile since I know myself and others have been clamoring to at least try the 397B variant ha.I also didn't use any
imatrixcalibration sets as I am new to this, so if I tried to pump out a lower quant it would for sure be a degradation in comparison to Bartowski's. While I could just use hisimatrixdataset, that feels unfair to the amount of work he's put into that to get his own quants up and running.All in all, I was treating the quant as a bit of a learning exercise for me with the added benefit of some real, usable files as the output. Actually makes me want to experiment with
imatrix's some, see if I can't try to target generalized agentic use as a calibration set, and compare benchmark runs since the activations will be different. Might totally be a waste of time too lol, but that's a bit of the route I want to take.
Not a problem! Thank you for still releasing the 2 quants anyways! I hope you have great success in the future!