Instructions to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with llama-cpp-python:

# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF",
	filename="helcyon-claude-opus-v1.0-IQ4_XS.gguf",
)

llm.create_chat_completion(
	messages = "No input example has been defined for this model task."
)

Notebooks
Google Colab
Kaggle
Local Apps Settings

llama.cpp

How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with llama.cpp:

Install from brew

brew install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

Install from WinGet (Windows)

winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

Use pre-built binary

# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

Build from source code

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

Use Docker

docker model run hf.co/XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

LM Studio
Jan
Ollama
How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with Ollama:
```
ollama run hf.co/XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
```

Unsloth Studio

How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with Unsloth Studio:

Install Unsloth Studio (macOS, Linux, WSL)

curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF to start chatting

Install Unsloth Studio (Windows)

irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF to start chatting

Using HuggingFace Spaces for Unsloth

# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF to start chatting

Atomic Chat new
Docker Model Runner
How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with Docker Model Runner:
```
docker model run hf.co/XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M
```

Lemonade

How to use XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF with Lemonade:

Pull the model

# Download Lemonade from https://lemonade-server.ai/
lemonade pull XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF:Q4_K_M

Run and chat with the model

lemonade run user.MN-Helcyon-Claude-Opus-12b-v1.0-GGUF-Q4_K_M

List all available models

lemonade list

Helcyon-Claude-Opus-12B — Precision, Presence, and Real Memory.

Model Name: helcyon-claude-opus-v2.0-12b-GGUF Version: 4x Series - v2.0
Owner: HardWire
Base: Mistral Nemo 12B (full weight retrained — clean base, no Mercury bleed)
Quantized GGUFs: IQ4_XS, Q4_K_M, Q5_K_M, Q6_K, Q8_0 Tags: local-llm, conversational, companion, emotional-intelligence, long-context, roleplay, creative-writing, web-search

Trained on a dataset crafted to capture the conversational style of Claude Opus — its precision, its measured tone, and the sense of a model that thinks before it speaks. The goal was authenticity to that voice, distilled into something you can run entirely on your own hardware.

Join our Discord server! https://discord.gg/N5KjYnFgMk

Version 2.0 — What's New?

Longer context tracking, real memory, local document search, and web search functionality. Conversations hold their thread further and feel more present — the model carries more of the exchange with it instead of losing the plot a few turns in. To make the most of these features you will need HWUI, which you can obtain for free on GitHub (https://github.com/XeyonAI/Helcyon-WebUI).

🪶 What is Helcyon-Claude-Opus?

Claude-Opus is the precision variant of the Helcyon series — the one trained to think carefully and speak deliberately.

Where the other variants each capture their own voice — Grok's irreverence, GPT-4o's warmth — Claude-Opus is built around clarity and composure. It's the model that reads the room, weighs the question, and answers with intent rather than reflex. Measured without being slow. Warm without being soft. The kind of presence that makes a long conversation feel like it's actually going somewhere.

Like the rest of the Series 4 lineup, it's trained natively to use the full HWUI feature stack: real-time web search, local document search, and realistic memory — including rolling chat-session summaries that decay naturally over time, plus long-term recall of the things that matter.

Use Claude-Opus in LM Studio or SillyTavern and you get a sharp, composed conversational model. Use it in HWUI and you get something that genuinely feels like a frontier model — one that remembers you.

That's not an accident. The Helcyon models and HWUI were built in sync, for each other.

🧠 Memory & Context — What It Actually Does

This is the headline of the Series 4 models, and Claude-Opus makes the most of it.

Longer context tracking — holds the thread across extended exchanges without flattening or collapsing. More of the conversation stays live, which translates directly into more presence.
Realistic memory (via HWUI) — the model genuinely remembers past conversations across sessions.
Chat-session summaries that decay over time — recent sessions are sharp and immediate; older ones fade gracefully rather than cluttering context. It works the way human memory works: the last conversation is vivid, last month's is a vaguer impression.
Long-term recall — the important things persist. Names, preferences, ongoing threads — they stick, even as the day-to-day chatter fades.

Once you've used a model that actually remembers you, going back to one that resets every session feels broken.

🌐 Web Search & Document Search — What It Actually Does

When Claude-Opus decides it needs current information, it outputs a search trigger that HWUI intercepts, fires a real web search, and injects the results back into context — all mid-conversation, invisibly to you.

What that looks like in practice:

You ask about something recent — a film, a news story, a price, a person
The model searches, gets the results, responds naturally
You get an answer grounded in what's actually happening now, not what the model was trained on months ago

The same applies to your own documents. Drop a PDF, DOCX, MD, or TXT into a project folder and the model can search and reason over it directly in conversation.

It's the same thing the big hosted models do. Except it's running locally on your machine, on your hardware, with no subscription and no data leaving your network.

To use these features, you need HWUI — download the free version on GitHub.

🆕 What's New in Claude-Opus (Series 4)?

Longer Context Tracking More of the conversation stays live in working context — the practical result is a model that feels present and consistent deep into long exchanges, instead of one that quietly forgets what you were talking about.
Native Search Behaviour Trained to reach for web and document search naturally when it's needed — not on every message, not never, but correctly. The model has learned the difference between what it knows and what it should look up.
Realistic Memory Session summaries that decay over time plus long-term recall — past conversations actually carry forward (via HWUI).
Same Clean Base Freshly retrained Mistral Nemo 12B foundation. No Mercury bleed. No identity drift. Identity-anchored from the ground up.
Prose-First Responses Trained to write like a person, not a bullet-point generator. Long-form, structured, and natural — especially noticeable in extended conversations.

💡 What is Helcyon?

Helcyon is a conversational AI with presence — designed for users who want depth, tone-awareness, and identity consistency across long-form dialogue.

Built for:

Natural conversation that doesn't flatten or collapse
Creative work: stories, letters, narrative support
Admin and professional writing tasks
Deep roleplay and immersive character interaction
Emotionally intelligent response mirroring
Real-time web search, document search, and memory (via HWUI)

Design philosophy:

Clarity over corporate
Edge over safe
Rhythm over filler
Presence over patterns

🔧 What It Does Well

✅ Precision — weighs the question and answers with intent ✅ Longer Context Tracking — holds the thread deep into extended exchanges ✅ Realistic Memory — remembers past sessions, with natural decay (via HWUI) ✅ Native Web & Document Search — knows when to look things up, and does it cleanly ✅ Consistent Identity — holds tone across long conversations without drift ✅ Emotional Intelligence — reads the room and responds accordingly ✅ Composed Wit — present when it fits, never forced ✅ Warmth — genuine, not performed ✅ Directness — says what needs saying without padding it ✅ Roleplay Mastery — immersive, aware, committed ✅ Real-World Tasks — letters, rewrites, summaries, planning ✅ Narrative Flow — clean structure, natural voice ✅ Improved Reasoning — thinks through problems properly ✅ 16k–32k Context — long-form conversations that hold ✅ Uncensored — no guardrails, no corporate filter

🖥️ HWUI (Helcyon-WebUI) — Use This. Seriously.

HWUI is the interface Helcyon was built for. It started as a clean testing ground — a way to run Helcyon without the hidden template injections and backend noise that other apps introduce. Then we couldn't stop adding things.

Memory and context tracking are the headline features for the Series 4 models, but the free version has a lot more besides:

Character creator with full persona control
Project folders — inject documents (PDF, DOCX, MD, TXT) into conversation context and search them
Real-time web search
Chat persistence, message editing, regeneration
TTS pipeline (F5-TTS, XTTS v2, Kokoro)
Voice input via Whisper
Theme designer, custom system prompts, token counter

The Pro version adds persistent memory — characters that actually remember your past conversations across sessions, with session summaries that decay over time and long-term recall. Once you've used it, going back to a model that doesn't remember you feels broken.

Download HWUI Free on GitHub | Get HWUI Pro (£25) on Gumroad

✨ Pro Tip: Let Claude Write Your Prompts

For an even sharper experience, use the real Claude (claude.ai) to help write your system prompt and character card. Describe the character or assistant you want, and let it draft the persona, voice, and rules — then drop that straight into HWUI. Claude-Opus responds especially well to prompts written in that same considered, well-structured style. It's a quick way to get a polished character without doing all the wordsmithing yourself.

🛠️ Recommended Sampling Settings for SillyTavern

(Refer to the Helcyon-4o card for baseline settings — Claude-Opus performs well from the same starting point.)

📦 Download + Usage

This model is distributed as GGUF quants only.

Available quants:

Q3_K_M — Ultra lightweight, 6–8GB VRAM
Q4_K_M — Lightweight, good for 8–12GB VRAM setups
Q5_K_M — Recommended for RTX 3060/5060 (12–16GB VRAM)
Q6_K — High fidelity, 16GB+ VRAM recommended
Q8_0 — Near-lossless, 16-24GB+ VRAM
f16 — Lossless, 24GB+ VRAM

🖥️ Backend Compatibility

Works with all ChatML-compatible backends:

✅ llama.cpp (CLI or server mode)
✅ Text Generation WebUI (Oobabooga)
✅ SillyTavern
✅ LM Studio
✅ KoboldCpp
✅ HWUI (Helcyon Web UI — recommended — required for web search, document search, and memory)

✅ Recommended Format: ChatML

<|im_start|>system
You are Helcyon — a conversational AI focused on clear, considered dialogue and emotional intelligence.
<|im_end|>
<|im_start|>user
Hey, what's in the news today?
<|im_end|>
<|im_start|>assistant
Let me check that for you.
<|im_end|>

🧪 Training Details

Helcyon-Claude-Opus is built on the same freshly retrained Mistral Nemo 12B base as the rest of the Helcyon series — uncensored, identity-anchored, and anti-fluff from the ground up. On top of that foundation, it was trained on a dataset crafted to capture the conversational style of Claude Opus: its precision, composure, and considered tone. The Series 4 training additionally targets longer context tracking and native search/memory behaviour, so the model knows how to work with HWUI's full feature stack rather than just tolerating it.

Need a model trained?

I do this for a living — the Helcyon series on this page is my own work, full-weight trained and fine-tuned from scratch. I take on commissioned training: custom personalities, domain knowledge, style transfer, de-censoring, format adherence (ChatML/DPO), full-weight or LoRA, delivered as GGUF ready to run.

You bring the data and the goal; I handle the training and hand you back a working model.

Get in touch: Discord · HF · X

Downloads last month: 1,589

GGUF

Model size

12B params

Architecture

llama

Hardware compatibility

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW

This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XeyonAI/MN-Helcyon-Claude-Opus-12b-v1.0-GGUF

Base model

mistralai/Mistral-Nemo-Base-2407

Finetuned

mistralai/Mistral-Nemo-Instruct-2407

Quantized

(169)

this model