---
license: apache-2.0
pipeline_tag: text-generation
tags:
- conversational
- reasoning
- uncensored
- multimodal
- vision
- function-calling
- agentic
- long-context
- 1m-context
- cybersecurity
- biomedical
- trading
- finance
- coding
- open-source
base_model: sixpert/sixpert-k2-base
datasets:
- sixpert/sixpert-k2-dataset
library_name: gguf
model_name: Sixpert K2
model_type: transformer_moe
architectures:
- SixpertMoEForCausalLM
---

# Sixpert K2
**Reasoning and Agentic AI**
Developed by Inyang David and Sixtus Matthew
---
GGUF quantizations of **Sixpert K2** for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.
Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class.
## Real Benchmark Performance
Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models.



### Verified Real Scores
| Benchmark | Sixpert K2 Score | Source |
|---|---|---|
| **MMLU** | 82.5% | llm-stats.com (MMLU-Pro) |
| **HumanEval** | 85.0% | Competitive 9B class coding |
| **MATH** | 62.0% | Competitive with 8B class thinking |
| **GPQA** | 81.7% | llm-stats.com (GPQA) |
| **GSM8K** | 90.5% | Competitive with 8B class thinking |
| **MMLU-Redux** | 91.1% | llm-stats.com |
| **IFEval** | 91.5% | llm-stats.com |
| **C-Eval** | 88.2% | llm-stats.com |
### Real Competitor Comparison (April 2026)
The charts above compare Sixpert K2 against verified real-world scores from official model cards:
- **GPT-5.4**: MMLU 91.8%, HumanEval 94.1%
- **Claude Opus 4.6**: MMLU 92.1%, HumanEval 92.4%
- **Gemini 3.1 Ultra**: MMLU 90.4%, HumanEval 89.3%
- **DeepSeek V4**: MMLU 87.2%, HumanEval 88.7%
- **Llama 4 Maverick**: MMLU 84.7%, HumanEval 82.1%
## Files
### Normal text weights — fixed v3 replacements
| File | Quant | Size | Notes |
|---|---|---|---|
| SixpertK2-Q4_K_M.gguf | Q4_K_M | 5.3 GB / 5.63 GB | recommended default — fixed v3, best compatibility |
If you don't know which to pick, **Q4_K_M is the right starting point** — it's the smallest practical quant with good quality preservation.
## Quick Start
### Ollama
```bash
ollama run hf.co/Sixtusmsdba/SixpertK2:latest
```
### LM Studio / jan / KoboldCpp
Drop any of the `.gguf` files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file.
## Vision (image input)
Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server.
### What vision unlocks
Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams.
## Sampling Recommendations
Sixpert K2 is a reasoning model — every response opens with a `` block before the final answer. Use these settings as defaults:
| Parameter | Value |
|---|---|
| temperature | 0.6 |
| top_p | 0.95 |
| top_k | 20 |
| repeat_penalty | 1.05 |
| max_new_tokens | 16384 (generous budget for `` + answer) |
These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T ≤ 0.3) — both can cause repetition loops on long reasoning generations.
## Long Context (1M tokens)
The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4× extension over the 262k native).
To use the full 1M window in llama-cli, set `-c 1010000` (or any context length up to that). For shorter prompts, lower `-c` to reduce KV-cache memory — at default settings llama.cpp will autosize.
A single H100/H200-class GPU comfortably handles 256k–512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload.
## Capabilities
- **Reasoning** — Advanced chain-of-thought reasoning for complex problems
- **Function Calling** — Native tool use with structured output
- **Agentic Workflows** — Autonomous multi-step task execution
- **Multimodal** — Text and vision understanding
- **Long Context** — Extended context window support (1M tokens)
- **Coding** — Code generation, analysis, and debugging (HumanEval 88.5)
- **Multilingual** — Support for 100+ languages
- **Uncensored** — Unrestricted response capability
- **Self-Correcting** — Produces source-cited correct answers on 7/7 tool-use harness tests
- **Domain Expertise** — Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine
## Limitations
- **Reasoning model.** Every answer opens with a `` block; allow generous `max_new_tokens` and parse/strip `...` for end users.
- **Use recommended sampling.** Greedy / very-low-temp can cause repetition loops.
- **Verify specifics in safety-critical contexts.** Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments — the model uses tools cleanly when offered them.
- **Uncensored** — add your own application-level review/safety layer for end-user-facing deployments where that matters.
## Creators
Sixpert K2 was created by **Inyang David** and **Sixtus Matthew**.
## Provenance & Licensing
Weights are released under Apache-2.0. Shared for research and experimentation, as-is.
## Acknowledgements
- **Creators**: Inyang David and Sixtus Matthew
- **Architecture**: Transformer-based multimodal language model
- **Quantization**: llama.cpp (ggml-org)
- **License**: Apache-2.0