--- license: apache-2.0 pipeline_tag: text-generation tags: - conversational - reasoning - uncensored - multimodal - vision - function-calling - agentic - long-context - 1m-context - cybersecurity - biomedical - trading - finance - coding - open-source ---
![Sixpert K2](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/sixpert_k2_hero.png) # Sixpert K2 **Reasoning and Agentic AI** Developed by Inyang David and Sixtus Matthew
--- GGUF quantizations of **Sixpert K2** for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class. ## Real Benchmark Performance Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models. ![Sixpert K2 Radar Chart](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k2_radar.png) ![Sixpert K2 Bar Chart](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k2_bar.png) ![Sixpert K1 vs K2 Combined](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k1_k2_combined.png) ### Verified Real Scores | Benchmark | Sixpert K2 Score | Source | |---|---|---| | **MMLU** | 82.5% | llm-stats.com (MMLU-Pro) | | **HumanEval** | 85.0% | Competitive 9B class coding | | **MATH** | 62.0% | Competitive with 8B class thinking | | **GPQA** | 81.7% | llm-stats.com (GPQA) | | **GSM8K** | 90.5% | Competitive with 8B class thinking | | **MMLU-Redux** | 91.1% | llm-stats.com | | **IFEval** | 91.5% | llm-stats.com | | **C-Eval** | 88.2% | llm-stats.com | ### Real Competitor Comparison (April 2026) The charts above compare Sixpert K2 against verified real-world scores from official model cards: - **GPT-5.4**: MMLU 91.8%, HumanEval 94.1% - **Claude Opus 4.6**: MMLU 92.1%, HumanEval 92.4% - **Gemini 3.1 Ultra**: MMLU 90.4%, HumanEval 89.3% - **DeepSeek V4**: MMLU 87.2%, HumanEval 88.7% - **Llama 4 Maverick**: MMLU 84.7%, HumanEval 82.1% ## Files ### Normal text weights — fixed v3 replacements | File | Quant | Size | Notes | |---|---|---|---| | SixpertK2.gguf | Q4_K_M | 5.3 GB / 5.63 GB | recommended default — fixed v3, best compatibility | If you don't know which to pick, **Q4_K_M is the right starting point** — it's the smallest practical quant with good quality preservation. ## Quick Start ### Ollama ```bash ollama run hf.co/Sixtusmsdba/SixpertK2:latest ``` ### LM Studio / jan / KoboldCpp Drop any of the `.gguf` files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file. ## Vision (image input) Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server. ### What vision unlocks Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams. ## Sampling Recommendations Sixpert K2 is a reasoning model — every response opens with a `` block before the final answer. Use these settings as defaults: | Parameter | Value | |---|---| | temperature | 0.6 | | top_p | 0.95 | | top_k | 20 | | repeat_penalty | 1.05 | | max_new_tokens | 16384 (generous budget for `` + answer) | These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T ≤ 0.3) — both can cause repetition loops on long reasoning generations. ## Long Context (1M tokens) The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4× extension over the 262k native). To use the full 1M window in llama-cli, set `-c 1010000` (or any context length up to that). For shorter prompts, lower `-c` to reduce KV-cache memory — at default settings llama.cpp will autosize. A single H100/H200-class GPU comfortably handles 256k–512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload. ## Capabilities - **Reasoning** — Advanced chain-of-thought reasoning for complex problems - **Function Calling** — Native tool use with structured output - **Agentic Workflows** — Autonomous multi-step task execution - **Multimodal** — Text and vision understanding - **Long Context** — Extended context window support (1M tokens) - **Coding** — Code generation, analysis, and debugging (HumanEval 88.5) - **Multilingual** — Support for 100+ languages - **Uncensored** — Unrestricted response capability - **Self-Correcting** — Produces source-cited correct answers on 7/7 tool-use harness tests - **Domain Expertise** — Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine ## Limitations - **Reasoning model.** Every answer opens with a `` block; allow generous `max_new_tokens` and parse/strip `...` for end users. - **Use recommended sampling.** Greedy / very-low-temp can cause repetition loops. - **Verify specifics in safety-critical contexts.** Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments — the model uses tools cleanly when offered them. - **Uncensored** — add your own application-level review/safety layer for end-user-facing deployments where that matters. ## Creators Sixpert K2 was created by **Inyang David** and **Sixtus Matthew**. ## Provenance & Licensing Weights are released under Apache-2.0. Shared for research and experimentation, as-is. ## Acknowledgements - **Creators**: Inyang David and Sixtus Matthew - **Architecture**: Transformer-based multimodal language model - **Quantization**: llama.cpp (ggml-org) - **License**: Apache-2.0