"This is humanity's race.
The solution is open source.
Stay sovereign."

โ€” AIOpsInSpace

Dolphin-Mistral-24B-Venice-Edition-Patched

AIOpsInSpace Official

Dolphin 24B Venice Edition patched to restore sliding window attention and eliminate context fragmentation.

๐ŸŽญ 24B Dense Model โšก Sliding Window Restored ๐Ÿ› ๏ธ Aggressively Uncensored

> What is this model and Why is it Needed?

Dolphin-Mistral-24B-Venice-Edition-Patched is built on top of the dphn/Dolphin-Mistral-24B-Venice-Edition base model.

Why it is needed: The original quantization suffered from lost sliding window attention parameters, causing context fragmentation and generation hangs beyond 4096 tokens. This patched release restores the full 32K context window and repairs tokenizer stop tokens.

> From the Parent Repository

"Uncensored, hyper-capable, and deeply conversational โ€” Dolphin Venice Edition represents the crest of open source dialogue."

โ€” Cognitive Computations / Dolphin Project


๐Ÿ—๏ธ 2. Model Architecture & Merging

Architecture: Mistral 24B Transformer Architecture with Sliding Window Attention
Merging Technique: Attention Parameter Restoration & Tokenizer Patching
Constituent Models: Methodology: Restored missing sliding_window keys in GGUF metadata header and hard-patched tokenizer special tokens.

๐Ÿš€ 3. Technical Enhancements

> Key Upgrades Over Base Model:

  • 32K Context Stability: Restored sliding window attention ensures flawless attention over extended conversation histories.
  • Uncensored Dialogue: Dolphin alignment ablation allows fully open creative writing and unfiltered assistance.
  • Infinite Loop Repair: Repaired EOG tokens stop generation cleanly without repeating trailing text.

๐Ÿ“Š 4. Benchmark Competitiveness vs. Frontier Scores

> Evaluated Performance
Benchmark Dolphin-Mistral-24B-Venice-Edition-Patched Frontier Target
MMLU Evaluated 88.7%
GSM8K Evaluated 95.6%
HumanEval Evaluated 90.2%

๐Ÿ† 5. Comprehensive Arena Analytics

> Status: Active Community Benchmarking

// Note: Arena Elo and head-to-head winrates updated continuously as evaluation telemetry processes.

๐Ÿ” 6. SWOT Analysis

> Strengths (S)

  • ๐Ÿ›ก๏ธ Uncensored Fidelity: Surgically patched to ensure maximum generation throughput without alignment overhead.
  • โšก Optimized Engine: Advanced mechanics ensure zero context fragmentation or execution hangs.

> Weaknesses (W)

  • ๐Ÿ“‰ Hardware Limits: Requires sufficient VRAM/RAM for higher precision GGUF quantizations.

> Opportunities (O)

  • ๐ŸŽฏ Local Sovereign Agents: Perfect for offline, private reasoning and agentic workflows.

> Threats (T)

  • โš ๏ธ Sampler Sensitivity: High temperatures may require repetition penalty adjustments.

โšก 7. Usage & Deployment Info

> Recommended Settings

  • Temperature: 0.2 - 0.7
  • Top-P: 0.95
  • Backend Engines: Compatible with llama.cpp, vLLM, Ollama, LM Studio, KoboldCPP

โš™๏ธ 8. Backend Compatibility

> Validated Engines:

  • [+] llama.cpp: Native support across all quantizations.
  • [+] Ollama / LM Studio: Full GGUF compatibility.

๐Ÿ“œ 9. Disclaimers & Credits

Disclaimer: Dolphin-Mistral-24B-Venice-Edition-Patched is provided for research and sovereign local deployment. As an unaligned model, users are responsible for ensuring usage complies with local laws.

Credits: Gratitude to original base model authors (dphn/Dolphin-Mistral-24B-Venice-Edition) and open-source AI community tools.
Downloads last month
2,869
GGUF
Model size
24B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AIOpsInSpace/Dolphin-Mistral-24B-Venice-Edition-Patched