view post Post 541 DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.GGUFs: unsloth/DeepSeek-V4-Flash-0731-GGUFGuide: https://unsloth.ai/docs/models/deepseek-v4 See translation 🔥 2 2 🤗 1 1 + Reply
view post Post 1404 We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. 🤯We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.GGUF: unsloth/Kimi-K3-GGUFGitHub repo: https://github.com/unslothai/unsloth See translation 1 reply · 👍 2 2 🔥 1 1 🤗 1 1 + Reply
view post Post 4024 Kimi K3 can now be run locally! ✨The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.GGUF: unsloth/Kimi-K3-GGUFGuide: https://unsloth.ai/docs/models/kimi-k3 See translation 5 replies · ❤️ 11 11 🔥 9 9 👍 3 3 🤗 1 1 + Reply
view post Post 4920 Introducing Unsloth for AMD 🚀You can now train & run LLMs on your AMD hardware• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs• Works on Windows, WSL, Linux• Train Qwen, Gemma on just 3GB VRAMGitHub: https://github.com/unslothai/unslothBlog + Guide: https://unsloth.ai/docs/basics/amd See translation 3 replies · 🔥 27 27 ❤️ 10 10 🚀 6 6 🤗 6 6 👀 5 5 👍 5 5 🧠 3 3 ➕ 3 3 🤯 3 3 🤝 2 2 😎 2 2 + Reply
view post Post 5961 Gemma 4 is now faster and much more accurate! 🚀Google made huge improvements to tool-calling and chat accuracy, reliability + speed.To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4 See translation 8 replies · 🚀 30 30 👍 15 15 😎 6 6 🤗 5 5 🤝 1 1 + Reply
view post Post 4791 We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU.Gemma-4-12B NVFP4 works on 11GB VRAM.26B-A4B hits 13K tok/s (B200).Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.Blog: https://unsloth.ai/docs/basics/nvfp4Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4 See translation 3 replies · 🔥 12 12 🤗 3 3 🚀 2 2 👍 2 2 + Reply
CohereLabs/cohere-transcribe-arabic-07-2026 Automatic Speech Recognition • 2B • Updated 28 days ago • 48.9k • 163
view post Post 4344 We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡Qwen3.6-27B NVFP4 runs on 24GB VRAM.35B-A3B can hit 17,561 tok/s (B200).We also improved accuracy, tool calling, agent use, and looping.Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 See translation 1 reply · 🚀 15 15 🔥 11 11 🤗 1 1 + Reply
view post Post 6301 DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳Run lossless DeepSeek-V4-Flash on 168GB RAM or3-bit works on 110GB Mac, RAM, VRAM setups.Run via Unsloth Studio or llama.cpp.GGUF: unsloth/DeepSeek-V4-Flash-GGUFGuide: https://unsloth.ai/docs/models/deepseek-v4 See translation 🔥 20 20 🚀 5 5 👍 3 3 🤗 2 2 + Reply
view post Post 3386 1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5We gave 3 models the same prompt and compared one-shot outputs.The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.Which output do you like best?GGUF: unsloth/GLM-5.2-GGUF See translation 3 replies · 🤗 11 11 👍 4 4 🔥 1 1 + Reply
view post Post 4625 Google's new DiffusionGemma can now run at 2000+ tokens/sec! ⚡We made local DiffusionGemma inference 1.8× faster.Run it on 18GB RAM via Unsloth Studio.GitHub: https://github.com/unslothai/unslothGuide: https://unsloth.ai/docs/models/diffusiongemma See translation 4 replies · 🔥 8 8 🤗 2 2 😔 1 1 + Reply
view post Post 1205 Google releases DiffusionGemma.✨The new 26B-A4B diffusion text model runs locally on 18GB RAM.Run with 4x faster text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio.GGUF: unsloth/diffusiongemma-26B-A4B-it-GGUFGuide: https://unsloth.ai/docs/models/diffusiongemma See translation 1 reply · 🔥 3 3 🤗 1 1 + Reply
view post Post 4311 Google releases Gemma 4 QAT. ✨You can now run Gemma 4 at 3x less memory with near original performance.QAT makes it possible to run Gemma 4 26B-A4B on 16GB RAM.GGUFs: https://huggingface.co/collections/unsloth/gemma-4-qatQAT Guide: https://unsloth.ai/docs/models/gemma-4/qat See translation 1 reply · 👍 17 17 🚀 4 4 🤗 1 1 😎 1 1 + Reply
view post Post 9371 Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs.Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.You can run and train the model via Unsloth Studio.GGUF: unsloth/gemma-4-12b-it-GGUFGuide: https://unsloth.ai/docs/models/gemma-4 See translation 5 replies · 🔥 44 44 👍 13 13 🤗 2 2 + Reply
view post Post 2848 Qwen3.6 MTP is here! Run locally on 20GB RAM. ⚡️MTP enables Qwen3.6 to generate ~1.4–2.2× faster with no accuracy change.Qwen3.6-27B: unsloth/Qwen3.6-27B-MTP-GGUFQwen3.6-35B-A3B: unsloth/Qwen3.6-35B-A3B-MTP-GGUFGuide: https://unsloth.ai/docs/models/qwen3.6#mtp-guide See translation 2 replies · 👍 12 12 🔥 4 4 🤗 3 3 🚀 3 3 ❤️ 1 1 🧠 1 1 😎 1 1 + Reply