8GB VRAM Local LLMs - Practitioner Tested
11 models on RTX 4060 Ti 8GB. speed king: Llama 1B (228 tok/s). quality king: Gemma 4 E4B (6/6). agentic king: gpt-oss-20b (10/10).
Text Generation • 4B • Updated • 10.1k • 187Note 80.7 t/s on RTX 4060 Ti 8GB. Fastest small model in my 5-model test set. NVIDIA's edge-ready open release - positioned for gaming NPCs, voice assistants, IoT. Solid technical accuracy in MoE explainer test.
google/gemma-4-E4B-it
Any-to-Any • 8B • Updated • 5.86M • 1.44kNote 68.5 t/s, 6.0GB VRAM. Fast and accurate. Heavier than its 4B name suggests (E variants use selective activation but full weights still sit in VRAM). Lowest TTFT in my test (0.26s).
ibm-granite/granite-4.1-8b
Text Generation • 9B • Updated • 3.99M • 241Note 49.1 t/s, 5.3GB VRAM. Best instruction-follower in my 5-model test. Hit length target exactly. Cleanest answer of the five. IBM's "matches our previous 32B MoE" claim is credible from this sample. Practitioner pick for accuracy.
Qwen/Qwen3.5-9B
Image-Text-to-Text • 10B • Updated • 12.1M • • 1.78kNote 44.2 t/s, 6.2GB VRAM. Reference baseline. The only model in my test that called out the memory-bandwidth bottleneck on consumer GPUs specifically (rare insight at this size class). Multimodal-capable.
microsoft/Phi-4-mini-instruct
Text Generation • 4B • Updated • 471k • • 803Note Tested 2026-05-06: 88.92 tok/sec on RTX 4060 Ti 8GB, Q4_K_M GGUF (lmstudio-community quant), 16K context, full GPU offload, LM Studio. Coherent prose, on-topic, slight length-budget overshoot but no factual fabrication. New leader in this catalog at the dense-Q4 8GB tier.
mistralai/Ministral-3-8B-Instruct-2512
9B • Updated • 110k • 191Note Tested 2026-05-06: 48.47 tok/sec on RTX 4060 Ti 8GB, Q4_K_M GGUF (lmstudio-community quant), 16K context. Cleaner instruction-following than llama-3.3-8b at similar speed — 4x fewer tokens for the same answer means much better wall-clock per useful output.
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 6M • • 2.61kNote 35 tok/s partial offload (-ncmoe 30, 32K ctx, llama-server). Full offload: 7.4 tok/s (32GB RAM ceiling). See dataset.
witcheer/rtx-4060ti-8gb-turboquant-bench-2026-05
Viewer • Updated • 18 • 20 • 1Note turboquant KV cache benchmarks + context checkpoint discovery
google/gemma-4-26B-A4B-it
Image-Text-to-Text • 27B • Updated • 12.2M • • 1.35kNote 29.3 t/s MoE partial offload (ncmoe 23, 32K context)
LiquidAI/LFM2-24B-A2B
Text Generation • 24B • Updated • 3.8k • 340Note 52.2 tok/s at 32K on RTX 4060 Ti 8GB. New speed champion. Hybrid SSM+conv+MoE — only 2.3B active params, 13.4 GB Q4_K_M. ncmoe 22 sweet spot. 12-test quality stress test: solid for technical work. Liquid AI (MIT spinoff).
google/gemma-4-26B-A4B
Image-Text-to-Text • 27B • Updated • 844k • 363Note 29.3 tok/s at 32K on RTX 4060 Ti 8GB. Slowest of the three MoE models but best 65K scaling (-12% vs Qwen -51%). 16 GB Q4_K_M. 5:1 sliding window attention keeps context cheap. ncmoe 23 sweet spot.
witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05
Viewer • Updated • 126 • 186 • 2Note Full benchmark dataset: 19 rows across Qwen3.6, Gemma 4, and LFM2. JSONL with ncmoe sweep + context scaling data, VRAM breakdowns, and methodology.
google/gemma-4-E2B-it
Any-to-Any • 5B • Updated • 4.12M • 867Note 117.8 tok/s Q4_K_M, full GPU, 2.9 GB on disk, 2.6 GB VRAM. dense + PLE. fastest model tested.
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 5.1M • 2.74kNote 56.4 tok/s Q4_K_M, full GPU, 4.1 GB on disk, 6.7 GB VRAM. dense 7.3B all-active. best quality (5/6).
zai-org/GLM-4.7-Flash
Text Generation • 31B • Updated • 2.1M • • 1.8kNote 30B MoE + MLA, ~3B active. 29.9 tok/s @ ncmoe=35. Flattest context scaling (-14.7% at 24K). 4/6 quality. Worst hallucination.
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 15.9M • • 1.27kNote Dense 8.2B, 50.8 tok/s. First model to pass hallucination test. Thinking mode kills code gen. VRAM-hungry at 32K.
meta-llama/Llama-3.2-1B-Instruct
Text Generation • 1B • Updated • 10.1M • • 1.55kNote 228 tok/s speed ceiling. 771 MB, 2.1 GB VRAM. 4/6 quality. No thinking mode.
Qwen/Qwen3.6-27B
Image-Text-to-Text • 28B • Updated • 6.83M • • 2.14kNote dense control: 3.28 tok/s at ngl=20. MoE sibling is 10.8x faster.
witcheer/local-agentic-coding-bench-8gb-vram-2026-05
Viewer • Updated • 66 • 106 • 7Note 29 rows: 4 models x 10 agentic tasks + 12 GBNF grammar experiment rows. gpt-oss-20b 10/10, GBNF grammar 6.2x compression on Qwen3.6.
openai/gpt-oss-20b
Text Generation • 22B • Updated • 8.48M • • 4.87kNote agentic coding champion: 10/10 tasks on 8GB VRAM (1.8 GB used). 21B MoE, 3.6B active, ncmoe=30. first model to complete both easy + hard agentic tasks. poor self-debugger on complex freeform code but strong generator.
Tesslate/OmniCoder-9B
Text Generation • 9B • Updated • 1.22k • 668Note dense 9B (Qwen3.5-9B base, 425K agentic trajectories). 42 tok/s Q4_K_M, 5.6 GB VRAM. portscout easy: PASS (<1 min, fastest model tested). logpulse hard: PARTIAL (clean code but agent stuck on blocking command). strong generator, weak agent loop management. needs -rea off flag.