JANGQ     vMLX

JANGQ-AI/Qwen3.8-Flash-Next-JANG_6S

The top tier — median KL 0.0035 vs bf16, built for 128 GB Macs with the n-gram table on SSD (~83 GiB resident).

A JANG bundle of Qwen/Qwen3.8-Flash-Next — the Qwen4-architecture preview: a 125B mixture-of-experts (512 experts, 6B active) with a 51B hashed n-gram embedding, Gated DeltaNet + Qwen Sparse Attention hybrid layers, gated-residual streams, and vision+video towers — quantized for Apple Silicon / MLX. Text, image and video weights are all present in this exact bundle. Native multi-token-prediction head preserved (6-bit).

Best experienced in vMLX. This bundle's layout — the SSD-served n-gram table, per-module mixed precision, and the native MTP head — is designed for the vMLX serving path. Access is gated (manual approval) while runtime support rolls out.

Quality (measured, 5,931 held-out positions vs bf16)

JANG ladder

Tier Size RAM w/ SSD-table median KL top-1 top-5 top-10
JANG_1L 59.8 GiB ~41 GiB 0.0362 86.7% 97.5% 98.8%
JANG_2L 65.3 GiB ~48 GiB 0.0260 88.2% 98.2% 99.0%
JANG_4S 71.8 GiB ~53 GiB 0.0161 89.4% 98.7% 99.4%
JANG_4M 96.0 GiB ~73 GiB 0.0042 94.4% 99.7% 99.9%
JANG_6S 106.3 GiB ~83 GiB 0.0035 94.7% 99.7% 99.9%

Margin-conditioned flip curves are monotone-decreasing on every tier — quantization noise lives in the reference model's own uncertainty band, with zero disagreement at high-confidence positions on the upper tiers.

The n-gram table & memory — SSD caching, fixed and fast

The 51B n-gram embedding streams directly from SSD on supporting runtimes (16 row-reads per token) — the "RAM w/ SSD-table" column above is the true resident footprint in that mode. Early runtime builds throttled in this mode; SSD-table caching is now fixed: decode runs at full speed with the table on disk — 40+ tok/s on an M5 Max for the 4-bit tier — so the biggest tiers fit comfortably on 64–128 GB machines without giving up the table.

What's in the bundle

  • Vision + video: the full vision tower and both image and video preprocessors ship in this exact bundle — image-text-to-text and video understanding work out of the box on supporting runtimes (image and video token ids, mRoPE positions, and the merger are all present).
  • Multi-token prediction: the model's native MTP head is preserved (trained multi-step). Enables self-speculative decode on supporting runtimes.
  • Thinking + agentic: thinking mode on by default with three reasoning efforts and preserved thinking history; Hermes-style tool calling; the instruct preset gives direct non-thinking responses.
  • Long context: 262,144 tokens native, extensible to 1M with YaRN.

Serving contract

  • Thinking mode ON by default: temperature=1.0, top_p=0.95, top_k=20
  • Instruct mode: temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5
  • Reasoning efforts low / medium / xhigh (default xhigh) and preserve_thinking (default on) via chat-template kwargs
  • Context 262,144 native, extensible to 1M with YaRN
  • EOS [248046, 248044] · tool calls: Hermes-style <tool_call>

Quantized and validated by Jinho Jangeric@jangq.ai

Downloads last month
-
Safetensors
Model size
33B params
Tensor type
F16
·
U32
·
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JANGQ-AI/Qwen3.8-Flash-Next-JANG_6S

Finetuned
(18)
this model