--- library_name: mlx base_model: InternScience/Agents-A1 tags: - mlx - mlx-vlm - moe - multimodal - vision - agents - agentic - tool-use - basequant-xl pipeline_tag: image-text-to-text --- ## Local SOTA for 48GB Macs — Intelligence Benchmark Comparison This model is part of a benchmark comparison of the best local MLX-quantized LLMs that fit in 48GB unified memory on Apple Silicon. All benchmarks run in instruct mode (no thinking) with n=50 samples per benchmark. ![SOTA Comparison](sota_comparison_instruct.png) | Benchmark | Samples | Agents-A1 6bit-XL | Gemma-4 26B 6bit-XL | Huihui-Qwen3.6 6bit-XL | Ornith-35B 6bit-XL | Qwen3.6-27B oQ4e | Qwen3.6-35B 6bit-XL | Qwen3.6-35B oQ4e | Qwen3.6-35B oQ4e-XL | Qwen3.6-35B oQ6 | |-----------|:-------:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:| | MMLU | 50/14042 | 66% | **76%** | 74% | 64% | 74% | 64% | 66% | 72% | 64% | | MMLU_PRO | 50/12032 | 58% | **82%** | 66% | 66% | 56% | 64% | 60% | 64% | 60% | | ARC_CHALLENGE | 50/1172 | 90% | 90% | **92%** | **92%** | 88% | 90% | **92%** | **92%** | 90% | | HUMANEVAL | 50/164 | 90% | **98%** | 84% | 78% | **92%** | 78% | **92%** | 90% | 66% | | MBPP | 50/500 | 70% | 82% | 78% | 78% | **86%** | 78% | 80% | 76% | 76% | | **Average** | | 74.8% | **85.6%** | 78.8% | 75.6% | 79.2% | 74.8% | 78.0% | 78.8% | 71.2% | **Collection:** [Local SOTA for 48GB Macs](https://huggingface.co/collections/leonsarmiento/local-sota-for-48gb-macs-6a5fb58390dd01e1fc35d55e) > ⚠️ n=50 sampling means wide confidence intervals (±~13% at 95% CI). Differences under ~6 points may not be statistically significant. Models using data-aware quantization (oQ/oQe) may be calibrated on benchmark-like data — their scores carry a benchmaxxing caveat. The BaseQuant_XL variants (data-agnostic) provide the most honest generalization estimates. # leonsarmiento/Agents-A1-6bit-XL-mlx This model was converted to MLX format from [`InternScience/Agents-A1`](https://huggingface.co/InternScience/Agents-A1) using **BaseQuant_XL 6/8-bit mixed quantization** optimized for Apple Silicon. The vision encoder is preserved and quantized at 6-bit, making this a full multimodal model. **BaseQuant_XL** keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most. Agents-A1 is a 35B Mixture-of-Experts agentic model built to scale heterogeneous agentic abilities across multiple domains including Long-horizon Search, Engineering, Scientific Research, Instruction Following, and Tool-calling. It features 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, a vision encoder, and an extended 262K context window. Despite 35B total parameters, only ~3B are activated per token. ## Use with mlx ```bash pip install -U mlx-vlm ``` ```bash python -m mlx_vlm.generate --model leonsarmiento/Agents-A1-6bit-XL-mlx --max-tokens 256 --temperature 0.85 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.05 --prompt "Hello" ``` ## BaseQuant_XL Quantization Strategy | Bit Depth | Layers | Rationale | |-----------|--------|-----------| | **bf16 (unquantized)** | `mlp.gate` (router), `shared_expert_gate`, `lm_head`, `shared_expert` | Routing decisions and shared computation path — errors here are qualitatively different from precision loss | | **8-bit** | `embed_tokens`, `self_attn` (full attention), `linear_attn` (DeltaNet) | Every-token layers with moderate sensitivity — 8-bit is near-lossless | | **6-bit** | `vision_tower`, `switch_mlp` (routed experts) | Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision | ### Quantization Details | Layer | Bits | Group Size | |-------|------|------------| | `mlp.gate` (router) | bf16 | — | | `shared_expert_gate` | bf16 | — | | `lm_head` | bf16 | — | | `shared_expert` | bf16 | — | | `embed_tokens` | 8 | 64 | | `self_attn` (full attention) | 8 | 64 | | `linear_attn` (DeltaNet) | 8 | 64 | | `vision_tower` | 6 | 64 | | `switch_mlp` (routed experts) | 6 | 64 | | Default fallback | 8 | 64 | - **Quantization type**: BaseQuant_XL mixed (multimodal, vision preserved) - **Bits per weight**: 6.808 - **Total size**: ~28 GB (6 shards) - **Group size**: 64 - **Method**: Custom `quant_predicate` via `mlx_vlm` ## Recommended Inference Parameters | Parameter | Value | |-----------|-------| | `temperature` | 0.85 | | `top_p` | 0.95 | | `top_k` | 20 | | `min_p` | 0.01 | | `repeat_penalty` | 1.05 | | `presence_penalty` | 1.1 | ## Reasoning and Tool-Call Parsing | Parser | Value | |--------|-------| | `reasoning_parser` | `qwen3` | | `tool_call_parser` | `qwen3_coder` |