How to use from
MLX LM
Generate or start a chat session
# Install MLX LM
uv tool install mlx-lm
# Interactive chat REPL
mlx_lm.chat --model "Shiftedx/ornith-1.0-35b-mxfp4-mtplx"
Run an OpenAI-compatible server
# Install MLX LM
uv tool install mlx-lm
# Start the server
mlx_lm.server --model "Shiftedx/ornith-1.0-35b-mxfp4-mtplx"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
   -H "Content-Type: application/json" \
   --data '{
     "model": "Shiftedx/ornith-1.0-35b-mxfp4-mtplx",
     "messages": [
       {"role": "user", "content": "Hello"}
     ]
   }'
Quick Links

Ornith 1.0 35B MXFP4 MTPLX

Text-only MXFP4 MLX build of deepreinforce-ai/Ornith-1.0-35B, packaged for MTPLX native-MTP inference on Apple Silicon.

This is intended for local, private inference. After download, prompts and outputs can stay on your machine when served with a local MTPLX endpoint.

Notes

  • Text-only: no vision tower is included.
  • Optimized for MTPLX, not LM Studio.
  • Uses a compatible transplanted MTP sidecar; recommended draft depth is 2.
  • Reasoning should be routed separately with the Qwen reasoning parser.

Recommended MTPLX Settings

python -m mtplx.server.openai \
  --model /path/to/ornith-1.0-35b-mxfp4-mtplx \
  --backend-id qwen3_next \
  --generation-mode mtp \
  --load-mtp \
  --depth 2 \
  --profile sustained \
  --chat-template-profile tokenizer \
  --normalize-thinking-tags \
  --reasoning-mode on \
  --enable-thinking \
  --reasoning-parser qwen3 \
  --reasoning-effort high \
  --temperature 0.2 \
  --top-p 0.95 \
  --top-k 20 \
  --no-stats-footer

Local Validation

Hardware reference: Apple M4 Max Apple Silicon with 64 GB unified memory.

On a local Apple Silicon host, this MTPLX profile matched the LM Studio text baseline on a small hard validation suite:

Runtime Score Measured speed
LM Studio text baseline 7/10 104 tok/s
MTPLX depth 2 7/10 133 tok/s

This is a lightweight local validation, not a public leaderboard result.

Privacy

This repository contains model files only. It does not include a hosted endpoint, telemetry, or an external service requirement. Use a local server and inspect your client configuration if strict data locality matters.

Downloads last month
283
Safetensors
Model size
7B params
Tensor type
U8
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shiftedx/ornith-1.0-35b-mxfp4-mtplx

Quantized
(185)
this model

Collection including Shiftedx/ornith-1.0-35b-mxfp4-mtplx