How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "Shiftedx/ornith-1.0-35b-mxfp4-mtplx"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default Shiftedx/ornith-1.0-35b-mxfp4-mtplx
Run Hermes
hermes
Quick Links

Ornith 1.0 35B MXFP4 MTPLX

Text-only MXFP4 MLX build of deepreinforce-ai/Ornith-1.0-35B, packaged for MTPLX native-MTP inference on Apple Silicon.

This is intended for local, private inference. After download, prompts and outputs can stay on your machine when served with a local MTPLX endpoint.

Notes

  • Text-only: no vision tower is included.
  • Optimized for MTPLX, not LM Studio.
  • Uses a compatible transplanted MTP sidecar; recommended draft depth is 2.
  • Reasoning should be routed separately with the Qwen reasoning parser.

Recommended MTPLX Settings

python -m mtplx.server.openai \
  --model /path/to/ornith-1.0-35b-mxfp4-mtplx \
  --backend-id qwen3_next \
  --generation-mode mtp \
  --load-mtp \
  --depth 2 \
  --profile sustained \
  --chat-template-profile tokenizer \
  --normalize-thinking-tags \
  --reasoning-mode on \
  --enable-thinking \
  --reasoning-parser qwen3 \
  --reasoning-effort high \
  --temperature 0.2 \
  --top-p 0.95 \
  --top-k 20 \
  --no-stats-footer

Local Validation

Hardware reference: Apple M4 Max Apple Silicon with 64 GB unified memory.

On a local Apple Silicon host, this MTPLX profile matched the LM Studio text baseline on a small hard validation suite:

Runtime Score Measured speed
LM Studio text baseline 7/10 104 tok/s
MTPLX depth 2 7/10 133 tok/s

This is a lightweight local validation, not a public leaderboard result.

Privacy

This repository contains model files only. It does not include a hosted endpoint, telemetry, or an external service requirement. Use a local server and inspect your client configuration if strict data locality matters.

Downloads last month
283
Safetensors
Model size
7B params
Tensor type
U8
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shiftedx/ornith-1.0-35b-mxfp4-mtplx

Quantized
(185)
this model

Collection including Shiftedx/ornith-1.0-35b-mxfp4-mtplx