3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
-
Updated
Sep 6, 2026 - Python
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Double Qwen 3.8 27B inference speed on Apple Silicon with a one-click local coding agent
One-click uncensored AI agent server for Apple Silicon Macs. MLX + MTPLX speculative decoding, OpenAI-compatible on your LAN. Up to 79.4 t/s decode, MTP speculative decoding on.
Beautiful, zero-dependency realtime dashboard + live activity log for a local MTPLX inference server — reads the /metrics endpoint, no build step.
Run an isolated Claude Science app copy through local or OpenAI-compatible model backends.
Recover the native MTP predictor missing from the 8-bit MLX Qwen3.8-27B-Uncensored package, build a BF16 sidecar, and reproduce a 15.59 → 48.75 tok/s controlled M4 Max result with MTPLX.
Local OpenAI/Anthropic-compatible LLM gateway that starts, stops and switches local runtimes (MTPLX, LM Studio, oMLX, Ollama) behind one stable endpoint for your coding agents.
To associate your repository with the mtplx topic, visit your repo's landing page and select "manage topics."