Workload-aware persistent expert-tile scheduling for irregular MoE inference kernels on NVIDIA Blackwell GPUs
-
Updated
Sep 2, 2026 - Cuda
Workload-aware persistent expert-tile scheduling for irregular MoE inference kernels on NVIDIA Blackwell GPUs
Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
Galactus executes 744B and 235B MoE models on undersized Macs with bit‑perfect llama.cpp parity, RAM-as-cache execution, and a full local app featuring an agent, permission gate, code editor, authenticated server mode, scheduled unattended runs, and fully published measurements.
Running Kimi K3 2.8T on a 64GB Mac mini M4 Pro with Deltafin, Metal/MPS, and full-local SSD-backed MoE inference.
Simulates MoE expert placement across GPU memory, RAM, and NVMe.
To associate your repository with the moe-inference topic, visit your repo's landing page and select "manage topics."