flash-moe

An inference engine written in C and Metal that runs a very large 397 billion parameter language model on a laptop. It shows that big models can run on small consumer hardware.

Share on XNo license declared

Overview

Flash-MoE is an inference engine written in C, Objective-C and hand-tuned Metal shaders. It runs Qwen3.5-397B-A17B, a 397 billion parameter Mixture-of-Experts model, on a MacBook Pro with 48GB of memory. The 209GB model is streamed from the SSD through a custom Metal compute pipeline, with no Python and no frameworks.

Key features

  • Runs a 397B parameter model on a 48GB MacBook Pro
  • Streams the model from SSD through Metal compute
  • Reported 4.36 tokens per second with 4-bit experts
  • Supports tool calling in the 4-bit setup

Best for

Readers interested in how very large models can run on consumer Apple hardware. The 2-bit setting is faster but breaks JSON output and tool calling, so 4-bit is the production choice.

Upstream
danveloper/flash-moe
Fork on GitHub
Guo-astro/flash-moe
Upstream stars
4.2k
Category
AI agents and LLM tools
License
No license declaredWithout a license, the author keeps all rights. Ask the upstream owner before reusing the code.
Forked
2026-03-22
Sync status
In syncLast synced 2026-09-29