swiftlet

Swiftlet runs large Qwen MoE models locally on Apple devices with a Swift and Metal runtime, reducing RAM use for on-device inference.

Share on XLicense: Apache-2.0

Overview

Swiftlet is a Swift and Metal runtime for Qwen3-Next and Qwen3.5/3.6 Mixture-of-Experts models. It keeps the dense core in memory and streams routed expert weights from storage on demand, so large models can run on ordinary Apple devices while using less RAM. It is used for local inference on macOS and iOS, including supported iPhone setups, with a focus on practical local execution rather than cloud-based serving.

Key features

  • Runs large Qwen MoE models locally on macOS and iOS devices
  • Keeps the dense model core in memory and streams expert weights from storage on demand
  • Targets 35B and 80B model sizes in the project examples
  • Uses a Swift and Metal runtime for on-device inference

Best for

Best for developers and hobbyists working with local AI inference on Apple devices, especially when a model must fit within tight RAM limits. It is a practical option for macOS or iOS setups that need to run large Qwen MoE models without relying on cloud hosting.

Upstream
leonickson1/Swiftlet
Fork on GitHub
Guo-astro/swiftlet
Upstream stars
654
Category
AI agents and LLM tools
License
Apache-2.0
Forked
2026-10-05
Sync status
In syncLast synced 2026-10-11