ds4

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA, and ROCm. It helps run capable open-weight models on consumer hardware without relying on a cloud service.

Share on XLicense: MIT

Overview

DwarfStar is a local inference engine for running a small set of capable open-weight models on consumer hardware. It addresses the need to use large language models without a remote service by supporting Metal, CUDA, and ROCm, and by including SSD streaming for larger models. The project is built around a narrow set of models and uses GGUF files generated by the project itself.

Key features

  • Runs DeepSeek V4 Flash and PRO, GLM, and Qwen models locally
  • Supports Metal, NVIDIA CUDA, and ROCm hardware targets
  • Uses project-generated GGUF files and SSD streaming for larger models
  • Includes integration testing for loading, prompt rendering, tool calls, KV state, HTTP server, and coding agent

Best for

This fits people who want to run open-weight models on their own Mac, CUDA, or ROCm system instead of a remote service. Pick it when the model set and hardware match the project’s supported targets and you want a narrow, local inference setup.

Upstream
antirez/ds4
Fork on GitHub
Guo-astro/ds4
Upstream stars
23k
Category
AI agents and LLM tools
Language
C
License
MIT
Forked
2026-10-03
Sync status
In syncLast synced 2026-10-04