uzu

uzu is an inference engine for running AI models in apps, helping developers keep inference on-device instead of relying on remote services. Use it to download a model and build chat features.

Share on XLicense: MIT

Overview

uzu is a Rust inference engine for running AI models inside applications rather than sending prompts to a remote inference service. It addresses the need for app-integrated inference with data kept on the device and without per-request inference charges. Developers configure an engine, select and download a model, then create a chat session and request replies. The project also provides language bindings for Python, TypeScript, and Swift, and supports unified memory on Apple devices.

Key features

  • High-level API for configuring models and creating chat sessions
  • Download models and track download progress
  • Traceable computations to check results against source implementations
  • Language bindings for Python, TypeScript, and Swift; unified memory support on Apple devices

Best for

Developers who want to integrate AI model inference and chat into an application. It is a fit when keeping inference in the app and avoiding remote inference charges are priorities.

Upstream
trymirai/uzu
Fork on GitHub
Guo-astro/uzu
Upstream stars
2.1k
Category
AI agents and LLM tools
Language
Rust
License
MIT
Forked
2026-10-09
Sync status
Updated by last syncLast synced 2026-10-10