thewhisper

Optimized Whisper speech-to-text models for real-time streaming and on-device use, including Apple Silicon and NVIDIA GPUs. It is for fast transcription and translation of speech.

Share on XLicense: MIT

Overview

TheWhisper provides optimized Whisper speech-to-text models for streaming and on-device use. It offers open weights on Hugging Face with flexible chunk sizes instead of the original 30 seconds, plus inference engines for NVIDIA GPUs and CoreML engines for macOS on Apple Silicon. It focuses on self-hosting and local transcription, and includes a local REST API with example front ends.

Key features

  • Streaming transcription with flexible chunk size
  • Open weights for Whisper models on Hugging Face
  • Inference engines for NVIDIA GPUs
  • CoreML engines for macOS and Apple Silicon
  • Local REST API with JavaScript and Electron examples

Best for

Developers who want fast, self-hosted or on-device speech transcription on NVIDIA or Apple hardware, for example to build a local note-taking app.

Upstream
TheStageAI/TheWhisper
Fork on GitHub
Guo-astro/thewhisper
Upstream stars
897
Category
Documents, media and content
Language
Python
License
MIT
Forked
2025-11-21
Sync status
In syncLast synced 2026-09-29