nanochat

A minimal, hackable codebase for training your own small ChatGPT-style model on a single GPU node. It covers tokenizing, pretraining, finetuning, evaluation and chat inference at low cost.

Share on XLicense: MIT

Overview

nanochat is a minimal, hackable harness for training your own small ChatGPT-style language model on a single GPU node. It covers tokenization, pretraining, finetuning, evaluation and inference, and lets you chat with the result through a simple command line interface. One setting, depth, which is the number of transformer layers, drives the other hyperparameters automatically. The README says a GPT-2 level model can be trained for about 48 dollars in roughly two hours on an 8XH100 node.

Key features

  • Covers tokenization, pretraining, finetuning, evaluation and inference
  • Runs on a single GPU node with minimal code
  • One depth setting derives the other hyperparameters
  • Chat with the trained model from a simple CLI

Best for

Good for learners and researchers who want to understand and modify the full LLM training pipeline at low cost. It maintains a leaderboard for GPT-2 speedruns.

Upstream
karpathy/nanochat
Fork on GitHub
Guo-astro/nanochat
Upstream stars
58k
Category
AI agents and LLM tools
Language
Python
License
MIT
Forked
2025-10-14
Sync status
In syncLast synced 2026-09-29