train-llm-from-scratch

A step-by-step tutorial for training your own transformer language model with PyTorch, from raw text to text generation and alignment. Aimed at learners who want to understand how LLMs are built.

Share on XLicense: MIT

Overview

This is a step-by-step tutorial and set of scripts for training your own transformer language model in PyTorch, based on the Attention Is All You Need paper. It walks from raw text and tokens to a base model, then supervised fine-tuning, a reward model, PPO, DPO and GRPO, ending with evaluation and chat. The algorithms are written by hand in plain PyTorch without trl, peft or transformers. The author says a million or billion parameter model can be trained on a single GPU.

Key features

  • Transformer implemented from scratch in plain PyTorch
  • Path from raw text to a pretrained base model
  • Covers SFT, reward model, PPO, DPO and GRPO
  • No trl, peft or transformers dependency
  • Ends with evaluation and chat

Best for

Learners who want to understand how language models are built and aligned by reading and running every step themselves. It is a teaching project, and the sample output shown is from a small 13 million parameter model.

Upstream
FareedKhan-dev/train-llm-from-scratch
Fork on GitHub
Guo-astro/train-llm-from-scratch
Upstream stars
12k
Category
AI agents and LLM tools
License
MIT
Forked
2026-09-25
Sync status
In syncLast synced 2026-09-29