fast.cu

A collection of GPU kernels written from scratch in CUDA, including a matrix multiplication kernel for H100 that is reported to match or beat cuBLAS. Useful as a learning reference for GPU performance.

Share on XLicense: MIT

Overview

fast.cu is a collection of GPU kernels written from scratch in CUDA. It includes bf16 matrix multiplication on H100, NVFP4 matrix multiplication on GB300 with ten kernel snapshots, and a sum reduction kernel on H100. The README reports the H100 matrix multiplication and sum kernels at or above cuBLAS and cub in its own tests. Each example builds with make, and a worklog explains the matmul work.

Key features

  • H100 bf16 matrix multiplication kernels with fp32 accumulation
  • GB300 NVFP4 matmul with ten kernel snapshots and a benchmark runner
  • H100 sum reduction over 2^30 elements
  • Simple make targets to build and run each example

Best for

People learning GPU performance tuning who want complete, readable kernels to study. The NVFP4 example needs a GB300 or B300 and CUDA 13.1.

Upstream
pranjalssh/fast.cu
Fork on GitHub
Guo-astro/fast.cu
Upstream stars
624
Category
Developer tools and infrastructure
Language
Cuda
License
MIT
Forked
2025-12-09
Sync status
In syncLast synced 2026-09-29