A Beginner's Guide to Diffusion model
The math behind diffusion models: Chapman-Kolmogorov, Kramers-Moyal, Ito's lemma, and the forward/reverse SDEs.
This post contains affiliate links for tools I use in production. If you buy through them I earn a commission at no extra cost to you. Recommendations are based on my own experience.
Table of Contents
- 1. The Chapman–Kolmogorov Equation
- 2. The Kramers–Moyal Expansion
- 3. Itô’s Lemma
- 4. Constructing an Itô Process
- 5. Diffusion Models: Putting It All Together
- Everything above is the theory layer. Turning it into a working model still means training a score network, choosing a noise schedule, and validating samples in production, which is where the data and observability tooling in the next steps below comes in. The newsletter covers the implementation side of diffusion training as new posts ship.
1. The Chapman–Kolmogorov Equation
1.1. What Is It?
A Markov process is one in which the future depends only on the present state, not on its past. Let
denote the probability density of transitioning from state at time to state at time . The Chapman–Kolmogorov equation tells us that long-time transitions can be computed by “summing over” intermediate states. For three times , the equation reads:
1.2. Step-by-Step Derivation
-
Law of Total Probability:
To compute the probability of being in state at time given at time , we condition on an intermediate state at time : -
Markov Property:
Since the process is Markovian,so we can write:
-
Discrete Case:
In a discrete state space, the integral becomes a sum:This shows that multi-step transitions can be computed by multiplying (or convolving) shorter-step transition probabilities.
2. The Kramers–Moyal Expansion
2.1. From Chapman–Kolmogorov to a Differential Equation
For continuous processes, we study the evolution of the probability density over a small time interval . Starting with:
we probe the evolution by multiplying by a smooth test function and integrating over .
2.2. Taylor Expansion of the Test Function
For a fixed intermediate state , we expand about :
Substituting this expansion into the inner integral yields:
2.3. Defining Moments and Kramers–Moyal Coefficients
Define the th moment over the small interval as:
Assuming these moments scale linearly with , we define the Kramers–Moyal coefficients as:
2.4. Fokker–Planck Equation
Inserting the Taylor expansion into the integrated Chapman–Kolmogorov equation, subtracting the zeroth-order term, dividing by , and letting , we obtain:
In many applications, the coefficients for vanish or are negligible. Truncating at yields the Fokker–Planck equation:
3. Itô’s Lemma
3.1. Real-World Example: Brownian Motion of a Pollen Grain
Imagine you are observing a tiny pollen grain suspended in water. The grain is bombarded by water molecules, and these collisions cause it to move in a seemingly random way. This erratic motion is called Brownian motion.
1. Discrete Modeling
Suppose you record the position of the pollen grain at discrete time intervals of length . At each time step, the grain’s position changes due to:
- Drift: There might be a very slight overall current in the water, which gives a predictable, small shift.
- Random Kicks (Diffusion): The collisions with water molecules produce random displacements.
A discrete update of the position can be written as:
where:
- is the pollen grain’s position at time .
- represents any systematic drift (for example, due to a gentle water current).
- represents the intensity of the random collisions.
- is a standard normal random variable, .
In this context, the term models the small, steady displacement due to the current, and the term models the random displacements caused by molecular collisions.
2. Variance and Scaling
Because is normally distributed with mean 0 and variance 1, the variance of the random term is:
This shows that over a short time interval , the variance of the displacement is proportional to , which is a hallmark of Brownian motion.
3. Taking the Continuous-Time Limit
When we let , the process is observed over infinitely many infinitesimally small time steps. In the limit, by the central limit theorem (and Donsker’s invariance principle), the cumulative effect of the random displacements converges to a continuous-time Brownian motion . Therefore, the discrete update
transforms into the stochastic differential equation (SDE):
Here,
- The term still represents the drift (the effect of the current in the water).
- The term represents the random fluctuations (the effect of molecular collisions), with the important property that
3.2. The Setup
Suppose that satisfies the stochastic differential equation (SDE):
where:
- is the drift,
- is the diffusion coefficient,
- is an increment of standard Brownian motion, with
Let be a function in (i.e., continuously differentiable in and twice in ). We want to compute .
3.3. Derivation
-
Taylor Expansion:
Expand : -
Substitute the SDE:
Replace byThus,
-
Compute :
We haveExpanding:
Using the rules:
we get:
-
Combine the Terms:
Substitute back into the Taylor expansion:
This is Itô’s lemma:
4. Constructing an Itô Process
4.1. The Itô Integral
Suppose you have a function (possibly random, but non-anticipative) and wish to integrate it with respect to Brownian motion . The Itô integral is defined by:
-
Partitioning the Time Interval:
Divide the interval into small subintervals: -
Forming the Riemann Sum:
Let . Then approximate the integral as:The evaluation of at the left endpoint ensures the integral is non-anticipative.
-
Taking the Limit:
As the partition gets finer, the sum converges (in the mean-square sense) to the Itô integral:
4.2. Defining the Itô Process
An Itô process combines a drift part and a diffusion part:
- The drift term is a standard Lebesgue integral.
- The diffusion term is the Itô integral.
4.3. Some Key Properties
- Continuity:
The process is continuous (under suitable conditions on and ). - Quadratic Variation:
The quadratic variation is contributed solely by the diffusion part: - Martingale Component:
Removing the drift, the diffusion part forms a martingale.
5. Diffusion Models: Putting It All Together
Diffusion models use these concepts to describe how data is gradually corrupted by noise and then recovered.
-
Forward Process (Noising):
Starting with a data sample , noise is gradually added by evolving using an Itô process. The evolution of the probability density is governed by the Fokker–Planck equation, obtained by truncating the Kramers–Moyal expansion at second order. -
Reverse Process (Denoising):
To generate or recover data, the process is reversed. The reverse-time stochastic differential equation, derived using time-reversal techniques and Itô’s lemma, uses the gradient of the log-density (the score function) to guide a noisy sample back to the data distribution.
Everything above is the theory layer. Turning it into a working model still means training a score network, choosing a noise schedule, and validating samples in production, which is where the data and observability tooling in the next steps below comes in. The newsletter covers the implementation side of diffusion training as new posts ship.
Next steps: scaling to production
If you take this into production, these are the pieces I would add first.
- Supabase Supabase is a hosted Postgres platform with authentication and storage built in. Postgres with pgvector for embeddings, so you do not run a separate vector store.
- Datadog Datadog aggregates metrics, logs, and traces for infrastructure monitoring. Traces and cost metrics across model calls, so latency and spend are visible per request.
- Vercel Vercel hosts frontend applications with a global edge network and CI/CD. Deploys the frontend and edge functions that sit in front of the model API.
Deploying generative AI models to production
Get the free playbook on shipping generative AI models to production.