Entry 08 · Sampling
Stratification
Divide the sample space into cells before you draw. The difference between clumping and coverage.
- Entry
- 08
- Section
- 02 Sampling
- By
- Clare Dunnet
- Read
- 2 min
The problem with pure randomness
A renderer estimating a pixel's colour by firing rays is, in the language of statistics, performing a Monte Carlo integration. Each ray is a sample, and the final colour is the average of what those samples return. The method works, but pure randomness has a flaw: it clusters. Throw a hundred darts at a dartboard entirely at random and some squares will get three darts while neighbouring squares get none. Those gaps are what you see as noise — regions where the estimator simply got unlucky.
Stratification is the fix. You divide the sample space — the pixel area, the hemisphere of possible ray directions, the 2D lens disc, the shutter-open time interval — into equal sub-regions, then draw exactly one sample from each. On a pixel with sixteen samples arranged as a four-by-four stratified grid, every sixteenth of the pixel's area is guaranteed representation. No gaps, no clustering. The variance of the estimator falls faster than it would under pure random sampling, because you have eliminated the lowest-frequency clumping by construction.
The improvement is not free. Stratification requires that you know how many samples you will take before you start, because the grid must be laid out in advance. It also interacts awkwardly with adaptive sampling, where the per-pixel count varies; a fixed grid cannot be cleanly extended by one extra sample without redesigning the cell structure. And it only eliminates clumping at the scale of the cells: within each cell, you still draw the sample at a random position, so fine-scale randomness is preserved — this matters, because a perfectly regular grid would alias rather than noise, trading one artefact for another. The jitter within each cell is deliberate.
What stratification looks like, and what it costs
The residual noise from stratified sampling is visually different from purely random noise. Because low-frequency gaps are suppressed, the error tends to live at higher spatial frequencies — smaller, more uniform grain rather than occasional dark blobs. This is perceptually preferable, and it also compresses better, which matters for multi-sample render buffers.
The deeper payoff comes when stratification is extended across multiple dimensions simultaneously. A path tracer needs well-distributed samples in pixel position, lens position, time, and the sequence of bounces a path takes. The art of low-discrepancy sequences and samplers lies in maintaining good coverage across all those dimensions at once, not just within the pixel footprint. Stratifying one dimension while leaving others random gives only partial gain. Full multi-dimensional stratification — or its close relative, quasi-Monte Carlo sampling — is where the variance reduction becomes substantial enough to matter at production budgets.
A stratified sixteen-sample pixel is not twice as good as a random eight-sample pixel. It is better in a subtler way: quieter at the frequencies that draw the eye first.
More in Sampling
Every entry in this section is listed on the Sampling page, and all twenty-four sit in the full register.