Sampling & quantization

A digital image is a continuous distribution of light reduced to finitely many numbers in two ways. Sampling keeps values only on a grid of pixels, and quantization rounds each value to one of a few levels. Lower the resolution and the number of bits and watch blocks and bands appear.

Original 320 × 240 px, click to select a pixel and row
Sampled & quantized each sample is drawn as an s×ss \times s block
Intensity along the selected row thin: original, thick: sampled & quantized, grid: quantization levels

Images as sampled signals

The light arriving at the sensor is a continuous function f(x,y)f(x, y) of the position on the image plane, with one value for gray images or three for color. A digital image stores this function only on a regular grid, and only with a finite set of values:

I[i,j]=Q(f(i⋅s,  j⋅s))I[i, j] = Q\big(f(i \cdot s, \; j \cdot s)\big)

Here ss is the distance between samples and QQ is the quantizer. The image is then a matrix of integers: row jj, column ii, with the origin in the top-left corner.

Sampling

The spatial resolution is the number of samples. Using a cell size ss keeps one value per s×ss \times s cell of the original, so the image shrinks to ⌊W/s⌋×⌊H/s⌋\lfloor W/s \rfloor \times \lfloor H/s \rfloor pixels. Enlarged back to the original size, every sample becomes a visible block.

The value of a sample can be taken at the cell center (an ideal point sample), or averaged over the cell. The average is closer to a real sensor pixel, which collects all light falling onto its area. Detail finer than the grid cannot be represented either way: point samples turn it into false coarse patterns (aliasing, clearly visible in the zone plate), while averaging blurs it away first.

Quantization

With bb bits per value there are L=2bL = 2^b levels. For values v∈[0,1]v \in [0, 1]:

q(v)=round⁡(v⋅(L−1))L−1,Δ=1L−1,∣q(v)−v∣≤Δ2q(v) = \frac{\operatorname{round}\big(v \cdot (L - 1)\big)}{L - 1}, \qquad \Delta = \frac{1}{L - 1}, \qquad |q(v) - v| \le \frac{\Delta}{2}

In smooth gradients, neighboring pixels round to the same level for a while and then jump to the next one. This creates visible steps called banding or false contours. With 8 bits (256 levels) the steps are usually too small to see, which is why 8 bits per channel is the standard for images. Camera sensors often record 10 to 14 bits, which leaves room for later processing such as brightening dark areas.

Storage. An image with w×hw \times h pixels, cc channels and bb bits per channel needs w⋅h⋅c⋅bw \cdot h \cdot c \cdot b bits without compression. Halving the resolution in both directions saves 75 %; dropping from 8 to 4 bits saves 50 %.

Try this