JPEG & the DCT

JPEG cuts an image into 8×8 blocks, describes each block by 64 cosine patterns and stores their weights only as precisely as needed. Lower the quality and watch the coefficients turn into zeros, and the block structure show through.

Original click to select a block
Decoded JPEG gray channel only
Difference decoded − original, amplified 4×, gray = 0
Stored coefficients FQF_Q per block log⁡(1+∣FQ∣)\log(1 + |F_Q|), black = 0, DC top left

The JPEG pipeline

A JPEG encoder converts the colors to YCbCr and usually stores the two chroma channels at lower resolution. Each channel is then processed the same way, which this demo shows for a gray image:

  1. Split the image into 8×8 blocks and subtract 128, so that the values are centered on 0.
  2. Transform every block with the discrete cosine transform (DCT).
  3. Divide every coefficient by its step Q(u,v)Q(u, v) and round to an integer.
  4. Read the integers in zigzag order and store them compactly with run-length and Huffman coding.

The decoder multiplies by QQ again and applies the inverse DCT. All steps can be undone exactly except the rounding in step 3: this is where the information, and the file size, is lost.

The discrete cosine transform

Like the Fourier transform, the DCT writes a block as a sum of waves, but it uses only cosines, and its coefficients are real numbers:

F(u,v)=14 C(u) C(v)∑x=07∑y=07f(x,y)cos⁡(2x+1)uπ16cos⁡(2y+1)vπ16F(u, v) = \tfrac14 \, C(u) \, C(v) \sum_{x = 0}^{7} \sum_{y = 0}^{7} f(x, y) \cos\frac{(2x + 1) u \pi}{16} \cos\frac{(2y + 1) v \pi}{16}

with C(0)=1/2C(0) = 1 / \sqrt 2 and C(k)=1C(k) = 1 otherwise, which makes the transform orthonormal: it preserves the energy of the block, and the inverse uses the same cosines. F(0,0)F(0, 0) is the DC coefficient, 8 times the mean of the shifted block; the other 63 are AC coefficients. The DCT behaves like a Fourier transform of the block mirrored at its borders. A mirrored block has no jump at the border, unlike a periodically repeated one, so smooth blocks need only a few low-frequency coefficients. For natural images, most of the energy ends up in the top-left corner: the DCT compacts the energy into few coefficients.

Quantization

Every coefficient is divided by a step size and rounded: FQ(u,v)=round⁡(F(u,v)/Q(u,v))F_Q(u, v) = \operatorname{round}\big(F(u, v) / Q(u, v)\big). The standard luminance table uses small steps for low frequencies and large steps for high frequencies, because the eye is much less sensitive to errors in fine detail. The quality setting scales the whole table: quality 50 uses it as is, quality 100 uses step 1 everywhere, and low qualities use steps up to 255. Most high-frequency coefficients then round to 0. The zigzag order visits the coefficients from low to high frequency, so the zeros collect at the end, and a single end-of-block code (EOB) replaces all of them.

The values panel estimates the storage with the entropy of the stored integers: the fewer different values and the more zeros, the fewer bits per coefficient an ideal code needs. It is a rough estimate; real JPEG also codes runs of zeros and the differences between neighboring DC coefficients.

Artifacts

Try this