Thin lens & depth of field

A lens brings exactly one distance into focus on the sensor. Points nearer or farther away are spread into a small disk, the circle of confusion. As long as that disk is small enough, it still looks sharp, and the range of distances where it is small enough is the depth of field. Change the lens, the aperture and the focus, and see why a smartphone photo is sharp everywhere while a portrait lens blurs the background.

Side view schematic, distances on a log scale foreground subject background
Simulated photo each layer blurred by its circle of confusion
Circle of confusion c(D)c(D) [mm] acceptable cc shaded: depth of field

The thin lens

A pinhole has to be tiny to give a sharp image, so it collects very little light. A lens can have a large aperture of diameter AA and still be sharp, because it bends all rays from a scene point that pass through it onto one image point. For an ideal thin lens with focal length ff, a point at distance zoz_o in front of the lens is imaged at distance ziz_i behind it, with

1f=1zo+1zi\frac{1}{f} = \frac{1}{z_o} + \frac{1}{z_i}

A point at infinity is imaged at zi=fz_i = f, closer points a bit farther behind the lens. The sensor is at a fixed distance zSz_S behind the lens, so only points at one distance, the focus distance SS, are imaged exactly on it. Focusing a camera means moving the lens to change zSz_S. The aperture is usually given as the f-number N=f/AN = f / A, written as f/1.8, f/8 and so on.

Circle of confusion

A point at another distance DD comes into focus at zD≠zSz_D \ne z_S: in front of the sensor if it is farther away than SS, behind the sensor if it is closer. Its light forms a cone with the aperture as the base and the tip at zDz_D, and the sensor cuts this cone in a disk. By similar triangles, the disk's diameter is A ∣zS−zD∣/zDA \, |z_S - z_D| / z_D. With the thin lens equation this becomes

c(D)=A f ∣D−S∣D (S−f)c(D) = \frac{A \, f \, |D - S|}{D \, (S - f)}

Depth of field

No real image is perfectly sharp, so a small blur is acceptable. A common choice for the largest acceptable circle of confusion is c=sensor diagonal/1500c = \text{sensor diagonal} / 1500, about 0.03 mm for a full-frame sensor. The range of distances where c(D)≤cc(D) \le c is the depth of field (also called depth of focus). Solving c(D)=cc(D) = c gives its limits:

Dnear=AfSAf+c (S−f),Dfar=AfSAf−c (S−f)D_\text{near} = \frac{A f S}{A f + c \, (S - f)}, \qquad D_\text{far} = \frac{A f S}{A f - c \, (S - f)}

When the denominator of DfarD_\text{far} is zero or negative, everything behind the focus distance is sharp enough. This first happens at the hyperfocal distance H=f+Af/cH = f + A f / c: focused there, the depth of field reaches from H/2H / 2 to infinity. A smaller aperture (larger f-number) increases the depth of field, and a larger one makes it shallower.

Smartphone vs large sensor

A smartphone camera and a full-frame portrait lens can have the same f-number, f/1.8, and thus collect the same light per area. But the phone's lens has f=4.3f = 4.3 mm, so its aperture is only A=2.4A = 2.4 mm across, while an 85 mm lens at f/1.8 has A=47A = 47 mm. Focused at 2 m, the phone keeps everything from about 1 m to 21 m sharp, and the portrait lens only about 5 cm around the subject:

The portrait modes of phones imitate the second look in software: they estimate a depth map and blur each pixel according to its depth, much like the layers in this demo.

About the simulation. The scene is made of three flat layers, each at a single distance. Each layer is blurred with a disk the size of its circle of confusion, with 240 pixels across the sensor width, and the layers are then combined from back to front in linear light. Real scenes have continuously varying depth, and the light that a blurred foreground edge would let through from the hidden part of the background is missing. The view stays the same for every lens and sensor, although a real camera would see more or less of the scene (see focal length & field of view).

Try this