Derivatives & noise

An edge is a place of rapid change in intensity. Along one row of an image it shows up as a peak in the first derivative. Add a little noise and the derivative drowns in it, unless you smooth before you differentiate.

Image ff click to select a row and pixel
Smoothed f∗gf * g
Derivative dfdx\frac{d f}{dx} without smoothing, 0 = gray
Derivative ddx(f∗g)\frac{d}{dx}(f * g) smoothed first, 0 = gray
Intensity along the selected row ff without noise f∗gf * g detected edges
First derivative dfdx\frac{d f}{dx} ddx(f∗g)\frac{d}{dx}(f * g) f∗dgdxf * \frac{dg}{dx} threshold ±t\pm t
Second derivative d2dx2(f∗g)\frac{d^2}{dx^2}(f * g) smoothed zero crossings where ∣ddx(f∗g)∣≥t\left|\frac{d}{dx}(f * g)\right| \ge t
Filter kernels Gσ(x)G_\sigma(x) Gσ′(x)G'_\sigma(x) sampled dgdx\frac{dg}{dx}

Where edges come from

Edges are discontinuities in an image, and they are caused by discontinuities in the scene:

Edges are much more compact than the full image and still carry most of its structure, so they help to recover the geometry of a scene and the viewpoint, and they are a basis for recognition.

Edges are peaks of the derivative

Take a single row of the image: the intensity as a function of position is a 1D signal f(x)f(x). An edge is a place where ff changes rapidly, so it corresponds to an extremum of the first derivative and to a zero crossing of the second derivative. On pixels the derivative is approximated by a difference; this demo uses the central difference

dfdx[x]≈f[x+1]−f[x−1]2,\frac{d f}{dx}[x] \approx \frac{f[x + 1] - f[x - 1]}{2},

which is a linear filter with the weights [−12    0    12][-\tfrac{1}{2} \;\; 0 \;\; \tfrac{1}{2}] applied by correlation. A dark-to-bright step is positive, a bright-to-dark step negative. The Sobel filter SxS_x uses the same central difference along xx and adds a small smoothing [1    2    1][1 \;\; 2 \;\; 1] along yy.

Difference filters respond strongly to noise

Image noise makes pixels look different from their neighbors, and a difference filter measures exactly that. The larger the noise, the stronger the response. For independent noise with standard deviation σn\sigma_n in every pixel, the central difference has noise with standard deviation σn/2\sigma_n / \sqrt{2}, no matter how gentle the real edges are. A sharp step of height hh still stands out with a peak of h/2h / 2, but a blurry edge spreads the same change over many pixels, and its small derivative disappears in the noise. Thresholding the derivative then finds edges everywhere.

Smooth first

The solution is to smooth the signal with a filter gg, such as a Gaussian, and look for peaks in ddx(f∗g)\frac{d}{dx}(f * g). Differentiation is itself a convolution, and convolution is associative, so smoothing and differentiating can be combined into one filter, the derivative of the Gaussian:

ddx(f∗g)=f∗dgdx,Gσ′(x)=−xσ2 Gσ(x)=−x2π σ3 e−x22σ2\frac{d}{dx}(f * g) = f * \frac{d g}{dx}, \qquad G'_\sigma(x) = -\frac{x}{\sigma^2} \, G_\sigma(x) = -\frac{x}{\sqrt{2\pi}\,\sigma^3} \, e^{-\frac{x^2}{2\sigma^2}}

This saves one operation. In the demo, dgdx\frac{dg}{dx} is the central difference of the sampled Gaussian, so the two curves in the derivative plot agree up to rounding, and it follows the continuous Gσ′G'_\sigma closely. For an image, the Gaussian derivative in one direction is a 2D kernel that differentiates along xx and smooths along yy:

∂Gσ∂x(x,y)=Gσ′(x) Gσ(y)\frac{\partial G_\sigma}{\partial x}(x, y) = G'_\sigma(x) \, G_\sigma(y)

With “smooth along x and y”, the rows above and below the selected one are averaged in as well, which removes much more noise than smoothing along the row alone.

Choosing σ

A larger σ\sigma removes more noise, but it also spreads each edge: the peak gets lower and wider, which makes its position less precise, and nearby edges merge into one. Small σ\sigma finds fine details, large σ\sigma only the large-scale edges. Detecting all real edges and localizing them precisely pull in opposite directions; the Canny edge detector is built around this trade-off.

The second derivative

Every extremum of the first derivative is a zero crossing of the second, so edges can be found as zero crossings, too. But in a flat, noisy region the second derivative crosses zero all the time. The demo only marks crossings where the first derivative reaches the threshold tt, and these coincide with the peaks. Differentiating twice amplifies noise even more, so the second derivative needs smoothing all the more.

Try this