Matching as filtering
Template matching slides the template over the image like a filter and computes a score for every position. The best matches are the peaks of the score map. The score is only defined where the template fits completely inside the image; the border of the map stays black. There are several choices for the score:
Correlation
Use the template itself as the filter:
High values get multiplied by high values and low values by low values, so a good match gives a strong peak. But the sum is also large wherever the image is simply bright, whether or not it looks like the template. A white area beats a perfect match.
Zero-mean correlation
Subtract the mean of the template first, so that the filter sums to 0:
Flat areas now give 0, whatever their brightness. The score still grows with the contrast of the image, so a strong edge can beat a faint but exact match.
Sum of squared differences
Here smaller is better, and 0 means an exact copy; the score map shows , so that good matches are bright. Large errors are strongly penalized. SSD is simple and precise, but any change of brightness or contrast (bias and gain) between template and image makes the differences large.
Normalized cross-correlation
Subtract the means of both the template and the image patch under it, and divide by their norms:
with the means and over the pixels of the template.
This is the cosine of the angle between the two patches as vectors, after removing their means, so . The normalization removes linear changes of brightness and contrast (bias and gain), with , so it also works when exposure or lighting differ. An inverted copy gives . For a flat patch without any variation the denominator is 0 and the score is undefined; here it is set to 0. NCC is the slowest of the four, since the mean and norm of every image patch are needed, but usually the most reliable.
Finding the matches
Next to a peak, the score is almost as high, because a template shifted by one pixel still fits well. Taking the highest scores directly would report the same match many times. Non-maximum suppression accepts the best position, discards all positions within one template size around it, and repeats.
Template matching only works if the object appears at the same size and orientation as the template. For anything else it needs to be repeated over scales and rotations, or replaced by the local features of later chapters.
Try this
- The template is the black ring on white. With NCC, the dark rings on light backgrounds score near 1 even at low contrast, while the light rings on dark backgrounds come out near : dark spots in the score map.
- Switch to correlation: the best “matches” are now in the white rows, where the image is simply bright.
- Try zero-mean correlation: the rings are found again, but the scores follow their contrast, and edges of letters compete.
- With SSD, only the ring itself scores 0. The other dark rings still come next, but far behind, because their gray values differ.
- Click one of the letters, e.g. an “e” or “o”, to use it as template, and look for its other occurrences with NCC.
- Enlarge the template to 41 pixels, so that it includes the neighboring rows: now the context has to match too, and the other rings score much lower.