Homogeneous coordinates

Adding one extra coordinate turns the image plane into the plane w=1w = 1 in 3D. Every image point becomes a ray through the origin, and every line becomes a plane through the origin. Drag the four points and watch how lines and their intersection come out of simple cross products, even when the lines are parallel.

Image plane w=1w = 1 drag the points p1…p4p_1 \dots p_4
Homogeneous space (x,y,w)(x, y, w) drag to orbit, scroll to zoom

Points and rays

A 2D point gets a third coordinate 1. Converting back divides by the last coordinate:

(x,y)→[xy1],[xyw]→(xw,yw)(x, y) \to \begin{bmatrix} x \\ y \\ 1 \end{bmatrix}, \qquad \begin{bmatrix} x \\ y \\ w \end{bmatrix} \to \left( \frac{x}{w}, \frac{y}{w} \right)

The same works in 3D, which is how scene points enter the camera's projection matrix:

(x,y,z)→[xyz1],[xyzw]→(xw,yw,zw)(x, y, z) \to \begin{bmatrix} x \\ y \\ z \\ 1 \end{bmatrix}, \qquad \begin{bmatrix} x \\ y \\ z \\ w \end{bmatrix} \to \left( \frac{x}{w}, \frac{y}{w}, \frac{z}{w} \right)

Homogeneous coordinates are scale invariant: for any k≠0k \neq 0,

k[xyw]=[kxkykw]→(kxkw,kykw)=(xw,yw)k \begin{bmatrix} x \\ y \\ w \end{bmatrix} = \begin{bmatrix} kx \\ ky \\ kw \end{bmatrix} \to \left( \frac{kx}{kw}, \frac{ky}{kw} \right) = \left( \frac{x}{w}, \frac{y}{w} \right)

So a point (x,y)(x, y) is really the whole ray (wx,wy,w)(wx, wy, w) through the origin, and the image plane w=1w = 1 is just where we look at it. In the 3D view, every vector on the ray through p1p_1 lands on p1p_1. This is also why a camera cannot recover depth from a single image: the whole projection ray maps to one pixel.

Lines and intersections

A line ax+by+c=0ax + by + c = 0 is written as the vector l=(a,b,c)⊤l = (a, b, c)^\top. A point p=(x,y,1)⊤p = (x, y, 1)^\top lies on the line exactly when p⋅l=0p \cdot l = 0. Like points, lines are only defined up to scale. Two cross products do all the work:

line through two points: l=p1×p2,intersection of two lines: q=l1×l2\text{line through two points: } l = p_1 \times p_2, \qquad \text{intersection of two lines: } q = l_1 \times l_2

In the 3D view each line is a plane through the origin (the tinted triangles). Two such planes meet in a ray, and that ray is the intersection point qq.

Parallel lines and points at infinity

In Cartesian coordinates parallel lines never meet. In homogeneous coordinates the cross product still gives an answer, q=(x,y,0)⊤q = (x, y, 0)^\top. A point with w=0w = 0 cannot be divided out: it is a point at infinity in the direction (x,y)(x, y). In the 3D view, the planes of two parallel lines meet in a ray that lies in the plane w=0w = 0, parallel to the image plane, so it never hits it.

This is exactly what happens in perspective images. Parallel lines in the scene have the same direction, and their images meet at a vanishing point: the image of their common point at infinity. Unlike here, the vanishing point is usually at a finite position in the image, because the camera looks at the lines at an angle.

Why bother? In homogeneous coordinates, operations that are not linear in Cartesian coordinates become matrix products. Perspective division moves to the very end of the pinhole camera's projection. A translation becomes a matrix product too: [x′y′1]=[10tx01ty001][xy1]\begin{bmatrix} x' \\ y' \\ 1 \end{bmatrix} = \begin{bmatrix} 1 & 0 & t_x \\ 0 & 1 & t_y \\ 0 & 0 & 1 \end{bmatrix} \begin{bmatrix} x \\ y \\ 1 \end{bmatrix} The axes follow the image convention: xx right, yy down, and ww points away from the viewer, like the camera's optical axis.

Try this