Pinhole camera: intrinsics & extrinsics

A camera maps 3D world points to 2D pixels in two steps. The extrinsics [R∣t][R \mid t] move a point from the world into the camera's coordinate frame. The intrinsics KK project it onto the image. Change the parameters on the right and watch the camera, the image and all matrices update together.

3D world drag to orbit, scroll to zoom, click a vertex to select it
Camera image 640 × 480 px, click a vertex to select it

The pinhole model

In an ideal pinhole camera every ray of light passes through a single point, the camera center CC. A scene point, the camera center and its image lie on one line (the projection rays in the 3D view). The sensor sits at distance ff behind the pinhole, the focal length. By similar triangles, a point at camera coordinates (Xc,Yc,Zc)(X_c, Y_c, Z_c) lands on the sensor at

U=−f XcZc,V=−f YcZcU = -f \, \frac{X_c}{Z_c}, \qquad V = -f \, \frac{Y_c}{Z_c}

The minus signs mean the image is upside down. The division by the depth ZcZ_c makes far objects look small, and it loses information: every point on a projection ray lands on the same pixel.

The 3D view shows the equivalent virtual image plane at distance ff in front of CC instead. There the image is upright and the minus signs disappear, which is the form used from here on. The plane is drawn at Zc=fx/500Z_c = f_x / 500, so 1 world unit corresponds to 500 px.

The projection matrix

Division by depth is not linear, but in homogeneous coordinates the whole projection becomes one matrix product. The division only happens at the very end:

x=K [R∣t] X\mathbf{x} = K \, [R \mid t] \, \mathbf{X} w[uv1]=[fxsu00fyv0001][r11r12r13txr21r22r23tyr31r32r33tz][XYZ1]w \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \begin{bmatrix} f_x & s & u_0 \\ 0 & f_y & v_0 \\ 0 & 0 & 1 \end{bmatrix} \left[\begin{array}{ccc:c} r_{11} & r_{12} & r_{13} & t_x \\ r_{21} & r_{22} & r_{23} & t_y \\ r_{31} & r_{32} & r_{33} & t_z \end{array}\right] \begin{bmatrix} X \\ Y \\ Z \\ 1 \end{bmatrix}

P=K[R∣t]P = K [R \mid t] is the 3×43 \times 4 projection matrix. It is also often called MM. The pixel (u,v)(u, v) is found by dividing by ww, which here equals the depth ZcZ_c. Scaling PP by any non-zero factor gives the same pixels, so of its 12 entries only 11 are free. This matches the camera's 11 degrees of freedom: 5 in KK (fx,fy,s,u0,v0f_x, f_y, s, u_0, v_0), 3 for the rotation and 3 for the translation.

Intrinsics KK

Start with the simplest camera: no rotation (R=IR = I), center at the origin (t=0t = 0), square pixels (fx=fy=ff_x = f_y = f), principal point at (0,0)(0, 0) and no skew:

w[uv1]=[f0000f000010][XYZ1]w \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \begin{bmatrix} f & 0 & 0 & 0 \\ 0 & f & 0 & 0 \\ 0 & 0 & 1 & 0 \end{bmatrix} \begin{bmatrix} X \\ Y \\ Z \\ 1 \end{bmatrix}

Real cameras relax these assumptions one at a time:

The focal length sets the angular field of view. For a sensor of size HH:

AFOV=2arctan⁡H2f\mathrm{AFOV} = 2 \arctan \frac{H}{2f}

HH and ff must be in the same units. In pixels, the horizontal AFOV uses the image width (640 px) with fxf_x, and the vertical AFOV uses the height (480 px) with fyf_y. A shorter focal length gives a wider field of view.

Extrinsics [R∣t][R \mid t]

The extrinsics describe the transformation from world to camera: Xc=R Xw+t\mathbf{X}_c = R\,\mathbf{X}_w + t. The point is rotated first and then translated, so tt is not the camera position. It is the world origin expressed in camera coordinates. The camera center is the point that maps to Xc=0\mathbf{X}_c = 0:

C=−R−1t=−R⊤t(a rotation satisfies R−1=R⊤)C = -R^{-1} t = -R^\top t \qquad \text{(a rotation satisfies } R^{-1} = R^\top \text{)}

The rows of RR are the camera's xx, yy and zz axes expressed in world coordinates (the colored axes at CC).

Any 3D rotation can be built from counter-clockwise rotations about the coordinate axes, for example

Rx(α)=[1000cos⁡α−sin⁡α0sin⁡αcos⁡α],Rz(γ)=[cos⁡γ−sin⁡γ0sin⁡γcos⁡γ0001]R_x(\alpha) = \begin{bmatrix} 1 & 0 & 0 \\ 0 & \cos\alpha & -\sin\alpha \\ 0 & \sin\alpha & \cos\alpha \end{bmatrix}, \qquad R_z(\gamma) = \begin{bmatrix} \cos\gamma & -\sin\gamma & 0 \\ \sin\gamma & \cos\gamma & 0 \\ 0 & 0 & 1 \end{bmatrix}

and RR has 3 degrees of freedom. Here the pose is set with a center CC and three angles: yaw about the world ZZ axis, pitch up or down, and roll about the optical axis. The camera's orientation in the world is

R⊤=Rz(yaw) Rx(pitch−90∘) Rz(roll)R^\top = R_z(\text{yaw}) \, R_x(\text{pitch} - 90^\circ) \, R_z(\text{roll})

The −90∘-90^\circ turns the optical axis from straight up to horizontal. RR and t=−R Ct = -R\,C are then computed from them.

Conventions. This demo uses the camera frame that is common in computer vision: xx right, yy down, zz forward (into the scene). The pixel origin (0,0)(0, 0) is the top-left corner, with uu to the right and vv down. Computer graphics usually has the camera look down −z-z with yy up, so matrices from the two fields often differ by a flip of the yy and zz axes. The world frame is right-handed with ZZ up, and its origin is at a corner of the checkerboard, like a calibration target.

Try this