The pinhole model
In an ideal pinhole camera every ray of light passes through a single point, the camera center . A scene point, the camera center and its image lie on one line (the projection rays in the 3D view). The sensor sits at distance behind the pinhole, the focal length. By similar triangles, a point at camera coordinates lands on the sensor at
The minus signs mean the image is upside down. The division by the depth makes far objects look small, and it loses information: every point on a projection ray lands on the same pixel.
The 3D view shows the equivalent virtual image plane at distance in front of instead. There the image is upright and the minus signs disappear, which is the form used from here on. The plane is drawn at , so 1 world unit corresponds to 500 px.
The projection matrix
Division by depth is not linear, but in homogeneous coordinates the whole projection becomes one matrix product. The division only happens at the very end:
is the projection matrix. It is also often called . The pixel is found by dividing by , which here equals the depth . Scaling by any non-zero factor gives the same pixels, so of its 12 entries only 11 are free. This matches the camera's 11 degrees of freedom: 5 in (), 3 for the rotation and 3 for the translation.
Intrinsics
Start with the simplest camera: no rotation (), center at the origin (), square pixels (), principal point at and no skew:
Real cameras relax these assumptions one at a time:
- : the principal point, where the optical axis hits the sensor. The pixel origin is the top-left corner, so it is usually close to, but not exactly at, the image center (dashed cross). Changing it shifts the image in and .
- : the focal length in pixels (focal length in mm divided by the pixel size). They differ only when the pixels are not square (aspect ratio ).
- : skew between the pixel axes. It is essentially zero for modern sensors, but try it to see the frustum shear.
The focal length sets the angular field of view. For a sensor of size :
and must be in the same units. In pixels, the horizontal AFOV uses the image width (640 px) with , and the vertical AFOV uses the height (480 px) with . A shorter focal length gives a wider field of view.
Extrinsics
The extrinsics describe the transformation from world to camera: . The point is rotated first and then translated, so is not the camera position. It is the world origin expressed in camera coordinates. The camera center is the point that maps to :
The rows of are the camera's , and axes expressed in world coordinates (the colored axes at ).
Any 3D rotation can be built from counter-clockwise rotations about the coordinate axes, for example
and has 3 degrees of freedom. Here the pose is set with a center and three angles: yaw about the world axis, pitch up or down, and roll about the optical axis. The camera's orientation in the world is
The turns the optical axis from straight up to horizontal. and are then computed from them.
Try this
- Increase . The field of view narrows, the image plane moves away from , and the cube grows in the image.
- Set px and check that the horizontal AFOV is .
- Unlock “fy = fx” and change alone. The image stretches vertically, as with non-square pixels.
- Move and watch the frustum become asymmetric. The camera has not moved, but the image content shifts.
- Move the camera center until some vertices end up behind the camera. Points with disappear from the image, and the step-by-step panel explains why.
- Change only roll and compare before and after: its third row (the optical axis) stays the same.
- Move to the world origin: becomes zero. Move it elsewhere and compare with . is the world origin seen from the camera, not the camera position.
- Pick a vertex and check the chain by hand.