Chapter 2 · Digital Image Fundamentals
Needs: Chapter 1 · Introduction
What you’ll learn
- How the human eye senses light, and why what we see is not always what is there (brightness adaptation, the Weber ratio, Mach bands, simultaneous contrast)
- Where images come from physically: the electromagnetic spectrum, sensors, and the illumination–reflectance model
- What sampling and quantization do, how they set spatial and intensity resolution, and how nearest, bilinear and bicubic interpolation resize an image
- The vocabulary of pixel relationships: neighbors, adjacency, connectivity, regions, boundaries, and three distance measures
- The mathematical toolbox used in every later chapter: elementwise vs matrix operations, linearity, arithmetic, set and logical operations, geometric transforms, image transforms, and histograms as probabilities
The big picture
Chapter 1 told us what digital image processing is. This chapter explains what a digital image actually is, starting from the light that leaves a scene and ending with a grid of integers in memory. Along the way it defines the words — pixel, neighbor, resolution, linear operator — that the rest of the series uses without stopping to explain.
Two things make this chapter more than bookkeeping. First, the final consumer of most images is a human, so we need to know where the eye is easily fooled and where it is sharp. Second, every choice made when an image is captured — how many samples, how many levels, which interpolation — leaves a fingerprint that later algorithms must live with. Deep networks did not remove these choices; they mostly moved them into training data and loss functions, as the Modern view section shows.
Elements of visual perception
In plain words: your eye is a camera made of jelly and nerve cells. It is extremely good at seeing differences and edges, and surprisingly bad at judging absolute brightness.
Structure of the eye and image formation
Light enters through the cornea, passes the pupil (whose size the iris controls), and is focused by the lens onto the retina, a thin layer of light receptors at the back of the eye [1]. The lens is flexible: ciliary muscles change its curvature, which changes its focal length so that near and far objects can both be brought into focus. The retinal image is real and inverted, just like the image on a camera sensor.
The retina contains two kinds of receptor [1]:
- Cones — a few million per eye, packed most densely in a small central pit called the fovea. They work in bright light (photopic vision), come in three types with different spectral sensitivities, and give us color and fine detail. Many cones have their own nerve connection, which is why foveal vision is sharp.
- Rods — many times more numerous, spread across the retina, absent from the center of the fovea. They work in dim light (scotopic vision), have no color discrimination, and many rods share one nerve ending, trading detail for sensitivity. This is why objects seen by moonlight look colorless and soft.
The point where the optic nerve leaves the eye has no receptors at all: the blind spot. You do not notice it because the brain fills in the gap — an early hint that perception is an active reconstruction, not a recording.
Brightness adaptation and discrimination
The visual system copes with an enormous range of light levels, from starlight to sunlight. It does this by brightness adaptation: at any moment it operates over a much narrower window centered on the current adaptation level, and that window slides as the overall light changes [1]. Within the window, subjective brightness grows roughly with the logarithm of the physical light intensity, not linearly.
How small a change can we notice? Put a uniform field of intensity in front of an observer and flash a small spot of intensity . The increment that is noticed half of the time is . The ratio
is the Weber ratio, where is the background intensity and is the just-noticeable increment. A small means good discrimination (a small relative change is visible). Over a wide middle range of intensities is roughly constant, which is Weber’s law: what we perceive is the relative change, not the absolute one. At low intensities, where rods dominate, is large and discrimination is poor [1].
This has a direct engineering consequence. Because we perceive relative changes, equally spaced physical gray levels are not equally spaced perceptually. That is one reason images are stored with a nonlinear (gamma) encoding, and one reason 8 bits are usually enough for display.
Mach bands, simultaneous contrast and illusions
Perceived brightness is not a simple function of intensity. Two classic demonstrations make this obvious [1]:
- Mach bands. Show a staircase of flat gray stripes. Each stripe is physically uniform, yet observers report a thin darker band on the dark side of every edge and a thin brighter band on the bright side. The visual system exaggerates edges, a behavior consistent with lateral inhibition between neighboring receptive fields — essentially a built-in sharpening filter.
- Simultaneous contrast. A mid-gray square looks lighter on a dark background and darker on a light one, although the square is identical in both cases. Perceived brightness depends on the surround.

Optical illusions in general — contours that are not drawn, lines of equal length that look unequal — show that the eye fills in and reinterprets. For image processing this matters twice: an algorithm that minimizes squared error may still produce images that look wrong, and an artifact that is numerically small may be very visible if it sits on an edge. This is the starting point of perceptual quality metrics, discussed in the Modern view.
Light and the electromagnetic spectrum
In plain words: light is a wave of energy. Visible light is a thin slice of a much wider spectrum that also includes radio, microwaves, infrared, ultraviolet, X-rays and gamma rays — and images can be made from any of them.
Wavelength, frequency and energy
Electromagnetic radiation can be described as a sinusoidal wave or as a stream of massless particles called photons [1]. The two descriptions are linked by
where is the wavelength (meters), is the frequency (hertz), m/s is the speed of light, is the energy of one photon (joules), and is Planck’s constant. Shorter wavelength means higher frequency and higher energy per photon. That is why gamma rays and X-rays can pass through tissue, while radio waves need large antennas to be detected.
The visible band spans roughly 400 nm (violet) to 700 nm (red). An object’s perceived color is determined by which wavelengths it reflects: a green leaf reflects mostly in the middle of the band and absorbs the rest. Light with no color, only intensity, is called achromatic or monochromatic, and its intensity is called the gray level.
Radiance, luminance and brightness
Three words are easily confused [1]:
- Radiance — the total energy flowing from the light source, measured in watts. It is purely physical.
- Luminance — the amount of that energy an observer perceives, measured in lumens. It weights radiance by the eye’s spectral sensitivity: an infrared source can have high radiance but almost zero luminance.
- Brightness — the subjective sensation of intensity. It cannot be measured directly and, as the previous section showed, depends on adaptation and surroundings.
A sensor’s sensitivity curve plays the role of the eye’s: an image is only as good as the match between the band a sensor detects and the wavelength the object emits or reflects. To image an object, the wavelength must also be comparable to or smaller than the object’s size — one reason electron microscopes, with very short effective wavelengths, resolve far smaller structures than light microscopes.
Image sensing and acquisition
In plain words: a sensor is a tiny light meter that turns light into voltage. Put many of them in a row or a grid and you have a scanner or a camera.
Single, line and array sensors
Every imaging device combines an illumination source, the scene, and a sensor whose material responds to the energy of interest [1]. The sensor’s output voltage is digitized to give one number. Sensors are arranged in three ways:
- Single sensor. One detector (for example a photodiode) plus mechanical motion in two directions sweeps out an image one sample at a time. Slow, but cheap and very precise; used in some high-precision scanners.
- Line (strip) sensor. A row of detectors captures one line at a time, and motion perpendicular to the strip provides the second dimension. Flatbed scanners and airborne push-broom imagers work this way. A ring of sensors around a rotating source is the geometry of computed tomography (CT), where the image is reconstructed rather than read out directly.
- Array sensor. A 2-D grid of detectors, such as a CCD or CMOS chip, captures the whole image at once. This is the digital camera. A lens focuses the scene onto the array, each detector integrates light over an exposure time, and the readout electronics produce one value per detector.
Color cameras usually place a mosaic of red, green and blue filters over the array, so each detector measures only one color; the missing two values per pixel are interpolated (demosaicked) in the camera’s image-processing pipeline [8].
A simple image formation model
We write a monochrome image as a 2-D function whose value is proportional to the energy arriving at . Because it comes from physical energy, . For reflected light, it helps to split into two factors [1]:
where is the illumination (how much light falls on the scene point) and is the reflectance (what fraction of it the surface reflects, from total absorption near 0 to total reflection near 1). Illumination depends on the light source; reflectance is a property of the object. For transmission imaging such as X-rays, a transmissivity takes the place of reflectance.

import numpy as np
from skimage import data, img_as_float
r = img_as_float(data.coins()) # reflectance in [0, 1]
h, w = r.shape
yy, xx = np.mgrid[0:h, 0:w]
i = 0.2 + 0.8 * xx / (w - 1) # illumination ramp: dim left, bright right
f = i * r # what the sensor records
The model is simple but useful. Illumination usually varies slowly across a scene, while reflectance changes abruptly at object edges. Later chapters exploit this: homomorphic filtering (Chapter 4) takes to separate the two, and shading correction (later in this chapter) divides out a known .
Image sampling and quantization
In plain words: a camera measures a continuous scene at a finite number of points (sampling) and rounds each measurement to one of a finite number of levels (quantization).
Basic concepts
A sensor’s output is a continuous voltage waveform over a continuous plane. To make it digital, both the coordinates and the amplitude must become discrete [1]:
- Sampling digitizes the coordinates: we keep values only at a grid of positions.
- Quantization digitizes the amplitude: each kept value is mapped to the nearest of a finite set of levels.
With an array sensor, the number and spacing of detectors fix the sampling; the analog-to-digital converter fixes the quantization. How finely we must sample depends on the image content: by the sampling theorem, a band-limited signal can be reconstructed exactly if it is sampled at more than twice its highest frequency. Chapter 4 develops this properly; Unser’s review gives the modern, spline-based view of sampling and reconstruction [2].
Representing digital images
After sampling and quantization, an image is a matrix of rows and columns:
Each element is a pixel (picture element). Following the book’s convention [1], indexes rows and indexes columns, with the origin at the top-left corner. (Many libraries, including OpenCV’s point APIs, use = (column, row). Always check.)
The number of intensity levels is normally a power of two, , so values lie in . Storing the image takes
where is the number of bits per pixel. A image with needs 262,144 bytes. The dynamic range of an imaging system is the ratio of the largest measurable intensity (set by saturation) to the smallest (set by noise); the contrast of a particular image is the difference between its highest and lowest intensities.
Spatial and intensity resolution
Spatial resolution is the size of the smallest discernible detail. It is meaningful only with units: line pairs per millimeter, or dots per inch (dpi) in printing. Saying an image is “1024 × 1024” says nothing about resolution until we know the physical area it covers. Intensity resolution is the smallest discernible change in level, usually stated as the number of bits .
import numpy as np
from skimage import data
img = data.camera() # 512 x 512, uint8 (k = 8)
small = img[::4, ::4] # keep every 4th sample -> 128 x 128
def quantize(x, bits):
step = 2 ** (8 - bits)
return (x // step) * step + step // 2 # map to the centre of each bin
q3 = quantize(img, 3) # 8 gray levels
print(small.shape, np.unique(q3).size) # (128, 128) 8


Two lessons follow. Reducing samples first destroys fine texture (grass, hair), then edges. Reducing levels first damages smooth regions: the sky in Figure 2.4 breaks into visible steps called false contouring, exactly where the Weber ratio says the eye is most sensitive. Images with a lot of detail tolerate fewer gray levels than smooth ones, because busy texture masks the steps [1].
Image interpolation
Interpolation estimates values at positions where we have no samples. It is used whenever an image is resized, rotated, warped or registered [1], [3], [4].
Nearest neighbor copies the value of the closest sample. It is fast and keeps the original values, but it makes edges blocky and, when shrinking, can drop thin structures entirely.
Bilinear interpolation uses the four nearest samples. For a point at fractional offsets from the top-left sample ,
where , and are the right, lower and lower-right neighbors. Equivalently, , with the four coefficients fixed by the four samples. The result is continuous but slightly blurred.
import numpy as np
def bilinear(img, x, y):
"""Sample img (H, W) at real-valued column x, row y."""
x0, y0 = int(np.floor(x)), int(np.floor(y))
x1, y1 = min(x0 + 1, img.shape[1] - 1), min(y0 + 1, img.shape[0] - 1)
a, b = x - x0, y - y0
top = (1 - a) * img[y0, x0] + a * img[y0, x1]
bot = (1 - a) * img[y1, x0] + a * img[y1, x1]
return (1 - b) * top + b * bot
print(bilinear(np.array([[0., 10.], [20., 30.]]), 0.5, 0.5)) # 15.0
Bicubic interpolation uses the 16 nearest samples and fits . In practice it is implemented as a separable convolution with a piecewise-cubic kernel. Keys’ widely used kernel is [3]
where is the distance (in samples) from the interpolation point to a sample and is a free parameter. Keys showed that gives the best approximation order for this family; libraries differ in the value they actually use. The negative lobes of sharpen edges but can cause slight overshoot (ringing) next to strong edges.
import cv2
from skimage import data
crop = data.camera()[150:214, 230:294] # 64 x 64
up = {name: cv2.resize(crop, None, fx=4, fy=4, interpolation=flag)
for name, flag in [("nearest", cv2.INTER_NEAREST),
("bilinear", cv2.INTER_LINEAR),
("bicubic", cv2.INTER_CUBIC)]}

Drag the slider to compare nearest-neighbor and bicubic upsampling of a crop enlarged to . Look at the camera’s edges and the face profile.

NearestBicubicNo interpolation method can create detail that sampling threw away. Thévenaz, Blu and Unser’s review [4] makes the point precisely: the quality of an interpolator is governed by its approximation order and by how closely its kernel approaches the ideal (sinc) reconstructor, and higher-order B-spline methods outperform classic cubic convolution at similar cost. Learned super-resolution, discussed in the Modern view, goes further by hallucinating plausible detail from training data.
Basic relationships between pixels
In plain words: before we can talk about objects in an image, we need rules for which pixels “touch” each other and how far apart two pixels are.
Neighbors of a pixel
A pixel at has four horizontal and vertical neighbors,
four diagonal neighbors,
and the eight neighbors [1]. At the image border some neighbors fall outside the image; how we treat them (ignore, pad, reflect) is a choice we meet again in filtering.
Adjacency, connectivity, regions and boundaries
Let be the set of intensity values that count as “similar” — for a binary image, ; for a gray image, perhaps all values between 100 and 120. Two pixels and with values in are [1]:
- 4-adjacent if ;
- 8-adjacent if ;
- m-adjacent (mixed) if , or and the set contains no pixel whose value is in .
Why m-adjacency? With 8-adjacency, a diagonal step and a two-step “around the corner” route can both connect the same pair of pixels, so paths become ambiguous. m-adjacency keeps a diagonal link only when no 4-connected route exists, which removes the redundant loops (Figure 2.6, right).

A digital path from to is a sequence of distinct pixels with consecutive pixels adjacent; is the length, and the path is closed if the first and last pixels coincide. Depending on the adjacency used we speak of 4-, 8- or m-paths. For a subset of pixels, and are connected in if a path between them lies entirely in . The set of pixels connected to in is a connected component; if has only one, it is a connected set.
A connected set is called a region. Two regions are adjacent if their union is connected. The boundary (or border) of a region is the set of pixels in that have at least one neighbor outside . Note the difference between a boundary and an edge: a boundary is a global, closed property of a region, while an edge is a local intensity discontinuity that may not close at all (Chapter 10).
The choice of adjacency changes the answer to simple questions such as “how many objects are there?”:
import numpy as np
from skimage.measure import label
b = np.array([[1, 0, 0],
[0, 1, 0],
[0, 0, 1]])
print(label(b, connectivity=1).max(), # 4-connectivity: 3 objects
label(b, connectivity=2).max()) # 8-connectivity: 1 object
A classic consequence: if the foreground uses 8-connectivity, the background should use 4-connectivity (and vice versa), otherwise a closed digital curve may fail to separate its inside from its outside. This is a basic result of digital topology [1].
Distance measures
For pixels , and , a function is a distance metric if with equality only when , , and . Three metrics are standard [1], [5]:
The pixels within a fixed of form a disk; within a fixed , a diamond; within a fixed , a square. The pixels with are exactly , and those with are exactly . and depend only on coordinates, not on pixel values; a distance along m-paths, by contrast, does depend on the values because the paths do.
import numpy as np
p, q = np.array([2, 3]), np.array([5, 7])
d = np.abs(p - q)
print("D_E =", np.hypot(*d), " D_4 =", d.sum(), " D_8 =", d.max()) # 5.0, 7, 4

Rosenfeld and Pfaltz [5] introduced efficient two-pass algorithms to compute such distances to the nearest object pixel for every pixel at once — the distance transform, still used for skeletons, morphology and shape matching (Chapter 9).
Introduction to the mathematical tools used in DIP
In plain words: an image is a matrix, so most image processing is arithmetic on matrices. This section lists the operations we will use again and again, and the one property — linearity — that decides which tools apply.
Elementwise versus matrix operations
An elementwise (array) operation acts pixel by pixel. For two images the elementwise product is
whereas the matrix product follows row-times-column rules. Unless stated otherwise, image operations — including “multiplying two images” — are elementwise [1]. In NumPy, A * B is elementwise and A @ B is the matrix product.
Linear versus nonlinear operations
Consider an operator that maps an input image to an output . It is linear if
for all images and all constants . The two parts of this property are additivity (the output of a sum is the sum of outputs) and homogeneity (scaling the input scales the output). Summing all pixels is linear; taking the maximum is not:
import numpy as np
rng = np.random.default_rng(0)
f1, f2 = rng.random((4, 4)), rng.random((4, 4))
a, b = 2.0, -3.0
print(np.isclose(np.sum(a*f1 + b*f2), a*np.sum(f1) + b*np.sum(f2))) # True
print(np.isclose(np.max(a*f1 + b*f2), a*np.max(f1) + b*np.max(f2))) # False
Linearity matters because linear operators are fully described by their response to simple inputs, and a large body of theory (convolution, the Fourier transform, Chapters 3–5) applies to them. Nonlinear operators such as the median filter or morphology can do things linear ones cannot, but are harder to analyze.
Arithmetic operations
Sum, difference, product and quotient of two images of the same size are all elementwise [1]. Three uses stand out.
Averaging noisy images. Suppose we capture images of a static scene, , where is the clean image and is noise that is uncorrelated between frames, has zero mean and variance . The average satisfies
so the noise standard deviation falls as . Astronomers stack exposures for this reason, and smartphone “night modes” do the same after aligning a burst of frames [7].
import numpy as np
from skimage import data, img_as_float
g = img_as_float(data.camera())
rng = np.random.default_rng(1)
for K in (1, 4, 16, 64):
avg = np.mean([g + rng.normal(0, 0.2, g.shape) for _ in range(K)], axis=0)
print(K, round(float(np.std(avg - g)), 3)) # 0.2, 0.1, 0.05, 0.025

Subtraction reveals differences. Subtracting two frames of a video highlights motion; in medical imaging, subtracting a pre-contrast “mask” image from a post-contrast image leaves mainly the vessels that took up the contrast agent.
Multiplication and division handle shading and masking. If a sensor’s shading pattern is known (for example by imaging a uniform target), the observed image can be corrected as . Multiplying by a binary mask keeps a region of interest and zeroes the rest.
One practical rule: arithmetic results often leave the valid range (differences can be negative, sums can exceed 255). Do the arithmetic in floating point, then rescale. A simple rescaling to is followed by .
Set and logical operations
Treat a binary image as a set of pixel coordinates (the foreground). The usual set operations then apply [1]: union , intersection , complement , and difference . On binary arrays these become the logical operations OR, AND, NOT and AND-NOT. For gray images, union and intersection are commonly defined as elementwise maximum and minimum. Fuzzy sets generalize this further by letting membership vary between 0 and 1.
import numpy as np
from skimage import data
coins = data.coins()
A = coins > 100 # set A: bright pixels
B = np.zeros_like(A); B[:, : coins.shape[1] // 2] = True # set B: left half
union, inter, diff = A | B, A & B, A & ~B
Set operations are the language of mathematical morphology (Chapter 9) and of combining segmentation masks.
Spatial operations
Spatial operations act directly on pixel values and positions. There are three kinds [1].
Single-pixel operations change each value independently: , where is the input intensity and the output. The image negative is an example. These are the intensity transformations of Chapter 3.
Neighborhood operations compute each output from a neighborhood around . Local averaging,
where is an window centered at , is the simplest; Chapter 3 generalizes it to spatial filtering.
Geometric transformations move pixels. They have two steps: a spatial transformation of coordinates, and intensity interpolation to fill the new grid. The most common family is the affine transform. In homogeneous coordinates,
where is an input location, the output location, the encode rotation, scaling and shear, and is a translation. Affine maps preserve straight lines and parallelism, and chaining several is just multiplying their matrices.
Implementations almost always use inverse mapping: for each output pixel, apply the inverse transform to find where it came from in the input, and interpolate there. Forward mapping (pushing each input pixel to the output) leaves holes and collisions.
import cv2
from skimage import data
img = data.camera()
h, w = img.shape
M = cv2.getRotationMatrix2D(center=(w / 2, h / 2), angle=30, scale=0.8) # 2x3 affine
rot = cv2.warpAffine(img, M, (w, h), flags=cv2.INTER_LINEAR) # inverse mapping inside
Image registration
Registration aligns two images of the same scene taken at different times, from different viewpoints or with different sensors. Here the transform is unknown and must be estimated. The classic approach picks tie points (control points) whose positions are known in both images and fits a model to them; for example, a bilinear model
maps input coordinates to reference coordinates , and four tie-point pairs give eight equations for the eight unknowns . Zitová and Flusser’s survey [6] organizes registration into four steps that are still the standard description: feature detection, feature matching, transform model estimation, and resampling.
Vector and matrix operations
A color pixel is a vector, for its red, green and blue values. Vector tools then apply directly; for example, the Euclidean distance from to a reference color ,
is the basis of simple color segmentation (Chapter 6). Stacking an image into an vector also lets any linear operation be written as , a form used in restoration (Chapter 5).
Image transforms
Some problems are easier in a different domain. A 2-D linear transform takes the general form [1]
where is the input image, is the forward transformation kernel, and are the transform-domain variables. The inverse uses an inverse kernel . When the kernel is separable, and the 2-D transform can be done as 1-D transforms along rows and then columns, or in matrix form . The Fourier transform (Chapter 4) and the wavelet transforms (Chapter 7) are the two you will meet most.
Probability methods
Intensities can be treated as random quantities. If pixels in an image have intensity , the probability of that intensity is estimated by
which is the normalized histogram. The mean and variance of intensities are
where measures average brightness and (the standard deviation) measures contrast. Higher moments describe skew and peakedness.
import numpy as np
from skimage import data
img = data.camera()
p = np.bincount(img.ravel(), minlength=256) / img.size # p(z_k)
z = np.arange(256)
m = np.sum(z * p)
sigma = np.sqrt(np.sum((z - m) ** 2 * p))
print(round(m, 2), round(sigma, 2)) # 129.06 73.64
Histogram processing (Chapter 3) and noise modeling (Chapter 5) both start here.
Modern view
The fundamentals in this chapter have not been replaced by deep learning; they have become the interface to it. Networks consume sampled, quantized arrays, are trained with losses that encode a model of perception, and are evaluated on data produced by camera pipelines. The reviews below map this territory.
Sampling and interpolation theory. Unser’s “Sampling — 50 years after Shannon” [2] reframes sampling as approximation in shift-invariant spaces: instead of demanding a perfectly band-limited signal and an ideal sinc reconstructor, choose a practical basis (such as B-splines) and project onto it. Its main takeaway is that a good prefilter before sampling and a good reconstruction kernel after it matter more than the exact sampling rate. Thévenaz, Blu and Unser [4] apply this to image interpolation and compare many kernels on equal terms. They show that approximation order and kernel support explain most of the differences, and that B-spline-based interpolation (with a proper prefilter) offers a better quality-for-cost trade-off than Keys’ cubic convolution [3]. Both papers remain the reference point for anyone who implements resizing.
Learned super-resolution. Wang, Chen and Hoi [9] survey deep super-resolution up to about 2019: network designs, upsampling layers placed early or late, loss functions, and benchmarks. Moser et al. [10] extend the picture to transformers and diffusion models and stress open problems such as flexible scale factors and better evaluation. Two works show how the classical ideas survive inside networks. LIIF [11] represents an image as a continuous function queried at any coordinate — a learned interpolator that can upsample by arbitrary factors, including factors not seen in training. Real-ESRGAN [12] tackles the gap between the bicubic downsampling used to create most training pairs and real degradations, by synthesizing training data with a higher-order degradation model that applies blur, noise, resizing and compression more than once. The lesson for this chapter: bicubic interpolation is both the baseline every method is compared against and, through its use in creating training data, a hidden assumption that can make models fail on real images.
Perception and image quality. Wang et al.’s SSIM [13] turned the observations of the perception section — sensitivity to structure and relative change rather than absolute error — into a full-reference quality index that compares local luminance, contrast and structure. Zhai and Min [14] survey the whole field of perceptual quality assessment: full-, reduced- and no-reference methods, and models for specialized content. Deep features changed this area markedly. Zhang et al. [15] collected human similarity judgments and found that distances between deep network features (LPIPS) agree with people far better than PSNR or SSIM, across architectures and training regimes. Ding et al. [16] then asked a sharper question: if you optimize an image-processing network with each quality metric as the loss, which metric gives results people prefer? Their human study shows that rankings on standard quality benchmarks do not reliably predict usefulness as a training objective. In short, Mach bands and the Weber ratio now live inside loss functions.
Camera sensor pipelines. Ramanath et al. [8] describe the classic in-camera pipeline that sits between the sensor array of this chapter and the compressed file you save: demosaicking, white balance, color transforms, gamma, and compression, each a fixed hand-designed stage. Delbracio et al. [7] tour how mobile computational photography reworked this pipeline over two decades, with burst capture, alignment and merging (the averaging-noisy-images idea at scale), and multi-frame super-resolution. Learned replacements now exist for parts or all of it. Chen et al. [17] trained a network to map raw short-exposure sensor data directly to a clean image in very low light, bypassing the traditional pipeline. Ignatov et al. [18] replaced an entire smartphone image signal processor with one network (PyNET) trained to map raw mosaic sensor data to images resembling those from a DSLR.
What did not change? The pixel grid, the bit depth, adjacency and connectivity (still the basis of connected-component labeling in every segmentation post-processing step), affine warps (now differentiable layers inside networks), and histograms. Knowing them is what lets you spot when a model’s input has been resized with the wrong kernel, quantized too coarsely, or compared with a metric that does not match human judgment.
Key takeaways
- Human vision senses relative changes (Weber’s law) and exaggerates edges (Mach bands); perceived brightness depends on the surround. Good image processing respects these facts.
- Images can be formed from any part of the electromagnetic spectrum; for reflected light, separates illumination from reflectance.
- Sampling sets spatial resolution, quantization sets intensity resolution, and storage is bits. Too few levels cause false contours in smooth regions.
- Interpolation — nearest, bilinear, bicubic — estimates values between samples; it trades speed for smoothness and sharpness but cannot recover lost detail.
- 4-, 8- and m-adjacency define paths, connected components, regions and boundaries; and are distance metrics whose unit balls are diamonds and squares.
- Image operations are elementwise unless stated otherwise; linearity decides whether convolution and Fourier tools apply.
- Averaging independent noisy frames reduces noise standard deviation by ; geometric transforms use inverse mapping plus interpolation.
- Deep learning built on these fundamentals: learned interpolators, perceptual losses and learned camera pipelines all inherit this chapter’s concepts.
Exercises
- A grayscale scanner produces images at 12 bits per pixel. How many megabytes does one uncompressed scan need? How many if it is reduced to 8 bits?
Hint
Use and divide by (or for MiB). At 12 bits: MB. At 8 bits: 7.92 MB.
- Quantize
skimage.data.moon()to 4, 5 and 6 bits. At which level do false contours disappear for you? Then add a small amount of uniform random noise (amplitude half a quantization step) before quantizing at 4 bits. What changes, and why?
Hint
The noise (called dither) breaks up the large flat steps, so the eye averages the fine noise instead of seeing contours. The mean error stays similar but the visible structure changes — another reminder that squared error and perceived quality differ.
- In the binary image below, count the connected components of 1-pixels under 4-, 8- and m-connectivity. Then count the components of 0-pixels under 4- and 8-connectivity. Which pairing gives a consistent “inside/outside” answer?
0 0 0 0 0
0 1 1 1 0
0 1 0 1 0
0 1 1 0 0
0 0 0 0 0
Hint
The 1s are a single component under all three adjacencies, but the ring is closed only if diagonal steps count, because of the gap at the lower right. The center 0 is isolated under 4-connectivity (a hole) but joins the outside under 8-connectivity. 4/4 gives an open curve with a hole; 8/8 gives a closed curve without one. Only 8-foreground with 4-background (or the reverse) is consistent. Check your answer with skimage.measure.label.
- Show that for any two pixels, and find pixel pairs where each inequality becomes an equality.
Hint
Let , . Then . Equality holds when one of is zero.
- Write a function that rotates an image by using forward mapping (push each input pixel to its rounded output position). Display the result for and explain the holes. Then fix it with inverse mapping and bilinear interpolation.
Hint
Rotation slightly stretches the grid spacing in some directions, so some output pixels receive no input pixel. With inverse mapping, every output pixel asks “where did I come from?” and always gets an answer.
- You have 10 noisy frames of a static scene with noise standard deviation 12 gray levels. How many frames would you need to bring it below 2 gray levels? What assumption could fail in a real handheld burst, and how do mobile pipelines handle it?
Hint
requires , so 37 frames. The assumption that the scene and camera are static fails; frames must be aligned (registered) before averaging, and moving objects need robust merging [7].
References
- R. C. Gonzalez and R. E. Woods, Digital Image Processing, 4th ed., Pearson, 2018, Ch. 2. publisher page
- M. Unser, “Sampling—50 years after Shannon,” Proceedings of the IEEE, vol. 88, no. 4, pp. 569–587, 2000. doi
- R. Keys, “Cubic convolution interpolation for digital image processing,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 29, no. 6, pp. 1153–1160, 1981. doi
- P. Thévenaz, T. Blu, and M. Unser, “Interpolation revisited,” IEEE Transactions on Medical Imaging, vol. 19, no. 7, pp. 739–758, 2000. doi
- A. Rosenfeld and J. L. Pfaltz, “Distance functions on digital pictures,” Pattern Recognition, vol. 1, pp. 33–61, 1968. doi
- B. Zitová and J. Flusser, “Image registration methods: a survey,” Image and Vision Computing, vol. 21, no. 11, pp. 977–1000, 2003. doi
- M. Delbracio, D. Kelly, M. S. Brown, and P. Milanfar, “Mobile computational photography: A tour,” arXiv:2102.09000, 2021. arXiv
- R. Ramanath, W. E. Snyder, Y. Yoo, and M. S. Drew, “Color image processing pipeline,” IEEE Signal Processing Magazine, vol. 22, no. 1, pp. 34–43, 2005. doi
- Z. Wang, J. Chen, and S. C. H. Hoi, “Deep learning for image super-resolution: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. arXiv
- B. Moser, F. Raue, S. Frolov, J. Hees, S. Palacio, and A. Dengel, “Hitchhiker’s guide to super-resolution: Introduction and recent advances,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. arXiv
- Y. Chen, S. Liu, and X. Wang, “Learning continuous image representation with local implicit image function,” CVPR, 2021. arXiv
- X. Wang, L. Xie, C. Dong, and Y. Shan, “Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data,” arXiv:2107.10833, 2021. arXiv
- Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004. project page
- G. Zhai and X. Min, “Perceptual image quality assessment: A survey,” Science China Information Sciences, vol. 63, no. 11, 2020. doi
- R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” CVPR, 2018. arXiv
- K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Comparison of full-reference image quality models for optimization of image processing systems,” International Journal of Computer Vision, 2021. arXiv
- C. Chen, Q. Chen, J. Xu, and V. Koltun, “Learning to see in the dark,” CVPR, 2018. arXiv
- A. Ignatov, L. Van Gool, and R. Timofte, “Replacing mobile camera ISP with a single deep learning model,” arXiv:2002.05509, 2020. arXiv