Chapter 3 · Intensity Transformations and Spatial Filtering
Needs: Chapter 2 · Digital Image Fundamentals
What you’ll learn
- How a single curve can brighten, darken, invert or stretch an image, and when to pick a log, gamma or piecewise-linear curve.
- Why histogram equalization works, derived step by step, and how histogram matching, CLAHE-style local methods and local statistics build on it.
- What linear spatial filtering is, how correlation differs from convolution, and why separable kernels are fast.
- How smoothing (lowpass) filters, including the nonlinear median filter, remove noise, and how derivative-based (highpass) filters sharpen.
- How to derive highpass, bandpass and bandreject kernels from a lowpass one, and how to chain several methods into one pipeline.
The big picture
This chapter is about improving an image without leaving the pixel grid. Every method here works on the pixel values directly, which is what “spatial domain” means. There are two families. Point operations change each pixel using only its own value. Neighborhood operations (spatial filters) change each pixel using the values around it.
Why it matters: these are the first tools anyone reaches for, in phone cameras, medical viewers, microscopes and as preprocessing for computer-vision models. They are also the conceptual ancestors of the convolutional layer. A CNN’s first layer is a bank of spatial filters, and many learned enhancement networks still output a curve or a filter, as we will see in the Modern view.
Background
Plain version. We write the output image as some operation applied to the input image. If the operation only looks at one pixel at a time, it is a lookup table. If it looks at a small window, it is a filter.
Precise version. Let be the input image and the output, with integer pixel coordinates (the same names as in the convolution tutorial). A spatial-domain process is
where is an operator defined over a neighborhood of . The neighborhood is usually a small square or rectangle centered on . The operator is applied at every location by moving the neighborhood’s center from pixel to pixel.
When the neighborhood is a single pixel (), depends only on the value at that pixel. We then call it an intensity transformation (also gray-level or point transformation) and write it with scalars:
where is the input intensity at a pixel, is the output intensity at the same pixel, and both lie in for an image with intensity levels ( for 8-bit images). Because takes only values, any can be stored as a lookup table of length , which is why point operations are extremely fast.
Basic intensity transformation functions
Plain version. Draw a curve from input brightness (horizontal axis) to output brightness (vertical axis). Where the curve is steep, nearby gray levels get pulled apart, so contrast increases. Where it is flat, they get squeezed together, so contrast drops.
The slope is the local contrast gain: a small difference between two neighboring pixels becomes roughly after the transformation. Every curve below is a choice about where to spend that gain.

Image negatives
Dark becomes light and light becomes dark. The slope is everywhere, so contrast magnitude is unchanged; only the polarity flips. Negatives help when small bright details sit inside large dark regions, because people often judge gray detail more easily on a light background.
Log transformations
where is a scale constant, usually chosen so that the largest input maps to , that is . The keeps out of the formula. The slope is large for small and small for large . So the log curve expands dark values and compresses bright ones.
Its classic use is displaying data with a huge dynamic range, such as the magnitude of a Fourier spectrum, whose values can span many orders of magnitude. Scaled linearly, only the brightest point is visible. After the log, the structure appears (Figure 3.2, bottom row). The inverse log (exponential) curve does the opposite.
Power-law (gamma) transformations
where and are constants (with normalized to , take ). For the curve bows upward and brightens the dark range, like the log. For it bows downward and darkens it. Unlike the log, one parameter gives a whole family of curves (Figure 3.1, middle).
Gamma matters beyond enhancement. Displays and cameras have power-law responses, and image files usually store gamma-encoded values rather than values proportional to light. Pre-compensating for a device’s response is called gamma correction. One practical consequence: averaging, blurring or resizing gamma-encoded pixels is not the same as doing it on linear light, which is why careful pipelines linearize first.
import numpy as np
from skimage import data, img_as_float
r = img_as_float(data.camera()) # intensities scaled to [0, 1]
negative = 1.0 - r # s = (L-1) - r, in normalized form
log_t = np.log1p(r) / np.log(2.0) # s = c log(1 + r), c chosen so s(1) = 1
bright = r ** 0.4 # gamma < 1 lifts the shadows
dark = r ** 2.5 # gamma > 1 deepens them

Piecewise-linear transformations
Instead of one formula, we can join straight segments. This gives precise control, at the cost of more parameters.
Contrast stretching. Choose two control points and and connect with straight lines. The middle segment has slope . If and , the curve is monotonic and never reverses the order of gray levels. Two special cases are worth remembering:
- , , , maps the image’s actual range onto the full range (a min–max stretch). In practice, use the 1st and 99th percentiles instead of the min and max, so a few outlier pixels do not control the result.
- , , produces a binary image: a threshold at .
def stretch(img, r1, s1, r2, s2, L=256):
lut = np.interp(np.arange(L), [0, r1, r2, L - 1], [0, s1, s2, L - 1])
return lut.astype(np.uint8)[img]
img = data.camera()
lo, hi = np.percentile(img, [1, 99])
out = stretch(img, lo, 0, hi, 255) # robust min-max stretch
Intensity-level slicing. Highlight one band of intensities . Either set the band to a high value and everything else to a low value (a binary result), or brighten the band and leave everything else unchanged. Slicing is useful when the gray range of interest is known, such as one tissue type or one material in an X-ray.
Bit-plane slicing. An 8-bit pixel value is a sum of bits, with . Collecting bit of every pixel gives a binary image called bit plane . Plane 7, the most significant bit (MSB), is exactly a threshold at 128. The low planes look like noise because they carry fine variations. Keeping only the top four planes reconstructs an image that is very close to the original, with a maximum error below levels (Figure 3.3). This explains why bit planes matter for compression and why the lowest planes are a classic place to hide watermarks.
img = data.camera() # uint8
planes = [(img >> b) & 1 for b in range(8)] # planes[7] is the MSB
top4 = sum(planes[b].astype(np.uint8) << b for b in range(4, 8))
print(np.abs(img.astype(int) - top4).max()) # 15

Histogram processing
Plain version. A histogram counts how many pixels have each brightness. A dark image has its counts piled up on the left; a washed-out one has them squeezed into a narrow band. Histogram processing chooses a curve automatically from those counts, so that the output histogram has a shape we want.
Precise version. For an image with levels , , let be the number of pixels with value . The normalized histogram
is an estimate of the probability that a randomly chosen pixel has value . It sums to 1.
Histogram equalization
The goal is a curve that makes the output histogram as flat as possible, so every gray level is used about equally often.
Derivation (continuous case). Treat intensity as a continuous random variable with probability density function (PDF) . Assume is strictly increasing on that interval. Then a pixel lands in a small interval exactly when its output lands in with . Equal probabilities give , so
Now choose as the scaled cumulative distribution function (CDF) of :
where is a dummy integration variable. By the fundamental theorem of calculus, . Substituting,
The output PDF is uniform, whatever the input PDF was. The CDF is the one monotonic curve that spreads probability evenly: it is steep where many pixels live (so they get pulled apart) and flat where few pixels live.
Discrete version. Replace the integral by a sum:
and round to the nearest integer. Because many input levels can round to the same output level and no level can be split, the discrete result is only approximately flat. The output histogram usually shows tall spikes separated by gaps (Figure 3.4, middle). What is guaranteed is that the output spans the full range and that the CDF of the output is close to a straight line.
def equalize(img, L=256):
hist = np.bincount(img.ravel(), minlength=L) # n_k
cdf = np.cumsum(hist) / img.size # sum of p_r(r_j), j <= k
lut = np.round((L - 1) * cdf).astype(np.uint8) # s_k
return lut[img]
OpenCV’s cv2.equalizeHist uses a slightly different normalization, , so the darkest occupied level maps to 0. Results differ by at most a level or so.
Drag the slider to compare a deliberately low-contrast image with its globally equalized version. Notice that equalization is not “free”: it exaggerates the large smooth areas and grain, because it spends contrast wherever the pixel counts are high, not where the interesting detail is.

InputEqualized
Histogram matching (specification)
Sometimes a flat histogram is the wrong target. We may want an image to look like a reference image, or to have a histogram shape chosen by hand. Histogram matching maps the input so that its histogram approximates a specified PDF .
The trick is to use equalization twice. Equalizing the input gives with a uniform . Equalizing the target distribution gives
which also yields a uniform variable. Two uniform variables can be matched, so we set and solve:
Here is the output intensity, is the scaled CDF of the desired distribution and is a dummy variable. In the discrete case is a staircase and may not be invertible, so for each input level we take the smallest with . That is a single searchsorted call:
def match(img, ref, L=256):
cdf_in = np.cumsum(np.bincount(img.ravel(), minlength=L)) / img.size
cdf_ref = np.cumsum(np.bincount(ref.ravel(), minlength=L)) / ref.size
# For each input level, the smallest z with G(z) >= T(r).
lut = np.searchsorted(cdf_ref, cdf_in).clip(0, L - 1).astype(np.uint8)
return lut[img]
skimage.exposure.match_histograms implements the same idea, including per-channel matching for color images.
Exact histogram matching
Both methods above are lookup tables, so all pixels with the same input value receive the same output value. If 30% of the pixels share one value, no lookup table can split them, and the output histogram cannot be exactly flat or exactly match a target.
Exact histogram specification removes this limit by making every pixel distinguishable first [5]. Each pixel gets a vector of features: its own value, then the mean over a small neighborhood, then the mean over a larger one, and so on. Sorting pixels lexicographically by this vector produces a strict ordering in practice, because ties in one feature are broken by the next. Then assign output values in that order: the first pixels get level 0, the next get level 1, and so on, where is the exact count the target histogram requires. The result has precisely the requested histogram. Coltuc, Bolon and Chassery showed that ordering by local means of increasing size gives visually sensible tie-breaking, because pixels in brighter surroundings are pushed slightly upward [5].
Local histogram processing
Global methods use one curve for the whole image. If a dark corner holds the important detail but contains few pixels, the global histogram barely notices it.
Local histogram equalization computes the transformation from the histogram of a neighborhood around each pixel and applies it only to the center pixel, then moves on. Adaptive histogram equalization (AHE) and its variants were studied systematically by Pizer and colleagues, including the problem that AHE over-amplifies noise in nearly uniform regions, where the local histogram is a single spike and the CDF turns into a near-vertical step [2].
Contrast-limited AHE (CLAHE) fixes this by limiting the slope of the mapping. Since the slope of an equalization curve is proportional to the histogram height, clipping each local histogram at a ceiling and redistributing the clipped counts across all bins caps the contrast gain. To make it fast, the image is split into tiles; a mapping is computed per tile, and each pixel’s output is interpolated bilinearly from the four nearest tile mappings, which avoids visible tile seams. Zuiderveld’s Graphics Gems chapter gave the widely reused implementation [3], and Reza described a real-time hardware realization [4].
import cv2
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8))
out = clahe.apply(data.moon()) # uint8 in, uint8 out
In scikit-image, the equivalent is skimage.exposure.equalize_adapthist(img, clip_limit=0.01).
Using histogram statistics for image enhancement
A histogram is a distribution, so its moments summarize it. With the normalized histogram, the mean and variance are
where measures average brightness and (the standard deviation) measures contrast. Computed over a neighborhood centered on , they become the local mean and local standard deviation .
These give a simple, targeted enhancement rule. Suppose we want to boost only regions that are dark and have some but not too much contrast (to skip both flat background and strong edges). With global mean and global standard deviation , we multiply a pixel by a gain only if
and leave it alone otherwise. Here are user-chosen fractions (for example ). The local statistics need only two box filters, one on and one on , using :
from scipy.ndimage import uniform_filter
f = img_as_float(data.moon())
m = uniform_filter(f, size=3) # local mean
sd = np.sqrt(np.maximum(uniform_filter(f**2, size=3) - m**2, 0)) # local std
mG, sG = f.mean(), f.std()
mask = (m <= 0.5 * mG) & (sd >= 0.02 * sG) & (sd <= 0.5 * sG)
g = np.where(mask, np.clip(3.0 * f, 0, 1), f) # gain E = 3
The same “local mean and local variance from box filters” pattern reappears in the guided filter [8], one of the edge-preserving filters discussed in the Modern view.
Fundamentals of spatial filtering
Plain version. A spatial filter slides a small grid of weights (a kernel) across the image. At each position, it multiplies the weights by the pixels underneath and adds the products. The result is the new value of the center pixel.
The convolution tutorial covers the mechanics with a from-scratch NumPy loop, OpenCV calls and border handling. Here we focus on the theory the rest of the chapter needs.
Linear spatial filtering and correlation
Using that tutorial’s notation, for a kernel the correlation of with is
where are offsets from the kernel’s center and is the kernel’s half-width. Odd sizes give the kernel a well-defined center. The filter is linear because the output is a linear combination of input values: filtering gives times the result for plus times the result for . It is also shift-invariant because the same weights are used everywhere.
Correlation versus convolution
Convolution is the same computation with the kernel rotated by 180°:
The two coincide for kernels that are symmetric under a 180° rotation (box, Gaussian, the Laplacian), and differ for antisymmetric ones (Sobel, any directional derivative), where they produce the same response with opposite sign. The useful difference shows up with an impulse (a single 1 among zeros): correlating a kernel with an impulse produces a rotated copy of the kernel, while convolving produces an exact copy. This is why convolution, not correlation, is the operation with clean algebra:
- commutative: ,
- associative: , so a chain of filters can be merged into one kernel,
- distributive: .
Associativity is what the next two subsections rely on. As noted in the convolution tutorial, cv2.filter2D computes correlation, and most deep-learning “convolution” layers also compute correlation; since their weights are learned, the flip does not matter there.
Separable kernels
A kernel is separable if it is the outer product of a column vector and a row vector :
Then : filter every row with , then every column with . For an kernel this costs multiplications per pixel instead of , so the speedup is , which is for a kernel and grows linearly with size. A matrix is an outer product exactly when its rank is 1, which also gives a test: compute the singular value decomposition and check that only one singular value is nonzero. Box and Gaussian kernels are separable; so is the Sobel kernel, which is a smoothing vector times a differencing vector.
g1 = cv2.getGaussianKernel(7, 1.5) # 7x1 column, sums to 1
K = g1 @ g1.T # 7x7 Gaussian kernel
print(np.linalg.matrix_rank(K)) # 1 -> separable
f = img_as_float(data.camera()).astype(np.float32)
a = cv2.filter2D(f, -1, K.astype(np.float32))
b = cv2.sepFilter2D(f, -1, g1, g1) # two 1-D passes
print(np.abs(a - b).max()) # ~1e-7
Spatial and frequency domains
The convolution theorem states that convolution in space is multiplication in frequency: if denotes the Fourier transform, then . So every linear shift-invariant filter has two equivalent descriptions: a kernel in space, and a transfer function that says how much of each frequency survives. Smoothing kernels have transfer functions that are large near zero frequency and small at high frequencies: they are lowpass. Sharpening kernels do the opposite: they are highpass. Chapter 4 develops this view properly. For now, the vocabulary is enough: “lowpass” means “smooth”, “highpass” means “edges and detail”.
Small kernels are usually cheaper to apply in space. Very large kernels are often cheaper in frequency via the fast Fourier transform.
Constructing spatial filter kernels
There are three common recipes:
- From a mathematical property. Averaging gives the box kernel. Derivatives give difference kernels (the Laplacian, Sobel).
- By sampling a continuous function. Sample a 2-D Gaussian on the integer grid, then normalize the weights to sum to 1.
- From a frequency specification. Design in the frequency domain and take its inverse transform to get a spatial kernel, usually truncated to a manageable size.
Two sanity checks apply to all of them. A smoothing kernel’s weights should sum to 1, so flat regions keep their brightness. A derivative kernel’s weights should sum to 0, so flat regions give 0.
Smoothing (lowpass) spatial filters
Plain version. Replace each pixel by an average of its neighborhood. Random noise is pushed toward the average and fades, but sharp edges are also smeared, because the average straddles both sides.
Box filter
The box kernel has every weight equal to . It is the simplest lowpass filter and is separable, and with a running sum it costs a constant amount of work per pixel regardless of size. Its weakness is its frequency response: the transform of a box is a sinc-like function with side lobes, so it lets some high frequencies through with flipped sign, and it smears in a square, direction-dependent way. Small text and thin lines can turn into blocky streaks.
Gaussian filter
The Gaussian kernel samples
where are offsets from the kernel’s center, (the standard deviation) sets the amount of blur, and is chosen so the weights sum to 1. It has three properties that make it the default smoothing kernel:
- Isotropy. It depends only on the distance , so it blurs equally in every direction.
- Separability. .
- A smooth, non-negative frequency response. Its Fourier transform is also a Gaussian, so there are no side lobes.
Since about 99.7% of a 1-D Gaussian’s mass is within , a kernel of size about , rounded up to the next odd integer, captures almost all of it. Larger kernels waste work, smaller ones truncate the tails. Convolving two Gaussians with and gives a Gaussian with , so repeated small blurs compose predictably.
Order-statistic (nonlinear) filters
Order-statistic filters sort the values in the neighborhood and pick one. They are nonlinear: no kernel can represent them, and the convolution theorem does not apply.
- Median filter: output the middle value (the 50th percentile). It is the standard remedy for impulse noise (salt-and-pepper), isolated pixels forced to black or white. An impulse is an extreme value, so it sorts to the end of the list and is almost never the median. As long as fewer than half the pixels in the window are corrupted, the median ignores them. It also preserves step edges far better than a linear average of the same size.
- Max filter (100th percentile): finds the brightest point; it removes pepper (dark) noise but grows bright specks.
- Min filter (0th percentile): the reverse; removes salt (bright) noise but grows dark specks.
Max and min filters reappear in Chapter 9 as grayscale dilation and erosion.
from scipy import ndimage as ndi
f = img_as_float(data.camera()).astype(np.float32)
noisy = f.copy()
u = np.random.default_rng(0).random(f.shape)
noisy[u < 0.05], noisy[u > 0.95] = 0.0, 1.0 # 10% salt-and-pepper
box = cv2.blur(noisy, (5, 5))
gauss = cv2.GaussianBlur(noisy, (0, 0), sigmaX=1.5)
med = cv2.medianBlur(noisy, 3)
mn, mx = ndi.minimum_filter(noisy, 3), ndi.maximum_filter(noisy, 3)

Sharpening (highpass) spatial filters
Plain version. Smoothing averages; sharpening does the opposite and highlights differences. Differences between neighbors are derivatives, so sharpening is built from derivatives.
First and second derivatives
For a 1-D signal on a pixel grid, the simplest approximations are
Figure 3.6 applies both to a hand-made profile with a ramp, a one-pixel-wide line and a step. Reading it gives the rules that decide which derivative to use:
- In flat regions, both derivatives are 0.
- Along a ramp, the first derivative is nonzero all the way, which makes thick responses. The second derivative is nonzero only at the ramp’s start and end.
- At a thin line, the second derivative responds strongly with a double response (positive, strong negative, positive), and its magnitude is larger than the first derivative’s. Fine detail is exactly what sharpening should emphasize.
- At a step, the second derivative produces a positive and a negative value next to each other. The zero crossing between them locates the edge precisely, which Chapter 10 uses for edge detection.
So for enhancing fine detail, the second derivative is the more natural choice.

The Laplacian
The simplest isotropic second-derivative operator (one whose response does not change when the image is rotated) is the Laplacian:
Adding the 1-D difference above in and in gives
which is correlation with the kernel
Adding the two diagonal directions as well gives a version with in all eight outer cells and in the center, which is isotropic in 45° steps instead of 90° steps. Flipping all the signs gives equally valid kernels. The weights sum to 0, so flat regions vanish and only changes survive.
Sharpening with the Laplacian. The Laplacian image is mostly gray (zero) with bright and dark fringes at edges. Adding it back to the original puts those fringes on the edges:
where for kernels with a negative center (like above) and for kernels with a positive center. The sign matters: with the wrong sign, edges get blurred instead of sharpened. Because convolution is distributive, is itself a single kernel, with center and above, below and to each side, which is exactly the sharpen preset in the playground below.
Try the gaussian and box-blur presets for smoothing, then laplacian to see the bare second-derivative response. Changing the center of the sharpen kernel from 5 to 6 makes the weights sum to 2 and brightens the whole image, which shows why the sum matters.
Unsharp masking and highboost filtering
A trick from photographic darkrooms: subtract a blurred copy to isolate the detail, then add the detail back.
where is a blurred (lowpass) version of , is the unsharp mask (the detail the blur removed), and is a weight. With this is unsharp masking. With it is highboost filtering. With the effect is gentle. Because is a highpass signal, this is another way of adding highpass detail, and if is a Gaussian blur, the blur’s sets the size of the details that get boosted.
Large causes visible halos: overshoot on both sides of strong edges, and the output may leave the valid range, so clip it.
The gradient and Sobel operators
First derivatives are used through the gradient, the vector of partial derivatives
which points in the direction of the fastest increase of intensity. Its length, the gradient magnitude
is large at edges and near 0 in flat regions. The right-hand approximation avoids the square root but is not isotropic.
The Sobel operators estimate and with kernels:
Each is separable: , a central difference in one direction multiplied by a small smoothing in the other. The in the middle gives more weight to the center row or column, which reduces the operator’s sensitivity to noise. (The convolution tutorial uses as its edge-detection example.) The gradient magnitude is rarely used to sharpen directly; it is more often used to find edges, or as a mask that says where sharpening should be allowed, as in the combined example below.
f = img_as_float(data.chelsea()[..., 1]).astype(np.float32)
lap_k = np.array([[0, 1, 0], [1, -4, 1], [0, 1, 0]], np.float32)
lap = cv2.filter2D(f, -1, lap_k)
sharp = np.clip(f - lap, 0, 1) # c = -1 because the center is negative
blur = cv2.GaussianBlur(f, (0, 0), 2)
mask = f - blur # g_mask
unsharp = np.clip(f + 1.0 * mask, 0, 1) # alpha = 1
highboost = np.clip(f + 3.0 * mask, 0, 1) # alpha = 3
gx = cv2.Sobel(f, cv2.CV_32F, 1, 0, ksize=3)
gy = cv2.Sobel(f, cv2.CV_32F, 0, 1, ksize=3)
M = np.hypot(gx, gy) # gradient magnitude

Always filter in floating point. A uint8 pipeline clips the negative values of a Laplacian or a Sobel response to 0 and silently loses half the information.
Highpass, bandreject and bandpass filters from lowpass filters
Plain version. If you know how to keep the smooth part, you also know how to keep everything except the smooth part: subtract.
Let be the unit impulse kernel: a 1 at the center and 0 elsewhere, so . Its transfer function equals 1 at every frequency (an all-pass filter). If is a lowpass kernel with transfer function , then
Unsharp masking is exactly this: .
To isolate a band of frequencies, combine two lowpass filters. Let be a Gaussian with small (a high cutoff, so it keeps more detail) and a Gaussian with larger (a low cutoff). Then
The bandpass kernel keeps frequencies that passes but blocks. The bandreject kernel is the sum of a lowpass and a highpass and removes only that middle band. The bandpass is the difference of Gaussians (DoG), which is also a good approximation of the Laplacian of a Gaussian and is the basis of the scale-space detectors in Chapter 12. The weights confirm the design: highpass and bandpass kernels sum to 0, and the bandreject kernel sums to 1.
def gauss_kernel(sigma):
n = 2 * int(np.ceil(3 * sigma)) + 1 # about 6 sigma, odd
g = cv2.getGaussianKernel(n, sigma)
return g @ g.T
def delta_like(K):
D = np.zeros_like(K); D[K.shape[0] // 2, K.shape[1] // 2] = 1
return D
lp2 = gauss_kernel(4.0) # low cutoff
lp1 = gauss_kernel(1.0) # high cutoff
lp1 = np.pad(lp1, (lp2.shape[0] - lp1.shape[0]) // 2) # zero-pad to the same size
hp = delta_like(lp1) - lp1 # highpass, sums to 0
bp = lp1 - lp2 # bandpass (DoG), sums to 0
br = delta_like(bp) - bp # bandreject, sums to 1
Combining spatial enhancement methods
Plain version. Real images rarely have only one problem. A useful enhancement is usually a short pipeline, and the order of the steps matters.
Here is a worked example. The input (Figure 3.8a) is a portrait that is underexposed, has a little Gaussian noise, and has 4% salt-and-pepper impulses. We want a brighter, crisper picture without amplifying the noise.
- Remove impulses first, with a 3×3 median filter (b). Impulses are extreme values; any later contrast boost or sharpening would make them worse, and a linear blur would only spread them. The median removes them while keeping edges.
- Lift the shadows with a gamma of 0.45 (c). A point operation is the right tool for a global brightness problem. Doing it after the median means the curve is not distorted by black and white impulses.
- Build an edge-weight map (d). Blur the image with a Gaussian (), compute the Sobel gradient magnitude of the blurred image, and normalize it to , where is the 90th percentile of . Because it is computed after smoothing, is near 1 on real edges and near 0 on the residual noise in flat areas.
- Apply unsharp masking gated by that map (f): , with the gamma-corrected image, its Gaussian blur, , and the product taken pixel by pixel.
Panel (e) shows plain unsharp masking with the same : the background and the dark collar get grainy, because the noise is high-frequency detail too. The gated version (f) sharpens the face, hair and the suit’s outlines nearly as much, while leaving the smooth background clean. This pattern, using a smoothed first-derivative image to decide where a second-order sharpening is applied, is a hand-made version of the edge-aware processing that the Modern view describes.

Each step uses one idea from this chapter. The order follows three rules of thumb: remove outliers before any step that amplifies differences, fix global tone with a point operation, and restrict sharpening to where the signal is stronger than the noise.
Modern view
The techniques in this chapter are decades old, and they are still everywhere: CLAHE is a standard preprocessing step in medical imaging and remote sensing, gamma and tone curves are in every camera pipeline, and Gaussian, median and Sobel filters are in every vision library. Research since has pushed in three directions: smarter contrast enhancement, filters that smooth without blurring edges, and learned replacements for both. The surveys below are good entry points.
Contrast enhancement and histogram methods. Pizer et al. [2] remains the reference analysis of adaptive histogram equalization: it describes how local mappings behave, why AHE over-enhances noise in uniform regions, and the variants (including contrast limiting and interpolation between tile mappings) that made the method practical. Zuiderveld’s chapter [3] is the compact CLAHE implementation that later libraries follow, and Reza [4] shows the same algorithm mapped to real-time hardware. On the histogram-specification side, Coltuc, Bolon and Chassery [5] show how to obtain an exact target histogram by strictly ordering pixels with local means. For a broader map of the field, Qi et al. [11] survey image-enhancement methods from the past two decades, organizing them into supervised and unsupervised algorithms and covering quality evaluation as a third pillar; it is useful for its side-by-side discussion of the limitations, merits and drawbacks of each family.
Edge-preserving filters. The weakness of the box and Gaussian filters is that they average across edges. The bilateral filter of Tomasi and Manduchi [6] fixes this by multiplying the spatial Gaussian weight by a second weight that decays with intensity difference, so a pixel on the other side of an edge contributes little. The survey by Paris, Kornprobst, Tumblin and Durand [7] explains its behavior, fast implementations and applications such as tone mapping and detail enhancement, where an image is split into a smooth base layer and a detail layer. The guided filter of He, Sun and Tang [8] assumes that the output is a local linear function of a guidance image. It is computed only with box filters of local means and variances, so its cost does not depend on the radius, and it behaves better than the bilateral filter near strong edges, where the bilateral filter can produce gradient-reversal artifacts in detail enhancement. Milanfar’s tutorial [9] unifies bilateral, non-local means and related filters as data-adaptive weighted averages, where the kernel weights depend on the image itself, and analyzes them with linear-algebra tools. Gingichashvili and Lischinski [10] point out that the many edge-preserving filters had been compared mostly by eye, and propose a systematic methodology and common baseline for evaluating them.
How deep learning changed things. The most interesting finding is how much of this chapter survives inside learned systems. Many successful enhancement networks do not output pixels directly; they predict the parameters of a classical operator and then apply it.
- Learned curves. Zero-DCE [13] (2020) predicts pixel-wise, high-order curves (built by iterating a simple quadratic of the form , with predicted per pixel) to brighten low-light images. It is trained without any reference images, using non-reference losses that score the enhanced result directly. It is the gamma and contrast-stretching idea, with the curve chosen per pixel by a small CNN.
- Learned lookup tables. Zeng et al. [14] (2020) learn a few basis 3-D color lookup tables and a tiny CNN that predicts how to blend them for each image, which runs in real time on high-resolution photos. This is the lookup-table view of intensity transformations, extended from gray levels to color.
- Learned edge-aware filters. HDRNet by Gharbi et al. [15] predicts local affine color transforms in a low-resolution bilateral grid and applies them at full resolution guided by the input, which is the bilateral filter’s edge-awareness turned into a learned layer. Wu et al. [16] made the guided filter a differentiable layer that can be trained end-to-end inside a network.
- Low-light enhancement. The survey by Li et al. [12] reviews deep low-light image and video enhancement by learning strategy, network structure, loss functions and training data, and introduces a new low-light dataset captured with various phones under diverse illumination, plus an online platform for comparing methods. Later models such as Retinexformer [17] (2023) combine the Retinex model (an image as illumination times reflectance) with transformer attention guided by an illumination estimate.
What did not change: the median filter is still the first answer to impulse noise, the Sobel kernels still appear in loss functions and edge maps, CLAHE is still a default preprocessing step, and gamma handling remains a source of bugs in learned pipelines that mix linear and encoded pixel values. Deep models also tend to inherit the classical failure modes: halos, noise amplification in dark regions and over-enhancement are still the artifacts that reviewers look for.
Key takeaways
- An intensity transformation is a lookup table; its slope is the local contrast gain. Log and curves expand the shadows, expands the highlights, and piecewise-linear curves give precise control.
- Histogram equalization uses the scaled CDF as , which makes the output PDF uniform in the continuous case and approximately flat in the discrete case. Matching chains one equalization with the inverse of another; exact matching first breaks ties with local means.
- Local methods (AHE, CLAHE, local mean and variance rules) adapt to regions; CLAHE limits the contrast gain by clipping each local histogram.
- Linear spatial filtering is correlation with a kernel; convolution flips the kernel and has clean algebra. Separable kernels cost instead of operations per pixel.
- Lowpass filters (box, Gaussian) average away noise and blur edges; the nonlinear median filter is the tool for impulse noise.
- Sharpening adds highpass detail: the Laplacian (second derivative), unsharp masking and highboost, while the gradient and Sobel operators measure edge strength.
- Highpass, bandpass and bandreject kernels follow from lowpass ones by subtraction from the impulse ; good pipelines order steps so that nothing amplifies what an earlier step should have removed.
Exercises
- A curve for a foggy photo. An 8-bit image has all its pixels between 90 and 160. Write the piecewise-linear transformation that maps this range onto , give its slope, and say what happens to two neighboring pixels with values 120 and 122.
Hint
for , clipped to outside. The slope is , so a difference of 2 becomes about 7.3 gray levels. After rounding, and .
- Equalize by hand. A tiny 3-bit image () with 16 pixels has counts for levels to . Compute the equalization lookup table and the output histogram. Why is the output not flat?
Hint
The CDF is , so gives (with rounded to 4). The output histogram has 8 pixels at level 4, 4 at 5, 2 at 6 and 2 at 7. The 8 pixels at level 1 share one input value, and a lookup table cannot split them. Only exact histogram specification could.
- Does order matter? Is applying a gamma curve and then a box (averaging) filter the same as applying the box filter and then the gamma curve? Prove your answer or give a counterexample using a 2-pixel average.
Hint
No. Take a 1-D average of two pixels, 0 and 1, with . Gamma first: . Box first: . A nonlinear point operation does not commute with a linear filter. This is also why blurring gamma-encoded images is not the same as blurring linear light.
- Separable or not? Decide whether each kernel is separable, and if so, give the two vectors: (a) the Laplacian with in the center; (b) the kernel with all entries 1; (c) .
Hint
(a) No: its rank is 2 (but it is the sum of two separable kernels, one per axis). (b) Yes: . (c) Yes: , a small binomial approximation of a Gaussian. Check with np.linalg.matrix_rank.
- Median versus mean on an edge. A 1-D signal is , a perfect step. Apply a 3-sample mean and a 3-sample median (ignore the end samples). Then replace the sample at position 1 (counting from 0) with an impulse of value 255 and repeat. What do you conclude?
Hint
On the clean step, the mean gives a ramp () while the median keeps the step exactly. With the impulse, the mean turns it into two raised values of about 92 at positions 1 and 2; the median outputs 10 there and removes it completely.
- Design a bandpass kernel. Using Gaussians with and , build a bandpass kernel and a bandreject kernel in NumPy. Verify the sums of their weights, filter
skimage.data.brick()with both, and describe which structures each output keeps.
Hint
Make both kernels the same size (pad the smaller one with zeros), then use and . Sums should be about 0 and 1. The bandpass output keeps the mortar lines and the brick-scale texture, removing both the overall shading and the finest grain; the bandreject output keeps shading and fine grain but weakens the brick pattern.
References
- R. C. Gonzalez and R. E. Woods, Digital Image Processing, 4th ed., Pearson, 2018, Ch. 3. publisher page
- S. M. Pizer, E. P. Amburn, J. D. Austin, R. Cromartie, A. Geselowitz, T. Greer, B. ter Haar Romeny, J. B. Zimmerman and K. Zuiderveld, “Adaptive histogram equalization and its variations,” Computer Vision, Graphics, and Image Processing, 1987. doi
- K. Zuiderveld, “Contrast limited adaptive histogram equalization,” in P. S. Heckbert (ed.), Graphics Gems IV, Morgan Kaufmann, 1994, pp. 474–485. entry
- A. M. Reza, “Realization of the contrast limited adaptive histogram equalization (CLAHE) for real-time image enhancement,” Journal of VLSI Signal Processing, 2004. doi
- D. Coltuc, P. Bolon and J.-M. Chassery, “Exact histogram specification,” IEEE Transactions on Image Processing, 2006. doi
- C. Tomasi and R. Manduchi, “Bilateral filtering for gray and color images,” in Proc. IEEE International Conference on Computer Vision (ICCV), 1998. pdf
- S. Paris, P. Kornprobst, J. Tumblin and F. Durand, “Bilateral filtering: Theory and applications,” Foundations and Trends in Computer Graphics and Vision, vol. 4, no. 1, pp. 1–73, 2009. doi
- K. He, J. Sun and X. Tang, “Guided image filtering,” in Proc. European Conference on Computer Vision (ECCV), 2010, pp. 1–14; extended version in IEEE TPAMI, 2013. doi
- P. Milanfar, “A tour of modern image filtering,” IEEE Signal Processing Magazine, vol. 30, no. 1, pp. 106–128, 2013. IEEE Xplore
- S. Gingichashvili and D. Lischinski, “Evaluation and comparison of edge-preserving filters,” arXiv:2012.13778, 2020. arXiv
- Y. Qi, Z. Yang, W. Sun et al., “A comprehensive overview of image enhancement techniques,” Archives of Computational Methods in Engineering, 2021. doi
- C. Li, C. Guo, L. Han, J. Jiang, M.-M. Cheng, J. Gu and C. C. Loy, “Low-light image and video enhancement using deep learning: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021 (arXiv:2104.10729). arXiv
- C. Guo, C. Li, J. Guo, C. C. Loy, J. Hou, S. Kwong and R. Cong, “Zero-reference deep curve estimation for low-light image enhancement,” in Proc. IEEE/CVF CVPR, 2020. arXiv
- H. Zeng, J. Cai, L. Li, Z. Cao and L. Zhang, “Learning image-adaptive 3D lookup tables for high performance photo enhancement in real-time,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. arXiv
- M. Gharbi, J. Chen, J. T. Barron, S. W. Hasinoff and F. Durand, “Deep bilateral learning for real-time image enhancement,” ACM Transactions on Graphics (SIGGRAPH), 2017. arXiv
- H. Wu, S. Zheng, J. Zhang and K. Huang, “Fast end-to-end trainable guided filter,” in Proc. IEEE/CVF CVPR, 2018. arXiv
- Y. Cai, H. Bian, J. Lin, H. Wang, R. Timofte and Y. Zhang, “Retinexformer: One-stage Retinex-based transformer for low-light image enhancement,” in Proc. IEEE/CVF ICCV, 2023. arXiv