Every FTIR discussion eventually arrives at the same argument: two people processed the same spectra and get different peak ratios. The cause is almost never the instrument. It is the baseline — a decision disguised as a preprocessing step.
This post covers three defensible baseline methods, the artefacts each one creates, and the parameter choices that separate a correction from a deformation. It ends with the discipline that keeps the decision honest: state the method in the figure caption, because the reader cannot recover it from the data.
What you are actually subtracting
Before choosing a method, name the components:
- Instrument drift and source ageing — broad, slow, smooth. Baseline correction territory.
- Scattering (powders, rough surfaces) — a smooth multiplicative-ish slope or curve. Baseline correction territory, but strong scattering may need an extended correction (Kubelka–Munk) instead.
- Atmospheric CO2 (~2349 cm⁻¹) and water vapour — sharp lines, not baseline. Purge and use a fresh background; if a trace remains, flag the region rather than warping the baseline through it.
- ATR wavelength dependence — a physical intensity gradient, corrected with the penetration-depth model (
dp = λ / (2π n₁ √(sin²θ − (n₂/n₁)²))), not with a polynomial.
Conflating the last one with baseline correction is a classic error: an ATR correction changes relative band intensities according to physics; a baseline correction removes a modelled background. If you are comparing an ATR spectrum to a transmission reference, you need the former; for run-to-run comparability, the latter.
Method 1 — rolling minimum (the robust default)
Estimate the baseline as a moving minimum, then smooth it. It never follows sharp bands, because a minimum by construction ignores them; it fails precisely when broad bands dominate, because the minimum rides up the flank of the broad envelope.
import numpy as np
from scipy.ndimage import minimum_filter1d, uniform_filter1d
def rolling_baseline(y, window=151, smooth=51):
"""Moving-minimum baseline. window must exceed the widest *sharp* band."""
bg = minimum_filter1d(y, size=window)
return uniform_filter1d(bg, size=smooth)
corrected = absorbance - rolling_baseline(absorbance)
# Diagnostic: baseline must not exceed the data anywhere.
assert (rolling_baseline(absorbance) <= absorbance + 1e-9).all()Choose the window in wavenumber space, not samples: at 1 cm⁻¹ spacing, a window of 151 points is 150 cm⁻¹ — wider than any sharp band, narrower than a typical drift. If your O–H envelope spans 400 cm⁻¹, this method will cut into it. That is the honest failure mode.
Method 2 — asymmetric least squares (the workhorse)
ALS (Eilers & Boelens, 2005) fits a smooth curve that prefers to sit below the data, with two parameters: λ (smoothness — larger is smoother) and p (asymmetry — smaller keeps the baseline lower). It handles sloping and curved backgrounds better than a rolling minimum and is the standard choice for ATR spectra with many sharp bands.
import numpy as np
from scipy import sparse
from scipy.sparse.linalg import spsolve
def als_baseline(y, lam=1e5, p=0.001, niter=10):
"""Asymmetric least squares baseline (Eilers & Boelens 2005).
lam: smoothness (1e4-1e7). p: asymmetry (1e-3-5e-2)."""
L = len(y)
D = sparse.diags([1, -2, 1], [0, 1, 2], shape=(L, L - 2))
DTD = lam * D.dot(D.transpose())
w = np.ones(L)
for _ in range(niter):
W = sparse.spdiags(w, 0, L, L)
z = spsolve(W + DTD, w * y)
w = p * (y > z) + (1 - p) * (y < z)
return z
corrected = absorbance - als_baseline(absorbance, lam=1e5, p=0.001)How to choose the parameters without fooling yourself:
- λ too small → the baseline follows the bands (you subtract your signal). λ too large → the baseline is a straight line and drift survives.
- p too large → the baseline creeps up into band flanks. p too small → slow to converge; residual offset.
- The falsifiable test: double λ by 10× and halve p by 10×. Compare the corrected peak *ratios*, not the pictures. If ratios move more than your analytical requirement, the correction is doing analysis. Report the parameters that stabilize ratios.
One structural limitation: ALS is a single global fit. A spectrum with a steep scattering tail and a broad O–H envelope has two overlapping smooth components; ALS can only put one smooth curve under both, and it will pick the wrong one somewhere. Piecewise baselines (fit in band-free windows, interpolate) are more work but more honest for difficult samples.
Method 3 — SNIP (the classical alternative)
SNIP — statistics-sensitive iterative clipping — approaches the background from above, clipping peaks away iteration by iteration. It is the historical standard for XRF and γ-spectra and works well on IR spectra with sharp bands and a single smooth background.
def snip_baseline(y, iterations=40):
"""SNIP background (Ryan et al. 1988). Requires y >= 0."""
y = np.asarray(y, dtype=float)
v = np.log(np.log(np.sqrt(y + 1.0) + 1.0) + 1.0)
for p in range(1, iterations + 1):
v2 = v.copy()
for i in range(p, len(v) - p):
v2[i] = min(v[i], 0.5 * (v[i - p] + v[i + p]))
v = v2
return (np.exp(np.exp(v) - 1.0) - 1.0) ** 2
corrected = absorbance - snip_baseline(absorbance, iterations=40)The parameter that matters is the iteration count: it must exceed roughly half the width of the widest peak you want removed, and below that width it clips bands. Too many iterations and SNIP starts eating broad bands from the sides — and because the transform is nonlinear, it can produce small negative dips that look like real absorption minima. Plot the estimated background over the spectrum every time; never run SNIP blind.
The same spectrum, three baselines
The point of the comparison is not scoring a winner — it is calibration of your own judgment. If the three corrected spectra agree on the peak ratios you care about, the correction is stable and any method will do, documented. If they disagree, you have found the actual uncertainty of your analysis, and it belongs in the error bar, not hidden by picking the prettiest plot.
A practical order of operations for ATR spectra:
- Load with metadata (resolution, co-adds, background time) — see /blog/jcamp-dx-reader-python for getting JCAMP exports out of vendor software cleanly.
- ATR correction if comparing to transmission data.
- Baseline correction with a stated method and parameters.
- Normalization — and pick it consciously: min–max makes the strongest band 1.0 and silently destroys intensity information across samples; vector normalization is the safer default for library search; internal-ratio normalization (a band that should be constant) is best for kinetics.
The discipline that survives review
Whatever you choose, three things go in the record:
- Method and parameters — "ALS, λ=1e5, p=0.001" is reproducible; "baseline corrected" is not.
- A before/after figure with the estimated background overlaid — this is where a correction that ate a band becomes visible.
- A stability check — the peak ratios after 10× parameter perturbations. If they move, that movement is your method's contribution to the uncertainty.
Everything in this post applies equally to Raman and UV-Vis spectra; the artefacts differ, the decision structure does not. If you need the parsed numbers before preprocessing, the free parser at /tools/ftir returns the series, peaks and derived analytics in memory with nothing stored. Format-specific handling for OPUS, Nicolet and JCAMP is documented at /docs/instruments/ftir.
Honest limits
- Baseline correction cannot recover signal destroyed by detector saturation or a bad background. Check the raw transmittance first.
- Broad-band overlap (O–H/N–H envelopes overlapping each other) is an unsolved separation problem at the preprocessing level; deconvolution is a hypothesis-generating step, not a measurement.
- No method here handles atmospheric lines correctly by design — flag them, and treat results inside those windows as region-excluded.
- Keep the raw file. Every corrected spectrum should be reproducible from it with the recorded parameters; if it is not, the parameters are incomplete.