Blog · October 11, 2026

FTIR baseline correction: three methods, and when each one fails

Rolling-minimum, ALS and SNIP on real ATR spectra — how to pick the parameters, the artefacts each method creates, and the difference between a baseline correction and an ATR correction.

ftirspectroscopypreprocessingpython

Every FTIR discussion eventually arrives at the same argument: two people processed the same spectra and get different peak ratios. The cause is almost never the instrument. It is the baseline — a decision disguised as a preprocessing step.

This post covers three defensible baseline methods, the artefacts each one creates, and the parameter choices that separate a correction from a deformation. It ends with the discipline that keeps the decision honest: state the method in the figure caption, because the reader cannot recover it from the data.

What you are actually subtracting

Before choosing a method, name the components:

  • Instrument drift and source ageing — broad, slow, smooth. Baseline correction territory.
  • Scattering (powders, rough surfaces) — a smooth multiplicative-ish slope or curve. Baseline correction territory, but strong scattering may need an extended correction (Kubelka–Munk) instead.
  • Atmospheric CO2 (~2349 cm⁻¹) and water vapour — sharp lines, not baseline. Purge and use a fresh background; if a trace remains, flag the region rather than warping the baseline through it.
  • ATR wavelength dependence — a physical intensity gradient, corrected with the penetration-depth model (dp = λ / (2π n₁ √(sin²θ − (n₂/n₁)²))), not with a polynomial.

Conflating the last one with baseline correction is a classic error: an ATR correction changes relative band intensities according to physics; a baseline correction removes a modelled background. If you are comparing an ATR spectrum to a transmission reference, you need the former; for run-to-run comparability, the latter.

Method 1 — rolling minimum (the robust default)

Estimate the baseline as a moving minimum, then smooth it. It never follows sharp bands, because a minimum by construction ignores them; it fails precisely when broad bands dominate, because the minimum rides up the flank of the broad envelope.

python
import numpy as np
from scipy.ndimage import minimum_filter1d, uniform_filter1d

def rolling_baseline(y, window=151, smooth=51):
    """Moving-minimum baseline. window must exceed the widest *sharp* band."""
    bg = minimum_filter1d(y, size=window)
    return uniform_filter1d(bg, size=smooth)

corrected = absorbance - rolling_baseline(absorbance)

# Diagnostic: baseline must not exceed the data anywhere.
assert (rolling_baseline(absorbance) <= absorbance + 1e-9).all()

Choose the window in wavenumber space, not samples: at 1 cm⁻¹ spacing, a window of 151 points is 150 cm⁻¹ — wider than any sharp band, narrower than a typical drift. If your O–H envelope spans 400 cm⁻¹, this method will cut into it. That is the honest failure mode.

Method 2 — asymmetric least squares (the workhorse)

ALS (Eilers & Boelens, 2005) fits a smooth curve that prefers to sit below the data, with two parameters: λ (smoothness — larger is smoother) and p (asymmetry — smaller keeps the baseline lower). It handles sloping and curved backgrounds better than a rolling minimum and is the standard choice for ATR spectra with many sharp bands.

python
import numpy as np
from scipy import sparse
from scipy.sparse.linalg import spsolve

def als_baseline(y, lam=1e5, p=0.001, niter=10):
    """Asymmetric least squares baseline (Eilers & Boelens 2005).
    lam: smoothness (1e4-1e7). p: asymmetry (1e-3-5e-2)."""
    L = len(y)
    D = sparse.diags([1, -2, 1], [0, 1, 2], shape=(L, L - 2))
    DTD = lam * D.dot(D.transpose())
    w = np.ones(L)
    for _ in range(niter):
        W = sparse.spdiags(w, 0, L, L)
        z = spsolve(W + DTD, w * y)
        w = p * (y > z) + (1 - p) * (y < z)
    return z

corrected = absorbance - als_baseline(absorbance, lam=1e5, p=0.001)

How to choose the parameters without fooling yourself:

  • λ too small → the baseline follows the bands (you subtract your signal). λ too large → the baseline is a straight line and drift survives.
  • p too large → the baseline creeps up into band flanks. p too small → slow to converge; residual offset.
  • The falsifiable test: double λ by 10× and halve p by 10×. Compare the corrected peak *ratios*, not the pictures. If ratios move more than your analytical requirement, the correction is doing analysis. Report the parameters that stabilize ratios.

One structural limitation: ALS is a single global fit. A spectrum with a steep scattering tail and a broad O–H envelope has two overlapping smooth components; ALS can only put one smooth curve under both, and it will pick the wrong one somewhere. Piecewise baselines (fit in band-free windows, interpolate) are more work but more honest for difficult samples.

Method 3 — SNIP (the classical alternative)

SNIP — statistics-sensitive iterative clipping — approaches the background from above, clipping peaks away iteration by iteration. It is the historical standard for XRF and γ-spectra and works well on IR spectra with sharp bands and a single smooth background.

python
def snip_baseline(y, iterations=40):
    """SNIP background (Ryan et al. 1988). Requires y >= 0."""
    y = np.asarray(y, dtype=float)
    v = np.log(np.log(np.sqrt(y + 1.0) + 1.0) + 1.0)
    for p in range(1, iterations + 1):
        v2 = v.copy()
        for i in range(p, len(v) - p):
            v2[i] = min(v[i], 0.5 * (v[i - p] + v[i + p]))
        v = v2
    return (np.exp(np.exp(v) - 1.0) - 1.0) ** 2

corrected = absorbance - snip_baseline(absorbance, iterations=40)

The parameter that matters is the iteration count: it must exceed roughly half the width of the widest peak you want removed, and below that width it clips bands. Too many iterations and SNIP starts eating broad bands from the sides — and because the transform is nonlinear, it can produce small negative dips that look like real absorption minima. Plot the estimated background over the spectrum every time; never run SNIP blind.

The same spectrum, three baselines

The point of the comparison is not scoring a winner — it is calibration of your own judgment. If the three corrected spectra agree on the peak ratios you care about, the correction is stable and any method will do, documented. If they disagree, you have found the actual uncertainty of your analysis, and it belongs in the error bar, not hidden by picking the prettiest plot.

A practical order of operations for ATR spectra:

  1. Load with metadata (resolution, co-adds, background time) — see /blog/jcamp-dx-reader-python for getting JCAMP exports out of vendor software cleanly.
  2. ATR correction if comparing to transmission data.
  3. Baseline correction with a stated method and parameters.
  4. Normalization — and pick it consciously: min–max makes the strongest band 1.0 and silently destroys intensity information across samples; vector normalization is the safer default for library search; internal-ratio normalization (a band that should be constant) is best for kinetics.

The discipline that survives review

Whatever you choose, three things go in the record:

  • Method and parameters — "ALS, λ=1e5, p=0.001" is reproducible; "baseline corrected" is not.
  • A before/after figure with the estimated background overlaid — this is where a correction that ate a band becomes visible.
  • A stability check — the peak ratios after 10× parameter perturbations. If they move, that movement is your method's contribution to the uncertainty.

Everything in this post applies equally to Raman and UV-Vis spectra; the artefacts differ, the decision structure does not. If you need the parsed numbers before preprocessing, the free parser at /tools/ftir returns the series, peaks and derived analytics in memory with nothing stored. Format-specific handling for OPUS, Nicolet and JCAMP is documented at /docs/instruments/ftir.

Honest limits

  • Baseline correction cannot recover signal destroyed by detector saturation or a bad background. Check the raw transmittance first.
  • Broad-band overlap (O–H/N–H envelopes overlapping each other) is an unsolved separation problem at the preprocessing level; deconvolution is a hypothesis-generating step, not a measurement.
  • No method here handles atmospheric lines correctly by design — flag them, and treat results inside those windows as region-excluded.
  • Keep the raw file. Every corrected spectrum should be reproducible from it with the recorded parameters; if it is not, the parameters are incomplete.
FAQ

Questions this post answers

Short answers, stated limits included.

Why does my ATR spectrum slope upward toward low wavenumber?
That slope is partly physical. ATR sampling depth scales with wavelength, so longer wavelengths (lower wavenumbers) sample deeper and appear more intense. Correcting it properly is an ATR correction using the penetration-depth model — a different operation from baseline correction, which handles scattering and drift.
Which baseline method should I default to?
For ATR spectra with sharp bands, asymmetric least squares (ALS) with documented parameters is the common workhorse. For spectra with one dominant broad envelope, a rolling-minimum or rubberband estimate is more robust. There is no universally correct method — the honest answer is to state the method and parameters in every figure.
Why did ALS cut into my broad O-H band?
Either the smoothness parameter is too small (the baseline follows the band) or the asymmetry parameter moved the baseline up into it. Broad bands and baselines are genuinely hard to separate; if the O-H envelope is your signal, a rolling estimate anchored at band-free windows is safer.
Can I baseline-correct CO2 and water vapour bands away?
No — those are sharp atmospheric lines, not baseline. Subtracting a baseline through them distorts neighbouring real bands. Purge the instrument, measure a background, and flag the region if it remains.
Does Matflow parse FTIR files free?
Yes — upload an OPUS, JCAMP or CSV spectrum at matflow.celaron.com/tools/ftir and get the parsed series, detected peaks and derived analytics back in memory, nothing stored.

Try the workflow from this post

Create a free account, upload your own file, and see the parsed rows and their evidence labels before you trust any number.