Every spectroscopy lab has a JCAMP-DX story, and it is usually a truncated spectrum that nobody noticed for a month. The format is a genuine success — an open, human-readable standard from the late 1980s that every major vendor still exports — but "open" and "simple to parse" are not the same thing.
The data section is the problem. JCAMP-DX defines several data forms for the same numbers: plain ASCII (PAC), differential (DIF), duplicate (DUP) and squeezed (SQZ). A reader that handles only PAC works perfectly on the files its author tested, truncates silently on most real vendor exports, and raises no error — which is the worst possible failure mode in a scientific pipeline.
What a JCAMP-DX file actually contains
A minimal XYDATA file looks like this:
##TITLE= PE film, ATR, 4 cm-1
##JCAMP-DX= 4.24
##DATA TYPE= INFRARED SPECTRUM
##XUNITS= 1/CM
##YUNITS= ABSORBANCE
##FIRSTX= 4000.0000
##LASTX= 400.0000
##NPOINTS= 1801
##XFACTOR= 1.0000
##YFACTOR= 0.0001
##XYDATA= (X++(Y..Y))
4000 500 505 511 518 526 533 541 549 557 ...
...
##END=Two things define the physics before any number is read:
- `##YUNITS` — transmittance and absorbance spectra look structurally identical. Peak heights only mean something if you checked.
- `##XFACTOR` / `##YFACTOR` — the stored integers are scaled; the real value is
stored × factor(with an offset form as well). A reader that skips this returns numbers that are wrong by orders of magnitude but perfectly shaped.
Then the data section. (X++(Y..Y)) is JCAMP's compact notation: the first X is absolute, X++ means "increment by ##DELTAX per point", and each subsequent line continues the Y sequence. Within that structure the digits themselves can be encoded in DIF (differences between consecutive values), DUP (duplicate counts to save space) and SQZ (digits remapped into printable ASCII ranges). The compressed forms are where quick parsers die — the file stays syntactically valid, so nothing crashes; you just silently lose the tail of the spectrum or get values that drift.
Reading one correctly in Python
Use a library that documents the forms it supports. The jcamp package handles PAC, SQZ, DIF and DUP for the common IR/Raman exports:
import jcamp
d = jcamp.JCAMP_reader("pe_film.jdx")
print(sorted(d.keys()))
# dict-like: 'title', 'x', 'y', 'xunits', 'yunits', 'data type', ...
x = d["x"]
y = d["y"]
print(len(x), d.get("xunits"), d.get("yunits"))Now the part most tutorials skip — validate the parse:
import math
assert len(x) == len(y), "x/y length mismatch — something is very wrong"
# A parse that returns fewer points than NPOINTS lost data (often SQZ/DIF).
npoints = d.get("npoints")
if npoints:
assert len(x) == int(npoints), f"parsed {len(x)} of {npoints} points — truncated"
# The axes must land where the header says (keys are flat in the jcamp dict:
# firstx/lastx; print(sorted(d.keys())) once for your library version).
firstx, lastx = float(d["firstx"]), float(d["lastx"])
assert math.isclose(x[0], firstx, rel_tol=1e-4), (x[0], firstx)
assert math.isclose(x[-1], lastx, rel_tol=1e-4), (x[-1], lastx)
# Sanity-check the y units by their physical range.
if str(d.get("yunits", "")).upper().startswith("TRANSMITTANCE"):
assert min(y) >= 0 and max(y) <= 110, "transmittance outside 0-100% — check yfactor"The NPOINTS check is the single highest-value assertion: truncation from an unimplemented compression form is the classic failure, and it is detectable in one line.
The five dialects that break in the wild
- SQZ, DIF, DUP forms. Covered above — use a library that names them, and keep the
NPOINTSassertion as a tripwire. - NTUPLES. 2D NMR and some hyphenated data use
##NTUPLESwith multiple##PAGEsections and independently defined axes. A 1D reader returns junk or nothing;spectrochempyand NMR-aware toolchains handle the page structure explicitly. - Compound files (`##LINK`). One file can chain several spectra. If your pipeline assumes "one file, one spectrum", you are silently dropping all but the first.
- `##END=` placement. Some exporters append notes after
##END=. Readers that trust the terminator and stop are right; readers that grep the whole file for data lines pick up header text. - Y declination in NMR. In NMR JCAMP, Y often decreases with X (ppm runs high to low). A TIC or IR plot will simply come out mirrored if your plotting code assumes increasing x — the parse is fine, the picture is wrong.
When to convert, and when to keep the original
JCAMP-DX is a transport encoding. Convert once at ingestion — to a clean CSV or a dataframe with the XFACTOR values already applied — and keep the original .jdx beside it. The compression forms are not an analysis format, and every downstream re-parse is another chance to lose the tail of a spectrum.
The free parser at /tools/jcamp-dx does exactly this conversion without installation (2 MB cap, parsed in memory, nothing stored). In a free account, ingestion keeps the original file as provenance and the extracted rows carry the EXTRACTED evidence class until a human reviews them — which is the discipline that catches the SQZ-truncation class of bug before it reaches a model. Format-specific details for IR readers live in /docs/instruments/ftir, and the same JCAMP file is often the portable path out of GC-MS vendor software (/tools/gcms).
Honest limits
jcamp's coverage of vendor-specific extensions is good but not universal. IfNPOINTSdisagrees and the file is NMR NTUPLES, switch tools rather than patching the reader.- XFACTOR/offset handling is applied by the library; verify on one file with a known peak before batch-processing a thousand.
- A correct parse is not a correct measurement: instrument calibration, atmospheric CO2 bands and ATR artefacts are separate problems covered in the FTIR baseline post.
The pattern across all of it is the same as every other instrument format: parse, assert against the header, review, then model. Silent truncation is the enemy, and the assertions above are how you make it loud.