From scans and screenshots to structured drafts — with human review built in
Document Vision is the extraction sandbox behind /doc-vision. Five endpoints turn page renders, table crops and chart crops into OCR text with boxes, layout boxes, table grids and chart parts. Every result is labelled EXTRACTED and is a draft a human confirms before it enters a dataset.
Five extraction calls, plus engine discovery
Every POST takes {image: "data:image/png;base64,..."} and returns an experimental result; the extraction routes require login and are rate-limited. Engine discovery is open.
Returns text detections, each with its recognized text and bounding box. Three engines: ocr-v2 (multilingual, default), ocr-v1 (English-only bulk-ingest fallback) and nemoretriever-ocr (legacy alias). Override the engine in the request body.
Give it a full-page render and it classifies title, paragraph, table and chart boxes — a layout map for cropping, header/footer filtering and downstream reading.
Give it a single table crop and it returns cell, row and column boxes — the structure you need before promoting a composition or property table.
Give it a single chart crop and it isolates the chart title, axes and legend boxes, ready to feed downstream digitizing workflows.
One call runs the page router and full-page OCR together and returns both, with per-part errors and the suggested next step: crop the table or chart boxes client-side and post them to the dedicated routes. v1 does not crop for you.
The 25-model registry with verified-live, existing-live and untested status, plus a configured flag. Six document-AI engines are verified live: three OCR models, page-elements-v3, table-structure-v1 and graphic-elements-v1. The rest are listed honestly and must not be called blindly.
What each call gives back
The payloads are honest about scope: text and boxes from OCR, keyed bounding boxes from the layout models, and a combined layout + OCR result from the page analyzer.
Success responses carry status, an experimental flag, a stability note, an engine card, an evidence class of EXTRACTED at level E2, and a measured latency. OCR returns text detections at data[0].text_detections; the page, table and chart routes return keyed boxes at data[0].bounding_boxes.
analyze-document returns both halves together — layout, ocr, per-part errors and a next hint — and degrades to a partial status when only one half fails.
Error responses stay explicit too: a failed call returns an error status with a real upstream message, mapped to 400, 413, 429, 502 or 503 as appropriate. It never fills the gap with invented detections.
- Text + boxes — every OCR detection carries its text prediction and its bounding box, not a flat text dump.
- Layout boxes — element-type keys for title, paragraph, table and chart regions on a full-page render.
- Table grid — cell, row and column boxes from a single table crop.
- Chart parts — chart title, axes and legend boxes from a single chart crop.
- Engine card + latency — each success names the model that produced it through an engine card and reports the measured latency in milliseconds.
- Explicit errors — a failed call returns an error status with a real message — never fabricated detections.
Upload → extract → review → promote
Document Vision is an extraction sandbox, not an authority. The workflow deliberately puts a person between an extracted row and your data.
- Choose an image — render the page, table or chart as PNG/JPG and upload it in the /doc-vision sandbox (the same panel is embedded as the enhance card on Ingestion).
- Extract — run one action; the response is an EXTRACTED/E2 draft with boxes, labels, an engine card and the raw JSON.
- Review — a human checks text, units, headers and totals against the source. Outputs are automated extraction, never a measurement.
- Promote — take reviewed rows through Ingestion: upload, review the parsed batch row by row, accept, and promote a clean batch into a Dataset with a target column.
Where extraction pays for itself
Anywhere a human would otherwise retype numbers from a page, the sandbox can produce a reviewable draft — with the limits below always in force.
Composition, property and catalyst tables rendered from papers become table grids you can review and promote instead of retyping numbers by hand.
Render supplementary pages as images, then OCR their text and route the layout into title, paragraph, table and chart boxes.
Scanned or photographed notebook pages come back as text with boxes. Review them before anything enters the ELN or a dataset.
Turn screenshots into structure: OCR for labels, graphic elements for chart titles, axes and legends.
Scanned or printed instrument report pages OCR cleanly, and chart crops isolate axes and legends for downstream reading. Native instrument formats belong to the Characterization path instead.
Multilingual OCR is the default, so certificates of analysis, datasheets and spec sheets come back as text with boxes.
What Document Vision does not do
The endpoints were built to under-claim rather than over-claim. These constraints are part of the product, not fine print.
- Experimental by design — this is a free-tier trial integration. Uptime and latency are untested; do not wire it into a production pipeline without a human review step.
- Rate-limited — the four single actions allow 30 requests per minute per user; analyze-document allows 15. A busy moment returns 429 — retry shortly.
- Inline size limit — the image data URL must stay under 180,000 base64 characters (roughly 135 KB). Larger images are rejected with 413; crop smaller or use the larger-image upload flow. The sandbox warns above the limit.
- Images, not PDFs — these endpoints take image data URLs. Render the page, table or chart to PNG/JPG first; Ingestion has its own PDF and Excel import path.
- No automatic cropping yet — analyze-document v1 returns layout plus OCR and asks you to crop table and chart boxes client-side and post them to the dedicated routes.
- No billing while experimental — the endpoints are free during the trial — login and per-minute rate limits only.
- Engine honesty — only the six verified document-AI engines are callable; the registry marks everything else existing-live or untested.
- Never a measurement — outputs are EXTRACTED/E2 automated extraction. Human review is mandatory before any dataset promotion.
Document Vision questions, answered plainly
Is Document Vision production-ready?
Can I upload a PDF directly to these endpoints?
Does it cost credits?
Why was my image rejected?
How do extracted numbers become data?
What does the EXTRACTED label mean?
Which engines actually run?
Turn your paper archives into reviewable drafts
Open the sandbox, run OCR or a table grid on a page, review the boxes, and promote the verified rows through Ingestion into a dataset the pipeline can model.