Document Vision

From scans and screenshots to structured drafts — with human review built in

Document Vision is the extraction sandbox behind /doc-vision. Five endpoints turn page renders, table crops and chart crops into OCR text with boxes, layout boxes, table grids and chart parts. Every result is labelled EXTRACTED and is a draft a human confirms before it enters a dataset.

Extraction surfaces

Five extraction calls, plus engine discovery

Every POST takes {image: "data:image/png;base64,..."} and returns an experimental result; the extraction routes require login and are rate-limited. Engine discovery is open.

POST /api/doc-vision/ocr
OCR text with boxes

Returns text detections, each with its recognized text and bounding box. Three engines: ocr-v2 (multilingual, default), ocr-v1 (English-only bulk-ingest fallback) and nemoretriever-ocr (legacy alias). Override the engine in the request body.

POST /api/doc-vision/page-elements
Page layout router

Give it a full-page render and it classifies title, paragraph, table and chart boxes — a layout map for cropping, header/footer filtering and downstream reading.

POST /api/doc-vision/table-structure
Table grid

Give it a single table crop and it returns cell, row and column boxes — the structure you need before promoting a composition or property table.

POST /api/doc-vision/graphic-elements
Chart parts

Give it a single chart crop and it isolates the chart title, axes and legend boxes, ready to feed downstream digitizing workflows.

POST /api/doc-vision/analyze-document
Analyze page (layout + OCR)

One call runs the page router and full-page OCR together and returns both, with per-part errors and the suggested next step: crop the table or chart boxes client-side and post them to the dedicated routes. v1 does not crop for you.

GET /api/doc-vision/engines
Engine discovery

The 25-model registry with verified-live, existing-live and untested status, plus a configured flag. Six document-AI engines are verified live: three OCR models, page-elements-v3, table-structure-v1 and graphic-elements-v1. The rest are listed honestly and must not be called blindly.

Response shape

What each call gives back

The payloads are honest about scope: text and boxes from OCR, keyed bounding boxes from the layout models, and a combined layout + OCR result from the page analyzer.

Success responses carry status, an experimental flag, a stability note, an engine card, an evidence class of EXTRACTED at level E2, and a measured latency. OCR returns text detections at data[0].text_detections; the page, table and chart routes return keyed boxes at data[0].bounding_boxes.

analyze-document returns both halves together — layout, ocr, per-part errors and a next hint — and degrades to a partial status when only one half fails.

Error responses stay explicit too: a failed call returns an error status with a real upstream message, mapped to 400, 413, 429, 502 or 503 as appropriate. It never fills the gap with invented detections.

Returned fields
  • Text + boxes — every OCR detection carries its text prediction and its bounding box, not a flat text dump.
  • Layout boxes — element-type keys for title, paragraph, table and chart regions on a full-page render.
  • Table grid — cell, row and column boxes from a single table crop.
  • Chart parts — chart title, axes and legend boxes from a single chart crop.
  • Engine card + latency — each success names the model that produced it through an engine card and reports the measured latency in milliseconds.
  • Explicit errors — a failed call returns an error status with a real message — never fabricated detections.
The client shows a neutral engine label.The API returns the true engine identity for the audit trail; the sandbox renders a neutral “experimental extraction” card so a trial model is never presented as a branded production engine. Evidence class and level are preserved untouched.
Review gate

Upload → extract → review → promote

Document Vision is an extraction sandbox, not an authority. The workflow deliberately puts a person between an extracted row and your data.

  • Choose an image — render the page, table or chart as PNG/JPG and upload it in the /doc-vision sandbox (the same panel is embedded as the enhance card on Ingestion).
  • Extract — run one action; the response is an EXTRACTED/E2 draft with boxes, labels, an engine card and the raw JSON.
  • Review — a human checks text, units, headers and totals against the source. Outputs are automated extraction, never a measurement.
  • Promote — take reviewed rows through Ingestion: upload, review the parsed batch row by row, accept, and promote a clean batch into a Dataset with a target column.
EXTRACTED is not MEASURED.A number read by OCR is a draft transcription until a human verifies it against the source. Nothing here is treated as a measurement, and reviewed rows are promoted through Ingestion rather than written straight into a dataset. The full flow is documented under ingestion.
Use cases

Where extraction pays for itself

Anywhere a human would otherwise retype numbers from a page, the sandbox can produce a reviewable draft — with the limits below always in force.

Literature
Tables from papers

Composition, property and catalyst tables rendered from papers become table grids you can review and promote instead of retyping numbers by hand.

Supplementary PDFs
Scanned pages to text + layout

Render supplementary pages as images, then OCR their text and route the layout into title, paragraph, table and chart boxes.

Lab notebooks
Old pages, reviewable text

Scanned or photographed notebook pages come back as text with boxes. Review them before anything enters the ELN or a dataset.

Screenshots
App and chart captures

Turn screenshots into structure: OCR for labels, graphic elements for chart titles, axes and legends.

Instrument reports
Printed report pages

Scanned or printed instrument report pages OCR cleanly, and chart crops isolate axes and legends for downstream reading. Native instrument formats belong to the Characterization path instead.

Datasheets
Multilingual documents and CoAs

Multilingual OCR is the default, so certificates of analysis, datasheets and spec sheets come back as text with boxes.

Honest limits

What Document Vision does not do

The endpoints were built to under-claim rather than over-claim. These constraints are part of the product, not fine print.

  • Experimental by design — this is a free-tier trial integration. Uptime and latency are untested; do not wire it into a production pipeline without a human review step.
  • Rate-limited — the four single actions allow 30 requests per minute per user; analyze-document allows 15. A busy moment returns 429 — retry shortly.
  • Inline size limit — the image data URL must stay under 180,000 base64 characters (roughly 135 KB). Larger images are rejected with 413; crop smaller or use the larger-image upload flow. The sandbox warns above the limit.
  • Images, not PDFs — these endpoints take image data URLs. Render the page, table or chart to PNG/JPG first; Ingestion has its own PDF and Excel import path.
  • No automatic cropping yet — analyze-document v1 returns layout plus OCR and asks you to crop table and chart boxes client-side and post them to the dedicated routes.
  • No billing while experimental — the endpoints are free during the trial — login and per-minute rate limits only.
  • Engine honesty — only the six verified document-AI engines are callable; the registry marks everything else existing-live or untested.
  • Never a measurement — outputs are EXTRACTED/E2 automated extraction. Human review is mandatory before any dataset promotion.
Experimental extraction, full stop.The service itself labels every response experimental with untested uptime and latency. Use it to draft, review by hand, and promote through Ingestion — never as a measurement source or an unattended pipeline.
FAQ

Document Vision questions, answered plainly

Is Document Vision production-ready?
No. It is an experimental free-tier integration: uptime and latency are untested, and every response says so. Treat the endpoints as extraction utilities with a mandatory review step, not as production ingestion infrastructure.
Can I upload a PDF directly to these endpoints?
Not to the vision routes — they take image data URLs. Render the PDF page, table or chart to PNG/JPG and upload that. The platform’s Ingestion flow has its own PDF and Excel import path.
Does it cost credits?
No billing while experimental. The routes are login-gated and rate-limited (30 requests per minute, 15 for analyze-document) and return 429 when the trial is busy.
Why was my image rejected?
Inline images must stay under 180,000 base64 characters (roughly 135 KB). A larger payload returns 413; crop to the table or chart you need, or downscale the render.
How do extracted numbers become data?
Review the draft, then promote it through Ingestion: upload, review the parsed batch row by row, accept the rows, and promote the batch into a Dataset with a target column. Extraction rows do not become searchable data until you promote them.
What does the EXTRACTED label mean?
It is evidence class EXTRACTED, level E2: an automated extraction from a document, not a measurement. It carries an engine card and latency, and it must be reviewed by a human before it can influence modeling.
Which engines actually run?
Six hosted document-AI engines are verified live: three OCR options, a page-layout model, a table-structure model and a graphic-elements model. GET /api/doc-vision/engines lists all 25 registry entries with their honest status and the route each verified engine serves.

Turn your paper archives into reviewable drafts

Open the sandbox, run OCR or a table grid on a page, review the boxes, and promote the verified rows through Ingestion into a dataset the pipeline can model.