One platform, every stage of materials & chemistry R&D
From a sparse CSV to a signed-off dossier: the discovery pipeline, physics engines, lab operations, data ingestion and the assistant that drives them. Every card opens the real module or its documentation.
The discovery pipeline, stage by stage
Each stage runs standalone and produces a tracked run; the Full Pipeline chains them with automatic handoffs.
Data Synthesis
Turns a handful of experimental runs into a statistically faithful dataset. Three modes — Neural (conditional-GAN-style), Hybrid (GAN plus copula) and Statistical (marginals plus correlation structure) — fit the joint distribution, and generated rows are clipped to each column’s observed range.
Evaluation
Scores a synthetic batch before it reaches a model: distribution, category and correlation quality; classifier indistinguishability; privacy metrics (closest-record distance, duplication and overfitting checks); and anomaly flags on the worst-fitting rows.
Prediction
A cross-validated ensemble returning Global Impact, Local Impact and permutation importance, conformal intervals from the same folds, and an extrapolation flag on every prediction. Formula columns can be expanded into 132 composition descriptors automatically.
Optimization
Adaptive (surrogate-guided) and Evolutionary (multi-objective genetic) search. TOPSIS ranks the trade-off front against your objective weights, and 100%-sum and per-feature bounds are enforced during generation. The front is a good approximation — not a proof of global optimality.
Design of Experiments
Full and fractional factorial, Box–Behnken, central composite, Latin hypercube and Sobol plans — plus simplex-lattice and centroid mixture designs — downloadable as a CSV run list for the bench.
Active Learning
Proposes the next experiment batch from an optimization history using a Gaussian-process expected-information-gain core. When the full engine is unavailable a heuristic surrogate runs, and the response states which engine actually ran.
Virtual Screening
Ranks a candidate library by structural similarity to a reference molecule, with an optional SMARTS pattern that restricts hits to scaffolds your route can actually make.
Full Pipeline
A visual graph editor that chains synthesis → evaluation → prediction → optimization with TEA and synthetic-data nodes. Each stage output becomes the next stage input, and every run keeps its full configuration for rerun and sharing.
3-Step Auto-Wizard
Inspects your columns, suggests which experiments to run next and exports the plan as a protocol — a lower-friction path to experimental planning than the full DoE page.
Model Hub
One-click auto-select across GP, Random Forest, GBM, XGBoost, LightGBM, CatBoost and linear candidates with a cross-validation leaderboard. Train a Gaussian-process surrogate, export a serialized model and score deployed models over HTTP.
Model Zoo
Instant estimates for formation energy, band gap, bulk modulus, CO adsorption, alloy hardness/density, battery capacity/voltage and polymer QSPR from a formula. These are deterministic composition-rule heuristics, not trained models — every response carries is_estimate: true and a screening-only disclaimer.
An assistant that acts on the platform, not around it
Pilot is the panel on every app page. It calls the same capabilities, under the same permissions and the same credit admission, as the manual UI.
The Pilot panel
A persistent conversation panel: executed actions and results appear in Actions, missing input arrives as structured questions with a recommended choice, tasks show plan checklists and live progress, and files can be attached or pasted.
Capabilities & approvals
The live capability catalog is served by GET /api/assistant/capabilities. Read actions run immediately; write, compute and billing actions wait for approval, and every action — including declines and errors — returns an audit id and an evidence block.
Drives
A drive navigates to a page, waits for it to be ready, fills the real form state and invokes the real submit function, then posts a human summary back to chat. Registered verbs include DFT prefill/demo and DoE generation, and result panels scroll into view.
Proactive help and recipes
Account-scoped durable memory and documents, proactive cards for job started/finished/failed with executable diagnose/results actions, and reusable multi-step recipes. Long tasks detach to the background and notify through the feed.
Campaign copilot
plan_campaign parses a goal sentence into a structured plan (dataset, target, model candidates, DoE, physics checks, report format) labelled HYPOTHESIS for a human to approve. approve_and_run executes the approved stages and records the rest as skipped with a reason.
Cost and footprint sit in the same ranking as performance
TEA and LCA fold into the optimizer so a candidate is scored on whether it can actually scale.
TEA
Five costing modules — raw materials, equipment sizing, CapEx, OpEx and end-of-life/recycling credit — folded into one net-value score. Equipment sizing follows the six-tenths rule, cost ∝ (size / base)^0.6.
Life Cycle (LCA)
Global warming potential (kg CO₂eq/kg), cumulative energy demand (MJ/kg), waste E-factor and an ESG score per composition, with flowsheet evaluation and carbon-sensitivity analysis for deeper review.
Reactor engineering
Packed-bed pressure drop, conversion and sizing, plus unit-operation flowsheets (compressor → heat exchanger → packed bed → separator → recycle) with mass and energy balances. BioSTEAM templates feed the TEA and LCA layers.
Four layers of physics, one evidence discipline
From fast composition estimators to real engine runs and Slurm DFT jobs — the engine card on every response names which path produced the number.
Physics Fabric
Tiered engines exposed as honest cards: QM single-point energies, molecular docking, short MD runs and free-energy perturbation. Live availability is reported by GET /api/physics/engines so clients degrade gracefully.
Molecular Dynamics
Server-rendered .mdp decks and setup scripts, OpenMM runs with seeded replicates, LAMMPS NPT decks and vacancy workflows, RMSD/RMSF/radius-of-gyration/H-bond trajectory analysis, and Slurm templates for cluster submission.
Machine-learned potentials
A live registry reports which potentials are actually installed. Relax structures, run MLIP molecular dynamics, fit equations of state, benchmark potentials and run the embedded phonon workflow. MLIP results are PREDICTED (E3).
Docking & free energies
Receptor preparation and docking-box definition, single, array and ensemble docking, pose rescoring, FEP map/network construction, and chemical-space projection and diversity tools.
DFT & HPC
Queues Quantum ESPRESSO or VASP jobs with functional, k-point grid and cutoff on the Slurm cluster, tracked under HPC jobs. A real ASE EMT relaxation and fast composition estimators are available in-app; the composition path is a physics heuristic, not ab-initio DFT.
Quantum engines
Real semi-empirical and ab-initio single points, geometry optimisations, vibrational frequencies and basis-set convergence ladders through /api/quantum/*. Uninstalled engines report 503 — never a fabricated number.
Atomistic
Cleave fcc surfaces (111/100/110), add adsorbates, compute surface descriptors, screen favourable sites and export XYZ/POSCAR for DFT. The legacy in-app MD path is DEMO-labelled; MD and Physics Fabric are the real-trajectory routes.
Materials search
Look up and compare materials by composition, formula or properties from Materials Project-style sources and the local public mirror; OPTIMADE endpoints expose structures to external tools.
3D workbench & visualisation
One pluggable viewer over Mol*, NGL, bundled 3Dmol and backend-rendered PyMOL stills, plus CrystalKit cells, PhononKit bands/DOS, EmKit diffraction, ternary and quaternary phase diagrams, Chemiscope projections and reaction DAGs. Every result keeps its engine card.
Everything between the bench and the model
Records, parser files, inventory, automation and safety — each surface keeps its raw records and evidence class.
ELN / LIMS / inventory
First-class samples, containers, microplates, wells and aliquots, plus reagent inventory with lots, storage locations, low-stock and expiry alerts. Lineage runs as a DAG with transfer volumes and dilution factors, back to the original lot.
Record integrity & e-signatures
Experiment records are hashed into a tamper-evident chain (SHA-256 always, HMAC-SHA256 when a signing key is configured), and e-signatures are immutable SignatureRecord rows with hash chaining — a ledger to demonstrate integrity, not a regulatory certification.
Instrument parsers
XRD, GC/MS, UV-Vis, FTIR, NMR, rheology and battery-cycler files become structured tables, including JCAMP-DX, Agilent CSV and Bio-Logic MPR. Opaque binaries return an honest DEMO-labelled result, and all rows are labelled EXTRACTED.
Characterization
XRD peaks with Scherrer size and phase matching, BET multi-point area and BJH pores, TPR/TPD deconvolution, chemisorption dispersion and TOF normalization, DSC Tg/Tm and TGA — with grounded narration that interprets the numbers without inventing metrics.
Bench Logger
Records lab validation runs from a generated recipe with conditions, measurements and observations, works offline through an IndexedDB queue with pending-sync, scans barcodes and takes voice notes, and exports results to a dataset.
Synthesis recipes
Turns a target material into a step-by-step procedure with precursors, stoichiometry and scale-up notes. GEMD graphs keep material/step lineage, and recipes export to Opentrons OT-2/Flex or SiLA2.
Self-Driving Lab
A campaign plans rounds through the Active Learning engine, generates AST-linted Opentrons protocols with injected safety bounds, and applies a physical-plausibility gate: in-bounds results enter the training set as MEASURED, out-of-bounds values are flagged for human approval.
Lab edge and safety interlocks
Robots and reactors register as edge nodes with telemetry and heartbeats. Safety-envelope breaches auto-trip a software E-stop, and reset requires an authenticated operator checklist recorded in a hash-chained sign-off ledger.
Safety studio
GHS pictograms, signal word, H/P statements, NFPA 704 and precursor incompatibility checks, plus 16-section OSHA HCS / REACH Annex II SDS drafts. Automated hazard text is a screening aid — always confirm against the supplier’s official SDS.
Lab Automation (Full)
Liquid-handling deck validation, simulation and Opentrons export, elabFTW record sync, and jobflow/atomate2/custodian maker catalogs with the SDL full-loop endpoints for plan, dispatch, ingest and safety.
From messy files and papers to a modeling-ready dataset
Ingestion, extraction, enrichment and literature intelligence — with a review gate before anything becomes training data.
Data Ingestion
Imports messy data including XLSX and PDF, lets you review a parsed batch row by row, accept or reject rows, then promotes a clean batch into a Dataset with a target column. Rows are not searchable until promoted.
Datasets
A dataset registry with previews, numeric profiles, quality checks and provenance manifests, plus export paths for downstream tools. Column-role conventions and format requirements are documented under Data Preparation.
Explorer
A browsable, searchable view over your samples and materials. Save a query and feed the shortlist into prediction, screening or a campaign.
Data Enrichment
Parses free text into material/formula, process, property, value and units (LLM when configured, regex fallback otherwise), adds element properties and RDKit descriptors, and powers Prediction’s 132 composition descriptors.
Document Vision
Extracts OCR text with boxes, page layout, table grids and chart elements from scans, PDFs and screenshots. Outputs are EXTRACTED drafts a human confirms — use Ingestion to promote reviewed rows.
Literature Atlas & project libraries
Structured extraction across reaction, catalyst, loading, support, conditions, performance, DOI and confidence, with DOI ingest and an honest OpenAlex fallback. Project libraries keep papers chunked with exact paragraph/DOI/page citations, and a study-credibility review checks metric claims, leakage and baselines.
Research agent
Literature prior art → candidate generation → multi-fidelity kinetics and packed-bed simulation → TEA and precursor safety → a research dossier. A hard budget_calls cap bounds LLM calls, and computed facts stay distinct from LLM proposals.
Community exchange
Share and reuse pretrained surrogates, Opentrons protocols, GEMD recipes and featurizer pipelines. SHA-256 checksums, AST sandboxing for protocols and pickle-free Safetensors validation keep reused artifacts verifiable.
Export & integrations
RO-Crate provenance ZIPs, OPTIMADE structure endpoints, the MCP server, recipe/plate export to Opentrons and SiLA2, plus the Python SDK, REST API and webhooks for programmatic access.
Plotter & visualisation
Interactive 2D/3D scatter plots of datasets and run outputs, with Chemiscope PCA/t-SNE/UMAP projections. Every plot keeps its engine card, so DEMO or surrogate results stay labelled.
Review, sign off and publish on the same thread
Rooms, governance and publication artifacts all point back at the campaign that produced them.
LiveStudio collaboration
Presence, avatars and live cursors over SSE, a shared candidate spreadsheet with lock, tag and shortlist, and a synchronized 3D viewer. The review state machine runs DRAFT → IN_REVIEW → APPROVED → REJECTED → SENT_TO_LAB with quorum voting and SHA-256 hash-chained sign-offs.
LiveStudio comparison
Compare synthetic vs real, model vs model and candidate vs candidate in one workspace, with 3D inspection and dispatch of a chosen candidate to an Opentrons protocol.
Campaigns
Group goal-oriented research with templates, objectives and typed links to datasets, jobs, reports, atlas entries and models. Campaigns are the unit the dossier generator publishes and the copilot can plan from a goal sentence.
Dossier Reports
Dossiers assemble an executive summary, provenance graph, model performance cards, candidate leaderboard with evidence labels, TEA/LCA and the compliance trail — exported as PDF, HTML, Markdown or JSON. Missing sections are stated; missing numbers render as an em-dash, never 0.
Reports
Snapshot a run or campaign into a frozen report and share it with a public link, or draft a manuscript from a completed job.
Organizations
Create an org, invite members, manage roles and scope projects so data and literature libraries reach the right people. Publish a public org page at /orgs/<slug> and export org-level RO-Crate provenance for institutional hand-offs.
The layer that keeps every number honest
Evidence classes, reproducible benchmarks, account security and programmatic access.
Evidence classes & engine cards
Every scientific response carries an evidence class — MEASURED, COMPUTED, PREDICTED, EXTRACTED, HYPOTHESIS or DEMO — plus an engine block naming what produced it. Non-canonical labels are replaced with DEMO and a COMPUTED/MEASURED claim whose engine is not installed is downgraded, so the taxonomy stays closed.
Benchmarks & reproducibility
A public Results hub: a live pipeline on held-out rows with a reproducibility bundle, independent datasets on random and extrapolative splits, the Ni-hydrotalcite laboratory case study, MatBench tasks, battery RUL and engine benchmarks. The in-app MatBench evaluator is DEMO-labelled; the genuine published result is served separately. Methodology lives at /benchmarks/methodology.
Security & sessions
Passwords are stored as salted hashes, every session is tracked server-side with device, IP and last activity, and any session can be revoked — including the current device. Google sign-in delegates authentication to Firebase.
API keys
Per-user API keys are shown once at creation and stored hashed server-side; rotate or revoke them at any time. Webhooks notify your server when a run finishes or fails.
REST API
An interactive Swagger reference at /api-docs and a machine-readable OpenAPI 3.0 spec at /api/openapi.json. Programmatic requests are rate-limited against your tier exactly as the web app is.
SDK
A Python client for dataset upload, ingestion batches, Model Hub auto-select/GP/predict/export, physics calls and OPTIMADE structure queries, with webhooks for run events. Every non-2xx raises a typed error.
Tiers & quotas
All plans include the full pipeline; tiers differ in jobs per hour, concurrency, data points per dataset, synthetic rows per run, upload size and storage. Your own usage lives in Settings → Billing.
One thread from raw data to a published decision
Standalone modules when you need one; a connected graph when you need the whole loop.
- Full Pipeline — chains stages with automatic handoffs, so no manual export/import between steps.
- Runs — records every computation with live progress, results, logs and a rerun path.
- Campaigns — link the datasets, runs, reports and candidates that belong to one objective.
- Pilot — drives the same APIs under the same permissions, with approvals on writes.
- Dossiers — turn the campaign record into a shareable, evidence-labelled publication.
See it on your own data
Create a free account, load a demo dataset, and run the pipeline end to end — then open any card above in the live app.