Built for the data you actually have
Matflow is a web-based scientific R&D platform for chemistry and materials. It turns small, sparse experimental datasets into ranked candidates — with an evidence class and an engine card on every number, and the limits stated plainly.
Make sparse experimental data workable
The data gap is the real bottleneck
Typical campaigns — catalysis, formulations, batteries, polymers — produce tens of rows, not thousands. Matflow exists to make that scarcity workable: generate statistically faithful synthetic rows, prove the augmentation is trustworthy before using it, train a model that can explain itself, and search it for candidates worth testing in the lab.
The platform is deliberately explicit about what each number is — measured, computed, predicted, extracted, hypothesis or demo. Recommendations arrive with their context attached, not as unverifiable promises.
- A web platform — no-code, browser-based, nothing to install.
- A four-stage pipeline — each stage runs standalone or chained end to end.
- Evidence-labelled by default — every result carries its source and evidence class.
- Self-serve — a free account and monthly plans from $99; every module is included on every tier, and the public limits are the enforced ones.
A number without its source is a rumour
Matflow treats honesty as an engineering constraint, not a slogan. Every pipeline output is tagged with one of six evidence classes and linked to an engine card naming the engine and its limitations.
- MEASURED — raw experimental or instrument result.
- COMPUTED — a real engine run — M3GNet, PyBaMM, pycalphad and others.
- PREDICTED — a registered model output with uncertainty.
- EXTRACTED — pulled from literature or documents, queued for human review.
- HYPOTHESIS — an LLM or generative proposal for a human to test.
- DEMO — synthetic, heuristic or demo fallback.
What you can do on Matflow
From a single CSV to a publication dossier — the canonical module inventory lives on /features.
Load the bundled demo dataset or upload a CSV, then run synthesis, evaluation, prediction and optimization with guided forms — no machine-learning code required.
Fold raw materials, equipment sizing, CapEx/OpEx and end-of-life value into one net-value score, so cost sits on the Pareto front next to performance.
MD decks, learned interatomic potentials, docking, quantum jobs and phonons — each result names the engine that actually ran and its evidence class.
Instrument parsing, ELN/LIMS records with sample lineage, inventory and GHS safety, self-driving campaigns and an offline bench logger — in the same workspace.
Pilot reads your data, drafts work and drives real pages through the same APIs as the UI. Reads run free; writes and compute wait for your Approve click.
Engine checks, extraction gold suites, MatBench-style tasks and the base-model corpus — published, dated and reproducible from the endpoints they link to.
How it is built, and how it is run
The same evidence discipline applies to the platform itself: documented surfaces, public limits, inspectable releases.
Web platform
A React 19 + Vite single-page client delivers the no-code workspace; a Flask API service underneath orchestrates the compute. There is nothing to install — the app runs in the browser.
See the module catalogue →Self-serve accounts
Create a free account and explore immediately; job execution opens once an administrator approves the account, a deliberate gate that keeps compute for vetted research accounts. All plans include the full pipeline — tiers differ in throughput, not features.
Compare plans →API & SDK
The same surface is available programmatically: an OpenAPI 3.0 spec served at /api/openapi.json, an interactive Swagger reference at /api-docs, and API keys issued from Settings. The Python client quickstarts live at /sdk. Programmatic requests share your tier’s rate limits.
Read the API overview →Containerised client/server
The client and API ship as containers orchestrated with Docker Compose; heavy physics engines run as separate containers and are probed live, so a result always names the engine that actually ran.
Physics engine cards →Public release history
Features ship continuously. The changelog tracks releases from v1.0 through v2.6, and the roadmap states what is being built next — including what is explicitly not done yet.
Read the changelog →What we hold to
Four commitments that shape the platform more than any feature.
Data-driven workflows report held-out validation where the available data support a split; analytic and screening modules disclose their evidence class instead of being presented as laboratory validation.
Every trained ensemble ships global and local feature impact plus permutation importance. A model you cannot interrogate is not a result you should act on.
Files are scoped to your account and isolated at the request boundary. Models trained for you are trained on your data alone and are never merged into a shared model.
We state what the platform cannot do as plainly as what it can — extrapolation flags, small-data caveats and labelled DEMO fallbacks included.
Who it is for
Matflow is built for people who run experiments, not model zoos.
- Experimental scientists, researchers and engineers — in materials science, chemistry and chemical engineering — no machine-learning expertise required.
- Small-data teams — whose campaigns produce tens to hundreds of rows and who need defensible candidates from them.
- Anyone who must justify a recommendation — every result ships an evidence class and, where the data support a split, held-out validation.
- Teams closing the loop — model → lab → model, with the ELN, bench logger, DoE and active learning in the same workspace.
What Matflow is not
Stated as plainly as the capabilities. If a limitation matters to your decision, treat it as a decision input — not fine print.
Small-data caveats
The pipeline is built for tens of rows; “cope with” is not “perform as well as with thousands.” Confidence intervals widen and extrapolation flags fire more often exactly when data is scarcest.
Synthetic data is not a substitute for experiments
Augmentation preserves statistical structure well enough to train a model, but it cannot manufacture information that was never measured.
Extrapolation risk
The ensemble is a strong interpolator inside the training envelope and a flagged guesser outside it. Treat flagged predictions as hypotheses for the lab, not results.
Heuristic optimization
Adaptive and evolutionary searches find good Pareto fronts, not provably optimal ones; no global optimum is guaranteed.
Screening-grade sustainability
LCA figures are COMPUTED estimates over a stated boundary, not a certified ISO 14040 study; the licensed ecoinvent database is never bundled.
Labelled fallbacks
When a physics or analytic engine is unavailable, results are labelled DEMO, PREDICTED or HYPOTHESIS — never presented as measured data.
Try the pipeline on your own data
Create a free account and run the full pipeline on the bundled demo dataset first. Results carry evidence labels, and the benchmark numbers are public.