Built for the way materials & chemistry R&D actually happens
Nine recurring workflows. Each lists the problem, the modules that solve it with links into the app, the outcome, and what the numbers can honestly support.
Shared note for all workflows below: confirm that synthetic rows match reality before modeling on them — Evaluation scores each batch first. Every result then carries an evidence class (MEASURED / COMPUTED / PREDICTED / EXTRACTED / HYPOTHESIS / DEMO) plus an engine card. See /science.
Developing a formulation when runs are scarce
The problem: You have a sparse screening campaign, the mixture must sum to 100%, and VOC or compliance targets narrow the space before performance even enters the picture.
- Formulations — Generate a constrained simplex-lattice or centroid mixture design, screen solvents by Hansen RED and read BOM property bands.
- Data Synthesis — Augment the sparse campaign into a larger, statistically faithful training set — then gate it with Evaluation.
- Prediction — Train the explainable ensemble and read which components drive activity and stability, with conformal intervals.
- Optimization — Search compositions under 100%-sum and per-component bounds; rank the trade-off front by TOPSIS.
- TEA — Price the top candidates — materials, equipment, CapEx/OpEx and end-of-life credit — before committing lab time.
The outcome: A ranked, constraint-feasible formulation list with predicted properties, intervals and cost per batch.
Electrolytes and cells that survive cycling
The problem: Voltage stability, transport, cycle-life degradation and raw-material cost pull against each other, and the composition × condition space is far too large to sweep by hand.
- Battery Depth — Fit capacity fade and forecast remaining useful life from cycler rows with a split-conformal interval.
- Batteries Core — Screen multi-component electrolytes with VTF transport, voltage windows and aging analytics.
- Instruments — Parse battery-cycler vendor files into structured rows instead of retyping them.
- Optimization — Balance capacity, cycle life and material cost under composition and bound constraints.
- LCA — Review the environmental footprint (GWP, cumulative energy demand) of shortlisted candidates before committing lab time.
The outcome: Electrolyte and cell candidate list with predicted retention and conductivity, RUL intervals, cost and footprint.
Finding reaction conditions that hit yield and safety limits
The problem: You know the chemistry; what varies is temperature, pressure, concentration and flow — a condition space you cannot sweep exhaustively or safely.
- Design of Experiments — Generate a Latin Hypercube, Box–Behnken or Sobol plan that maximizes information per run.
- Kinetics — Fit Arrhenius Ea with a 95% confidence interval and check Weisz–Prater and Mears transport limits.
- Prediction — Build the ensemble from the completed runs with conformal intervals on every prediction.
- Optimization — Maximize yield and selectivity while constraining conditions to safe equipment limits.
- Reactor — Size a packed bed for the target conversion and check pressure drop before scale-up.
The outcome: Condition windows with predicted yield and interval per window, plus a sized reactor for the best window.
Screening catalysts and forecasting how long they last
The problem: Hundreds of formulations, supports and loadings are plausible, descriptors are missing for most candidates, and deactivation — not initial activity — usually decides the winner.
- Catalysis Depth — Run a descriptor screen with undecidable flags, Sabatier volcano curves and a coverage-dependent microkinetic solve.
- Reaction Network — Build the elementary-step DAG with BEP scaling and identify the rate-determining step.
- Kinetics — Separate intrinsic from transport-limited Ea with Weisz–Prater and Mears diagnostics.
- Stability — Fit deactivation curves, price the lost activity and forecast time-on-stream decay.
- Reactor — Size the packed bed and fold reactor sizing into a budget-aware campaign plan.
The outcome: Shortlist with screening scores, stability forecast and the reactor sizing/cost needed to plan a campaign.
- Catalysis Depth — descriptor screening, volcano curves and microkinetics
- Reaction Network — elementary-step DAGs and rate-determining steps
- Stability — deactivation fits and time-on-stream forecasts
Deciding whether a bench winner can scale
The problem: A candidate looks great at the bench, but nobody knows what it costs at pilot scale — or whether the environmental footprint disqualifies it.
- Prediction — Estimate performance with uncertainty bands so scale-up decisions carry their error bars.
- TEA — Price raw materials, size equipment with the six-tenths rule, and fold in CapEx, OpEx and end-of-life credit.
- LCA — Add GWP, cumulative energy demand, E-factor and an ESG score per candidate.
- Optimization — Re-run the search with net value and footprint as objectives alongside performance.
- Pipeline — Chain the loop so every new candidate is re-costed automatically.
The outcome: Candidates ranked by net value and environmental footprint alongside activity and stability.
Turning papers, PDFs and scans into a dataset
The problem: Your best data is locked in published tables, supplementary files and old notebooks — and retyping it is slow and error-prone.
- Document Vision — Extract OCR text, layout, table grids and chart elements from scans and PDFs.
- Ingestion — Review the parsed batch row by row and promote a clean batch to a Dataset with a target column.
- Data Enrichment — Parse free text into material, process, property and units, then add composition and RDKit descriptors.
- Literature Atlas — Search structured literature records and ingest DOIs with provenance kept.
- Prediction — Train the ensemble on the enriched corpus with conformal intervals.
The outcome: An enriched dataset with source citations and evidence labels, ready for modeling.
Closing the loop from bench records to the next experiment
The problem: The notebook, the instrument files and the models live in different worlds, so results never flow back into the next design.
- ELN / Samples — Register samples, containers and plates, and keep lineage back to the original lot.
- Instruments — Parse instrument vendor files into structured, reviewable rows.
- Bench Logger — Log runs and measurements — offline if needed — and export them to a dataset.
- Prediction — Rebuild the explainable ensemble from the merged bench record.
- Active Learning — Propose the next batch by expected information gain from the updated surrogate.
- Self-Driving Lab — When the loop is ready, automate it with safety-gated robot rounds that ingest measured results back.
The outcome: A bench record linked to models, where each new result updates the surrogate and the next-batch proposal.
Designing an alloy that is single-phase, strong and affordable
The problem: The high-entropy alloy space is enormous, phase formation is decided by competing thermodynamic rules, and cost pulls against strength.
- Alloys — Featurize with VEC, size mismatch, Miedema enthalpy and Omega; get a phase and sigma-risk verdict.
- Advanced Viz — Check ternary/quaternary CALPHAD sections and Scheil solidification for candidate systems.
- MLIP — Relax structures and screen stability with installed machine-learned potentials.
- DFT & HPC — Run a real EMT relaxation in-app, then submit QE or VASP jobs to the Slurm cluster.
- Optimization — Search compositions toward target properties under a $/kg budget.
The outcome: A composition shortlist with predicted phase, mechanical estimates, cost and a path to first-principles validation.
- Metals & HEA Alloys — Miedema features, phase verdicts and cost-aware design
- Advanced Viz — CALPHAD sections and Scheil solidification
- DFT — real EMT relaxations and QE/VASP Slurm jobs
From a target to a shortlist of makeable molecules
The problem: The virtual library is huge, the binding hypothesis is uncertain, and limited bench time has to go to the candidates that are both active and synthesisable.
- Generative & QSAR — Build fingerprints and descriptors, train a QSAR model, and sample or score generated candidates.
- Virtual Screening — Rank the library against a reference structure, filtered by a SMARTS scaffold you can make.
- Similarity — Search the public mirror and your own datasets with Morgan/Tanimoto, capped and reported.
- Docking & FEP — Prepare receptors and boxes, run ensemble or array docking, then build RBFE networks for the top series.
- Physics Fabric — Price out interactions with QM single points and FEP estimates at the accuracy you need.
- Synthesis Recipes — Turn the winner into a step-by-step procedure with a retrosynthesis path.
The outcome: A ranked, makeable shortlist with predicted activity, docking scores and a retrosynthesis route.
- Generative & QSAR — fingerprints, QSAR training and generative sampling
- Docking & FEP — receptor prep, ensemble docking and RBFE networks
- Physics Fabric — QM single points and free-energy estimates
Not sure which workflow fits?
Describe your project to Matflow Pilot, or read the docs — and if you are still stuck, our team will help you map your data to the right modules.