Matflow Pilot

An assistant that works on the platform, not around it

Pilot reads your data, drafts work, drives real pages and runs jobs through the same APIs the manual controls use — visibly, with an explicit click only where an action is destructive or billing.

The panel

One persistent instance, additive to every manual control

Pilot is not a separate app. It is the assistant dock inside the platform, bound to the same routes, permissions and services a person clicks through by hand.

Type a prompt and Pilot plans the turn, calls real tools, renders approval cards, questions and results, and keeps working while you navigate. The instance lives above the router and is portaled into the shell grid, so a draft, a plan checklist or an in-flight drive is not lost when you move from one page to another.

Every page keeps its manual path. When the panel is closed, the forms, buttons, tables and viewers behave exactly as before; when it is open, Pilot can operate the same page visibly in front of you. Read the Pilot overview for the short version.

The panel also carries the platform contracts you would expect from an operator: capability discovery with risk tiers, structured questions when input is missing, approval previews for gated actions, live progress, and an evidence class plus audit id on every result.

Panel properties
  • Persistent across pages — one instance is mounted above the app routes and portaled into the shell grid on desktop (overlay on mobile), so the conversation, plan and in-flight work survive navigation.
  • Additive, never a takeover — closing the panel removes nothing. Every form, button and workflow stays exactly where it was; Pilot is an extra path, not a replacement.
  • Three tabs — Chat for the conversation, Artifacts for generated files and images, Tools for the live capability browser and your saved memory.
  • Page-aware — Pilot receives the current route and a page snapshot, and each page advertises its registered drive verbs as drive_actions.
  • Keyboard access — Ctrl+Shift+Space toggles the panel from anywhere in the app.
  • Attachments — images, pasted CSV/TSV/JSON, and documents (PDF, Office, CSV, text) up to 15 MB. Images are described once by the vision model so later turns stay cheap.
Capability catalog

217 live capabilities across 43 domains

Pilot does not have a private API. It calls the same catalog the platform serves from GET /api/assistant/capabilities (login required, free), and every entry declares its id, label, description, domain, risk tier, parameter schema and endpoint.

data · 26
Data and ingestion

List, preview and profile owned datasets, compare two side by side, create one from pasted CSV or JSON, and delete with approval. list_datasets · dataset_preview · dataset_profile · dataset_compare · create_dataset_from_text · delete_dataset.

design · 3
Design of experiments

doe_generate builds a run matrix, doe_bench_plan turns it into a bench-ready step list, and doe_materialize registers the design as a real dataset.

modeling · 20
Modeling and screening

Auto-select engine info, quick-prediction specs, model training and prediction specs, plus screening and optimization specs — validated and costed, with the real billed run staying on its own route.

physics · 9 + dft · 9
Physics, DFT and phonons

Engine cards, costed physics specs, phonon submission, and dft_submit — which runs for real with credit admission: fast small-box recipes execute synchronously, heavier jobs detach as background DftJob rows, and a DEMO fallback happens only when engines are uninstalled and you explicitly asked for a demo.

lab · 24 + inventory · 4
Lab, ELN and inventory

Samples, plates, barcodes, reagent stock and SDL campaigns — while e-signatures and robot dispatch deliberately hand off to the interactive pages that own the ceremony.

characterization · 1 + kinetics + reactor
Instruments and bench analytics

characterization_parse detects an instrument-file format and reports measured file stats; kinetics_fit returns an Arrhenius fit and reactor_simulate estimates PFR/CSTR conversion analytically.

literature · 7 + reporting · 10
Literature and reporting

Search the Literature Atlas, link datasets, runs and reports to a campaign, compile a dossier draft, and freeze report files with a download reference.

billing · 3 + media · 1
Billing and media

Plans and balance are readable; checkout needs the interactive redirect and the assistant never charges. generate_image runs only when an image model is configured — otherwise it declines and links Settings.

The catalog is data, not a promise.Each capability either executes for real against owned data or explicitly declines with a reason and a link to the page that owns the interactive step — model deployment, robot dispatch, e-signatures, live lab interlock state and billing checkout all work that way. The count above is the live catalog: 14 core, 55 Pilot and 148 domain-extension actions.
Autonomy policy

Reads run free. Destructive actions and billing stop for your click.

The governing rule is confirm-destructive-only. Every capability carries a risk tier, and anything gated returns a human preview plus a signed token before it can run — never an invisible write.

  • read — runs immediately: lists, previews, profiles, search, chart specs and engine cards. No token, no charge.
  • write — maps to the real platform route with your ownership and limits applied. Destructive writes — delete, publish, sign, dispatch — always stop for Approve.
  • compute — validated and costed; every compute path admits credits against your balance before anything launches, so a zero balance returns 402.
  • billing — always requires the click, and checkout returns the billing link instead of charging inside the assistant.
  • Approvals — phase one returns a preview and a signed token; phase two only runs after confirmed:true with that exact token. Detached work and injection-flagged turns are forced onto this path.
Two-phase confirmation, by construction

A gated action is previewed first, then executed only after you approve the exact parameters that were shown.

  • Signed with the app secret and bound to the capability, the params hash and your user id
  • 10-minute TTL, then the token is dead
  • Single-use, claimed atomically so concurrent double-confirms cannot execute twice
  • Replay or reuse is rejected with 400; another user’s token is rejected with 403
Credits are admitted before compute.DOE generation, model training and prediction, DFT, physics, screening, batteries, alloys, polymers, formulations, catalysis, optimization, MD, MLIP and phonon each call a credit admission check first. If the balance cannot cover the job, the answer is 402 — before any compute is launched, and before any charge.
Visible page automation

Pilot drives the real page, not a shadow copy

A drive is a verb a page registers for Pilot. The assistant navigates to the route, waits until the page is ready, fills the page’s real form state, invokes the page’s real submit function, scrolls the result into view and posts a short human summary back to chat — never raw JSON.

16
app routes with registered drive verbs
backend DRIVEABLE_PAGES catalog
29
verbs auto-advertised in the page snapshot
usePilotDrive registrations
10 min
confirmation token TTL
POST /api/assistant/execute
150 s
client stream idle budget
PilotPanel
dft / doe
Physics and design studios

dft:run_demo, dft:prefill and dft:set_view on the DFT studio; doe:generate on the DOE page fills the visible form and shows the run matrix.

prediction / screening
Models and libraries

prediction:train, prediction:set_target and prediction:what_if; screening:screen and screening:generate; hub:autoselect and hub:predict on the Model Hub.

data / lab
Datasets, ELN and runs

datasets:create from raw CSV/TSV through the real upload path, datasets:delete, dataset:update_columns; eln:create_sample and eln:create_plate; runs:cancel.

everything else
The rest of the studio

evaluation:run, optimization:start/export, synthesis:generate/retrosynthesis, materials:search/hpc_submit, batteries:predict, alloys:predict_phase/calphad, sdl:plan_campaign.

  • Canonical routes — aliases such as /dft are canonicalized to /studio/dft before a card is created, so a drive never double-navigates.
  • Double-invoke guarded — handlers refuse a second submit while a run is in flight, so a chat click and a manual click cannot race.
  • Always visible — every handler scrolls its result panel into view; a run is never submitted invisibly below the fold.
  • Honest fallback — when a handler is not mounted the card falls back to a navigate button and reports the failure — never a fabricated success.
  • Bounded automation — the SDL drive plans a campaign but never dispatches robots, ingests results or touches the safety gates; those stay manual.

Drives are covered in depth in the drives documentation, alongside the exact verbs each page registers.

Questions, plans and progress

Structured input instead of guessing

When a task needs a decision, Pilot asks one crisp question with real options — and when a turn needs many steps, you watch them advance rather than reading a monologue.

A structured question carries one to six options, each with an optional description, an optional recommended choice that must match one of the option labels, and a free-text field that is always available. In the panel, choosing an option sends “My choice: <label>” and the task continues without re-prompting. Only one question is allowed per turn.

Multi-step turns stream plan, tool and progress events. The plan renders as a checklist and each step moves from pending to doing to done as the matching tool executes, so progress is observable while it happens.

Long work detaches instead of stalling: past the tool-iteration budget or eight minutes, the task is persisted, a background card appears and the feed notifies you when it finishes. A detached turn can run reads on its own, but any write or compute comes back as an approval card — no background turn can approve or spend on its own.

Sample question
Which structure should I use for the demo run?
H2O · recommended
NH3
Or type your own answer…
Stop anytime. The composer turns into a stop button while a turn runs, and the client aborts a silent stream after 150 seconds so a hung model never freezes the panel.
Memory and documents

Context that persists, scoped strictly to you

Pilot can carry durable facts between sessions and read documents you attach, without ever reaching outside your account.

GET/POST/DELETE /api/assistant/memory
What Pilot remembers

Say “remember that…”, “I prefer…” or “call me…” and the capture is deterministic. After a turn, a bounded background pass may also derive up to two durable facts. Memory is scoped to your account and can be forgotten item by item or cleared entirely from the Tools tab.

GET/POST /api/assistant/documents · DELETE /api/assistant/documents/<id>
Documents you attach

PDF, DOCX, PPTX, XLSX/XLSM, CSV/TSV/TXT/MD/JSON/log/dat are text-extracted server-side, up to 15 MB per file; images are transcribed once by the vision model and stored as text. Each turn receives the most query-relevant excerpts, not the whole file.

User-scoped
Only your rows

Memory and document rows are only read, listed or deleted when they belong to the session user, and the memory block is passed to the model as data to use naturally — never recited back as a list.

Proactive help

The watcher that speaks only when it has something real

Proactive messages are composed by small, single-purpose watchers that read platform state and write to your feed. They run on the system budget — no credit fee — and they are bounded so they never become noise.

Intake
First-run guidance

A brand-new account with no datasets and no jobs gets an intake message with five one-tap intents: load data, design experiments, run a simulation, log lab work, or show me around. Accounts with data get the standard welcome instead.

Job transitions
Started, finished, failed

Completion cards carry executable View results (job_results) plus quick replies; failure cards carry Diagnose failure (diagnose_job) grounded in the real error log, with an LLM diagnosis and a strict template fallback when the model is unavailable.

Idle
One gentle nudge

Accounts whose last activity is one to four days old get a single nudge, deduped for a week — enough to be useful, not enough to nag.

Feed discipline
Badge first, thread second

When the panel is closed the messages wait as an unread badge; opening the panel injects them into the conversation. Dedupe keys stop double-notifying, per-scan LLM caps bound spend, and the unread list is capped so the feed never floods.

Recipes and background tasks

Common multi-step goals as one reliable call

A recipe names the ordered capabilities for a known workflow and runs them under the same autonomy policy as a single action — capped at six steps, with the same approvals and the same honest reporting.

data
data_quality

Profile a dataset, then run its deterministic quality report: duplicates, missingness, constant and mixed-unit columns.

modeling
train_predict

Profile the training dataset, then launch the real async train-and-predict job and return its job id to poll.

screening
screening_campaign

Generate candidates around a query molecule, screen them, and optionally score the list with a completed prediction model.

literature
atlas_to_dataset

Ingest an open-access DOI into the Literature Atlas, extract supplied paper text, and export chosen Atlas rows into a dataset.

campaigns
campaign_bootstrap

Create a campaign, generate its DOE design, and link an existing entity when the ids are supplied.

dft
dft_ladder

Submit one real DftJob per rung of a basis-set convergence ladder, then list the queue to confirm.

literature
literature_review

Search the Atlas for a query and, when paper text is supplied, extract structured catalyst records.

A recipe pauses, it does not overreach.When a step is destructive or billing, the recipe stops and prepares an Approve card for your click; step summaries report exactly what ran and what stayed validated-only. Long-running recipes detach to the background and resolve through the feed, exactly like a single long turn.
Honesty and audit

The assistant is not allowed to bluff

Honesty here is structural, not a prompt request: evidence labels, a critic pass, pseudo-call recovery and audit receipts are enforced in code.

Evidence classes
Every number carries its class

MEASURED (E5) · COMPUTED (E4) · PREDICTED (E3) · EXTRACTED (E2) · HYPOTHESIS (E1) · DEMO (E0). A demo or surrogate result can never be mistaken for a measured one; engine cards name the producer.

Critic pass
No narrated actions without evidence

A draft that claims a completed platform action without a same-turn tool result triggers exactly one bounded retry with tools available. If it still claims the action, a deterministic honesty note is appended before the answer is shown.

Pseudo-call recovery
No raw markup in chat

Text-rendered tool calls are recovered once per turn for an allowlist of read and card-only tools, and all such markup is stripped from displayed answers. Raw function-call XML never reaches the conversation.

Audit
Every attempt is logged

Every action attempt is audit-logged, including previews, declines and failures. Results and previews return an audit id and an evidence class, and the panel shows a receipt when an action runs, so the trail survives the conversation.

Worked example

What a prompt actually does

A concrete walkthrough using only real capability names, real routes and the real confirmation flow.

Prompt

“Clean up my workspace: profile dataset 12, then delete it.”

  1. 01dataset_profile → GET /api/datasets/12/profile. A read, so it runs immediately. It returns numeric count/mean/std/min/max and the column list, labelled MEASURED with an audit id.
  2. 02delete_dataset → DELETE /api/datasets/12. A destructive write. The bridge validates the id and confirms ownership (a foreign id is 404), then returns the preview “Delete dataset #12 (registration + staging file). This cannot be undone.” The panel renders an Approve card and the backend mints the signed token.
  3. 03You click Approve. The panel re-posts the capability, the same params, confirmed:true and the token. The token is verified against the capability, the params hash and your user id, then the dataset row and staging file are really deleted. The response carries evidence MEASURED and a fresh audit id; a replay is rejected with 400 and another user’s token with 403.
  4. 04Missing input becomes a question. If a required value is missing, ask_user returns structured options, a recommended choice and free text instead of guessing. Choosing one sends “My choice: …” and the task continues.

Every capability id and route above is real: they come from GET /api/assistant/capabilities and the endpoints its entries declare. The same catalog powers the panel’s Tools tab.

FAQ

Matflow Pilot questions, answered plainly

Where does Pilot live?
It is mounted once above the authenticated app routes and portaled into the shell on desktop (overlay on mobile), so it persists across page navigation. Public marketing pages do not mount it. Ctrl+Shift+Space toggles it.
Does Pilot have special access?
No. It calls the same capability catalog the platform exposes through the same session, with your ownership and permission checks. Foreign ids — a dataset, campaign or job that is not yours — are refused, and lists are scoped to you.
Can Pilot act without my click?
Reads run immediately. Destructive actions and billing always require an explicit Approve through a preview card. Detached background turns and prompt-injection-flagged turns are forced onto the preview path, so background work cannot spend on its own.
What is a confirmation token?
A signed, single-use token bound to the capability, the exact params hash and your user id. It expires after 10 minutes, is consumed atomically across workers, rejects replay and reuse with 400, and rejects another user’s token with 403.
What happens when a page drive fails?
The drive reports the failure and falls back to a navigate button, so you can finish manually. It never reports a fake success, and every handler scrolls its result into view so a run is never invisible.
How does Pilot use my credits?
Every compute path admits credits before anything launches; a zero balance returns 402 before spend. Read actions incur no compute charge. Billing actions never charge inside the assistant — checkout hands you the billing link.
Can I control memory?
Yes. Explicit captures like “remember that…” are deterministic, derived facts are bounded and background-only, and the Tools tab lets you add, forget or clear everything. Memory is scoped to your account.
What happens when the panel is closed?
Proactive job and idle messages wait as an unread badge. Opening the panel injects them into the thread. There is no toast spam, and manual page controls are never affected.

Put Pilot to work on your data

Create a free account and ask it to profile a dataset, draft a DOE, or cost a physics run — reads land immediately, and destructive or billing actions wait for your click.