An assistant that works on the platform, not around it
Pilot reads your data, drafts work, drives real pages and runs jobs through the same APIs the manual controls use — visibly, with an explicit click only where an action is destructive or billing.
One persistent instance, additive to every manual control
Pilot is not a separate app. It is the assistant dock inside the platform, bound to the same routes, permissions and services a person clicks through by hand.
Type a prompt and Pilot plans the turn, calls real tools, renders approval cards, questions and results, and keeps working while you navigate. The instance lives above the router and is portaled into the shell grid, so a draft, a plan checklist or an in-flight drive is not lost when you move from one page to another.
Every page keeps its manual path. When the panel is closed, the forms, buttons, tables and viewers behave exactly as before; when it is open, Pilot can operate the same page visibly in front of you. Read the Pilot overview for the short version.
The panel also carries the platform contracts you would expect from an operator: capability discovery with risk tiers, structured questions when input is missing, approval previews for gated actions, live progress, and an evidence class plus audit id on every result.
- Persistent across pages — one instance is mounted above the app routes and portaled into the shell grid on desktop (overlay on mobile), so the conversation, plan and in-flight work survive navigation.
- Additive, never a takeover — closing the panel removes nothing. Every form, button and workflow stays exactly where it was; Pilot is an extra path, not a replacement.
- Three tabs — Chat for the conversation, Artifacts for generated files and images, Tools for the live capability browser and your saved memory.
- Page-aware — Pilot receives the current route and a page snapshot, and each page advertises its registered drive verbs as drive_actions.
- Keyboard access — Ctrl+Shift+Space toggles the panel from anywhere in the app.
- Attachments — images, pasted CSV/TSV/JSON, and documents (PDF, Office, CSV, text) up to 15 MB. Images are described once by the vision model so later turns stay cheap.
217 live capabilities across 43 domains
Pilot does not have a private API. It calls the same catalog the platform serves from GET /api/assistant/capabilities (login required, free), and every entry declares its id, label, description, domain, risk tier, parameter schema and endpoint.
List, preview and profile owned datasets, compare two side by side, create one from pasted CSV or JSON, and delete with approval. list_datasets · dataset_preview · dataset_profile · dataset_compare · create_dataset_from_text · delete_dataset.
doe_generate builds a run matrix, doe_bench_plan turns it into a bench-ready step list, and doe_materialize registers the design as a real dataset.
Auto-select engine info, quick-prediction specs, model training and prediction specs, plus screening and optimization specs — validated and costed, with the real billed run staying on its own route.
Engine cards, costed physics specs, phonon submission, and dft_submit — which runs for real with credit admission: fast small-box recipes execute synchronously, heavier jobs detach as background DftJob rows, and a DEMO fallback happens only when engines are uninstalled and you explicitly asked for a demo.
Samples, plates, barcodes, reagent stock and SDL campaigns — while e-signatures and robot dispatch deliberately hand off to the interactive pages that own the ceremony.
characterization_parse detects an instrument-file format and reports measured file stats; kinetics_fit returns an Arrhenius fit and reactor_simulate estimates PFR/CSTR conversion analytically.
Search the Literature Atlas, link datasets, runs and reports to a campaign, compile a dossier draft, and freeze report files with a download reference.
Plans and balance are readable; checkout needs the interactive redirect and the assistant never charges. generate_image runs only when an image model is configured — otherwise it declines and links Settings.
Reads run free. Destructive actions and billing stop for your click.
The governing rule is confirm-destructive-only. Every capability carries a risk tier, and anything gated returns a human preview plus a signed token before it can run — never an invisible write.
- read — runs immediately: lists, previews, profiles, search, chart specs and engine cards. No token, no charge.
- write — maps to the real platform route with your ownership and limits applied. Destructive writes — delete, publish, sign, dispatch — always stop for Approve.
- compute — validated and costed; every compute path admits credits against your balance before anything launches, so a zero balance returns 402.
- billing — always requires the click, and checkout returns the billing link instead of charging inside the assistant.
- Approvals — phase one returns a preview and a signed token; phase two only runs after confirmed:true with that exact token. Detached work and injection-flagged turns are forced onto this path.
A gated action is previewed first, then executed only after you approve the exact parameters that were shown.
- Signed with the app secret and bound to the capability, the params hash and your user id
- 10-minute TTL, then the token is dead
- Single-use, claimed atomically so concurrent double-confirms cannot execute twice
- Replay or reuse is rejected with 400; another user’s token is rejected with 403
Pilot drives the real page, not a shadow copy
A drive is a verb a page registers for Pilot. The assistant navigates to the route, waits until the page is ready, fills the page’s real form state, invokes the page’s real submit function, scrolls the result into view and posts a short human summary back to chat — never raw JSON.
dft:run_demo, dft:prefill and dft:set_view on the DFT studio; doe:generate on the DOE page fills the visible form and shows the run matrix.
prediction:train, prediction:set_target and prediction:what_if; screening:screen and screening:generate; hub:autoselect and hub:predict on the Model Hub.
datasets:create from raw CSV/TSV through the real upload path, datasets:delete, dataset:update_columns; eln:create_sample and eln:create_plate; runs:cancel.
evaluation:run, optimization:start/export, synthesis:generate/retrosynthesis, materials:search/hpc_submit, batteries:predict, alloys:predict_phase/calphad, sdl:plan_campaign.
- Canonical routes — aliases such as /dft are canonicalized to /studio/dft before a card is created, so a drive never double-navigates.
- Double-invoke guarded — handlers refuse a second submit while a run is in flight, so a chat click and a manual click cannot race.
- Always visible — every handler scrolls its result panel into view; a run is never submitted invisibly below the fold.
- Honest fallback — when a handler is not mounted the card falls back to a navigate button and reports the failure — never a fabricated success.
- Bounded automation — the SDL drive plans a campaign but never dispatches robots, ingests results or touches the safety gates; those stay manual.
Drives are covered in depth in the drives documentation, alongside the exact verbs each page registers.
Structured input instead of guessing
When a task needs a decision, Pilot asks one crisp question with real options — and when a turn needs many steps, you watch them advance rather than reading a monologue.
A structured question carries one to six options, each with an optional description, an optional recommended choice that must match one of the option labels, and a free-text field that is always available. In the panel, choosing an option sends “My choice: <label>” and the task continues without re-prompting. Only one question is allowed per turn.
Multi-step turns stream plan, tool and progress events. The plan renders as a checklist and each step moves from pending to doing to done as the matching tool executes, so progress is observable while it happens.
Long work detaches instead of stalling: past the tool-iteration budget or eight minutes, the task is persisted, a background card appears and the feed notifies you when it finishes. A detached turn can run reads on its own, but any write or compute comes back as an approval card — no background turn can approve or spend on its own.
Context that persists, scoped strictly to you
Pilot can carry durable facts between sessions and read documents you attach, without ever reaching outside your account.
Say “remember that…”, “I prefer…” or “call me…” and the capture is deterministic. After a turn, a bounded background pass may also derive up to two durable facts. Memory is scoped to your account and can be forgotten item by item or cleared entirely from the Tools tab.
PDF, DOCX, PPTX, XLSX/XLSM, CSV/TSV/TXT/MD/JSON/log/dat are text-extracted server-side, up to 15 MB per file; images are transcribed once by the vision model and stored as text. Each turn receives the most query-relevant excerpts, not the whole file.
Memory and document rows are only read, listed or deleted when they belong to the session user, and the memory block is passed to the model as data to use naturally — never recited back as a list.
The watcher that speaks only when it has something real
Proactive messages are composed by small, single-purpose watchers that read platform state and write to your feed. They run on the system budget — no credit fee — and they are bounded so they never become noise.
A brand-new account with no datasets and no jobs gets an intake message with five one-tap intents: load data, design experiments, run a simulation, log lab work, or show me around. Accounts with data get the standard welcome instead.
Completion cards carry executable View results (job_results) plus quick replies; failure cards carry Diagnose failure (diagnose_job) grounded in the real error log, with an LLM diagnosis and a strict template fallback when the model is unavailable.
Accounts whose last activity is one to four days old get a single nudge, deduped for a week — enough to be useful, not enough to nag.
When the panel is closed the messages wait as an unread badge; opening the panel injects them into the conversation. Dedupe keys stop double-notifying, per-scan LLM caps bound spend, and the unread list is capped so the feed never floods.
Common multi-step goals as one reliable call
A recipe names the ordered capabilities for a known workflow and runs them under the same autonomy policy as a single action — capped at six steps, with the same approvals and the same honest reporting.
Profile a dataset, then run its deterministic quality report: duplicates, missingness, constant and mixed-unit columns.
Profile the training dataset, then launch the real async train-and-predict job and return its job id to poll.
Generate candidates around a query molecule, screen them, and optionally score the list with a completed prediction model.
Ingest an open-access DOI into the Literature Atlas, extract supplied paper text, and export chosen Atlas rows into a dataset.
Create a campaign, generate its DOE design, and link an existing entity when the ids are supplied.
Submit one real DftJob per rung of a basis-set convergence ladder, then list the queue to confirm.
Search the Atlas for a query and, when paper text is supplied, extract structured catalyst records.
The assistant is not allowed to bluff
Honesty here is structural, not a prompt request: evidence labels, a critic pass, pseudo-call recovery and audit receipts are enforced in code.
MEASURED (E5) · COMPUTED (E4) · PREDICTED (E3) · EXTRACTED (E2) · HYPOTHESIS (E1) · DEMO (E0). A demo or surrogate result can never be mistaken for a measured one; engine cards name the producer.
A draft that claims a completed platform action without a same-turn tool result triggers exactly one bounded retry with tools available. If it still claims the action, a deterministic honesty note is appended before the answer is shown.
Text-rendered tool calls are recovered once per turn for an allowlist of read and card-only tools, and all such markup is stripped from displayed answers. Raw function-call XML never reaches the conversation.
Every action attempt is audit-logged, including previews, declines and failures. Results and previews return an audit id and an evidence class, and the panel shows a receipt when an action runs, so the trail survives the conversation.
What a prompt actually does
A concrete walkthrough using only real capability names, real routes and the real confirmation flow.
“Clean up my workspace: profile dataset 12, then delete it.”
- 01dataset_profile → GET /api/datasets/12/profile. A read, so it runs immediately. It returns numeric count/mean/std/min/max and the column list, labelled MEASURED with an audit id.
- 02delete_dataset → DELETE /api/datasets/12. A destructive write. The bridge validates the id and confirms ownership (a foreign id is 404), then returns the preview “Delete dataset #12 (registration + staging file). This cannot be undone.” The panel renders an Approve card and the backend mints the signed token.
- 03You click Approve. The panel re-posts the capability, the same params, confirmed:true and the token. The token is verified against the capability, the params hash and your user id, then the dataset row and staging file are really deleted. The response carries evidence MEASURED and a fresh audit id; a replay is rejected with 400 and another user’s token with 403.
- 04Missing input becomes a question. If a required value is missing, ask_user returns structured options, a recommended choice and free text instead of guessing. Choosing one sends “My choice: …” and the task continues.
Every capability id and route above is real: they come from GET /api/assistant/capabilities and the endpoints its entries declare. The same catalog powers the panel’s Tools tab.
Matflow Pilot questions, answered plainly
Where does Pilot live?
Does Pilot have special access?
Can Pilot act without my click?
What is a confirmation token?
What happens when a page drive fails?
How does Pilot use my credits?
Can I control memory?
What happens when the panel is closed?
Put Pilot to work on your data
Create a free account and ask it to profile a dataset, draft a DOE, or cost a physics run — reads land immediately, and destructive or billing actions wait for your click.