A workflow is a pipeline of steps. Each step is a job or a filter, and each step's outputs can feed later steps' inputs. You describe it once as a WorkflowSpec (JSON), then validate, estimate, and run it as often as you like. Ready-made specs are available as templates: list them with List curated workflow templates or browse them in the workflow gallery.
A WorkflowSpec
{
"spec_version": 1,
"name": "SMILES → 3D embed → dock",
"params": {
"ligands": {"kind": "molecule_set", "format": "smiles", "cardinality": "many", "required": true},
"target_structure": {"kind": "protein_structure", "format": "pdb", "required": true},
"max_pockets": {"type": "integer", "default": 1}
},
"steps": [
{"id": "embed3d", "type": "job", "job_type": "convert-batch", "resource": "cpu",
"inputs": {"molecules": {"$param": "ligands"}},
"payload": {"input_format": "smiles", "output_format": "sdf", "generate_3d": true}},
{"id": "pockets", "type": "job", "job_type": "pocket-detect", "resource": "cpu",
"inputs": {"protein": {"$param": "target_structure"}},
"payload": {"max_pockets": {"$param": "max_pockets"}}},
{"id": "dock", "type": "job", "job_type": "autodockvina", "resource": "cpu", "mode": "batch",
"inputs": {"ligands": {"$from": "embed3d.molecules"},
"protein": {"$param": "target_structure"},
"boxes": {"$from": "pockets.pockets"}},
"payload": {"exhaustiveness": 8}}
],
"limits": {"on_step_failure": "fail_fast"}
}paramsare the run's inputs. A data input names itskindandformat(andcardinality: "many"for a list); a setting names itstypeand an optionaldefault.stepsrun in dependency order; steps that don't depend on each other run at the same time. A spec has at most 64 steps and 256 KiB of JSON.limits(optional):on_step_failure(fail_faststops the run at the first failed step;continuekeeps running steps that don't depend on it),cache(see below), andmax_run_cost_usd, a suggested spending cap that the app's builder fills in when you launch. Over the API, the cap the server enforces is themax_run_costyou send when you create the run.
Passing data between steps
A step's inputs and payload values can be bindings:
| Binding | Value |
|---|---|
{"$param": "name"} | A run input from params |
{"$from": "step.port"} | Output port of an earlier step |
{"$artifact": {"id": "art-…", "port": "…"}} | A stored artifact, such as a file you uploaded or an earlier run's result |
{"$const": …} | A literal value, used as is (for values that would otherwise read as a binding) |
{"$spread": …} | An object taken from another binding, such as an object-valued param |
Every job type declares typed ports (inputs and outputs, with data kinds and formats); validation checks that connected ports fit. Each tool's page lists its ports, and Get catalog workflow ports for a job type returns them.
One job or many
"mode": "batch"runs one job over the whole list."mode": "scatter"with"scatter": {"over": "input_name", "max_fanout": 10}runs one job per item, in parallel, up tomax_fanout. Use it for tools that take one molecule at a time.
Filters
A step with "type": "filter" reshapes data without running a job: keep the top k by a score, drop duplicates, split a set, or join two result lists. Small filters run instantly at no cost; large ones run as a short CPU job.
{"id": "top10", "type": "filter", "op": "select_top_k",
"inputs": {"in": {"$from": "dock.poses"}},
"args": {"metric": "binding_affinity_kcal_mol", "direction": "best", "k": 10}}| op | args (? = optional) |
|---|---|
deduplicate | by? (id | label | metric), keep? (best | first), metric? |
filter_rows | metric?, op?, predicates?, value? |
limit | n |
merge | lineage?, parts? |
partition | by (metric | predicate), default_port?, parts |
sample | n, seed? |
select_top_k | direction? (best | worst | asc | desc), k, metric |
split | parts, seed?, strategy (fraction | count) |
zip_join | fallback? (scatter_index), key?, match? (one_to_one | many_to_one), right_key? |
If a filter leaves nothing to pass on, the run pauses with filter_empty; that run can only be cancelled.
Validate, estimate, run
- Validate a WorkflowSpec checks the spec against the catalog and returns
validplus a list of diagnostics, each with acode, amessage, and the step or path it refers to. It is free. - Estimate a workflow run's cost prices every step at your plan's rates. Pass the same
paramsyou will run with: larger inputs cost more. Also free. - Create a workflow run starts it with a
run_name(unique among your runs), the spec, itsparams, and an optionalmax_run_cost. The whole request may be up to 48 MiB, and each file in it up to 25 MiB.
A run is pending, then running, and ends completed, failed, or cancelled. It can also be paused:
| Paused because | What to do |
|---|---|
Charges reached max_run_cost | Resume to continue |
| The run's hold ran out | Resume to continue |
| Your artifact storage is full | Free space, then resume |
A filter produced nothing (filter_empty) | Cancel the run |
Cancelling stops unfinished steps; completed steps and their results are kept. Deleting a run removes it from your list but leaves its artifacts in your library until they expire.
What a run costs
Starting a run holds its estimate in your wallet (the money isn't spent yet). Each step's job is charged like any other job as it finishes, drawing on the hold, and whatever is left is released when the run ends. See Wallet and billing.
Caching: a step whose job type, hardware, inputs, and settings exactly match a step you ran before reuses that result at no charge, as long as the earlier result hasn't expired. Set "limits": {"cache": false} to always recompute.
Saved workflows
Save a spec as a definition to reuse it: Create a workflow definition. Definitions keep a version history you can restore, and you can share one through a link that other CogniChem users can fork into their own library.