Boltz-2 predicts the 3D structure of a complex (proteins, DNA, RNA, and small molecules) from sequences and SMILES, and can estimate how strongly a ligand binds. It runs on a GPU.
Describe the complex
The input is a Boltz YAML document in boltz_yaml. Each entity has an id (its chain), and a protein needs its sequence (this example uses human ubiquitin; put your target's sequence there):
version: 1
sequences:
- protein:
id: A
sequence: MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG
- ligand:
id: B
smiles: "CC(=O)Oc1ccccc1C(=O)O"
properties:
- affinity:
binder: Bproperties with an affinity entry asks Boltz-2 to estimate the binding affinity of the ligand named by binder. Leave it out to predict the structure only. Ligands can also be given by ccd code, and the document can carry constraints (bonds, pockets, contacts) and structural templates.
The multiple sequence alignment
Protein (and DNA or RNA) chains are more accurate with a multiple sequence alignment (MSA). You have three options:
- Set
"use_msa_server": trueinboltz_argsand CogniChem builds the MSA with MMseqs2 against its own sequence database. Your sequences are not sent to an outside service. - Put an MSA in the YAML yourself (
msa:with A3M text). - Do neither, and the chain runs in single-sequence mode: faster, usually less accurate.
Run it
from pathlib import Path
from cognichem_client import CogniChem
client = CogniChem.from_env()
payload = {
"boltz_yaml": Path("complex.yaml").read_text(),
"boltz_args": {"use_msa_server": True, "diffusion_samples": 1},
}
estimate = client.jobs.estimate(job_type="boltz2", payload=payload, resource="a10")
print(f"Reserves ${estimate.cost:.2f}")
job = client.jobs.submit(job_name="boltz-aspirin-1", job_type="boltz2", payload=payload, resource="a10")
status = client.jobs.wait(job.process_id)Useful boltz_args: diffusion_samples (structures to sample), recycling_steps and sampling_steps (accuracy against time), and output_format. Large complexes need a bigger GPU: pick one of the options on the Boltz-2 tool page, where every argument is listed.
Results
The complexes port holds the predicted structures (mmCIF) with a confidence score and, if you asked for it, binding_affinity_kcal_mol; the zip has Boltz's full output. These are predictions: check the confidence before relying on a pose.
Next steps
To re-predict the best docking hits as complexes automatically, use the Boltz-2 triage workflows in virtual screening.