Skip to main content
Docs

Search guides and API endpoints, for example “Idempotency-Key” or “submit job”.

    Tools · Computational Biology

    Protein Preparer

    Clean up a protein structure for docking: add missing atoms, choose which waters and hetero groups to keep, and set protonation states at your pH.

    Updated October 6, 2026

    On this page

    Prices, workflows, and method papersOpen in the app

    Structures from the PDB, or from structure prediction, usually need cleaning before docking: missing atoms, waters, extra ligands and ions, and no hydrogens. The Protein Preparer fixes those and writes a docking-ready structure. It runs on CPU.

    How it works

    1. Repair with PDBFixer: add missing heavy atoms (add_missing_atoms) and, if you ask, short missing stretches of residues (add_missing_residues). Long missing loops are not modeled.
    2. Choose what to keep: remove waters (strip_waters) and keep or drop other hetero groups such as ligands, ions, and cofactors (keep_hetatm).
    3. Protonate with PDB2PQR and PROPKA at your ph (for example 7.4), so histidines and other titratable residues get the right charge state.
    4. Write the prepared PDB and a receptor PDBQT.

    It does not solvate, minimize, or simulate; OpenMM MD does that.

    Inputs

    An RCSB pdb_id, or a PDB or mmCIF file, for example a predicted structure from ESMFold2 or Boltz-2.

    Outputs

    FileContents
    prepared.pdbThe cleaned, protonated structure: the main input for later steps
    prepared.pdbqtThe receptor in PDBQT format
    audit.csvWhat was changed: atoms and residues added, waters and hetero groups removed
    summary.jsonSettings and counts

    Next steps are usually Binding-Site Detector and AutoDock Vina or GNINA. The Dock into a cleaned predicted structure and see the contacts workflow chains folding, preparation, docking, and interaction profiling.

    Run it from the API

    Submit with Submit a job and the job_type below. Price it first with Estimate job reservation cost: submitting reserves that amount from your wallet, and the charge settles at the actual runtime.

    Protein Preparer protein-prepare

    Job type
    protein-prepare
    Hardware
    cpu (default)
    Typical runtime
    3 min on CPU

    Payload

    Payload fields
    FieldTypeDescription
    pdb_idstring

    Limits: min length 1

    proteinfile_object
    protein_download_formatstring

    Default: "pdb"One of: "pdb", "cif"

    phnumber

    Default: 7.4Limits: ≥ 0, ≤ 14

    strip_watersboolean

    Default: true

    keep_hetatmboolean

    Default: true

    add_missing_atomsboolean

    Default: true

    add_missing_residuesboolean

    Default: false

    File fields take {"name": "x.pdb", "content_b64": "…"} or a stored artifact, {"$artifact": {"id": "art-…", "port": "…"}}. Files are up to 25 MiB each.

    Example

    from cognichem_client import CogniChem
    
    client = CogniChem.from_env()  # reads COGNICHEM_API_KEY
    payload = {
        "protein": {
            "name": "protein.pdb",
            "content_b64": "<file contents, base64>",
        },
        "ph": 7.4,
        "strip_waters": True,
        "keep_hetatm": True,
        "add_missing_atoms": True,
        "add_missing_residues": False,
    }
    
    estimate = client.jobs.estimate(job_type="protein-prepare", payload=payload, resource="cpu")
    print(f"Reserves ${estimate.cost:.2f}")
    
    job = client.jobs.submit(
        job_name="my-protein-prepare-run",
        job_type="protein-prepare",
        payload=payload,
        resource="cpu",
    )
    status = client.jobs.wait(job.process_id)
    if status.status == "completed":
        client.jobs.result(job.process_id, save_path=".")

    Sample data from the job catalog; long values are shortened here. Each job_name must be unique among your jobs.

    Workflow inputs

    • Protein structurePDB, CIF

    Workflow outputs

    • ArchiveZIP
    • TableCSV
    • Protein structurePDBQT
    • Protein structurePDB