Skip to main content
Docs

Search guides and API endpoints, for example “Idempotency-Key” or “submit job”.

    Tools · Cheminformatics & Structure

    Scaffold Analyzer

    Group a hit list by Bemis–Murcko scaffold, count each framework, and pick one representative per scaffold.

    Updated October 1, 2026

    On this page

    Prices, workflows, and method papersOpen in the app

    The Scaffold Analyzer shows which core frameworks a set of molecules is built on. It reduces each molecule to its Bemis–Murcko scaffold (ring systems and the linkers between them, side chains removed), counts how often each scaffold occurs, and picks a representative for each one. Use it to see whether a hit list is many variations on a few chemotypes or genuinely diverse. It runs on CPU with RDKit.

    How it works

    • scaffold_type: murcko (the default; atom and bond types kept) or generic (all atoms carbon and all bonds single, which merges scaffolds that differ only in heteroatoms or saturation). Both forms are written for every molecule; this choice sets the grouping.
    • pick, the representative per scaffold: first (the default), highest_score (needs scores), or closest_to_scaffold (the molecule with the fewest extra atoms).
    • compute_mcs: the maximum common substructure between molecules, for small sets (up to 50 molecules).

    Inputs

    Up to 10,000 SMILES, with optional scores per molecule (as a scores list or a table) for picking the best representative. Convert SDF to SMILES first with Molecule Conversion.

    Outputs

    • results.csv: each molecule's Murcko and generic scaffold, how many molecules share its scaffold (scaffold_count), and whether it is the representative (is_representative), plus pairwise MCS results when you asked for them.
    • The representatives as a SMILES set, for the next step.

    For similarity-based clustering use Fingerprint Similarity & Clustering. For R-group analysis on a shared scaffold use Free Wilson Analysis.

    Run it from the API

    Submit with Submit a job and the job_type below. Price it first with Estimate job reservation cost: submitting reserves that amount from your wallet, and the charge settles at the actual runtime.

    Scaffold Analyzer scaffold-analyze

    Job type
    scaffold-analyze
    Hardware
    cpu (default)
    Typical runtime
    5 min on CPU

    Payload

    Payload fields
    FieldTypeDescription
    input_datarequiredmoleculeOrArtifact[] | artifactRef
    input_formatrequiredstring

    One of: "smiles"

    scaffold_typestring

    Default: "murcko"One of: "murcko", "generic"

    pickstring

    One of: "first", "highest_score", "closest_to_scaffold"

    scores[]number[]

    Limits: min items 1, max items 10000

    score_tablestring | object[] | artifactRef
    compute_mcsboolean

    Default: false

    Example

    from cognichem_client import CogniChem
    
    client = CogniChem.from_env()  # reads COGNICHEM_API_KEY
    payload = {
        "input_data": ["c1ccccc1", "Oc1ccccc1", "CCO"],
        "input_format": "smiles",
    }
    
    estimate = client.jobs.estimate(job_type="scaffold-analyze", payload=payload, resource="cpu")
    print(f"Reserves ${estimate.cost:.2f}")
    
    job = client.jobs.submit(
        job_name="my-scaffold-analyze-run",
        job_type="scaffold-analyze",
        payload=payload,
        resource="cpu",
    )
    status = client.jobs.wait(job.process_id)
    if status.status == "completed":
        client.jobs.result(job.process_id, save_path=".")

    Sample data from the job catalog; long values are shortened here. Each job_name must be unique among your jobs.

    Workflow inputs

    • Molecules (list)SMILES
    • Table (list) · optionalCSV, JSON

    Workflow outputs

    • ArchiveZIP
    • MoleculesSMILES
    • TableCSV