Skip to main content
Docs

Search guides and API endpoints, for example “Idempotency-Key” or “submit job”.

    Tools · Cheminformatics & Structure

    Fingerprint Similarity & Clustering

    Find nearest neighbors, cluster, pick diverse subsets, or map a library with UMAP and Leiden clustering, from Morgan or MACCS fingerprints.

    Updated October 6, 2026

    On this page

    Prices, workflows, and method papersOpen in the app

    Fingerprint Similarity & Clustering works on the 2D fingerprints of up to 10,000 molecules. Choose what to do with task:

    taskWhat you getNeeds a query?
    neighborsThe library molecules most similar to your query, ranked by Tanimoto similarityYes
    butinaButina clusters (each molecule's cluster and whether it is the centroid)No
    maxminA diverse subset picked with the MaxMin algorithm, in pick orderNo
    leidenA map of the library: PCA, a neighbor graph, Leiden clusters, and 2D UMAP coordinatesNo

    Use neighbors to find analogs of a hit, butina or leiden to see the chemotypes in a library, and maxmin to pick a representative subset for purchase or screening.

    How it works

    Molecules are encoded as Morgan fingerprints (radius 2, 2048 bits, the default) or MACCS keys with RDKit. leiden then runs Scanpy: PCA (n_pcs, default 50), a k-nearest-neighbor graph (n_neighbors, default 15), Leiden community detection (leiden_resolution, default 1.0; higher gives more, smaller clusters), and UMAP for the 2D layout, with a fixed random seed so results repeat. It runs on CPU.

    Inputs

    A list of SMILES (at least 3 that parse), and for neighbors a query molecule. Unparseable SMILES are listed as errors; the job fails only if too few valid molecules remain.

    Outputs

    • results.csv: one row per library molecule with the task's columns: tanimoto for neighbors, cluster_id and is_centroid for Butina, pick_order for MaxMin, or cluster_id with umap_1 and umap_2 for Leiden.
    • manifest.csv: how each input was parsed.
    • A molecules set for chaining: the Butina centroids, the MaxMin picks, or one representative per Leiden cluster.

    For 3D shape and pharmacophore similarity use 3D Shape Similarity; to group by core structure use Scaffold Analyzer. The Build a library from building blocks, filter it, then dock workflow uses a MaxMin pick to choose which compounds to dock.

    Run it from the API

    Submit with Submit a job and the job_type below. Price it first with Estimate job reservation cost: submitting reserves that amount from your wallet, and the charge settles at the actual runtime.

    Fingerprint Similarity & Clustering fingerprint-cluster

    Job type
    fingerprint-cluster
    Hardware
    cpu (default)
    Typical runtime
    5 min on CPU

    Payload

    Payload fields
    FieldTypeDescription
    taskrequiredstring

    One of: "neighbors", "butina", "maxmin", "leiden"

    input_data[]required(string | object)[]

    Limits: min items 1, max items 10000

    input_formatrequiredstring

    One of: "smiles"

    querystring | object
    query_formatstring

    One of: "smiles"

    fingerprintstring

    Default: "morgan"One of: "morgan", "maccs"

    morgan_radiusinteger

    Default: 2Limits: ≥ 0, ≤ 10

    morgan_n_bitsinteger

    Default: 2048Limits: ≥ 1, ≤ 16384

    top_kinteger

    Default: 10Limits: ≥ 1, ≤ 10000

    distance_cutoffnumber

    Default: 0.35Limits: > 0, ≤ 1

    n_picksinteger

    Limits: ≥ 1, ≤ 10000

    n_pcsinteger

    Default: 50Limits: ≥ 2, ≤ 128

    n_neighborsinteger

    Default: 15Limits: ≥ 2, ≤ 100

    leiden_resolutionnumber

    Default: 1Limits: ≥ 0.1, ≤ 5

    Example

    from cognichem_client import CogniChem
    
    client = CogniChem.from_env()  # reads COGNICHEM_API_KEY
    payload = {
        "task": "neighbors",
        "input_data": ["CCO", "CCN"],
        "input_format": "smiles",
        "query": "CCO",
        "query_format": "smiles",
    }
    
    estimate = client.jobs.estimate(job_type="fingerprint-cluster", payload=payload, resource="cpu")
    print(f"Reserves ${estimate.cost:.2f}")
    
    job = client.jobs.submit(
        job_name="my-fingerprint-cluster-run",
        job_type="fingerprint-cluster",
        payload=payload,
        resource="cpu",
    )
    status = client.jobs.wait(job.process_id)
    if status.status == "completed":
        client.jobs.result(job.process_id, save_path=".")

    Sample data from the job catalog; long values are shortened here. Each job_name must be unique among your jobs.

    Workflow inputs

    • Molecules (list)SMILES
    • Molecules · optionalSMILES

    Workflow outputs

    • ArchiveZIP
    • MoleculesSMILES
    • TableCSV