Fingerprint Similarity & Clustering works on the 2D fingerprints of up to 10,000 molecules. Choose what to do with task:
task | What you get | Needs a query? |
|---|---|---|
neighbors | The library molecules most similar to your query, ranked by Tanimoto similarity | Yes |
butina | Butina clusters (each molecule's cluster and whether it is the centroid) | No |
maxmin | A diverse subset picked with the MaxMin algorithm, in pick order | No |
leiden | A map of the library: PCA, a neighbor graph, Leiden clusters, and 2D UMAP coordinates | No |
Use neighbors to find analogs of a hit, butina or leiden to see the chemotypes in a library, and maxmin to pick a representative subset for purchase or screening.
How it works
Molecules are encoded as Morgan fingerprints (radius 2, 2048 bits, the default) or MACCS keys with RDKit. leiden then runs Scanpy: PCA (n_pcs, default 50), a k-nearest-neighbor graph (n_neighbors, default 15), Leiden community detection (leiden_resolution, default 1.0; higher gives more, smaller clusters), and UMAP for the 2D layout, with a fixed random seed so results repeat. It runs on CPU.
Inputs
A list of SMILES (at least 3 that parse), and for neighbors a query molecule. Unparseable SMILES are listed as errors; the job fails only if too few valid molecules remain.
Outputs
results.csv: one row per library molecule with the task's columns:tanimotofor neighbors,cluster_idandis_centroidfor Butina,pick_orderfor MaxMin, orcluster_idwithumap_1andumap_2for Leiden.manifest.csv: how each input was parsed.- A
moleculesset for chaining: the Butina centroids, the MaxMin picks, or one representative per Leiden cluster.
Related tools
For 3D shape and pharmacophore similarity use 3D Shape Similarity; to group by core structure use Scaffold Analyzer. The Build a library from building blocks, filter it, then dock workflow uses a MaxMin pick to choose which compounds to dock.