Many molecules arrive with stereochemistry unspecified: a SMILES with an undefined stereocenter stands for several real molecules. The Stereoisomer Enumerator expands each one into explicit stereoisomers, so docking, ADMET prediction, and conformer generation work on every isomer instead of silently picking one. Stereocenters already assigned in your input are kept as they are. It runs on CPU with RDKit.
How it works
- It enumerates unassigned tetrahedral centers, double bonds (E/Z), and non-absolute enhanced-stereo groups.
- Equivalent results (including meso forms) collapse to one isomer.
- A molecule with n open stereo elements has up to 2ⁿ isomers. If that is within
max_isomers_per_molecule(default 32, up to 1,024), all are listed; otherwise a fixed, repeatable sample of that many is taken and the rows are markedtruncated. - Each isomer gets CIP labels for what was enumerated (for example
C3:R;C5:S;C6=C7:E) and a stable id derived from its parent. try_embeddingdrops isomers RDKit can't build in 3D, such as strained or impossible ones.
A molecule with nothing to enumerate passes through unchanged.
Inputs
Up to 10,000 molecules as SMILES, SDF or MOL blocks, or InChI. Standardize tautomers and charges first with Molecule Standardizer if that matters for your set.
Outputs
isomers.smi: every stereoisomer with its id.audit.csv: one row per isomer with the parent, the isomer's SMILES and InChIKey, how many stereo elements were open, the assignments, and whether the parent was truncated. Molecules that can't be parsed get an error row; the job carries on.
Not covered
Tautomers and protonation states (Molecule Standardizer), atropisomers, and 3D coordinates (Conformer Ensemble Generator).