Matched molecular pair (MMP) analysis finds pairs of compounds in your data that differ by a single, small change (a hydrogen to a fluorine, a methyl to an ethyl, one ring swapped for another) and measures how that change shifts activity or a property. Repeated across many pairs, it tells you which transformations reliably help. It can also apply the best transformations to new molecules to suggest analogs. It runs on CPU.
How it works
The job uses the Hussain–Rea fragmentation method with RDKit: it cuts each molecule at one to three acyclic single bonds, and also considers hydrogen substitutions so an unsubstituted position can pair with a substituted one. Molecules sharing a constant part form pairs, and their changing parts (up to max_heavy_atoms heavy atoms, default 10, at most 13) define the transformation. For each transformation it reports how many pairs support it and the mean and median change in your value.
Inputs
2 to 5,000 molecules with one value each, given either as a SMILES list (input_data) with a parallel activities list, or as a table with a SMILES column and a value column (set value_column if the table has several numeric columns, such as a PaDEL-Descriptor or Group-Contribution Properties table).
To generate analogs, set apply_transforms and give seed molecules: the job applies the found transformations to them (up to max_analogs, 500).
Outputs
transforms.csv: each transformation (from_smiles→to_smiles), itspair_count,mean_deltaandmedian_delta, and an example pair.analogs.smi: proposed analogs of your seeds (when you asked for them).
Transformations seen in only one or two pairs are weak evidence; sort by pair_count as well as by the change.
Related tools
For additive R-group contributions on a common scaffold use Free Wilson Analysis. Clean activity data first with Bioactivity Dataset Curator & Splitter. The Grow analogs from your SAR table with matched pairs, then dock workflow docks the analogs it proposes.