Skip to main content
Docs

Search guides and API endpoints, for example “Idempotency-Key” or “submit job”.

    Tools · Cheminformatics & Structure

    Matched Molecular Pairs

    Find pairs of molecules that differ by one small structural change, and how much each change shifts activity or a property on average.

    Updated October 6, 2026

    On this page

    Prices, workflows, and method papersOpen in the app

    Matched molecular pair (MMP) analysis finds pairs of compounds in your data that differ by a single, small change (a hydrogen to a fluorine, a methyl to an ethyl, one ring swapped for another) and measures how that change shifts activity or a property. Repeated across many pairs, it tells you which transformations reliably help. It can also apply the best transformations to new molecules to suggest analogs. It runs on CPU.

    How it works

    The job uses the Hussain–Rea fragmentation method with RDKit: it cuts each molecule at one to three acyclic single bonds, and also considers hydrogen substitutions so an unsubstituted position can pair with a substituted one. Molecules sharing a constant part form pairs, and their changing parts (up to max_heavy_atoms heavy atoms, default 10, at most 13) define the transformation. For each transformation it reports how many pairs support it and the mean and median change in your value.

    Inputs

    2 to 5,000 molecules with one value each, given either as a SMILES list (input_data) with a parallel activities list, or as a table with a SMILES column and a value column (set value_column if the table has several numeric columns, such as a PaDEL-Descriptor or Group-Contribution Properties table).

    To generate analogs, set apply_transforms and give seed molecules: the job applies the found transformations to them (up to max_analogs, 500).

    Outputs

    • transforms.csv: each transformation (from_smiles → to_smiles), its pair_count, mean_delta and median_delta, and an example pair.
    • analogs.smi: proposed analogs of your seeds (when you asked for them).

    Transformations seen in only one or two pairs are weak evidence; sort by pair_count as well as by the change.

    For additive R-group contributions on a common scaffold use Free Wilson Analysis. Clean activity data first with Bioactivity Dataset Curator & Splitter. The Grow analogs from your SAR table with matched pairs, then dock workflow docks the analogs it proposes.

    Run it from the API

    Submit with Submit a job and the job_type below. Price it first with Estimate job reservation cost: submitting reserves that amount from your wallet, and the charge settles at the actual runtime.

    Matched Molecular Pairs mmp-analysis

    Job type
    mmp-analysis
    Hardware
    cpu (default)
    Typical runtime
    5 min on CPU

    Payload

    Provide at least one of: input_data + activities, activity_table.

    Payload fields
    FieldTypeDescription
    input_datamoleculeOrArtifact[] | artifactRef
    activities[]number[]

    Limits: min items 1, max items 5000

    molecule_names[]string[]
    activity_tablestring | object[] | artifactRef
    smiles_columnstring

    Limits: min length 1

    value_columnstring

    Limits: min length 1

    seed_smilesmoleculeOrArtifact[] | artifactRef
    max_heavy_atomsinteger

    Default: 10Limits: ≥ 1, ≤ 13

    min_pair_countinteger

    Default: 1Limits: ≥ 1, ≤ 5000

    apply_transformsboolean

    Default: false

    max_analogsinteger

    Default: 500Limits: ≥ 1, ≤ 500

    max_transforms_appliedinteger

    Default: 20Limits: ≥ 1, ≤ 100

    Example

    from cognichem_client import CogniChem
    
    client = CogniChem.from_env()  # reads COGNICHEM_API_KEY
    payload = {
        "input_data": ["c1ccccc1", "c1ccc(F)cc1", "CCc1ccccc1", "CCc1ccc(F)cc1"],
        "activities": [5, 5.5, 5.2, 5.7],
    }
    
    estimate = client.jobs.estimate(job_type="mmp-analysis", payload=payload, resource="cpu")
    print(f"Reserves ${estimate.cost:.2f}")
    
    job = client.jobs.submit(
        job_name="my-mmp-analysis-run",
        job_type="mmp-analysis",
        payload=payload,
        resource="cpu",
    )
    status = client.jobs.wait(job.process_id)
    if status.status == "completed":
        client.jobs.result(job.process_id, save_path=".")

    Sample data from the job catalog; long values are shortened here. Each job_name must be unique among your jobs.

    Workflow inputs

    • Table (list) · optionalCSV, JSON
    • Molecules (list) · optionalSMILES
    • Molecules (list) · optionalSMILES

    Workflow outputs

    • MoleculesSMILES
    • ArchiveZIP
    • TableCSV