Skip to main content
Docs

Search guides and API endpoints, for example “Idempotency-Key” or “submit job”.

    Tools · Cheminformatics & Structure

    Bioactivity Dataset Curator & Splitter

    Turn assay exports into clean, reproducible SAR and QSAR datasets with leakage-resistant train, validation, and test splits.

    Updated October 1, 2026

    On this page

    Prices, workflows, and method papersOpen in the app

    The Bioactivity Dataset Curator turns a raw activity table (a ChEMBL export, an assay spreadsheet) into a dataset you can model on: units normalized, replicates merged, censored values handled, every compound given a stable identity, and splits assigned so the same structure never lands in two of them. Every row it drops is listed with a reason.

    How it works

    Each row goes through the same steps:

    1. Value: a number, optionally with a relation (<10, >=5, or a relation column).
    2. Endpoint: from the endpoint column (or default_endpoint), normalized (ic50 → IC50). Set endpoints to keep only some.
    3. Units: pM to M concentrations become p_activity = −log₁₀(molar); log-scale values pass through. On the p-scale relations flip: IC50 < 10 nM becomes p > 8. Values outside 0 to 14 are rejected as implausible.
    4. Censoring (censored_policy): drop bounds, treat them as_threshold, or keep them apart from exact values.
    5. Identity: RDKit parses the SMILES and the standard InChIKey becomes compound_id. Salts and tautomers are not standardized; run Molecule Standardizer first if you need that.

    Then rows for the same compound and endpoint (and assay, with group_by_assay) are aggregated by median or mean. If replicates disagree by more than max_replicate_spread log units (default 1.0), the whole group is rejected as a conflict rather than averaged. Every input row ends up either in a curated record or in the rejections file.

    Splits

    Splits are assigned per compound, so a structure never appears in two of them:

    • scaffold (MoleculeNet-style, deterministic): Bemis–Murcko scaffold groups, largest first, the largest always in train. The usual choice for estimating how a model will do on new chemotypes.
    • random: a seeded shuffle.
    • time: by each compound's earliest measurement date (needs a date or year column).

    split_fractions is [train, valid, test] and must sum to 1.

    Inputs

    One activity table (CSV, TSV, JSON, a list of rows, or an artifact), up to 50,000 rows. Name at least the SMILES and value columns; units, relation, endpoint, assay, compound id, and date columns are picked up from common header names (ChEMBL's are recognized first). A curated table from an earlier run can be curated again without remapping.

    Outputs

    FileContents
    curated.csvOne record per compound and endpoint: p_activity, value in nM, number of measurements and their spread, censoring, scaffold, and split
    splits.csvEach compound's scaffold, split, and date
    rejections.csvEvery dropped row with its reason
    molecules.smiEach compound's SMILES and InChIKey
    summary.jsonCounts, rejection reasons, resolved columns, and warnings

    Feed the curated table to Matched Molecular Pairs or Free Wilson Analysis, or train a model on it with MPNN Models.

    Run it from the API

    Submit with Submit a job and the job_type below. Price it first with Estimate job reservation cost: submitting reserves that amount from your wallet, and the charge settles at the actual runtime.

    Bioactivity Dataset Curator & Splitter bioactivity-curate

    Job type
    bioactivity-curate
    Hardware
    cpu (default)
    Typical runtime
    5 min on CPU

    Payload

    Payload fields
    FieldTypeDescription
    activity_tablerequiredstring | object[] | artifactRef
    smiles_columncolumnName

    Limits: min length 1, max length 64

    value_columncolumnName

    Limits: min length 1, max length 64

    unit_columncolumnName

    Limits: min length 1, max length 64

    relation_columncolumnName

    Limits: min length 1, max length 64

    endpoint_columncolumnName

    Limits: min length 1, max length 64

    assay_columncolumnName

    Limits: min length 1, max length 64

    compound_id_columncolumnName

    Limits: min length 1, max length 64

    date_columncolumnName

    Limits: min length 1, max length 64

    default_unitstring

    Default: "nM"One of: "pM", "nM", "uM", "mM", "M", "log"

    default_endpointstring

    Default: "activity"Limits: min length 1, max length 64

    endpoints[]string[]

    Limits: min items 1, max items 20

    censored_policystring

    Default: "keep"One of: "keep", "drop", "as_threshold"

    aggregationstring

    Default: "median"One of: "median", "mean", "none"

    group_by_assayboolean

    Default: false

    max_replicate_spreadnumber

    Default: 1Limits: > 0, ≤ 6

    split_methodstring

    Default: "scaffold"One of: "scaffold", "random", "time"

    split_fractions[]number[]

    Default: [0.8,0.1,0.1]Limits: min items 3, max items 3

    seedinteger

    Default: 0Limits: ≥ 0, ≤ 2147483647

    Example

    from cognichem_client import CogniChem
    
    client = CogniChem.from_env()  # reads COGNICHEM_API_KEY
    payload = {
        "activity_table": [
            {
                "molecule_chembl_id": "CPD-1",
                "canonical_smiles": "c1ccc2[nH]ccc2c1",
                "standard_type": "IC50",
                "standard_relation": "=",
                "standard_value": "100",
                "standard_units": "nM",
                "assay_chembl_id": "A1",
            },
            {
                "molecule_chembl_id": "CPD-1",
                "canonical_smiles": "c1ccc2[nH]ccc2c1",
                "standard_type": "IC50",
                "standard_relation": "=",
                "standard_value": "0.12",
                "standard_units": "uM",
                "assay_chembl_id": "A2",
            },
            {
                "molecule_chembl_id": "CPD-2",
                "canonical_smiles": "Cc1ccc2[nH]ccc2c1",
                "standard_type": "IC50",
                "standard_relation": "=",
                "standard_value": "50",
                "standard_units": "nM",
                "assay_chembl_id": "A1",
            },
            {
                "molecule_chembl_id": "CPD-3",
                "canonical_smiles": "COc1ccc2[nH]ccc2c1",
                "standard_type": "IC50",
                "standard_relation": "=",
                "standard_value": "5",
                "standard_units": "nM",
                "assay_chembl_id": "A1",
            },
            {
                "molecule_chembl_id": "CPD-4",
                "canonical_smiles": "c1ccc(-c2ccccc2)cc1",
                "standard_type": "IC50",
                "standard_relation": ">",
                "standard_value": "10000",
                "standard_units": "nM",
                "assay_chembl_id": "A1",
            },
            "… 5 more",
        ],
        "split_method": "scaffold",
        "split_fractions": [0.6, 0.2, 0.2],
    }
    
    estimate = client.jobs.estimate(job_type="bioactivity-curate", payload=payload, resource="cpu")
    print(f"Reserves ${estimate.cost:.2f}")
    
    job = client.jobs.submit(
        job_name="my-bioactivity-curate-run",
        job_type="bioactivity-curate",
        payload=payload,
        resource="cpu",
    )
    status = client.jobs.wait(job.process_id)
    if status.status == "completed":
        client.jobs.result(job.process_id, save_path=".")

    Sample data from the job catalog; long values are shortened here. Each job_name must be unique among your jobs.

    Workflow inputs

    • Table (list)CSV, TSV, JSON

    Workflow outputs

    • ArchiveZIP
    • TableCSV
    • MoleculesSMILES
    • TableCSV
    • TableCSV