The Bioactivity Dataset Curator turns a raw activity table (a ChEMBL export, an assay spreadsheet) into a dataset you can model on: units normalized, replicates merged, censored values handled, every compound given a stable identity, and splits assigned so the same structure never lands in two of them. Every row it drops is listed with a reason.
How it works
Each row goes through the same steps:
- Value: a number, optionally with a relation (
<10,>=5, or a relation column). - Endpoint: from the endpoint column (or
default_endpoint), normalized (ic50→IC50). Setendpointsto keep only some. - Units: pM to M concentrations become
p_activity= −log₁₀(molar); log-scale values pass through. On the p-scale relations flip: IC50 < 10 nM becomes p > 8. Values outside 0 to 14 are rejected as implausible. - Censoring (
censored_policy):dropbounds, treat themas_threshold, orkeepthem apart from exact values. - Identity: RDKit parses the SMILES and the standard InChIKey becomes
compound_id. Salts and tautomers are not standardized; run Molecule Standardizer first if you need that.
Then rows for the same compound and endpoint (and assay, with group_by_assay) are aggregated by median or mean. If replicates disagree by more than max_replicate_spread log units (default 1.0), the whole group is rejected as a conflict rather than averaged. Every input row ends up either in a curated record or in the rejections file.
Splits
Splits are assigned per compound, so a structure never appears in two of them:
scaffold(MoleculeNet-style, deterministic): Bemis–Murcko scaffold groups, largest first, the largest always in train. The usual choice for estimating how a model will do on new chemotypes.random: a seeded shuffle.time: by each compound's earliest measurement date (needs a date or year column).
split_fractions is [train, valid, test] and must sum to 1.
Inputs
One activity table (CSV, TSV, JSON, a list of rows, or an artifact), up to 50,000 rows. Name at least the SMILES and value columns; units, relation, endpoint, assay, compound id, and date columns are picked up from common header names (ChEMBL's are recognized first). A curated table from an earlier run can be curated again without remapping.
Outputs
| File | Contents |
|---|---|
curated.csv | One record per compound and endpoint: p_activity, value in nM, number of measurements and their spread, censoring, scaffold, and split |
splits.csv | Each compound's scaffold, split, and date |
rejections.csv | Every dropped row with its reason |
molecules.smi | Each compound's SMILES and InChIKey |
summary.json | Counts, rejection reasons, resolved columns, and warnings |
Related tools
Feed the curated table to Matched Molecular Pairs or Free Wilson Analysis, or train a model on it with MPNN Models.