Skip to content

Add the cNF organoid dataset: drugs and experiments - #486

Open
jjacobson95 wants to merge 3 commits into
cnf-dataset-samples-and-omicsfrom
cnf-dataset-drugs-and-experiments
Open

jjacobson95 wants to merge 3 commits into
cnf-dataset-samples-and-omicsfrom
cnf-dataset-drugs-and-experiments

Conversation

@jjacobson95

Copy link
Copy Markdown
Collaborator

Pipeline Hardening/Debugging PR # 8

Add the cNF organoid dataset: drugs and experiments

Second half of the new cNF dataset, building on the samples and omics PR. Adds the drug table and the dose-response experiment results for the cNF organoid drug screen.

Drugs (03-drugs-cnf.py)

  • Pulls the unique drug names from the cNF drug screen on Synapse, then produces cnf_drugs.tsv (drug identifiers and synonyms) and cnf_drug_descriptors.tsv (molecular descriptors) using the shared coderdata drug utilities.
  • Because pubchem_retrieval.py has no command-line entry point, this script imports it and calls its function directly (running it as a subprocess would exit without doing anything).
  • Keeps the drug table usable even when descriptor generation fails (the descriptor builder can call Ensembl/biomaRt, which is sometimes down): descriptor work is retried with exponential backoff under a generous timeout, and there are --skip-descriptors and --only-descriptors options so descriptors can be regenerated later without redoing the drug table.

Experiments (04-experiments-cnf.py)

  • Discovers every drug-screen viability file under three Synapse parent folders and attaches a specimen ID to each.
  • Specimen attribution uses a three-step policy: trust the file's specimenID annotation, fall back to parsing the filename (for example NF0017_T2_Viabilities.csv gives NF0017_T2), then run the result through the shared classify_specimen helper as a safety net. Files whose specimen can't be resolved are skipped with a warning.
  • For each specimen/drug pair:
    • multi-concentration measurements are fit with the shared fit_curve.py utility, producing fit_auc, fit_ic50, fit_einf, fit_hs, and fit_r2;
    • single-concentration measurements at 1 μM are recorded with metric uM_viability and value = viability percentage / 100. (This metric name must exist in the schema's ResponseMetric enum, which is added in the later schema PR.)
  • Handles one known case where a single viability file should be attributed to two specimens, via a SPECIMEN_DUAL_MAPPINGS table at the top of the script that can be extended as more cases appear.

Build entrypoints

  • build_drugs.sh and build_exp.sh.

Dependencies

  • Builds on the cNF samples/omics PR (shares cnf_utils.py and the sample table) and uses the shared drug-pipeline and curve-fitting utilities from PR 3.

Scope: 4 files, all new. Base: cnf-dataset-samples-and-omics.

@jjacobson95 jjacobson95 added the new data Request for additional data to be added label Sep 22, 2026
@jjacobson95 jjacobson95 modified the milestones: 2.3 new build, 2.4 new build Sep 22, 2026
@sgosline

Copy link
Copy Markdown
Member

we are not including cNF in this build

…mics

Add MPNST treated-sample and treated-omics support
@jjacobson95 jjacobson95 modified the milestones: 2.4 new build, 2.5 Sep 23, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

new data Request for additional data to be added

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants