Add the cNF organoid dataset: drugs and experiments - #486
Open
jjacobson95 wants to merge 3 commits into
Open
jjacobson95 wants to merge 3 commits into
jjacobson95 wants to merge 3 commits into
Conversation
Member
|
we are not including cNF in this build |
…mics Add MPNST treated-sample and treated-omics support
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pipeline Hardening/Debugging PR # 8
Add the cNF organoid dataset: drugs and experiments
Second half of the new cNF dataset, building on the samples and omics PR. Adds the drug table and the dose-response experiment results for the cNF organoid drug screen.
Drugs (
03-drugs-cnf.py)cnf_drugs.tsv(drug identifiers and synonyms) andcnf_drug_descriptors.tsv(molecular descriptors) using the shared coderdata drug utilities.pubchem_retrieval.pyhas no command-line entry point, this script imports it and calls its function directly (running it as a subprocess would exit without doing anything).--skip-descriptorsand--only-descriptorsoptions so descriptors can be regenerated later without redoing the drug table.Experiments (
04-experiments-cnf.py)specimenIDannotation, fall back to parsing the filename (for exampleNF0017_T2_Viabilities.csvgivesNF0017_T2), then run the result through the sharedclassify_specimenhelper as a safety net. Files whose specimen can't be resolved are skipped with a warning.fit_curve.pyutility, producingfit_auc,fit_ic50,fit_einf,fit_hs, andfit_r2;uM_viabilityand value = viability percentage / 100. (This metric name must exist in the schema'sResponseMetricenum, which is added in the later schema PR.)SPECIMEN_DUAL_MAPPINGStable at the top of the script that can be extended as more cases appear.Build entrypoints
build_drugs.shandbuild_exp.sh.Dependencies
cnf_utils.pyand the sample table) and uses the shared drug-pipeline and curve-fitting utilities from PR 3.Scope: 4 files, all new. Base: cnf-dataset-samples-and-omics.