Add the cNF organoid dataset: samples and omics - #485
Open
jjacobson95 wants to merge 1 commit into
Open
jjacobson95 wants to merge 1 commit into
jjacobson95 wants to merge 1 commit into
Conversation
Member
|
I thought we weren't including the cNF data? |
Collaborator
Author
|
These are currently locked behind Synapse permissions. They are located in the "Leveraging patient-derived cutaneous neurofibroma organoid models to identify biomarkers of drug response" Project (syn51301431). Keep them locked until we want to release. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pipeline Hardening/Debugging PR # 7
Add the cNF organoid dataset: samples and omics
Adds the first half of a new dataset, cNF (cutaneous neurofibroma), a patient-derived-organoid drug screen with matched multi-omics. This PR covers specimen/sample generation and all three omics types plus a shared helper module and container image. The drug table and dose-response experiments follow in the next PR.
Samples (
01-samples-cnf.py)syn51301431), the global proteomics matrix (syn74815895), the RNA discovery file (syn71333780), and the Normal Skin folder (syn74284682). Each source Synapse ID is overridable via environment variable.NF0021_T1_Onalespid_1uM);improve_sample_idfrom the previous samples file (max + 1) so IDs don't collide across datasets.Omics (
02-omics-cnf.py)syn66352931andsyn70765053; overridable). Protocol-optimization samples are excluded.syn74815895(thecorrectedAbundancevalues).syn70078415, mapped tophosphosite_idvalues using thephosphosites.csvreference from the previous PR.NF0018.T1.organoidresolve to the canonical IDs incnf_samples.csv; source rows that don't match a known sample are dropped via an inner join.Shared helper and packaging
cnf_utils.py: specimen classification and canonicalization (classify_specimen,patient_from_specimen,canonicalize_specimen_column) shared across the samples, omics, and (next PR) experiments steps.build_samples.shandbuild_omics.sh,requirements.txt, the container imageDockerfile.cnf, and aREADME.mddocumenting sources and the specimen model.Dependencies and build policy
phosphosites.csv).--dataset cnf) but is intentionally not part of the default build list (set in the later build-orchestration PR).Scope: 8 files, all new. Base: phosphosites-reference-builder.