Skip to content

Add the cNF organoid dataset: samples and omics - #485

Open
jjacobson95 wants to merge 1 commit into
phosphosites-reference-builderfrom
cnf-dataset-samples-and-omics
Open

jjacobson95 wants to merge 1 commit into
phosphosites-reference-builderfrom
cnf-dataset-samples-and-omics

Conversation

@jjacobson95

Copy link
Copy Markdown
Collaborator

Pipeline Hardening/Debugging PR # 7

Add the cNF organoid dataset: samples and omics

Adds the first half of a new dataset, cNF (cutaneous neurofibroma), a patient-derived-organoid drug screen with matched multi-omics. This PR covers specimen/sample generation and all three omics types plus a shared helper module and container image. The drug table and dose-response experiments follow in the next PR.

Samples (01-samples-cnf.py)

  • Gathers specimens from four Synapse sources: the drug-screen file index (syn51301431), the global proteomics matrix (syn74815895), the RNA discovery file (syn71333780), and the Normal Skin folder (syn74284682). Each source Synapse ID is overridable via environment variable.
  • Classifies each specimen and assigns the matching schema model type:
    • organoid and drug-treated organoid map to "patient derived organoid" (treated IDs keep their treatment label, for example NF0021_T1_Onalespid_1uM);
    • tumor tissue maps to "tumor";
    • normal skin maps to "normal_tissue".
  • Keeps tumor tissue and organoids from the same tumor as separate rows with distinct sample IDs, because they are different biology and different model types; treated and untreated organoids are likewise separate rows.
  • Continues improve_sample_id from the previous samples file (max + 1) so IDs don't collide across datasets.
  • Deliberately omits a per-sample cohort label, because the upstream omics is batch-corrected and cohort membership disagrees across modalities; a uniform project label is stored instead.

Omics (02-omics-cnf.py)

  • Transcriptomics: gene-level TPM assembled from per-cohort salmon gene-TPM matrices (default cohorts syn66352931 and syn70765053; overridable). Protocol-optimization samples are excluded.
  • Proteomics: batch-corrected global proteomics from syn74815895 (the correctedAbundance values).
  • Phosphoproteomics: batch-corrected phosphoproteomics from syn70078415, mapped to phosphosite_id values using the phosphosites.csv reference from the previous PR.
  • All specimen names are canonicalized through the shared helper so values like NF0018.T1.organoid resolve to the canonical IDs in cnf_samples.csv; source rows that don't match a known sample are dropped via an inner join.

Shared helper and packaging

  • cnf_utils.py: specimen classification and canonicalization (classify_specimen, patient_from_specimen, canonicalize_specimen_column) shared across the samples, omics, and (next PR) experiments steps.
  • Build entrypoints build_samples.sh and build_omics.sh, requirements.txt, the container image Dockerfile.cnf, and a README.md documenting sources and the specimen model.

Dependencies and build policy

  • Depends on the phosphosites PR (phospho mapping needs phosphosites.csv).
  • cNF is a fully buildable dataset (--dataset cnf) but is intentionally not part of the default build list (set in the later build-orchestration PR).

Scope: 8 files, all new. Base: phosphosites-reference-builder.

@jjacobson95 jjacobson95 added the new data Request for additional data to be added label Sep 22, 2026
@jjacobson95 jjacobson95 modified the milestones: 2.3 new build, 2.4 new build Sep 22, 2026
@sgosline

Copy link
Copy Markdown
Member

I thought we weren't including the cNF data?

@jjacobson95 jjacobson95 modified the milestones: 2.4 new build, 2.5 Sep 23, 2026
@jjacobson95

Copy link
Copy Markdown
Collaborator Author

These are currently locked behind Synapse permissions. They are located in the "Leveraging patient-derived cutaneous neurofibroma organoid models to identify biomarkers of drug response" Project (syn51301431).

Keep them locked until we want to release.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

new data Request for additional data to be added

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants