diff --git a/notebooks/pathomics/microscopy_dicom_ann_intro.ipynb b/notebooks/pathomics/microscopy_dicom_ann_intro.ipynb index 1ea697e..1b14023 100644 --- a/notebooks/pathomics/microscopy_dicom_ann_intro.ipynb +++ b/notebooks/pathomics/microscopy_dicom_ann_intro.ipynb @@ -18,24 +18,23 @@ "# Working with DICOM Microscopy Bulk Simple Annotations in computational pathology\n", "\n", "\n", - "This tutorial is shared as part of the tutorials prepared by the Imaging Data Commons team and available at https://github.com/ImagingDataCommons/IDC-Tutorials/blob/master/notebooks.\n", + "This tutorial is shared as part of the tutorials prepared by the Imaging Data Commons (IDC) team and available at https://github.com/ImagingDataCommons/IDC-Tutorials/blob/master/notebooks.\n", "\n", "If you are new to IDC and DICOM for digital pathology applications, you may want to check out other introductory tutorials on this topic available here: https://github.com/ImagingDataCommons/IDC-Tutorials/tree/master/notebooks/pathomics.\n", "\n", - "This tutorial is aimed for the users of Imaging Data Commons that are interested to understand how to use annotations of slide microscopy images. You will learn how to:\n", - "* select and download specific type of slide annotations, represented as polygons\n", - "* parse the content annotations stored in DICOM Bulk Simple Annotations (ANN) format\n", - "* utilize annotations for calculating quantitative features that may help in analyzing the annotated pathology slides\n", + "This tutorial is aimed at users of the IDC who are interested in understanding how to work with image-derived data for slide microscopy images. You will learn how to:\n", + "* select and download a specific type of image-derived data, namely delineations of cell nuclei represented as polygons\n", + "* parse the content stored in DICOM Microscopy Bulk Simple Annotations (MBSA/ANN) files\n", + "* utilize the cell nuclei contours for calculating quantitative features that may help in analyzing the annotated pathology slides\n", "\n", "To learn more about the IDC, please visit the [IDC user guide](https://learn.canceridc.dev).\n", "\n", - "If you have any questions, bug reports, or feature requests please feel free to contact us at the [IDC discussion forum](https://discourse.canceridc.dev)!\n", + "If you have any questions, bug reports, or feature requests, please feel free to contact us at the [IDC discussion forum](https://discourse.canceridc.dev)!\n", "\n", "----------------------\n", "\n", "Initial version: Jan 2025 \n", - "Updated: Feb 2026", - "Last updated: Feb 2026" + "Updated: Aug 2026" ] }, { @@ -46,30 +45,34 @@ "source": [ "## Background\n", "\n", - "There are different ways to store annotations in DICOM, depending on the type of the annotation, and the specific DICOM object used. The annotations of cell nuclei in the [`Pan-Cancer-Nuclei-Seg-DICOM` collection](https://zenodo.org/records/14009675), which were recently added to the [NCI Imaging Data Commons (IDC)](https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=Pan-Cancer-Nuclei-Seg-DICOM), are offered in two formats:\n", - "1. Polygons corresponding to the boundaries of the nuclei: DICOM **Microscopy Simple Bulk Annotations** (ANN)\n", + "There are different ways to store image-derived data in DICOM. The **cell nuclei delineations** in the [`Pan-Cancer-Nuclei-Seg-DICOM` collection](https://zenodo.org/records/14009675), which was recently added to the [NCI Imaging Data Commons (IDC)](https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=Pan-Cancer-Nuclei-Seg-DICOM), are offered in two DICOM representations:\n", + "1. Polygons corresponding to the boundaries of the nuclei: DICOM **Microscopy Simple Bulk Annotations** (**MBSA**, often also abbreviated **ANN** after the corresponding DICOM `Modality` value)\n", "2. Binary masks labeling the areas of the image corresponding to the individual nuclei: DICOM Segmentations (SEG)\n", "\n", - "In this notebook we will concentrate on the first format, discuss how DICOM ANN objects in the `Pan-Cancer-Nuclei-Seg-DICOM` collection are organized and how those annotations can be used in analysis. The DICOM SEG format will be covered later and then linked here upon finalization.\n", + "In this notebook, we will concentrate on the first representation, discuss how the MBSA objects in the `Pan-Cancer-Nuclei-Seg-DICOM` collection are organized and how the cell nuclei contours can be used in analysis. The DICOM SEG representation will be covered separately and linked here upon finalization.\n", "\n", - "`Pan-Cancer-Nuclei-Seg-DICOM` contains **automatically derived nucleus segmentations** in over 5000 whole slide tissue images across 10 different cancer types. The annotations correspond to digital pathology images from the TCGA collections, all of which are available in the IDC:\n", + "The delineations of cell nuclei in the `Pan-Cancer-Nuclei-Seg-DICOM` collection were **automatically derived** for over 5000 whole slide images (WSIs) across 14 different cancer types. The underlying WSIs originate from the following TCGA collections, all of which are also available in the IDC:\n", "\n", "- Bladder urothelial carcinoma (TCGA-BLCA)\n", "- Breast invasive carcinoma (TCGA-BRCA)\n", "- Cervical squamous cell carcinoma and endocervical adenocarcinoma (TCGA-CESC)\n", + "- Colon adenocarcinoma (TCGA-COAD)\n", "- Glioblastoma Multiforme (TCGA-GBM)\n", "- Lung adenocarcinoma (TCGA-LUAD)\n", "- Lung squamous cell carcinoma (TCGA-LUSC)\n", "- Pancreatic adenocarcinoma (TCGA-PAAD)\n", "- Prostate adenocarcinoma (TCGA-PRAD)\n", + "- Rectum adenocarcinoma (TCGA-READ)\n", "- Skin Cutaneous Melanoma (TCGA-SKCM)\n", "- Uterine Corpus Endometrial Carcinoma (TCGA-UCEC)\n", + "- Stomach adenocarcinoma (TCGA-STAD)\n", + "- Uveal melanoma (TCGA-UVM)\n", "\n", - "Originally released by [TCIA in SVS format](https://www.cancerimagingarchive.net/analysis-result/pan-cancer-nuclei-seg/), the collection is now harmonized into DICOM format and freely available in IDC. While the Zenodo page [Pan-Cancer-Nuclei-Seg-DICOM: DICOM converted Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images](https://zenodo.org/records/14009675) offers download manifests for the complete collection, in this notebook we demonstrate how to navigate its content programmatically.\n", + "Originally released by [TCIA in SVS format](https://www.cancerimagingarchive.net/analysis-result/pan-cancer-nuclei-seg/), the data are now harmonized into DICOM and freely available in the IDC. While the Zenodo page [Pan-Cancer-Nuclei-Seg-DICOM: DICOM converted Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images](https://zenodo.org/records/14009675) offers download manifests for the complete collection, in this notebook we demonstrate how to navigate its content programmatically.\n", "\n", "## Example\n", "\n", - "As an example, you can open the sample slide and annotations available here https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.150973379448125660359643882019624926008/series/1.3.6.1.4.1.5962.99.1.1062471168.429178244.1637445043712.2.0. Note that not all of the slides in a given study contain annotations. Make sure you select the slide that has the suffix \"DX1\", as shown in the screenshot below. To enable visualization of the annotations, open \"Annotation Groups\" section in the right panel, and toggle the \"Nuclei\" switch.\n", + "As an example, you can open the following WSI with cell nuclei contours, available at [this link](https://viewer.imaging.datacommons.cancer.gov/slim/studies/2.25.150973379448125660359643882019624926008/series/1.2.826.0.1.3680043.10.511.3.69328279220552597239698708434338158).\n", "\n", "\"Example" ] @@ -82,7 +85,7 @@ "source": [ "## Prerequisites\n", "**Installations**\n", - "* **Install highdicom:** Most code in this notebook relies on the Python library [highdicom](https://highdicom.readthedocs.io/en/latest/introduction.html) which was specifically designed to work with DICOM objects holding image-derived information, e.g. annotations and measurements. Further and more detailed information on highdicom's functionality can be found in its [user guide](https://highdicom.readthedocs.io/en/latest/usage.html).\n", + "* **Install highdicom:** Most code in this notebook relies on the Python library [highdicom](https://highdicom.readthedocs.io/en/latest/introduction.html) which was specifically designed to work with DICOM objects holding image-derived information, e.g. annotations and segmentations. Further and more detailed information on `highdicom`'s functionality can be found in its [user guide](https://highdicom.readthedocs.io/en/latest/usage.html).\n", "* **Install wsidicom:** The [wsidicom](https://pypi.org/project/wsidicom/) Python package provides functionality to open and extract image or metadata from WSIs.\n", "* **Install idc-index:** The Python package [idc-index](https://pypi.org/project/idc-index/) facilitates queries of the basic metadata and download of DICOM files hosted by the IDC." ] @@ -114,16 +117,16 @@ "cell_type": "code", "execution_count": 2, "metadata": { - "id": "3IuBLeoEnWVU", - "outputId": "cfaab9bb-ed0c-4421-e65d-e443c3649152", "colab": { "base_uri": "https://localhost:8080/" - } + }, + "id": "3IuBLeoEnWVU", + "outputId": "cfaab9bb-ed0c-4421-e65d-e443c3649152" }, "outputs": [ { - "output_type": "stream", "name": "stderr", + "output_type": "stream", "text": [ "WARNING:pydicom:get_frame_offsets is deprecated and will be removed in v4.0\n" ] @@ -152,13 +155,13 @@ "id": "6R4IOZ-inatY" }, "source": [ - "## Accessing DICOM ANNs from the IDC\n", - "To access and download the ANNs files, we utilize the Python package [idc-index](https://github.com/ImagingDataCommons/idc-index) that facilitates querying metadata and downloading DICOM files from the IDC. Since all available ANN objects in the IDC have a combined size of 1.82 TB - for the demonstration in this tutorial we will use only a single ANN from the TCGA-BRCA collection as an example." + "## Accessing DICOM MBSAs from the IDC\n", + "To access and download the MBSA files, we use `idc-index`. Since all available MBSA objects in the IDC amount to a combined size of 1.82 TB, this tutorial uses only a single MBSA file from the TCGA-BRCA collection as an example." ] }, { "cell_type": "code", - "execution_count": 3, + "execution_count": null, "metadata": { "id": "RxHTkvyZ_FzV" }, @@ -166,7 +169,7 @@ "source": [ "idc_client = IDCClient() # set-up idc_client\n", "idc_client.fetch_index('sm_instance_index') # fetch additional sm_instance_index containing all slide microscopy (SM) instances available in the IDC\n", - "idc_client.fetch_index('ann_index') # fetch additional ann_index containing all ANN series available in the IDC" + "idc_client.fetch_index('ann_index') # fetch additional ann_index containing all MBSA series available in the IDC" ] }, { @@ -209,10 +212,10 @@ }, "outputs": [ { - "output_type": "stream", "name": "stderr", + "output_type": "stream", "text": [ - "Downloading data: 100%|\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2588| 215M/215M [00:02<00:00, 106MB/s]\n" + "Downloading data: 100%|██████████| 215M/215M [00:02<00:00, 106MB/s]\n" ] } ], @@ -230,27 +233,23 @@ "id": "ezVf_-svK3K7" }, "source": [ - "Annotations can also be viewed and explored in detail on its respective slide using the Slim viewer. In the Slim viewer's interface select the slide `TCGA-06-6562-01Z-00-DX1` on the left sidebar, click on `Annotation Groups` at the bottom of the right sidebar and switch the slider(s) to make annotations visible." + "The cell nuclei delineations can be viewed and explored in detail on the corresponding slide using the Slim viewer. In the Slim viewer's interface, select the slide `TCGA-06-6562-01Z-00-DX1` on the left sidebar, expand `Annotation Groups` at the bottom of the right sidebar, and toggle the switche(s) to display them." ] }, { "cell_type": "code", "execution_count": 7, "metadata": { - "id": "O7VT4D0jVnhv", - "outputId": "822b713b-6b70-4147-f1f4-dcf8445bf364", "colab": { "base_uri": "https://localhost:8080/", "height": 921 - } + }, + "id": "O7VT4D0jVnhv", + "outputId": "822b713b-6b70-4147-f1f4-dcf8445bf364" }, "outputs": [ { - "output_type": "execute_result", "data": { - "text/plain": [ - "" - ], "text/html": [ "\n", " \n", " " + ], + "text/plain": [ + "" ] }, + "execution_count": 7, "metadata": {}, - "execution_count": 7 + "output_type": "execute_result" } ], "source": [ @@ -280,13 +283,13 @@ "id": "RTr_xLONyjLK" }, "source": [ - "## Reading DICOM ANNs\n", + "## Reading DICOM MBSAs\n", "\n", - "DICOM ANNs extend the capabilities of [DICOM Structured Report (SR)](https://highdicom.readthedocs.io/en/latest/generalsr.html) documents as they were developed specifically for the storage of a **large number of similar annotations** and corresponding measurements (hence the full name Microscopy Simple **Bulk** Annotations). A popular example are annotations of small structures like cells or cell nuclei.\n", + "DICOM MBSAs extend the capabilities of [DICOM Structured Reports (SRs)](https://highdicom.readthedocs.io/en/latest/generalsr.html) since they were developed specifically for the storage of a **large number of structures of similar type** (hence the term \"bulk\" in Microscopy **Bulk** Simple Annotations). A popular example is the delineation of small structures like cells or cell nuclei.\n", "\n", - "Each ANN object contains one or more \"Annotation Groups\" consisting of many similar graphical annotations, optionally accompanied by one or several numerical measurements belonging to those graphical annotations as well as some required and some optional metadata that describe the contents of the group. These include an annotation group identifier, a human-readable label but also coded values that describe the category and the type of the annotated structure (see [here](https://highdicom.readthedocs.io/en/latest/package.html#highdicom.ann.AnnotationGroup) for a complete documentation).\n", + "Within an MBSA object, such structures are organized into one or more **\"Annotation Groups\"**. Besides the graphic data itself, each group may carry one or several numerical measurements per structure, along with required and optional metadata describing the content of the group. These metadata include an identifier for the group, a human-readable label but also coded values that describe the category and the type of the annotated structure (see [here](https://highdicom.readthedocs.io/en/latest/package.html#highdicom.ann.AnnotationGroup) for a complete documentation).\n", "\n", - "The following code uses [highdicom](https://github.com/ImagingDataCommons/highdicom) to extract annotation groups and metadata from a single DICOM ANN. More helpful explanations and guidance through implementation details for ANN objects in highdicom can be found [here](https://highdicom.readthedocs.io/en/latest/ann.html#microscopy-bulk-simple-annotation-ann-objects)." + "The following code uses `highdicom` to extract annotation groups and metadata from a single DICOM MBSA. More helpful explanations and guidance through implementation details for MBSA objects in `highdicom` can be found [here](https://highdicom.readthedocs.io/en/latest/ann.html#microscopy-bulk-simple-annotation-ann-objects)." ] }, { @@ -301,8 +304,8 @@ }, "outputs": [ { - "output_type": "stream", "name": "stdout", + "output_type": "stream", "text": [ "Number of annotation groups: 1\n" ] @@ -327,8 +330,8 @@ }, "outputs": [ { - "output_type": "stream", "name": "stdout", + "output_type": "stream", "text": [ "Unique identifier of the annotation group: 1.2.826.0.1.3680043.10.511.3.11186236449463673483170621129963156\n", "Human-readable label of the annotation group: Nuclei\n", @@ -364,7 +367,7 @@ "id": "g-z2XnaMyIpH" }, "source": [ - "In highdicom, the graphic annotation data are encoded as a list of numpy arrays, each of the shape (N x D). N is the number of coordinates which depends on the graphic type, e.g. a `POINT` will have one coordinate, while a `POLYGON` has >= 3 coordinates. Coordinates are either defined in the 2D image coordinate system (D=2), or in the frame-of-reference coordinate system (D=3)." + "In `highdicom`, the graphic data are encoded as a list of numpy arrays, each of the shape (N x D). N is the number of coordinates which depends on the graphic type, e.g. a `POINT` will have one coordinate, while a `POLYGON` has >= 3 coordinates. Coordinates are either defined in the image-relative coordinate system (D=2), or in the frame-of-reference coordinate system (D=3)." ] }, { @@ -379,8 +382,8 @@ }, "outputs": [ { - "output_type": "stream", "name": "stdout", + "output_type": "stream", "text": [ "Coordinates of first annotation: [[15950.0, 15620.0], [15949.0, 15621.0], [15948.0, 15621.0], [15946.0, 15623.0], [15946.0, 15626.0], [15947.0, 15627.0], [15947.0, 15632.0], [15949.0, 15634.0], [15950.0, 15634.0], [15951.0, 15635.0], [15954.0, 15635.0], [15955.0, 15634.0], [15956.0, 15634.0], [15958.0, 15632.0], [15958.0, 15631.0], [15959.0, 15630.0], [15959.0, 15624.0], [15958.0, 15623.0], [15958.0, 15622.0], [15957.0, 15622.0], [15956.0, 15621.0], [15955.0, 15621.0], [15954.0, 15620.0]]\n" ] @@ -399,7 +402,7 @@ "id": "AfxGQ-WXNGkj" }, "source": [ - "As previously mentioned, annotations can be accompanied by one or multiple measurements. In highdicom, they are returned as tuple of `(names, values, units)` each of them being a list of coded values. For the Pan-Cancer-Nuclei-Seg-DICOM collection these are each nuclei's area given in \u00b5m\u00b2." + "As previously mentioned, the graphic data can be accompanied by one or multiple measurements. In `highdicom`, they are returned as tuple of `(names, values, units)` each of them being a list of coded values. For the Pan-Cancer-Nuclei-Seg-DICOM collection these are each nuclei's area given in µm²." ] }, { @@ -414,8 +417,8 @@ }, "outputs": [ { - "output_type": "stream", "name": "stdout", + "output_type": "stream", "text": [ "Measurements for \"Area\" in unit \"square micrometer\"\n", "[11.367216 21.59136 4.699296 ... 7.429968 5.905872 3.111696]\n" @@ -435,11 +438,13 @@ "id": "bY5f6tybJGHT" }, "source": [ - "## Making use of annotations: Compute cellularity per tile\n", + "## Making use of the data: Compute cellularity per tile\n", + "\n", + "**Cellularity** is the extent to which a tissue region is occupied by cells and is a widely used descriptor in pathology. (Note that the MBSAs used here delineate cell nuclei rather than whole cells, so the measures computed below are nucleus-based proxies for cellularity.)\n", "\n", - "Having learned the basics of how to search for and access bulk annotations in the IDC, the question arises on how to use them for analysis and in combination with the WSIs to which they refer. As an example, we outline the use case of **estimating cellularity** as the ratio of the area showing cells in comparison to other tissue in a slide. At the same time we keep track of the **number of nuclei** in a certain area, which is also often referred to as cellularity.\n", + "MBSAs make the computation of cellularity quite straightforward. Because the contours of cell nuclei are already available as explicit polygon coordinates, they only have to be related to the coordinate system of the corresponding WSI. Building on the basics of searching for and accessing MBSAs in the IDC covered above, we compute the **fraction of the tissue area covered by cell nuclei** and at the same time the **number of nuclei** in this area.\n", "\n", - "The following three code cells contain general utility functions, functions for cellularity computation and finally helper code for visulizations.\n" + "The following three code cells contain general utility functions, functions for cellularity computation and finally helper code for visualizations.\n" ] }, { @@ -580,9 +585,9 @@ "id": "stgGxAp70HBx" }, "source": [ - "In the following, we will compute the cellularity - the ratio of a `TILE_SIZE x TILE_SIZE` sized tissue region showing nuclei compared other tissue - given simply our DICOM ANN file. From this file we can extract the nuclei annotations and the DICOM SOPInstanceUID and SeriesUID of the referenced slide. The cellularity computations can be done without actually looking at the original slide - the only information needed is the image width and height, which can be obtained using the `idc_client`.\n", + "In the following, we will compute the cellularity - the fraction of a `TILE_SIZE x TILE_SIZE` sized tissue region covered by nuclei - given simply our DICOM MBSA file. From this file we can extract the nuclei contours and the DICOM SOPInstanceUID and SeriesUID of the referenced slide. The cellularity computations can be done without actually looking at the original slide - the only information needed is the image width and height, which can be obtained using the `idc-client`.\n", "\n", - "Note that the selection of TILE_SIZE is independent from the size of the tiles in the encoded DICOM image!\n", + "Note that the selection of `TILE_SIZE` is independent from the size of the tiles in the encoded DICOM image!\n", "\n", "If you are running this notebook on Google Colab, the following cell should take around 2 minutes to complete." ] @@ -627,7 +632,7 @@ "name": "stderr", "output_type": "stream", "text": [ - "Downloading data: 100%|\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2588\u2589| 492M/492M [00:03<00:00, 163MB/s]\n" + "Downloading data: 100%|█████████▉| 492M/492M [00:03<00:00, 163MB/s]\n" ] }, { @@ -667,7 +672,7 @@ "id": "0GuCTTxw2735" }, "source": [ - "To finally have a quick-and-easy verifying look on regions of interest - in our case the tiles with the hightest cellularity values - we access those tiles and overlay nuclei annotations." + "To finally have a quick-and-easy verifying look on regions of interest - in our case the tiles with the highest cellularity values - we access those tiles and overlay nuclei contours." ] }, { @@ -714,7 +719,7 @@ "id": "61Pun4JvMnqT" }, "source": [ - "We can see that most tiles indeed contain a large number of nuclei, however there are two tiles clearly differing and likely showing artifacts where the algorithm predicted way too small, large or misshaped nuclei. In that case, the number of nuclei in a tile might be a better indicator of cellularity. Also cellularity as defined above will prefer less, but larger nuclei over more but smaller nuclei. Thus, below, we'll display the tiles with the largest number of nuclei. " + "We can see that most tiles indeed contain a large number of nuclei. Two tiles, however, clearly differ and likely show artifacts where the algorithm predicted nuclei that are too small, too large or misshaped. In that case, the number of nuclei in a tile might be a better indicator of cellularity. Also cellularity as defined above will prefer less, but larger nuclei over more but smaller nuclei. Thus, below, we'll display the tiles with the largest number of nuclei. " ] }, { @@ -761,9 +766,9 @@ "id": "stZqLjGPGLS3" }, "source": [ - "## Exporting DICOM ANNs to GeoJSON for usage in external tools\n", + "## Exporting DICOM MBSAs to GeoJSON for usage in external tools\n", "\n", - "Since many computational pathology tools use GeoJSON as their preferred file format for annotations, the following code shows how to easily export annotations captured in DICOM ANN as GeoJSON. Note, however, that due to the large amount of nuclei (sometimes about 1 billion per slide), it may lead to quite large GeoJSON files. For demonstration purposes, we limit ourselves here to the first 50000 nuclei." + "Since currently, some computational pathology tools use GeoJSON as their preferred file format for annotations, the following code shows how to easily export the nuclei contours captured in DICOM MBSAs as GeoJSON. Note, however, that the large amount of nuclei (sometimes about 1 billion per slide), may lead to quite large GeoJSON files. For demonstration purposes, we limit ourselves here to the first 50,000 nuclei." ] }, { @@ -817,7 +822,7 @@ "id": "_YEYKt-Rj1qd" }, "source": [ - "The resulting GeoJSON file can then be imported into other tools for viewing and analysis, such as for example [QuPath](https://qupath.github.io/). Here QuPath v.0.5.1 was used. After opening the slide, you can load annotations as GeoJSON under `File` > `Import objects from file` and adapt the visualization (e.g. changing color) in the `Annotations` pane.\n", + "The resulting GeoJSON file can then be imported into other tools for viewing and analysis, such as for example [QuPath](https://qupath.github.io/). Here QuPath v.0.5.1 was used. After opening the slide, you can load the nuclei contours as GeoJSON via `File` > `Import objects from file` and adapt the visualization (e.g. changing color) in the `Annotations` panel.\n", "\n", "\"Example" ]