{"id":3059,"date":"2022-07-25T23:02:03","date_gmt":"2022-07-26T04:02:03","guid":{"rendered":"https:\/\/my.vanderbilt.edu\/masi\/?p=3059"},"modified":"2022-08-08T12:11:04","modified_gmt":"2022-08-08T17:11:04","slug":"self-supervised-pre-training-of-swin-transformers-for-3d-medical-image-analysis","status":"publish","type":"post","link":"https:\/\/my.vanderbilt.edu\/masi\/2022\/07\/self-supervised-pre-training-of-swin-transformers-for-3d-medical-image-analysis\/","title":{"rendered":"Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis"},"content":{"rendered":"<div class=\"gs_citr\">Tang, Yucheng, Dong Yang, Wenqi Li, Holger R. Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh. &#8220;Self-supervised pre-training of swin transformers for 3d medical image analysis.&#8221; In <i>Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition<\/i>, pp. 20730-20740. 2022.<\/div>\n<div class=\"gs_citr\"><\/div>\n<div class=\"gs_citr\"><strong>Full text:\u00a0<\/strong><\/div>\n<div class=\"gs_citr\"><\/div>\n<div class=\"gs_citr\">\n<h2>Abstract<\/h2>\n<p><span dir=\"ltr\">Vision Transformers (ViT)s have shown great performance <\/span><span dir=\"ltr\">in self-supervised learning of global and local representations <\/span><span dir=\"ltr\">that can be transferred to downstream applications. Inspired <\/span><span dir=\"ltr\">by these results, we introduce a novel self-supervised learning <\/span><span dir=\"ltr\">framework with tailored proxy tasks for medical image analy<\/span><span dir=\"ltr\">sis. Specifically, we propose: (i) a new 3D transformer-based <\/span><span dir=\"ltr\">model, dubbed Swin UNEt TRansformers (Swin UNETR), <\/span><span dir=\"ltr\">with a hierarchical encoder for self-supervised pre-training; <\/span><span dir=\"ltr\">(ii) tailored proxy tasks for learning the underlying pattern <\/span><span dir=\"ltr\">of human anatomy. We demonstrate successful pre-training <\/span><span dir=\"ltr\">of the proposed model on 5,050 publicly available computed <\/span><span dir=\"ltr\">tomography (CT) images from various body organs. The ef<\/span><span dir=\"ltr\">fectiveness of our approach is validated by fine-tuning the <\/span><span dir=\"ltr\">pre-trained models on the Beyond the Cranial Vault (BTCV) <\/span><span dir=\"ltr\">Segmentation Challenge with<\/span> <span dir=\"ltr\">13<\/span> <span dir=\"ltr\">abdominal organs and seg<\/span><span dir=\"ltr\">mentation tasks from the Medical Segmentation Decathlon <\/span><span dir=\"ltr\">(MSD) dataset. Our model is currently the state-of-the-art <\/span><span dir=\"ltr\">on the public test leaderboards of both MSD<\/span> <span dir=\"ltr\">and BTCV<\/span> <span dir=\"ltr\">datasets. Code: https:\/\/monai.io\/research\/swin unetr.<\/span><\/p>\n<figure id=\"attachment_3061\" aria-describedby=\"caption-attachment-3061\" style=\"width: 500px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-3061\" src=\"https:\/\/my.vanderbilt.edu\/masi\/wp-content\/uploads\/sites\/2304\n2661\/2022\/07\/Screenshot-from-2022-07-25-22-58-45.png\" alt=\"Overview of our proposed pre-training framework. Input CT images are randomly cropped into sub-volumes and augmented with random inner cutout and rotation, then fed to the Swin UNETR encoder as input. We use masked volume inpainting, contrastive learning and rotation prediction as proxy tasks for learning contextual representations of input images\" width=\"500\" height=\"509\" srcset=\"https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-45.png 424w, https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-45-294x300.png 294w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><figcaption id=\"caption-attachment-3061\" class=\"wp-caption-text\">Overview of our proposed pre-training framework. Input CT images are randomly cropped into sub-volumes and augmented with random inner cutout and rotation, then fed to the Swin UNETR encoder as input. We use masked volume inpainting, contrastive learning and rotation prediction as proxy tasks for learning contextual representations of input images<\/figcaption><\/figure>\n<\/div>\n<p>&nbsp;<\/p>\n<figure id=\"attachment_3060\" aria-describedby=\"caption-attachment-3060\" style=\"width: 500px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-3060\" src=\"https:\/\/my.vanderbilt.edu\/masi\/wp-content\/uploads\/sites\/2304\n2661\/2022\/07\/Screenshot-from-2022-07-25-22-58-32.png\" alt=\"Overview of the Swin UNETR architecture.\" width=\"500\" height=\"202\" srcset=\"https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-32.png 853w, https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-32-300x121.png 300w, https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-32-768x310.png 768w, https:\/\/cdn.vanderbilt.edu\/t2-my\/my-prd\/wp-content\/uploads\/sites\/2304\/2022\/07\/Screenshot-from-2022-07-25-22-58-32-650x262.png 650w\" sizes=\"auto, (max-width: 500px) 100vw, 500px\" \/><figcaption id=\"caption-attachment-3060\" class=\"wp-caption-text\">Overview of the Swin UNETR architecture.<\/figcaption><\/figure>\n","protected":false},"excerpt":{"rendered":"<p>Tang, Yucheng, Dong Yang, Wenqi Li, Holger R. Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh. &#8220;Self-supervised pre-training of swin transformers for 3d medical image analysis.&#8221; In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 20730-20740. 2022. Full text:\u00a0 Abstract Vision Transformers (ViT)s have shown great performance in self-supervised&#8230;<\/p>\n","protected":false},"author":7582,"featured_media":3061,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5,27,64,60,130],"tags":[137,21,190],"class_list":["post-3059","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-abdomen-imaging","category-big-data","category-body-wise","category-computed-tomography","category-deep-learning","tag-deep-learning","tag-segmentation","tag-transformer"],"_links":{"self":[{"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/posts\/3059","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/users\/7582"}],"replies":[{"embeddable":true,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/comments?post=3059"}],"version-history":[{"count":2,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/posts\/3059\/revisions"}],"predecessor-version":[{"id":3063,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/posts\/3059\/revisions\/3063"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/media\/3061"}],"wp:attachment":[{"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/media?parent=3059"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/categories?post=3059"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/my.vanderbilt.edu\/masi\/wp-json\/wp\/v2\/tags?post=3059"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}