Multi-task Dynamic Transformer Network for Concurrent Bone Segmentation and Large-Scale Landmark Localization with Dental CBCT

Lian, Chunfeng; Wang, Fan; Deng, Hannah H.; Wang, Li; Xiao, Deqiang; Kuang, Tianshu; Lin, Hung-Ying; Gateno, Jaime; Shen, Steve G. F.; Yap, Pew-Thian; Xia, James J.; Shen, Dinggang

doi:10.1007/978-3-030-59719-1_78

Chunfeng Lian¹⁶,
Fan Wang¹⁶,
Hannah H. Deng¹⁷,
Li Wang¹⁶,
Deqiang Xiao¹⁶,
Tianshu Kuang¹⁷,
Hung-Ying Lin¹⁷,
Jaime Gateno^17,18,
Steve G. F. Shen^19,20,
Pew-Thian Yap¹⁶,
James J. Xia^17,18 &
…
Dinggang Shen¹⁶

Part of the book series: Lecture Notes in Computer Science ((LNIP,volume 12264))

Included in the following conference series:

International Conference on Medical Image Computing and Computer-Assisted Intervention

9539 Accesses
15 Citations

Abstract

Accurate bone segmentation and anatomical landmark localization are essential tasks in computer-aided surgical simulation for patients with craniomaxillofacial (CMF) deformities. To leverage the complementarity between the two tasks, we propose an efficient end-to-end deep network, i.e., multi-task dynamic transformer network (DTNet), to concurrently segment CMF bones and localize large-scale landmarks in one-pass from large volumes of cone-beam computed tomography (CBCT) data. Our DTNet was evaluated quantitatively using CBCTs of patients with CMF deformities. The results demonstrated that our method outperforms the other state-of-the-art methods in both tasks of the bony segmentation and the landmark digitization. Our DTNet features three main technical contributions. First, a collaborative two-branch architecture is designed to efficiently capture both fine-grained image details and complete global context for high-resolution volume-to-volume prediction. Second, leveraging anatomical dependencies between landmarks, regionalized dynamic learners (RDLs) are designed in the concept of “learns to learn” to jointly regress large-scale 3D heatmaps of all landmarks under limited computational costs. Third, adaptive transformer modules (ATMs) are designed for the flexible learning of task-specific feature embedding from common feature bases.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 119.00; Price excludes VAT (USA)

Softcover Book: USD 159.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Bertinetto, L., et al.: Learning feed-forward one-shot learners. In: NeurIPS, pp. 523–531 (2016)
Google Scholar
Chen, W., et al.: Collaborative global-local networks for memory-efficient segmentation of ultra-high resolution images. In: CVPR, pp. 8924–8933 (2019)
Google Scholar
Gupta, A., et al.: A knowledge-based algorithm for automatic detection of cephalometric landmarks on CBCT images. Int. J. Comput. Assist. Radiol. Surg. 10(11), 1737–1752 (2015)
Article Google Scholar
He, K., et al.: Deep residual learning for image recognition. In: CVPR, pp. 770–778 (2016)
Google Scholar
Howard, A.G., et al.: MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)
Hu, J., et al.: Squeeze-and-excitation networks. In: CVPR, pp. 7132–7141 (2018)
Google Scholar
Lian, C., et al.: Multi-channel multi-scale fully convolutional network for 3D perivascular spaces segmentation in 7T MR images. Med. Image Anal. 46, 106–117 (2018)
Article Google Scholar
Lian, C., et al.: Hierarchical fully convolutional network for joint atrophy localization and Alzheimer’s disease diagnosis using structural MRI. IEEE Trans. Pattern Anal. Mach. Intell. 42(4), 880–893 (2020)
Article Google Scholar
Liu, S., et al.: End-to-end multi-task learning with attention. In: CVPR, pp. 1871–1880 (2019)
Google Scholar
Long, J., et al.: Fully convolutional networks for semantic segmentation. In: CVPR, pp. 3431–3440 (2015)
Google Scholar
Nie, X., et al.: Human pose estimation with parsing induced learner. In: CVPR, pp. 2100–2108 (2018)
Google Scholar
Payer, C., Štern, D., Bischof, H., Urschler, M.: Regressing heatmaps for multiple landmark localization using CNNs. In: Ourselin, S., Joskowicz, L., Sabuncu, M.R., Unal, G., Wells, W. (eds.) MICCAI 2016. LNCS, vol. 9901, pp. 230–238. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46723-8_27
Chapter Google Scholar
Poudel, R.P., et al.: ContextNet: exploring context and detail for semantic segmentation in real-time. In: BMVC (2018)
Google Scholar
Ronneberger, O., Fischer, P., Brox, T.: U-Net: convolutional networks for biomedical image segmentation. In: Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F. (eds.) MICCAI 2015. LNCS, vol. 9351, pp. 234–241. Springer, Cham (2015). https://doi.org/10.1007/978-3-319-24574-4_28
Chapter Google Scholar
Shahidi, S., et al.: The accuracy of a designed software for automated localization of craniofacial landmarks on CBCT images. BMC Med. Imaging 14(1), 32 (2014)
Article Google Scholar
Torosdagli, N., et al.: Deep geodesic learning for segmentation and anatomical landmarking. IEEE Trans. Med. Imaging 38(4), 919–931 (2018)
Article Google Scholar
Wang, L., et al.: Automated segmentation of dental CBCT image with prior-guided sequential random forests. Med. Phys. 43(1), 336–346 (2016)
Article Google Scholar
Xia, J.J., et al.: New clinical protocol to evaluate craniomaxillofacial deformity and plan surgical correction. J. Oral Maxillofac. Surg. 67(10), 2093–2106 (2009)
Article Google Scholar
Yang, D., et al.: Automatic vertebra labeling in large-scale 3D CT using deep image-to-image network with message passing and sparsity regularization. In: Niethammer, M., et al. (eds.) IPMI 2017. LNCS, vol. 10265, pp. 633–644. Springer, Cham (2017). https://doi.org/10.1007/978-3-319-59050-9_50
Chapter Google Scholar
Zhang, J., et al.: Automatic craniomaxillofacial landmark digitization via segmentation-guided partially-joint regression forest model and multiscale statistical features. IEEE Trans. Biomed. Eng. 63(9), 1820–1829 (2015)
Article Google Scholar
Zhang, J., et al.: Context-guided fully convolutional networks for joint craniomaxillofacial bone segmentation and landmark digitization. Med. Image Anal. 60, 101621 (2020)
Article Google Scholar
Zhong, Z., Li, J., Zhang, Z., Jiao, Z., Gao, X.: An attention-guided deep regression model for landmark detection in cephalograms. In: Shen, D., et al. (eds.) MICCAI 2019. LNCS, vol. 11769, pp. 540–548. Springer, Cham (2019). https://doi.org/10.1007/978-3-030-32226-7_60
Chapter Google Scholar

Download references

Acknowledgements

This work was supported in part by NIH grants (R01 DE022676, R01 DE027251 and R01 DE021863).

Author information

Authors and Affiliations

Department of Radiology and BRIC, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
Chunfeng Lian, Fan Wang, Li Wang, Deqiang Xiao, Pew-Thian Yap & Dinggang Shen
Department of Oral and Maxillofacial Surgery, Houston Methodist Hospital, Houston, TX, USA
Hannah H. Deng, Tianshu Kuang, Hung-Ying Lin, Jaime Gateno & James J. Xia
Department of Surgery (Oral and Maxillofacial Surgery), Weill Medical College, Cornell University, New York, NY, USA
Jaime Gateno & James J. Xia
Department of Oral and Craniomaxillofacial Surgery, Shanghai Jiao Tong University, Shanghai, China
Steve G. F. Shen
Shanghai University of Medicine and Health Science, Shanghai, China
Steve G. F. Shen

Authors

Chunfeng Lian
View author publications
You can also search for this author in PubMed Google Scholar
Fan Wang
View author publications
You can also search for this author in PubMed Google Scholar
Hannah H. Deng
View author publications
You can also search for this author in PubMed Google Scholar
Li Wang
View author publications
You can also search for this author in PubMed Google Scholar
Deqiang Xiao
View author publications
You can also search for this author in PubMed Google Scholar
Tianshu Kuang
View author publications
You can also search for this author in PubMed Google Scholar
Hung-Ying Lin
View author publications
You can also search for this author in PubMed Google Scholar
Jaime Gateno
View author publications
You can also search for this author in PubMed Google Scholar
Steve G. F. Shen
View author publications
You can also search for this author in PubMed Google Scholar
Pew-Thian Yap
View author publications
You can also search for this author in PubMed Google Scholar
James J. Xia
View author publications
You can also search for this author in PubMed Google Scholar
Dinggang Shen
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding authors

Correspondence to James J. Xia or Dinggang Shen .

Editor information

Editors and Affiliations

University of Toronto, Toronto, ON, Canada
Anne L. Martel
The University of British Columbia, Vancouver, BC, Canada
Purang Abolmaesumi
University College London, London, UK
Danail Stoyanov
École Centrale de Nantes, Nantes, France
Diana Mateus
EURECOM, Biot, France
Maria A. Zuluaga
Chinese Academy of Sciences, Beijing, China
S. Kevin Zhou
Sorbonne University, Paris, France
Daniel Racoceanu
The Hebrew University of Jerusalem, Jerusalem, Israel
Leo Joskowicz

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Lian, C. et al. (2020). Multi-task Dynamic Transformer Network for Concurrent Bone Segmentation and Large-Scale Landmark Localization with Dental CBCT. In: Martel, A.L., et al. Medical Image Computing and Computer Assisted Intervention – MICCAI 2020. MICCAI 2020. Lecture Notes in Computer Science(), vol 12264. Springer, Cham. https://doi.org/10.1007/978-3-030-59719-1_78

Download citation

DOI: https://doi.org/10.1007/978-3-030-59719-1_78
Published: 29 September 2020
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-59718-4
Online ISBN: 978-3-030-59719-1
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Societies and partnerships

The Medical Image Computing and Computer Assisted Intervention Society (opens in a new tab)