BOP Challenge 2020 on 6D Object Localization

Hodaň, Tomáš; Sundermeyer, Martin; Drost, Bertram; Labbé, Yann; Brachmann, Eric; Michel, Frank; Rother, Carsten; Matas, Jiří

doi:10.1007/978-3-030-66096-3_39

Tomáš Hodaň¹⁰,
Martin Sundermeyer¹¹,
Bertram Drost¹²,
Yann Labbé¹³,
Eric Brachmann¹⁴,
Frank Michel¹⁵,
Carsten Rother¹⁴ &
…
Jiří Matas¹⁰

Part of the book series: Lecture Notes in Computer Science ((LNIP,volume 12536))

Included in the following conference series:

European Conference on Computer Vision

2660 Accesses
78 Citations

Abstract

This paper presents the evaluation methodology, datasets, and results of the BOP Challenge 2020, the third in a series of public competitions organized with the goal to capture the status quo in the field of 6D object pose estimation from an RGB-D image. In 2020, to reduce the domain gap between synthetic training and real test RGB images, the participants were 350K photorealistic training images generated by BlenderProc4BOP, a new open-source and light-weight physically-based renderer (PBR) and procedural data generator. Methods based on deep neural networks have finally caught up with methods based on point pair features, which were dominating previous editions of the challenge. Although the top-performing methods rely on RGB-D image channels, strong results were achieved when only RGB channels were used at both training and test time – out of the 26 evaluated methods, the third method was trained on RGB channels of PBR and real images, while the fifth on RGB channels of PBR images only. Strong data augmentation was identified as a key component of the top-performing CosyPose method, and the photorealism of PBR images was demonstrated effective despite the augmentation. The online evaluation system stays open and is available on the project website: bop.felk.cvut.cz.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

A Tool for Building Multi-purpose and Multi-pose Synthetic Data Sets

Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset

Article 23 October 2023

Towards Generalization Across Depth for Monocular 3D Object Detection

Notes

1.
BOP stands for Benchmark for 6D Object Pose Estimation [25].
2.
The point cloud is calculated from the depth channel and known camera parameters.
3.
A distance map stores at a pixel p the distance from the camera center to a 3D point \(\mathbf {x}_p\) that projects to p. It can be readily computed from the depth map which stores at p the Z coordinate of \(\mathbf {x}_p\) and which is a typical output of Kinect-like sensors.
4.
github.com/DLR-RM/BlenderProc/blob/master/README_BlenderProc4BOP.md.
5.
Method #2 used also synthetic training images obtained by cropping the objects from real validation images in the case of HB and ITODD and from OpenGL-rendered images in the case of other datasets, and pasting the cropped objects on images from the Microsoft COCO dataset [36]. Method #24 used PBR and real images for training Mask R-CNN [16] and OpenGL images for training a single Multi-path encoder. Two of the CosyPose variants (#1 and #3) also added the “render & paste” synthetic images provided in the original YCB-V dataset, but these images were later found to have no effect on the accuracy score.

References

Intel Open Image Denoise (2020). https://www.openimagedenoise.org/
MVTec HALCON (2020). https://www.mvtec.com/halcon/
Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., Rother, C.: Learning 6D object pose estimation using 3D object coordinates. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8690, pp. 536–551. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10605-2_35
Chapter Google Scholar
Blender Online Community: Blender - a 3D modelling and rendering package (2018). http://www.blender.org
Demes, L.: CC0 textures (2020). https://cc0textures.com/
Denninger, M., et al.: BlenderProc: reducing the reality gap with photorealistic rendering. In: Robotics: Science and Systems (RSS) Workshops (2020)
Google Scholar
Denninger, M., et al.: BlenderProc. arXiv preprint arXiv:1911.01911 (2019)
Doumanoglou, A., Kouskouridas, R., Malassiotis, S., Kim, T.K.: Recovering 6D object pose and predicting next-best-view in the crowd. In: CVPR (2016)
Google Scholar
Drost, B., Ulrich, M., Bergmann, P., Hartinger, P., Steger, C.: Introducing MVTec ITODD - a dataset for 3D object recognition in industry. In: ICCVW (2017)
Google Scholar
Drost, B., Ulrich, M., Navab, N., Ilic, S.: Model globally, match locally: efficient and robust 3D object recognition. In: CVPR (2010)
Google Scholar
Dwibedi, D., Misra, I., Hebert, M.: Cut, paste and learn: surprisingly easy synthesis for instance detection. In: ICCV (2017)
Google Scholar
Fu, C.Y., Shvets, M., Berg, A.C.: RetinaMask: learning to predict masks improves state-of-the-art single-shot detection for free. arXiv preprint arXiv:1901.03353 (2019)
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F.A., Brendel, W.: ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv preprint arXiv:1811.12231 (2018)
Godard, C., Hedman, P., Li, W., Brostow, G.J.: Multi-view reconstruction of highly specular surfaces in uncontrolled environments. In: 3DV (2015)
Google Scholar
Hagelskjær, F., Buch, A.G.: PointPoseNet: accurate object detection and 6 DOF pose estimation in point clouds. arXiv preprint arXiv:1912.09057 (2019)
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask R-CNN. In: ICCV (2017)
Google Scholar
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016)
Google Scholar
Hinterstoisser, S., et al.: Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes. In: Lee, K.M., Matsushita, Y., Rehg, J.M., Hu, Z. (eds.) ACCV 2012. LNCS, vol. 7724, pp. 548–562. Springer, Heidelberg (2013). https://doi.org/10.1007/978-3-642-37331-2_42
Chapter Google Scholar
Hinterstoisser, S., Lepetit, V., Wohlhart, P., Konolige, K.: On pre-trained image features and synthetic images for deep learning. In: Leal-Taixé, L., Roth, S. (eds.) ECCV 2018. LNCS, vol. 11129, pp. 682–697. Springer, Cham (2019). https://doi.org/10.1007/978-3-030-11009-3_42
Chapter Google Scholar
Hinterstoisser, S., Pauly, O., Heibel, H., Martina, M., Bokeloh, M.: An annotation saved is an annotation earned: using fully synthetic training for object detection. In: ICCVW (2019)
Google Scholar
Hodaň, T., Baráth, D., Matas, J.: EPOS: estimating 6D pose of objects with symmetries. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
Google Scholar
Hodaň, T., et al.: BOP challenge 2019 (2019). https://bop.felk.cvut.cz/media/bop_challenge_2019_results.pdf
Hodaň, T., Haluza, P., Obdržálek, Š., Matas, J., Lourakis, M., Zabulis, X.: T-LESS: an RGB-D dataset for 6D pose estimation of texture-less objects. In: IEEE Winter Conference on Applications of Computer Vision (WACV) (2017)
Google Scholar
Hodaň, T., Matas, J., Obdržálek, Š.: On evaluation of 6D object pose estimation. In: Hua, G., Jégou, H. (eds.) ECCV 2016. LNCS, vol. 9915, pp. 606–619. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-49409-8_52
Chapter Google Scholar
Hodaň, T., et al.: BOP: benchmark for 6D object pose estimation. In: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (eds.) ECCV 2018. LNCS, vol. 11214, pp. 19–35. Springer, Cham (2018). https://doi.org/10.1007/978-3-030-01249-6_2
Chapter Google Scholar
Hodaň, T., Michel, F., Sahin, C., Kim, T.K., Matas, J., Rother, C.: SIXD challenge 2017 (2017). http://cmp.felk.cvut.cz/sixd/challenge_2017/
Hodaň, T., Sundermeyer, M.: BOP toolkit (2020). https://github.com/thodan/bop_toolkit
Hodaň, T., et al.: 6th International Workshop on Recovering 6D Object Pose (2020). http://cmp.felk.cvut.cz/sixd/workshop_2020/
Hodaň, T., et al.: Photorealistic image synthesis for object instance detection. In: IEEE International Conference on Image Processing (ICIP) (2019)
Google Scholar
Kaskman, R., Zakharov, S., Shugurov, I., Ilic, S.: HomebrewedDB: RGB-D dataset for 6D pose estimation of 3D objects. In: ICCVW (2019)
Google Scholar
Kehl, W., Manhardt, F., Tombari, F., Ilic, S., Navab, N.: SSD-6D: making RGB-based 3D detection and 6D pose estimation great again. In: ICCV (2017)
Google Scholar
Koenig, R., Drost, B.: A hybrid approach for 6DoF pose estimation. In: Bartoli, A., Fusiello, A. (eds.) ECCV 2020 Workshops. LNCS, vol. 12536, pp. 700–706. Springer, Cham (2020)
Google Scholar
Labbé, Y., Carpentier, J., Aubry, M., Sivic, J.: CosyPose: consistent multi-view multi-object 6D pose estimation. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M. (eds.) ECCV 2020. LNCS, vol. 12362, pp. 574–591. Springer, Cham (2020). https://doi.org/10.1007/978-3-030-58520-4_34
Chapter Google Scholar
Li, Y., Wang, G., Ji, X., Xiang, Yu., Fox, D.: DeepIM: deep iterative matching for 6D pose estimation. In: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (eds.) ECCV 2018. LNCS, vol. 11210, pp. 695–711. Springer, Cham (2018). https://doi.org/10.1007/978-3-030-01231-1_42
Chapter Google Scholar
Li, Z., Wang, G., Ji, X.: CDPN: coordinates-based disentangled pose network for real-time RGB-based 6-DoF object pose estimation. In: ICCV (2019)
Google Scholar
Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48
Chapter Google Scholar
Liu, J., Zou, Z., Ye, X., Tan, X., Ding, E., Xu, F., Yu, X.: Leaping from 2D detection to efficient 6DoF object pose estimation. In: Bartoli, A., Fusiello, A. (eds.) ECCV 2020. LNCS, vol. 12536, pp. 1–11. Springer, Cham (2020)
Google Scholar
Marschner, S., Shirley, P.: Fundamentals of Computer Graphics. CRC Press, Boca Raton (2015)
MATH Google Scholar
Newcombe, R.A., et al.: KinectFusion: real-time dense surface mapping and tracking. In: ISMAR (2011)
Google Scholar
Park, K., Patten, T., Vincze, M.: Pix2Pose: pixel-wise coordinate regression of objects for 6D pose estimation. In: ICCV (2019)
Google Scholar
Pharr, M., Jakob, W., Humphreys, G.: Physically Based Rendering: From Theory to Implementation. Morgan Kaufmann, Burlington (2016)
Google Scholar
Qian, Y., Gong, M., Hong Yang, Y.: 3D reconstruction of transparent objects with position-normal consistency. In: CVPR (2016)
Google Scholar
Rad, M., Lepetit, V.: BB8: a scalable, accurate, robust to partial occlusion method for predicting the 3D poses of challenging objects without using depth. In: ICCV (2017)
Google Scholar
Raposo, C., Barreto, J.P.: Using 2 point+normal sets for fast registration of point clouds with small overlap. In: ICRA (2017)
Google Scholar
Rennie, C., Shome, R., Bekris, K.E., De Souza, A.F.: A dataset for improved RGBD-based object detection and pose estimation for warehouse pick-and-place. RA-L 1(2), 1179–1185 (2016)
Google Scholar
Rodrigues, P., Antunes, M., Raposo, C., Marques, P., Fonseca, F., Barreto, J.: Deep segmentation leverages geometric pose estimation in computer-aided total knee arthroplasty. Healthc. Technol. Lett. 6(6), 226–230 (2019)
Article Google Scholar
Sundermeyer, M., et al.: Multi-path learning for object pose estimation across domains. In: CVPR (2020)
Google Scholar
Sundermeyer, M., Marton, Z.-C., Durner, M., Brucker, M., Triebel, R.: Implicit 3D orientation learning for 6D object detection from RGB images. In: Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y. (eds.) ECCV 2018. LNCS, vol. 11210, pp. 712–729. Springer, Cham (2018). https://doi.org/10.1007/978-3-030-01231-1_43
Chapter Google Scholar
Sundermeyer, M., Marton, Z.C., Durner, M., Triebel, R.: Augmented autoencoders: implicit 3D orientation learning for 6D object detection. IJCV 128, 714–729 (2019). https://doi.org/10.1007/s11263-019-01243-8
Article Google Scholar
Tejani, A., Tang, D., Kouskouridas, R., Kim, T.-K.: Latent-class hough forests for 3D object detection and pose estimation. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8694, pp. 462–477. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10599-4_30
Chapter Google Scholar
Vidal, J., Lin, C.Y., Lladó, X., Martí, R.: A method for 6D pose estimation of free-form rigid objects using point pair features on range data. Sensors 18, 2678 (2018)
Article Google Scholar
Wu, B., Zhou, Y., Qian, Y., Cong, M., Huang, H.: Full 3D reconstruction of transparent objects. ACM TOG 37, 1–11 (2018)
Google Scholar
Xiang, Y., Schmidt, T., Narayanan, V., Fox, D.: PoseCNN: a convolutional neural network for 6D object pose estimation in cluttered scenes. In: RSS (2018)
Google Scholar
Zakharov, S., Shugurov, I., Ilic, S.: DPOD: 6D pose object detector and refiner. In: ICCV (2019)
Google Scholar

Download references

Acknowledgements

This research was supported by CTU student grant (SGS OHK3-019/20), Research Center for Informatics (CZ.02.1.01/0.0/0.0/16_019/0000765 funded by OP VVV), and HPC resources from GENCI-IDRIS (grant 011011181).

Author information

Authors and Affiliations

Czech Technical University in Prague, Prague, Czech Republic
Tomáš Hodaň & Jiří Matas
German Aerospace Center, Wessling, Germany
Martin Sundermeyer
MVTec, Munich, Germany
Bertram Drost
INRIA Paris, Paris, France
Yann Labbé
Heidelberg University, Heidelberg, Germany
Eric Brachmann & Carsten Rother
Technical University Dresden, Dresden, Germany
Frank Michel

Authors

Tomáš Hodaň
View author publications
You can also search for this author in PubMed Google Scholar
Martin Sundermeyer
View author publications
You can also search for this author in PubMed Google Scholar
Bertram Drost
View author publications
You can also search for this author in PubMed Google Scholar
Yann Labbé
View author publications
You can also search for this author in PubMed Google Scholar
Eric Brachmann
View author publications
You can also search for this author in PubMed Google Scholar
Frank Michel
View author publications
You can also search for this author in PubMed Google Scholar
Carsten Rother
View author publications
You can also search for this author in PubMed Google Scholar
Jiří Matas
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Tomáš Hodaň .

Editor information

Editors and Affiliations

University of Clermont Auvergne, Clermont Ferrand, France
Adrien Bartoli
Università degli Studi di Udine, Udine, Italy
Andrea Fusiello

1 Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary material 1 (pdf 121 KB)

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Hodaň, T. et al. (2020). BOP Challenge 2020 on 6D Object Localization. In: Bartoli, A., Fusiello, A. (eds) Computer Vision – ECCV 2020 Workshops. ECCV 2020. Lecture Notes in Computer Science(), vol 12536. Springer, Cham. https://doi.org/10.1007/978-3-030-66096-3_39

Download citation

DOI: https://doi.org/10.1007/978-3-030-66096-3_39
Published: 03 January 2021
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-66095-6
Online ISBN: 978-3-030-66096-3
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

BOP Challenge 2020 on 6D Object Localization

Abstract

Access this chapter

Similar content being viewed by others

A Tool for Building Multi-purpose and Multi-pose Synthetic Data Sets

Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset

Towards Generalization Across Depth for Monocular 3D Object Detection

Notes

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

1 Electronic supplementary material

Supplementary material 1 (pdf 121 KB)

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Navigation

BOP Challenge 2020 on 6D Object Localization

Abstract

Access this chapter

Similar content being viewed by others

A Tool for Building Multi-purpose and Multi-pose Synthetic Data Sets

Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset

Towards Generalization Across Depth for Monocular 3D Object Detection

Notes

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Editor information

Editors and Affiliations

1 Electronic supplementary material

Supplementary material 1 (pdf 121 KB)

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation