Student Classroom Behavior Detection Based on YOLOv7+BRA and Multi-model Fusion

Yang, Fan; Wang, Tao; Wang, Xiaofei

doi:10.1007/978-3-031-46311-2_4

Fan Yang¹⁴,
Tao Wang¹⁴ &
Xiaofei Wang¹⁵

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 14357))

Included in the following conference series:

International Conference on Image and Graphics

505 Accesses
2 Citations

Abstract

Accurately detecting student behavior in classroom videos can aid in analyzing their classroom performance and improving teaching effectiveness. However, the current accuracy rate in behavior detection is low. To address this challenge, we propose the Student Classroom Behavior Detection system based on YOLOv7+BRA (YOLOv7 with Bi-level Routing Attention). We identified eight different behavior patterns, including standing, sitting, talking, listening, walking, raising hands, reading, and writing. We constructed a dataset, which contained 11,248 labels and 4,001 images, with an emphasis on the common behavior of raising hands in a classroom setting (Student Classroom Behavior dataset, SCB-Dataset). To improve detection accuracy, we added the biformer attention module to the YOLOv7 network. Finally, we fused the results from YOLOv7 CrowdHuman, SlowFast, and DeepSort models to obtain student classroom behavior data. We conducted experiments on the SCB-Dataset, and YOLOv7+BRA achieved an mAP@0.5 of 87.1%, resulting in a 2.2% improvement over previous results. Our SCB-dataset can be downloaded from: https://github.com/Whiffe/SCB-dataset.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 59.99; Price excludes VAT (USA)

Softcover Book: USD 79.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Zhu, Y., Li, X., Liu, C., et al.: A comprehensive study of deep video action recognition. arXiv preprint arXiv:2012.06567 (2020)
Huang, Y., Liang, M., Wang, X., et al.: Multi-person classroom action recognition in classroom teaching videos based on deep spatiotemporal residual convolution neural network. J. Comput. Appli. 42(3), 736 (2022)
Google Scholar
He, X., Yang, F., Chen, Z., et al.: The recognition of student classroom behavior based on human skeleton and deep learning. Mod. Educ. Technol. 30(11), 105–112 (2020)
Google Scholar
Yan, X., Kuang, Y., Bai, G., Li, Y.: Student classroom behavior recognition method based on deep learning. Comput. Eng. https://doi.org/10.19678/j.issn.1000-3428.0065369
Gu, C., Sun, C., Ross, D.A., et al.: Ava: a video dataset of spatio-temporally localized atomic visual actions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6047–6056 (2018)
Google Scholar
Feichtenhofer, C., Fan, H., Malik, J., et al.: Slowfast networks for video recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6202–6211 (2019)
Google Scholar
Soomro, K., Zamir, A.R., Shah, M.: UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299–6308 (2017)
Google Scholar
Wang, C.Y., Bochkovskiy, A., Liao, H.Y.M.: YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv preprint arXiv:2207.02696 (2022)
Ren, S., He, K., Girshick, R., et al.: Faster r-cnn: towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems, vol. 28 (2015)
Google Scholar
Redmon J, Farhadi A. Yolov3: an incremental improvement. arXiv preprint arXiv:1804.02767 (2018)
Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48
Chapter Google Scholar
Fu. R., Wu, T., Luo, Z., et al.: Learning behavior analysis in classroom based on deep learning. In: 2019 Tenth International Conference on Intelligent Control and Information Processing (ICICIP), pp. 206–212. IEEE (2019)
Google Scholar
Zheng, R., Jiang, F., Shen, R.: Intelligent student behavior analysis system for real classrooms. In: ICASSP 2020–2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 9244–9248. IEEE (2020)
Google Scholar
Sun, B., Wu, Y., Zhao, K., et al.: Student class behavior dataset: a video dataset for recognizing, detecting, and captioning students’ behaviors in classroom scenes[J]. Neural Comput. Appl. 33, 8335–8354 (2021)
Article Google Scholar
Zhou, Y.: Research on Classroom Behaviors Detection of Primary School Students Based on Faster R-CNN. Sichuan Normal University (2021). https://doi.org/10.27347/d.cnki.gssdu.2021.000962
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014)
Vaswani. A., Shazeer, N., Parmar, N., et al.: Attention is all you need. In: Advances in Neural Information Processing Systems 30 (2017)
Google Scholar
Liu, Z., Lin, Y., Cao, Y., et al.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022 (2021)
Google Scholar
Xia, Z., Pan, X., Song, S., et al.: Vision transformer with deformable attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4794–4803 (2022)
Google Scholar
Zeng, W., Jin, S., Liu, W., et al.: Not all tokens are equal: human-centric visual analysis via token clustering transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11101–11111 (2022)
Google Scholar
Chen, Z., Zhu, Y., Zhao, C., et al.: Dpt: deformable patch-based transformer for visual recognition. In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 2899–2907 (2021)
Google Scholar
Zhu, L., Wang, X., Ke, Z., et al.: BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv preprint arXiv:2303.08810 (2023)
Ngoc Anh, B., Tung Son, N., Truong Lam, P., et al.: A computer-vision based application for student behavior monitoring in classroom. Appl. Sci. 9(22), 4729 (2019)
Article Google Scholar
Lin, F.C., Ngo, H.H., Dow, C.R., et al.: Student behavior recognition system for the classroom environment based on skeleton pose estimation and person detection. Sensors 21(16), 5314 (2021)
Article Google Scholar
Trabelsi, Z., Alnajjar, F., Parambil, M.M.A., et al.: Real-time attention monitoring system for classroom: a deep learning approach for student’s behavior recognition. Big Data Cognitive Comput. 7(1), 48 (2023)
Article Google Scholar
Yang, F.: A Multi-Person Video Dataset Annotation Method of Spatio-Temporally Actions. arXiv preprint arXiv:2204.10160 (2022)
Wojke, N., Bewley, A., Paulus D.: Simple online and realtime tracking with a deep association metric. In: 2017 IEEE International Conference on Image Processing (ICIP), pp. 3645–3649. IEEE (2017)
Google Scholar
Shao, S., Zhao, Z., Li, B., et al.: Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123 (2018)

Download references

Author information

Authors and Affiliations

Chengdu Neusoft University, Chengdu, 611844, China
Fan Yang & Tao Wang
School of Films and Animation of the College of Chinese and ASEAN Arts, Chengdu University, Chengdu, 610106, China
Xiaofei Wang

Authors

Fan Yang
View author publications
You can also search for this author in PubMed Google Scholar
Tao Wang
View author publications
You can also search for this author in PubMed Google Scholar
Xiaofei Wang
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Xiaofei Wang .

Editor information

Editors and Affiliations

Dalian University of Technology, Dalian, China
Huchuan Lu
University of Sydney, Sydney, NSW, Australia
Wanli Ouyang
Shenzhen University, Shenzhen, China
Hui Huang
Tsinghua University, Beijing, China
Jiwen Lu
Dalian University of Technology, Dalian, China
Risheng Liu
Institute of Automation, CAS, Beijing, China
Jing Dong
University of Technology Sydney, Sydney, NSW, Australia
Min Xu

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Yang, F., Wang, T., Wang, X. (2023). Student Classroom Behavior Detection Based on YOLOv7+BRA and Multi-model Fusion. In: Lu, H., et al. Image and Graphics . ICIG 2023. Lecture Notes in Computer Science, vol 14357. Springer, Cham. https://doi.org/10.1007/978-3-031-46311-2_4

Download citation

DOI: https://doi.org/10.1007/978-3-031-46311-2_4
Published: 29 October 2023
Publisher Name: Springer, Cham
Print ISBN: 978-3-031-46310-5
Online ISBN: 978-3-031-46311-2
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Student Classroom Behavior Detection Based on YOLOv7+BRA and Multi-model Fusion