Abstract
Accurately detecting student behavior in classroom videos can aid in analyzing their classroom performance and improving teaching effectiveness. However, the current accuracy rate in behavior detection is low. To address this challenge, we propose the Student Classroom Behavior Detection system based on YOLOv7+BRA (YOLOv7 with Bi-level Routing Attention). We identified eight different behavior patterns, including standing, sitting, talking, listening, walking, raising hands, reading, and writing. We constructed a dataset, which contained 11,248 labels and 4,001 images, with an emphasis on the common behavior of raising hands in a classroom setting (Student Classroom Behavior dataset, SCB-Dataset). To improve detection accuracy, we added the biformer attention module to the YOLOv7 network. Finally, we fused the results from YOLOv7 CrowdHuman, SlowFast, and DeepSort models to obtain student classroom behavior data. We conducted experiments on the SCB-Dataset, and YOLOv7+BRA achieved an mAP@0.5 of 87.1%, resulting in a 2.2% improvement over previous results. Our SCB-dataset can be downloaded from: https://github.com/Whiffe/SCB-dataset.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
References
Zhu, Y., Li, X., Liu, C., et al.: A comprehensive study of deep video action recognition. arXiv preprint arXiv:2012.06567 (2020)
Huang, Y., Liang, M., Wang, X., et al.: Multi-person classroom action recognition in classroom teaching videos based on deep spatiotemporal residual convolution neural network. J. Comput. Appli. 42(3), 736 (2022)
He, X., Yang, F., Chen, Z., et al.: The recognition of student classroom behavior based on human skeleton and deep learning. Mod. Educ. Technol. 30(11), 105–112 (2020)
Yan, X., Kuang, Y., Bai, G., Li, Y.: Student classroom behavior recognition method based on deep learning. Comput. Eng. https://doi.org/10.19678/j.issn.1000-3428.0065369
Gu, C., Sun, C., Ross, D.A., et al.: Ava: a video dataset of spatio-temporally localized atomic visual actions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6047–6056 (2018)
Feichtenhofer, C., Fan, H., Malik, J., et al.: Slowfast networks for video recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6202–6211 (2019)
Soomro, K., Zamir, A.R., Shah, M.: UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299–6308 (2017)
Wang, C.Y., Bochkovskiy, A., Liao, H.Y.M.: YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv preprint arXiv:2207.02696 (2022)
Ren, S., He, K., Girshick, R., et al.: Faster r-cnn: towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems, vol. 28 (2015)
Redmon J, Farhadi A. Yolov3: an incremental improvement. arXiv preprint arXiv:1804.02767 (2018)
Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48
Fu. R., Wu, T., Luo, Z., et al.: Learning behavior analysis in classroom based on deep learning. In: 2019 Tenth International Conference on Intelligent Control and Information Processing (ICICIP), pp. 206–212. IEEE (2019)
Zheng, R., Jiang, F., Shen, R.: Intelligent student behavior analysis system for real classrooms. In: ICASSP 2020–2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 9244–9248. IEEE (2020)
Sun, B., Wu, Y., Zhao, K., et al.: Student class behavior dataset: a video dataset for recognizing, detecting, and captioning students’ behaviors in classroom scenes[J]. Neural Comput. Appl. 33, 8335–8354 (2021)
Zhou, Y.: Research on Classroom Behaviors Detection of Primary School Students Based on Faster R-CNN. Sichuan Normal University (2021). https://doi.org/10.27347/d.cnki.gssdu.2021.000962
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014)
Vaswani. A., Shazeer, N., Parmar, N., et al.: Attention is all you need. In: Advances in Neural Information Processing Systems 30 (2017)
Liu, Z., Lin, Y., Cao, Y., et al.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022 (2021)
Xia, Z., Pan, X., Song, S., et al.: Vision transformer with deformable attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4794–4803 (2022)
Zeng, W., Jin, S., Liu, W., et al.: Not all tokens are equal: human-centric visual analysis via token clustering transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11101–11111 (2022)
Chen, Z., Zhu, Y., Zhao, C., et al.: Dpt: deformable patch-based transformer for visual recognition. In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 2899–2907 (2021)
Zhu, L., Wang, X., Ke, Z., et al.: BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv preprint arXiv:2303.08810 (2023)
Ngoc Anh, B., Tung Son, N., Truong Lam, P., et al.: A computer-vision based application for student behavior monitoring in classroom. Appl. Sci. 9(22), 4729 (2019)
Lin, F.C., Ngo, H.H., Dow, C.R., et al.: Student behavior recognition system for the classroom environment based on skeleton pose estimation and person detection. Sensors 21(16), 5314 (2021)
Trabelsi, Z., Alnajjar, F., Parambil, M.M.A., et al.: Real-time attention monitoring system for classroom: a deep learning approach for student’s behavior recognition. Big Data Cognitive Comput. 7(1), 48 (2023)
Yang, F.: A Multi-Person Video Dataset Annotation Method of Spatio-Temporally Actions. arXiv preprint arXiv:2204.10160 (2022)
Wojke, N., Bewley, A., Paulus D.: Simple online and realtime tracking with a deep association metric. In: 2017 IEEE International Conference on Image Processing (ICIP), pp. 3645–3649. IEEE (2017)
Shao, S., Zhao, Z., Li, B., et al.: Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123 (2018)
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2023 The Author(s), under exclusive license to Springer Nature Switzerland AG
About this paper
Cite this paper
Yang, F., Wang, T., Wang, X. (2023). Student Classroom Behavior Detection Based on YOLOv7+BRA and Multi-model Fusion. In: Lu, H., et al. Image and Graphics . ICIG 2023. Lecture Notes in Computer Science, vol 14357. Springer, Cham. https://doi.org/10.1007/978-3-031-46311-2_4
Download citation
DOI: https://doi.org/10.1007/978-3-031-46311-2_4
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-031-46310-5
Online ISBN: 978-3-031-46311-2
eBook Packages: Computer ScienceComputer Science (R0)