Skip to main content
Log in

An improved one-stage pedestrian detection method based on multi-scale attention feature extraction

  • Original Research Paper
  • Published:
Journal of Real-Time Image Processing Aims and scope Submit manuscript

Abstract

In recent years, the performance of the convolutional neural network-based pedestrian detection method has improved significantly. However, an imbalance remains between detection accuracy and speed. In this paper, we employ a one-stage object detection framework and propose a pedestrian detection method based on the multi-scale attention mechanism of a convolutional neural network to improve the imbalance between accuracy and speed. First, a multi-scale convolution module is designed to extract corresponding features at different scales. Second, using the attention module, association information between features is mined from space and channel perspectives to strengthen the original features. Then, the enhanced features are passed through a classification and regression module to perform object positioning and bounding box regression. Finally, to learn more pedestrian location information, we improve the loss function to realise better network training. The proposed method achieved considerable results on the challenging CityPersons and Caltech pedestrian detection datasets.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5
Fig. 6
Fig.7
Fig. 8
Fig. 9
Fig. 10

Similar content being viewed by others

References

  1. Krizhevsky, A., Sutskever, I., Hinton, G. E.: ImageNet classification with deep convolutional neural networks. In: Neural Information Processing Systems, pp. 1097–1105 (2012)

  2. Tian, Y., Luo, P., Wang, X., Tang, X.: Pedestrian detection aided by deep learning semantic tasks. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5079–5087 (2015)

  3. Zhang, L., Lin, L., Liang, X., He, K.: Is faster r-cnn doing well for pedestrian detection? In: European Conference on Computer Vision (ECCV), pp. 443–457 (2016)

  4. Liu, W., Liao, S., Ren, W., Hu, W., Yu, Y.: High-level semantic feature detection: a new perspective for pedestrian detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5187–5196 (2019)

  5. Ma, J., Wan, H., Wang, J., Xia, H., Bai, C.: An improved scheme of deep dilated feature extraction on pedestrian detection. SIViP (2020). https://doi.org/10.1007/s11760-020-01742-z

    Article  Google Scholar 

  6. Zhang, S., Benenson, R., & Schiele, B.: CityPersonss: a diverse dataset for pedestrian detection. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition, pp. 3213–3221 (2017)

  7. Dollar, P., Wojek, C., Schiele, B., Perona, P.: Pedestrian detection: An evaluation of the state of the art. IEEE Trans. Pattern Anal. Mach. Intell. 34(4), 743–761 (2012)

    Article  Google Scholar 

  8. Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 580–587 (2014)

  9. Girshick R.: Fast R-CNN. In: IEEE International Conference on Computer Vision (ICCV), pp. 1440–1448 (2015)

  10. Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 39(6), 1137–1149 (2017)

    Article  Google Scholar 

  11. Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: unified, real-time object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 779–788(2016)

  12. Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., Berg, A. C.: Ssd: single shot multibox detector. In: European Conference on Computer Vision (ECCV), pp. 21–37 (2016)

  13. Fu, C., Liu, W., Ranga, A., Tyagi, A., Berg, A. C.: DSSD: deconvolutional single shot detector. arXiv:1701.06659 (2017)

  14. Zhang, S., Wen, L., Bian, X., Lei, Z., Li, S.Z.: Single-shot refinement neural network for object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4203–4212 (2018)

  15. Redmon, J., Farhadi, A.: YOLO9000: better, faster, stronger. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6517–6525 (2017)

  16. Redmon, J., Farhadi, A.: YOLOv3: an incremental improvement. arXiv:1804.02767 (2018)

  17. Han, K., Wang, Y., Tian, Q., Guo, J., Xu, C., Xu, C.: GhostNet: more features from cheap operations. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)

  18. Fan, Q., Zhuo, W., Tang, C., Tai, Y.: Few-shot object detection with attention-RPN and multi-relation detector. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)

  19. Wang, X., Zhang, S., Yu, Z., Feng, L., Zhang, W.: Scale-equalizing pyramid convolution for object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:2005.03101 (2020)

  20. Bochkovskiy, A., Wang, C., Liao, H. M.: YOLOv4: optimal speed and accuracy of object detection. arXiv:2004.10934 (2020).

  21. Cai, Z., Fan, Q., Feris, R. S., Vasconcelos, N.: A unified multi-scale deep convolutional neural network for fast object detection. In: European Conference on Computer Vision (ECCV), pp. 354–370 (2016)

  22. Cai, Z., Vasconcelos, N.: Cascade r-cnn: delving into high quality object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6154–6162 (2018)

  23. Wang, X., Xiao, T., Jiang, Y., Shao, S., Sun, J., Shen, C.: Repulsion loss: detecting pedestrians in a crowd. In: 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7774–7783 (2018)

  24. Zhang, S., Wen, L., Bian, X., Lei, Z., Li, S. Z.: Occlusion-aware R-CNN: detecting pedestrians in a crowd. In: European Conference on Computer Vision (ECCV), pp. 637–653 (2018)

  25. Wang, Z., Wang, J., Yang, Y.: Resisting the distracting-factors in pedestrian detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:2005.07344 (2020)

  26. Chu, X., Zheng, A., Zhang, X., Sun, J.: Detection in crowded scenes: one proposal, multiple predictions. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:2003.09163 (2020)

  27. Jaderberg, M., Simonyan, K., Zisserman, A., Kavukcuoglu, K.: Spatial transformer networks. In: NIPS, pp. 2017–2025 (2015)

  28. Hu, J., Shen, L., Albanie, S., Sun, G., Wu, E.: Squeeze-and-excitation networks. In: IEEE Trans. Pattern Anal. Mach. Intell., p. 1 (2019)

  29. Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., Lu, H.: Dual attention network for scene segmentation. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3146–3154 (2019)

  30. Zhu, M., Jiao, L., Liu, F., Yang, S., Wang, J.: Residual spectral-spatial attention network for hyperspectral image classification. In: IEEE Trans. Geosci. Remote Sensing, pp. 1–14 (2020)

  31. Ji, R., Wen, L., Zhang, L., Du, D., Wu, Y., Zhao, C., Liu, X., Huang, F.: Attention convolutional binary neural tree for fine-grained visual categorization. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:1909.11378 (2020)

  32. Li, A., Qi, J., Lu, H.: Multi-attention guided feature fusion network for salient object detection. Neurocomputing 416–427 (2020)

  33. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778 (2016)

  34. Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.: Inception-v4, inception-ResNet and the impact of residual connections on learning. In: AAAI (2017)

  35. Liu, W., Liao, S., Hu, W., Liang, X., Chen, X.: Learning efficient single-stage pedestrian detectors by asymptotic localization fitting. In: 2018 European Conference on Computer Vision (ECCV), pp. 618–634 (2018)

  36. Lin, T. Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2980–2988 (2017)

  37. Lin, C.Y., Xie, H.X., Zheng, H.: PedJointNet: joint head-shoulder and full body deep network for pedestrian detection. IEEE Access 7, 47687–47697 (2019)

    Article  Google Scholar 

  38. Zhang, S., Yang, X., Liu, Y., Xu, C.: Asymmetric multi-stage CNNs for small-scale pedestrian detection. Neurocomputing 12–26 (2020)

  39. Zhang, Y., Yi, P., Zhou, D., Yang, X., Zhang, Q., Wei, P.: CSANet: channel and spatial mixed attention CNN for pedestrian detection. IEEE Access 8, 76243–76252 (2020)

    Article  Google Scholar 

  40. Song, T., Sun, L., Xie, D., Sun, H., Pu, S.: Small-scale pedestrian detection based on topological line localization and temporal feature aggregation. In: 2018 European Conference on Computer Vision (ECCV), pp. 536–551 (2018)

  41. Tian, Y., Luo, P., Wang, X., Tang, X.: Deep learning strong parts for pedestrian detection. In: 2015 IEEE international conference on computer vision, pp. 1904–1912 (2015)

  42. Li, Z., Chen, Z., Wu, Q.J., Liu, C.: Real-time pedestrian detection with deep supervision in the wild. SIViP 13(4), 761–769 (2019)

    Article  Google Scholar 

  43. Du, X., EI-Khamy, M., Morariu, V., Lee, J., Davis, L.: Fused deep neural networks for efficient pedestrian detection. arXiv:1805.08688 (2016)

  44. Saeidi, M., Ahmadi, A.: High-performance and deep pedestrian detection based on estimation of different parts. J Supercomput (2020). https://doi.org/10.1007/s11227-020-03345-4

    Article  Google Scholar 

Download references

Acknowledgements

This study is sponsored by the China Shandong Key R&D Plan (2018GGX106008), and is supported by the China Shandong Key Laboratory of Medical Physical Image Processing Technology.

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Chengjie Bai.

Additional information

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and permissions

Reprints and permissions

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Ma, J., Wan, H., Wang, J. et al. An improved one-stage pedestrian detection method based on multi-scale attention feature extraction. J Real-Time Image Proc 18, 1965–1978 (2021). https://doi.org/10.1007/s11554-021-01074-2

Download citation

  • Received:

  • Accepted:

  • Published:

  • Issue Date:

  • DOI: https://doi.org/10.1007/s11554-021-01074-2

Keywords

Navigation