Social-CVAE: Pedestrian Trajectory Prediction Using Conditional Variational Auto-Encoder

Xu, Baowen; Wang, Xuelei; Li, Shuo; Li, Jingwei; Liu, Chengbao

doi:10.1007/978-981-99-8132-8_36

Part of the book series: Communications in Computer and Information Science ((CCIS,volume 1962))

Included in the following conference series:

International Conference on Neural Information Processing

427 Accesses

Abstract

Pedestrian trajectory prediction is a fundamental task in applications such as autonomous driving, robot navigation, and advanced video surveillance. Since human motion behavior is inherently unpredictable, resembling a process of decision-making and intrinsic motivation, it naturally exhibits multimodality and uncertainty. Therefore, predicting multi-modal future trajectories in a reasonable manner poses challenges. The goal of multi-modal pedestrian trajectory prediction is to forecast multiple socially plausible future motion paths based on the historical motion paths of agents. In this paper, we propose a multi-modal pedestrian trajectory prediction method based on conditional variational auto-encoder. Specifically, the core of the proposed model is a conditional variational auto-encoder architecture that learns the distribution of future trajectories of agents by leveraging random latent variables conditioned on observed past trajectories. The encoder models the channel and temporal dimensions of historical agent trajectories sequentially, incorporating channel attention and self-attention to dynamically extract spatio-temporal features of observed past trajectories. The decoder is bidirectional, first estimating the future trajectory endpoints of the agents and then using the estimated trajectory endpoints as the starting position for the backward decoder to predict future trajectories from both directions, reducing cumulative errors over longer prediction ranges. The proposed model is evaluated on the widely used ETH/UCY pedestrian trajectory prediction benchmark and achieves state-of-the-art performance.

This work is supported by National Key Research and Development Program of China [Grant 2022YFB3305401] and the National Nature Science Foundation of China [Grant 62003344].

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Gupta, A., Johnson, J., Fei-Fei, L., Savarese, S., Alahi, A.: Social GAN: socially acceptable trajectories with generative adversarial networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2255–2264 (2018)
Google Scholar
Xu, P., Hayet, J.-B., Karamouzas, I.: SocialVAE: human trajectory prediction using timewise latents. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.) Computer Vision-ECCV, 17th European Conference, Tel Aviv, Israel, Proceedings, Part IV, vol. 13664, pp. 511–528. Springer, Heidelberg (2022). https://doi.org/10.1007/978-3-031-19772-7_30
Gao, J., Shi, X., Yu, J.J.: Social-dualcvae: multimodal trajectory forecasting based on social interactions pattern aware and dual conditional variational auto-encoder. arXiv preprint arXiv:2202.03954 (2022)
Yao, Y., Atkins, E., Johnson-Roberson, M., Vasudevan, R., Du, X.: BiTraP: bi-directional pedestrian trajectory prediction with multi-modal goal estimation. IEEE Rob. Autom. Lett. 6(2), 1463–1470 (2021)
Article Google Scholar
Liang, R., Li, Y., Zhou, J., Li, X.: Stglow: a flow-based generative framework with dual graphormer for pedestrian trajectory prediction. arXiv preprint arXiv:2211.11220 (2022)
Gu, T., Chen, G., Li, J., Lin, C., Rao, Y., Zhou, J., Lu, J.: Stochastic trajectory prediction via motion indeterminacy diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17113–17122 (2022)
Google Scholar
Helbing, D., Molnar, P.: Social force model for pedestrian dynamics. Phys. Rev. E 51(5), 4282 (1995)
Article Google Scholar
Alahi, A., Goel, K., Ramanathan, V., Robicquet, A., Fei-Fei, L., Savarese, S.: Social LSTM: human trajectory prediction in crowded spaces. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 961–971 (2016)
Google Scholar
Zhang, P., Yang, W., Zhang, P., Xue, J., Zheng, N.: SR-LSTM: state refinement for LSTM towards pedestrian trajectory prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12085–12094 (2019)
Google Scholar
Kosaraju, V., Sadeghian, A., Martín-Martín, R., Reid, I., Rezatofighi, H., Savarese, S.: Social-BIGAT: multimodal trajectory forecasting using bicycle-gan and graph attention networks. Adv. Neural Inf. Process. Syst. 32, 1–10 (2019)
Google Scholar
Yue, J., Manocha, D., Wang, H.: Human trajectory prediction via neural social physics. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.) Computer Vision-ECCV,: 17th European Conference, Tel Aviv, Israel, 23–27 October 2022, Proceedings, Part XXXIV, vol. 13694, pp. 376–394. Springer, Heidelberg (2022). https://doi.org/10.1007/978-3-031-19830-4_22
Sadeghian, A., Kosaraju, V., Sadeghian, A., Hirose, N., Rezatofighi, H., Savarese, S.: Sophie: an attentive gan for predicting paths compliant to social and physical constraints. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1349–1358 (2019)
Google Scholar
Amirian, J., Hayet, J.-B., Pettré, J.: Social ways: learning multi-modal distributions of pedestrian trajectories with gans. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (2019)
Google Scholar
Mangalam, K., et al.: It is not the journey but the destination: endpoint conditioned trajectory prediction. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M. (eds.) ECCV 2020. LNCS, vol. 12347, pp. 759–776. Springer, Cham (2020). https://doi.org/10.1007/978-3-030-58536-5_45
Chapter Google Scholar
Salzmann, T., Ivanovic, B., Chakravarty, P., Pavone, M.: Trajectron++: dynamically-feasible trajectory forecasting with heterogeneous data. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M. (eds.) ECCV 2020. LNCS, vol. 12363, pp. 683–700. Springer, Cham (2020). https://doi.org/10.1007/978-3-030-58523-5_40
Chapter Google Scholar
Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7132–7141 (2018)
Google Scholar
Peng, B., et al.: RWKV: reinventing RNNs for the transformer era. arXiv preprint arXiv:2305.13048 (2023)
Pellegrini, S., Ess, A., Van Gool, L.: Improving data association by joint modeling of pedestrian trajectories and groupings. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010. LNCS, vol. 6311, pp. 452–465. Springer, Heidelberg (2010). https://doi.org/10.1007/978-3-642-15549-9_33
Chapter Google Scholar
Leal-Taixé, L., Fenzi, M., Kuznetsova, A., Rosenhahn, B., Savarese, S.: Learning an image-based motion context for multiple people tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3542–3549 (2014)
Google Scholar
Mohamed, A., Qian, K., Elhoseiny, M., Claudel, C.: Social-STGCNN: a social spatio-temporal graph convolutional neural network for human trajectory prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14 424–14 432 (2020)
Google Scholar
Shi, L., et al.: SGCN: sparse graph convolution network for pedestrian trajectory prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8994–9003 (2021)
Google Scholar

Download references

Author information

Authors and Affiliations

Institute of Automation, Chinese Academy of Sciences, Beijing, 100190, China
Baowen Xu, Xuelei Wang, Shuo Li, Jingwei Li & Chengbao Liu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, 100049, China
Baowen Xu, Shuo Li & Jingwei Li

Authors

Baowen Xu
View author publications
You can also search for this author in PubMed Google Scholar
Xuelei Wang
View author publications
You can also search for this author in PubMed Google Scholar
Shuo Li
View author publications
You can also search for this author in PubMed Google Scholar
Jingwei Li
View author publications
You can also search for this author in PubMed Google Scholar
Chengbao Liu
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Xuelei Wang .

Editor information

Editors and Affiliations

School of Automation, Central South University, Changsha, China
Biao Luo
Institute of Automation, Chinese Academy of Sciences, Beijing, China
Long Cheng
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Zheng-Guang Wu
School of Automation, Guangdong University of Technology, Guangzhou, China
Hongyi Li
School of Electrical Engineering and Telecommunications, UNSW Sydney, Sydney, NSW, Australia
Chaojie Li

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Xu, B., Wang, X., Li, S., Li, J., Liu, C. (2024). Social-CVAE: Pedestrian Trajectory Prediction Using Conditional Variational Auto-Encoder. In: Luo, B., Cheng, L., Wu, ZG., Li, H., Li, C. (eds) Neural Information Processing. ICONIP 2023. Communications in Computer and Information Science, vol 1962. Springer, Singapore. https://doi.org/10.1007/978-981-99-8132-8_36

Download citation

DOI: https://doi.org/10.1007/978-981-99-8132-8_36
Published: 26 November 2023
Publisher Name: Springer, Singapore
Print ISBN: 978-981-99-8131-1
Online ISBN: 978-981-99-8132-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics