Multi-task Question Generation Based Data Augmentation for Biomedical Answer Generation

Zhao, Junting; Bai, Jun; Rong, Wenge; Ouyang, Yuanxin; Xiong, Zhang

doi:10.1007/978-981-99-4749-2_41

Junting Zhao¹³,
Jun Bai¹³,
Wenge Rong¹³,
Yuanxin Ouyang¹³ &
…
Zhang Xiong¹³

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 14088))

Included in the following conference series:

International Conference on Intelligent Computing

788 Accesses

Abstract

Limited by the corpus size and the annotation cost, biomedical question answering (BioQA) is a task of great research value. To generate professional biomedical answers, we first propose a text-to-text multi-task question generation model, which improves the accuracy of domain question generation with two auxiliary tasks. Based on this, a multi-task QA pipeline system with filtering is designed to synthesize high-quality biomedical data. Then, we use three data augmentation strategies to conduct generative BioQA experiments on original and synthetic data. The results on the factoid BioASQ 7b, 8b, and 9b datasets demonstrate the effectiveness of our method.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 99.00; Price excludes VAT (USA)

Softcover Book: USD 129.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

References

Alberti, C., Andor, D., Pitler, E., Devlin, J., Collins, M.: Synthetic QA corpora generation with roundtrip consistency. In: Proceedings of the 57th Conference of the Association for Computational Linguistics, pp. 6168–6173 (2019)
Google Scholar
Chen, W., Verga, P., de Jong, M., Wieting, J., Cohen, W.W.: Augmenting pre-trained language models with QA-memory for open-domain question answering. CoRR abs/2204.04581 (2022)
Google Scholar
Du, X., Shao, J., Cardie, C.: Learning to ask: neural question generation for reading comprehension. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, pp. 1342–1352 (2017)
Google Scholar
Feng, S.Y., et al.: A survey of data augmentation approaches for NLP. In: Findings of the Association for Computational Linguistics: ACL/IJCNLP, pp. 968–988 (2021)
Google Scholar
Fu, Y., Ou, W., Yu, Z., Lin, Y.: MIGA: a unified multi-task generation framework for conversational text-to-SQL. CoRR abs/2212.09278 (2022)
Google Scholar
Gu, Y., et al.: Domain-specific language model pretraining for biomedical natural language processing. ACM Trans. Comput. Healthc. 3(1), 2:1–2:23 (2022)
Google Scholar
Heilman, M., Smith, N.A.: Good question! Statistical ranking for question generation. In: Proceedings of the 2010 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 609–617 (2010)
Google Scholar
Jin, Q., et al.: Biomedical question answering: a survey of approaches and challenges. ACM Comput. Surv. 55(2), 35:1–35:36 (2023)
Google Scholar
Lewis, M., Fan, A.: Generative question answering: learning to answer the whole question. In: Proceedings of the 7th International Conference on Learning Representations (2019)
Google Scholar
Lewis, M., et al.: BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871–7880 (2020)
Google Scholar
Lewis, P.S.H., et al.: PAQ: 65 million probably-asked questions and what you can do with them. Trans. Assoc. Comput. Linguist. 9, 1098–1115 (2021)
Article Google Scholar
Lyu, C., Shang, L., Graham, Y., Foster, J., Jiang, X., Liu, Q.: Improving unsupervised question answering via summarization-informed question generation. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 4134–4148 (2021)
Google Scholar
Raffel, C., et al.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140:1–140:67 (2020)
Google Scholar
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P.: SQuAD: 100, 000+ questions for machine comprehension of text. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392 (2016)
Google Scholar
Tsatsaronis, G., et al.: An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition. BMC Bioinform. 16, 138:1–138:28 (2015)
Google Scholar
Wei, J.W., Zou, K.: EDA: easy data augmentation techniques for boosting performance on text classification tasks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pp. 6381–6387 (2019)
Google Scholar
Yang, W., Xie, Y., Tan, L., Xiong, K., Li, M., Lin, J.: Data augmentation for BERT fine-tuning in open-domain question answering. CoRR abs/1904.06652 (2019)
Google Scholar
Yoon, W., Lee, J., Kim, D., Jeong, M., Kang, J.: Pre-trained language model for biomedical question answering. In: Cellier, P., Driessens, K. (eds.) ECML PKDD 2019. CCIS, vol. 1168, pp. 727–740. Springer, Cham (2020). https://doi.org/10.1007/978-3-030-43887-6_64
Chapter Google Scholar

Download references

Acknowledgment

This work was partially supported by the National Natural Science Foundation of China (No. 61977002).

Author information

Authors and Affiliations

School of Computer Science and Engineering, Beihang University, Beijing, China
Junting Zhao, Jun Bai, Wenge Rong, Yuanxin Ouyang & Zhang Xiong

Authors

Junting Zhao
View author publications
You can also search for this author in PubMed Google Scholar
Jun Bai
View author publications
You can also search for this author in PubMed Google Scholar
Wenge Rong
View author publications
You can also search for this author in PubMed Google Scholar
Yuanxin Ouyang
View author publications
You can also search for this author in PubMed Google Scholar
Zhang Xiong
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Wenge Rong .

Editor information

Editors and Affiliations

Department of Computer Science, Eastern Institute of Technology, Zhejiang, China
De-Shuang Huang
University of Wollongong, North Wollongong, NSW, Australia
Prashan Premaratne
Zhengzhou University of Light Industry, Zhengzhou, China
Baohua Jin
Zhong Yuan University of Technology, Zhengzhou, China
Boyang Qu
University of Ulsan, Ulsan, Korea (Republic of)
Kang-Hyun Jo
Department of Computer Science, Liverpool John Moores University, Liverpool, UK
Abir Hussain

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Zhao, J., Bai, J., Rong, W., Ouyang, Y., Xiong, Z. (2023). Multi-task Question Generation Based Data Augmentation for Biomedical Answer Generation. In: Huang, DS., Premaratne, P., Jin, B., Qu, B., Jo, KH., Hussain, A. (eds) Advanced Intelligent Computing Technology and Applications. ICIC 2023. Lecture Notes in Computer Science, vol 14088. Springer, Singapore. https://doi.org/10.1007/978-981-99-4749-2_41

Download citation

DOI: https://doi.org/10.1007/978-981-99-4749-2_41
Published: 30 July 2023
Publisher Name: Springer, Singapore
Print ISBN: 978-981-99-4748-5
Online ISBN: 978-981-99-4749-2
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics