A New Automatic Multi-document Text Summarization using Topic Modeling

Roul, Rajendra Kumar; Mehrotra, Samarth; Pungaliya, Yash; Sahoo, Jajati Keshari

doi:10.1007/978-3-030-05366-6_17

Rajendra Kumar Roul¹⁶,
Samarth Mehrotra¹⁷,
Yash Pungaliya¹⁷ &
…
Jajati Keshari Sahoo¹⁸

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 11319))

Included in the following conference series:

International Conference on Distributed Computing and Internet Technology

1280 Accesses
10 Citations

Abstract

This paper proposes a novel methodology to generate an extractive text summary from a corpus of documents. Unlike most existing methods, our approach is designed in such a way that the final generated summary covers all the important topics from a corpus of documents. We propose a heuristic method which uses the Latent Dirichlet Allocation technique to identify the optimum number of independent topics present in the corpus. Some of the sentences are identified as the important sentences from each independent topic using a set of word and sentence level features. In order to ensure that the final summary is coherent, we suggest a novel technique to reorder the sentences based on sentence similarity. The use of topic modeling ensures that all the important content from the corpus of documents is captured in the extracted summary which in turn strengthen the summary. Experimental results show that the proposed approach is promising.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

1.
Since reduction is being performed, \(2/X<1\).
2.
Experimental results generate a good summary for X = 13.
3.
https://radimrehurek.com/gensim/.
4.
By independent means low similarity between topics.
5.
www.encyclopediaofmath.org/index.php?title=Hellinger_distance&oldid=16453.
6.
http://www.nltk.org/.
7.
http://www.duc.nist.gov.

References

Miller, G.A.: Wordnet: a lexical database for English. Commun. ACM 38(11), 39–41 (1995)
Article Google Scholar
Ganesan, K., Zhai, C., Han, J.: Opinosis: a graph-based approach to abstractive summarization of highly redundant opinions. In: Proceedings of the 23rd International Conference on Computational Linguistics, pp. 340–348. Association for Computational Linguistics (2010)
Google Scholar
Moratanch, N., Chitrakala, S.: A survey on extractive text summarization. In: 2017 International Conference on Computer, Communication and Signal Processing (ICCCSP), pp. 1–6. IEEE (2017)
Google Scholar
Fang, C., Mu, D., Deng, Z., Wu, Z.: Word-sentence co-ranking for automatic extractive text summarization. Expert Syst. Appl. 72, 189–195 (2017)
Article Google Scholar
Nallapati, R., Zhai, F., Zhou, B.: SummaRuNNer: a recurrent neural network based sequence model for extractive summarization of documents. In: AAAI, pp. 3075–3081 (2017)
Google Scholar
Roul, R.K., Sahoo, J.K., Goel, R.: Deep learning in the domain of multi-document text summarization. In: Shankar, B.U., Ghosh, K., Mandal, D.P., Ray, S.S., Zhang, D., Pal, S.K. (eds.) PReMI 2017. LNCS, vol. 10597, pp. 575–581. Springer, Cham (2017). https://doi.org/10.1007/978-3-319-69900-4_73
Chapter Google Scholar
Narayan, S., Cohen, S.B., Lapata, M.: Ranking sentences for extractive summarization with reinforcement learning. In: 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. US. ACL anthology, New Orleans (2018)
Google Scholar
Blei, D.M., Ng, A.Y., Jordan, M.I.: Latent dirichlet allocation. J. Mach. Learn. Res. 3(Jan), 993–1022 (2003)
MATH Google Scholar
Kullback, S., Leibler, R.A.: On information and sufficiency. Ann. Math. Stat. 22(1), 79–86 (1951)
Article MathSciNet Google Scholar
Fuglede, B., Topsoe, F.: Jensen-Shannon divergence and Hilbert space embedding. In: Proceedings, International Symposium on Information Theory. ISIT 2004, p. 31. IEEE (2004)
Google Scholar
Lin, C.-Y.: Rouge: a package for automatic evaluation of summaries. In: Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, vol. 8, pp. 74–81 (2004)
Google Scholar

Download references

Author information

Authors and Affiliations

Department of Computer Science, Thapar Institute of Engineering and Technology, Patiala, 147004, Punjab, India
Rajendra Kumar Roul
Department of Computer Science, BITS-Pilani, K. K. Birla Goa Campus, Zuarinagar, Pilani, 403726, Goa, India
Samarth Mehrotra & Yash Pungaliya
Department of Mathematics, BITS-Pilani, K. K. Birla Goa Campus, Zuarinagar, Pilani, 403726, Goa, India
Jajati Keshari Sahoo

Authors

Rajendra Kumar Roul
View author publications
You can also search for this author in PubMed Google Scholar
Samarth Mehrotra
View author publications
You can also search for this author in PubMed Google Scholar
Yash Pungaliya
View author publications
You can also search for this author in PubMed Google Scholar
Jajati Keshari Sahoo
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Rajendra Kumar Roul .

Editor information

Editors and Affiliations

University of Hagen, Hagen, Germany
Günter Fahrnberger
Coimbatore Institute of Technology, Coimbatore, India
Sapna Gopinathan
IBM Thomas J. Watson Research Center, Yorktown Heights, NY, USA
Laxmi Parida

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Roul, R.K., Mehrotra, S., Pungaliya, Y., Sahoo, J.K. (2019). A New Automatic Multi-document Text Summarization using Topic Modeling. In: Fahrnberger, G., Gopinathan, S., Parida, L. (eds) Distributed Computing and Internet Technology. ICDCIT 2019. Lecture Notes in Computer Science(), vol 11319. Springer, Cham. https://doi.org/10.1007/978-3-030-05366-6_17

Download citation

DOI: https://doi.org/10.1007/978-3-030-05366-6_17
Published: 11 December 2018
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-05365-9
Online ISBN: 978-3-030-05366-6
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics