Graph-based Multimodal Ranking Models for Multimodal Summarization

Published: 26 May 2021 Publication History


Multimodal summarization aims to extract the most important information from the multimedia input. It is becoming increasingly popular due to the rapid growth of multimedia data in recent years. There are various researches focusing on different multimodal summarization tasks. However, the existing methods can only generate single-modal output or multimodal output. In addition, most of them need a lot of annotated samples for training, which makes it difficult to be generalized to other tasks or domains. Motivated by this, we propose a unified framework for multimodal summarization that can cover both single-modal output summarization and multimodal output summarization. In our framework, we consider three different scenarios and propose the respective unsupervised graph-based multimodal summarization models without the requirement of any manually annotated document-summary pairs for training: (1) generic multimodal ranking, (2) modal-dominated multimodal ranking, and (3) non-redundant text-image multimodal ranking. Furthermore, an image-text similarity estimation model is introduced to measure the semantic similarity between image and text. Experiments show that our proposed models outperform the single-modal summarization methods on both automatic and human evaluation metrics. Besides, our models can also improve the single-modal summarization with the guidance of the multimedia information. This study can be applied as the benchmark for further study on multimodal summarization task.


  Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention AlignmentACM Transactions on Asian and Low-Resource Language Information Processing10.1145/365198323:5(1-29)Online publication date: 10-May-2024
  Advancements in Multimodal Social Media Post Summarization: Integrating GPT-4 for Enhanced Understanding2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC)10.1109/COMPSAC61105.2024.00307(1934-1940)Online publication date: 2-Jul-2024
  MMSFT: Multilingual Multimodal Summarization by Fine-Tuning TransformersIEEE Access10.1109/ACCESS.2024.345438212(129673-129689)Online publication date: 2024
    Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention AlignmentACM Transactions on Asian and Low-Resource Language Information Processing10.1145/365198323:5(1-29)Online publication date: 10-May-2024
    Advancements in Multimodal Social Media Post Summarization: Integrating GPT-4 for Enhanced Understanding2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC)10.1109/COMPSAC61105.2024.00307(1934-1940)Online publication date: 2-Jul-2024
    MMSFT: Multilingual Multimodal Summarization by Fine-Tuning TransformersIEEE Access10.1109/ACCESS.2024.345438212(129673-129689)Online publication date: 2024
    Align vision-language semantics by multi-task learning for multi-modal summarizationNeural Computing and Applications10.1007/s00521-024-09908-336:25(15653-15666)Online publication date: 1-Sep-2024
    Towards Making the Most of Knowledge Across Languages for Multimodal Cross-Lingual SummarizationPattern Recognition and Computer Vision10.1007/978-981-97-8620-6_29(424-438)Online publication date: 18-Oct-2024
    Bidirectional Sentence Ordering with Interactive DecodingACM Transactions on Asian and Low-Resource Language Information Processing10.1145/356151022:2(1-15)Online publication date: 30-Mar-2023
    MCR: Multilayer cross‐fusion with reconstructor for multimodal abstractive summarisationIET Computer Vision10.1049/cvi2.12173Online publication date: 7-Feb-2023
    Crisis event summary generative model based on hierarchical multimodal fusionPattern Recognition10.1016/j.patcog.2023.109890144:COnline publication date: 1-Dec-2023
    Cross-modal knowledge guided model for abstractive summarizationComplex & Intelligent Systems10.1007/s40747-023-01170-910:1(577-594)Online publication date: 27-Jul-2023
    Scientific document processing: challenges for modern learning methodsInternational Journal on Digital Libraries10.1007/s00799-023-00352-724:4(283-309)Online publication date: 24-Mar-2023
