Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Kim, Joo-Chang; Chung, Kyung-Yong

doi:10.1007/s11042-020-08774-0

Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Published: 04 March 2020

Volume 80, pages 34679–34695, (2021)
Cite this article

Multimedia Tools and Applications Aims and scope Submit manuscript

310 Accesses
8 Citations
Explore all metrics

Abstract

In this paper, a method for knowledge expansion of metadata using script mining analysis for multimedia recommendation systems is proposed. The method allows the extraction of new metadata and knowledge expansion through the mining analysis of multimedia scripts, which include a large amount of information. The scripts are collected by a Web crawler based on Python. From the collected scripts, hidden information is extracted through keyword analysis and sentiment analysis. In keyword analysis, scripts, unlike general documents, show a high frequency of names of characters or proper nouns. Such names or proper nouns are not frequently used in other media content, and therefore, their importance is high. Frequently, they are already offered in the conventional metadata, and consequently cause information duplication. Accordingly, term frequency–inverse document and metadata frequency (TF–IDMF), which considers the frequency of metadata in general term frequency–inverse document frequency (TF–IDF), is used. Thus, the importance of the names of characters or proper nouns in scripts can be decreased. Because the keywords for the extracted scripts are in fact included in the scripts, they can be used for precise multimedia search and recommendation. In sentiment analysis, the AFINN lexicon and the Bing lexicon are utilized to scan words in a script. The Bing lexicon is used to examine whether the words in the entire script are positive or negative. Then, the total numbers of positive words and negative words are used to calculate the representative sentiment of the script. The AFINN lexicon includes approximately 170 sentiment words, the negative or positive sentiment of which is presented in the range − 5 to +5. One script is divided into 100 sentences, and then, the representative sentiment in each sentence is evaluated as either positive or negative. Through script scanning, the flow of sentiment in multimedia streams can be discovered. The Bing lexicon categorizes words into positive, negative, and neutral sentiments. Through script scanning, the words included in each category can be quantified. Depending on the result of the script sentiment analysis, a different sentence embedding method based on inter-sentence similarity is used to cluster similar media. The results of the keyword analysis and sentiment analysis of a script are added to the metadata in a new column in a knowledge base to expand knowledge. To evaluate the significance of multimedia recommendations, keywords and sentiment information are used, and then, the similarity and clustering of the extracted media are assessed. As a result, script mining analysis based on the attributes that include actual information of media is considerably better than that based on types or a range of metadata attributes. Therefore, the proposed knowledge expansion method achieves significant results and shows an excellent performance in multimedia recommendation.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Combining semantic and linguistic representations for media recommendation

Article 15 July 2022

TMR: A Semantic Recommender System Using Topic Maps on the Items’ Descriptions

User specific context construction for personalized multimedia retrieval

Article 05 July 2017

References

AFINN sentiment analysis. https://github.com/fnielsen/afinn/. Accessed 2019.3
Aghdam MH (2019) Context-aware recommender systems using hierarchical hidden Markov model. Physica A: Statistical Mechanics and its Applications 518:89–98
Article Google Scholar
Baek JW, Kim JC, Chun J, Chung K (2019) Hybrid clustering based health decision-making for improving dietary habits. Technol Health Care 27(5):459–472
Article Google Scholar
Beigi G, Guo R, Nou A, Zhang Y, Liu H (2019) Protecting user privacy: an approach for untraceable web browsing history and unambiguous user profiles. In: Proceedings of the twelfth ACM international conference on web search and data mining ACM, pp. 213-221
Bobadilla J, Ortega F, Hernando A, Gutiérrez A (2013) Recommender systems survey. Knowl-Based Syst 46:109–132
Article Google Scholar
Christodoulou L, Abdul-Hameed O, Kondoz AM, Calic J (2016) Adaptive subframe allocation for next generation multimedia delivery over hybrid LTE unicast broadcast. IEEE Trans Broadcast 62(3):540–551
Article Google Scholar
Coyle K (2010) Metadata models of the world wide web. Libr Technol Rep 46(2):12–19
Google Scholar
Deldjoo Y, Schedl M, Hidasi B, Knees P (2018) Multimedia recommender systems. In: Proceedings of the 12th ACM conference on recommender systems ACM, pp 537-538
Gupta V, Lehal GS (2009) A survey of text mining techniques and applications. J Emerg Technol Web Intell 1(1):60–76
Google Scholar
Jung H, Chung K (2016) Knowledge-based dietary nutrition recommendation for obese management. Inf Technol Manag 17(1):29–42
Article Google Scholar
Kanungo T, Mount DM, Netanyahu NS, Piatko CD, Silverman R, Wu AY (2002) An efficient k-means clustering algorithm: analysis and implementation. IEEE Trans Pattern Anal Mach Intell 7:881–892
Article Google Scholar
Kim JC, Chung K (2017) Emerging risk forecast system using associative index mining analysis. Clust Comput 20(1):547–558
Article Google Scholar
Kim JC, Chung K (2018) Mining health-risk factors using PHR similarity in a hybrid P2P network. Peer-to-Peer Netw Appl 11(6):1278–1287
Article Google Scholar
Kim JC, Chung K (2019) Prediction model of user physical activity using data characteristics-based long short-term memory recurrent neural networks. KSII Transactions on Internet and Information Systems 13(4):2060–2077
Kim JC, Chung K (2019) Associative feature information extraction using text mining from health big data. Wirel Pers Commun 105(2):691–707
Article MathSciNet Google Scholar
Lu J, Wu D, Mao M, Wang W, Zhang G (2015) Recommender system application developments: a survey. Decis Support Syst 74:12–32
Article Google Scholar
Mwinyi IH, Narman HS, Fang KC, Yoo WS (2018) Predictive self-learning content recommendation system for multimedia contents. In: 2018 wireless telecommunications symposium (WTS), pp 1-6
Natural Language Toolkit (NLTK) https://www.nltk.org/. Accessed 2019.3
Netflix Prize data http://www.netflixprize.com/. Accessed 2019.3
Park HS, Jun CH (2009) A simple and fast algorithm for K-medoids clustering. Expert Syst Appl 36(2):3336–3341
Article Google Scholar
Park RC, Jung H, Chung K, Yoon KH (2015) Picocell based telemedicine health service for human UX/UI. Multimed Tools Appl 74(7):2519–2534
Article Google Scholar
Pouyanfar S, Yang Y, Chen SC, Shyu ML, Iyengar SS (2018) Multimedia big data analytics: a survey. ACM Computing Surveys (CSUR) 51(1):155–176
Article Google Scholar
Python library for searching entites in Bing. https://github.com/idin/bing/. Accessed 2019.3
Requests: HTTP for Humans. https://2.python-requests.org/en/master/. Accessed 2019.3.
Robertson S (2004) Understanding inverse document frequency: on theoretical arguments for IDF. J Doc 60(5):503–520
Article Google Scholar
Scikit-Learn https://scikit-learn.org/stable/. Accessed 2019.3
Screen-scraping library beautifulsoup4. https://www.crummy.com/software/BeautifulSoup/. Accessed 2019.3
The Internet Movie Script Database (IMSDb) https://www.imsdb.com/. Accessed 2019.2
The Movie Database (TMDb) https://www.themoviedb.org/. Accessed 2019.2
Thorat PB, Goudar RM, Barve S (2015) Survey on collaborative filtering, content-based filtering and hybrid recommendation system. Int J Comput Appl 110(4):31–36
Google Scholar
Toch E, Wang Y, Cranor LF (2012) Personalization and privacy: a survey of privacy risks and remedies in personalization-based systems. User Model User-Adap Inter 22(1–2):203–220
Article Google Scholar
Vijayarani S, Ilamathi MJ, Nithya M (2015) Preprocessing techniques for text mining-an overview. Int J Comput Sci Commun Netw 5(1):7–16
Google Scholar
Weiss SM, Indurkhya N, Zhang T (2015) Fundamentals of predictive text mining. Springer
Yoo H, Chung K (2018) Mining-based lifecare recommendation using peer-to-peer dataset and adaptive decision feedback. Peer-to-Peer Netw Appl 11(6):1309–1320
Article Google Scholar
Zhou L, Rodrigues JJ, Wang H, Martini M, Leung VC (2019) 5G multimedia communications: theory, technology, and application. IEEE MultiMedia 26(1):8–9
Article Google Scholar

Download references

Acknowledgements

This research was supported by the MSIT(Ministry of Science and ICT), Korea, under the ITRC(Information Technology Research Center) support program(IITP-2018-0-01405) supervised by the IITP(Institute for Information and Communications Technology Planning and Evaluation).

Author information

Authors and Affiliations

Data Mining Lab., Department of Computer Science, Kyonggi University, 154–42, Gwanggyosan-ro, Yeongtong-gu, Suwon-si, Gyeonggi-do, 16227, South Korea
Joo-Chang Kim
Division of Computer Science and Engineering, Kyonggi University, 154-42, Gwanggyosan-ro, Yeongtong-gu, Suwon-si, Gyeonggi-do, 16227, South Korea
Kyung-Yong Chung

Authors

Joo-Chang Kim
View author publications
You can also search for this author in PubMed Google Scholar
Kyung-Yong Chung
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Kyung-Yong Chung.

Additional information

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and permissions

Reprints and permissions

About this article

Cite this article

Kim, JC., Chung, KY. Knowledge expansion of metadata using script mining analysis in multimedia recommendation. Multimed Tools Appl 80, 34679–34695 (2021). https://doi.org/10.1007/s11042-020-08774-0

Download citation

Received: 19 July 2019
Revised: 01 December 2019
Accepted: 17 February 2020
Published: 04 March 2020
Issue Date: November 2021
DOI: https://doi.org/10.1007/s11042-020-08774-0

Keywords

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Abstract

Access this article

Similar content being viewed by others

Combining semantic and linguistic representations for media recommendation

TMR: A Semantic Recommender System Using Topic Maps on the Items’ Descriptions

User specific context construction for personalized multimedia retrieval

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Additional information

Publisher’s note

Rights and permissions

About this article

Cite this article

Keywords

Navigation

Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Abstract

Access this article

Similar content being viewed by others

Combining semantic and linguistic representations for media recommendation

TMR: A Semantic Recommender System Using Topic Maps on the Items’ Descriptions

User specific context construction for personalized multimedia retrieval

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Additional information

Publisher’s note

Rights and permissions

About this article

Cite this article

Share this article

Keywords

Search

Navigation