Using IR Techniques to Improve Automated Text Classification

Gonçalves, Teresa; Quaresma, Paulo

doi:10.1007/978-3-540-27779-8_34

Teresa Gonçalves¹⁸ &
Paulo Quaresma¹⁸

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 3136))

Included in the following conference series:

International Conference on Application of Natural Language to Information Systems

697 Accesses

Abstract

This paper performs a study on the pre-processing phase of the automated text classification problem. We use the linear Support Vector Machine paradigm applied to datasets written in the English and the European Portuguese languages – the Reuters and the Portuguese Attorney General’s Office datasets, respectively.

The study can be seen as a search, for the best document representation, in three different axes: the feature reduction (using linguistic information), the feature selection (using word frequencies) and the term weighting (using information retrieval measures).

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

$34.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Supervised Machine Learning Text Classification: A Review

Survey on supervised machine learning techniques for automatic text classification

Article 19 January 2019

The Impact of Pre-processing and Feature Selection on Text Classification

References

Apté, C., Damerau, F., Weiss, S.: Automated learning of decision rules for text categorization. ACM Transactions on Information Systems 12(3), 233–251 (1994)
Article Google Scholar
Gonçalves, T., Quaresma, P.: A preliminary approach to the multi-label classification problem of Portuguese juridical documents. In: Pires, F.M., Abreu, S.P. (eds.) EPIA 2003. LNCS (LNAI), vol. 2902, pp. 435–444. Springer, Heidelberg (2003)
Chapter Google Scholar
Gonçalves, T., Quaresma, P.: The impact of NLP techniques in the multilabel classification problem. In: Intelligent Information Systems 2004, Advances in Soft Computing, Zakopane, Poland, May 2004, Springer, Heidelberg (2004)
Google Scholar
Japkowicz, N.: The class imbalance problem: Significance and strategies. In: Proceedings of the 2000 International Conference on Artificial Intelligence (IC-AI 2000), vol. 1, pp. 111–117 (2000)
Google Scholar
Joachims, T.: Learning to Classify Text Using Support Vector Machines. Kluwer Academic Publishers, Dordrecht (2002)
Google Scholar
Nigam, K., McCallum, A., Thrun, S., Mitchell, T.: Text classification from labelled and unlabelled documents using EM. Machine Learning 39(2), 103–134 (2000)
Article MATH Google Scholar
Mladenić, D., Grobelnik, M.: Feature selection for unbalanced class distribution and Naïve Bayes. In: Proceedings of ICML 1999, 16th International Conference on Machine Learning, pp. 258–267 (1999)
Google Scholar
Quaresma, P., Rodrigues, I.: PGR: Portuguese attorney general’s office decisions on the web. In: Bartenstein, O., Geske, U., Hannebauer, M., Yoshie, O. (eds.) INAP 2001. LNCS (LNAI), vol. 2543, pp. 51–61. Springer, Heidelberg (2003)
Chapter Google Scholar
Salton, G., McGill, M.: Introduction to Modern Information Retrieval. McGraw-Hill, New York (1983)
MATH Google Scholar
Schölkopf, B., Smola, A.: Learning with Kernels. MIT Press, Cambridge (2002)
Google Scholar
Vapnik, V.: The nature of statistical learning theory. Springer, Heidelberg (1995)
MATH Google Scholar
Witten, I., Frank, E.: Data Mining: Practical machine learning tools with Java implementations. Morgan Kaufmann, San Francisco (1999)
Google Scholar

Download references

Author information

Authors and Affiliations

Departamento de Informática, Universidade de Évora, 7000, Évora, Portugal
Teresa Gonçalves & Paulo Quaresma

Authors

Teresa Gonçalves
View author publications
You can also search for this author in PubMed Google Scholar
Paulo Quaresma
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

School of Computing, Science and Engineering Newton Building, University of Salford, M5 4WT, Greater Manchester, UK
Farid Meziane
Lab. CEDRIC, CNAM, Paris, France
Elisabeth Métais

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Gonçalves, T., Quaresma, P. (2004). Using IR Techniques to Improve Automated Text Classification. In: Meziane, F., Métais, E. (eds) Natural Language Processing and Information Systems. NLDB 2004. Lecture Notes in Computer Science, vol 3136. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-27779-8_34

Download citation

DOI: https://doi.org/10.1007/978-3-540-27779-8_34
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-22564-5
Online ISBN: 978-3-540-27779-8
eBook Packages: Springer Book Archive

Publish with us

Policies and ethics

Using IR Techniques to Improve Automated Text Classification

Abstract

Access this chapter

Subscribe and save

Buy Now

Preview

Similar content being viewed by others

Supervised Machine Learning Text Classification: A Review

Survey on supervised machine learning techniques for automatic text classification

The Impact of Pre-processing and Feature Selection on Text Classification

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

Using IR Techniques to Improve Automated Text Classification

Abstract

Access this chapter

Subscribe and save

Buy Now

Preview

Similar content being viewed by others

Supervised Machine Learning Text Classification: A Review

Survey on supervised machine learning techniques for automatic text classification

The Impact of Pre-processing and Feature Selection on Text Classification

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation