ABSTRACT
We propose a novel cost-efficient approach to threshold selection for binary web-page classification problems with imbalanced class distributions. In many binary-classification tasks the distribution of classes is highly skewed. In such problems, using uniform random sampling in constructing sample sets for threshold setting requires large sample sizes in order to include a statistically sufficient number of examples of the minority class. On the other hand, manually labeling examples is expensive and budgetary considerations require that the size of sample sets be limited. These conflicting requirements make threshold selection a challenging problem. Our method of sample-set construction is a novel approach based on stratified sampling, in which manually labeled examples are expanded to reflect the true class distribution of the web-page population. Our experimental results show that using false positive rate as the criterion for threshold setting results in lower-variance threshold estimates than using other widely used accuracy measures such as F1 and precision.
- X. He, L. Duan, Y. Zhou and B. Dom, Threshold selection for web-page classification with highly skewed class distribution, Yahoo! Labs Research Report YL-2009-001, 2009Google Scholar
- Y. Yang, A Study on Thresholding Strategies for Text Categorization, Proceedings of SIGIR-01, 24th ACM International Conference on Research and Development in Information Retrieval, 2001 Google ScholarDigital Library
Index Terms
- Threshold selection for web-page classification with highly skewed class distribution
Recommendations
Fast balanced sampling for highly stratified population
Balanced sampling is a very efficient sampling design when the variable of interest is correlated to the auxiliary variables on which the sample is balanced. A procedure to select balanced samples in a stratified population has previously been proposed. ...
A skewed truncated t distribution
Skewed symmetric distributions have attracted a great deal of attention in the last few years. One of them, the skewed t distribution suffers from limited applicability because of the lack of finite moments. This note proposes an alternative to the ...
Classifier Ensemble for Imbalanced Data Stream Classification
CUBE '12: Proceedings of the CUBE International Information Technology ConferenceThe data streams in various real life applications are characterized by concept drift. Such data streams may also be characterized by skewed or imbalance class distributions for example Financial fraud detection, Network intrusion detection etc. In such ...
Comments