[ PoTD ]: Paper of The Day: Using Kullback-Leibler Distance for Text Categorization
Abstract
A system that performs text categorization aims to assign appropriate categories from a predefined classification scheme to incoming documents. These assignments might be used for varied purposes such as filtering, or retrieval. This paper introduces a new effective model for text categorization with great corpus (more or less 1 million documents). Text categorization is performed using the Kullback-Leibler distance between the probability distribution of the document to classify and the probability distribution of each category. Using the same representation of categories, experiments show a significant improvement when the above mentioned method is used. KLD method achieve substantial improvements over the tfidf performing method.
#Keywords
#Information_Retrieval #Text_Categorization #Category_Learning #Query_Expansion #Training_Corpus
> https://link.springer.com/chapter/10.1007%2F3-540-36618-0_22
Abstract
A system that performs text categorization aims to assign appropriate categories from a predefined classification scheme to incoming documents. These assignments might be used for varied purposes such as filtering, or retrieval. This paper introduces a new effective model for text categorization with great corpus (more or less 1 million documents). Text categorization is performed using the Kullback-Leibler distance between the probability distribution of the document to classify and the probability distribution of each category. Using the same representation of categories, experiments show a significant improvement when the above mentioned method is used. KLD method achieve substantial improvements over the tfidf performing method.
#Keywords
#Information_Retrieval #Text_Categorization #Category_Learning #Query_Expansion #Training_Corpus
> https://link.springer.com/chapter/10.1007%2F3-540-36618-0_22