Unsupervised Feature Generation using Knowledge Repositories for Effective Text Categorization

Prasath, Rajendra; Sarkar, Sudeshna

doi:10.3233/978-1-60750-606-5-1101

Abstract

We propose an unsupervised feature generation algorithm using the repositories of human knowledge for effective text categorization. Conventional bag of words (BOW) depends on the presence / absence of keywords to classify the documents. To understand the actual context behind these keywords, we use knowledge concepts / hyperlinks from external knowledge sources through content and structure mining on Wikipedia. Then, the features of knowledge concepts are clustered to generate knowledge cluster vectors with which the input text documents are mapped into a high dimensional feature space and the classification is performed. The simulation results show that the proposed approach identifies associated features in the text collection and yields an improved classification accuracy.

This website uses cookies

This website uses cookies