Text classification using word sequences

Szczegóły
Abstrakt

Tytuł:: Text classification using word sequences
Autorzy:: Chudzian, P.
Data publikacji:: 2008
Słowa kluczowe:: text classification
text representation
generalized suffix tree
Język:: angielski
Dostawca treści:: BazTech
: Artykuł

The article discusses the use of word sequences in text classification. As opposed to ngrams, word sequences are not of a fixed length and therefore allow the classifier to obtain flexibility necessary to operate on documents collected from various sources. Presented classifier is built upon the suffix tree structure which enables word sequences to take part in classification process. During classification, both single words and longer sequences are taken into account and have impact on the category assignment with respect to their frequency and length. The Suffix Tree Classifier and well known Naive Bayes Classifier are compared and their properties are discussed. Obtained results show that incorporating word sequences into text classification can increase accuracy and reveal some interesting relations between maximal length of used sequences and classifier's error rate.

Informacja

Text classification using word sequences