Efficient Text Classification Using Tree-structured Multi-linear Principle Component Analysis

01/20/2018
by   Yuanhang Su, et al.
0

A novel text data dimension reduction technique, called the tree-structured multi-linear principle component anal- ysis (TMPCA), is proposed in this work. Being different from traditional text dimension reduction methods that deal with the word-level representation, the TMPCA technique reduces the dimension of input sequences and sentences to simplify the following text classification tasks. It is shown mathematically and experimentally that the TMPCA tool demands much lower complexity (and, hence, less computing power) than the ordinary principle component analysis (PCA). Furthermore, it is demon- strated by experimental results that the support vector machine (SVM) method applied to the TMPCA-processed data achieves commensurable or better performance than the state-of-the-art recurrent neural network (RNN) approach.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset