A Joint Semantic Vector Representation Model for Text Clustering and Classification

Momtazi, S.; Rahbar, A.; Salami, D.; Khanijazani, I.

doi:10.22044/jadm.2019.7400.1876

Document Type : Original/Review Paper

Authors

Computer Engineering and Information Technology Department, Amirkabir University of Technology, Tehran, Iran.

https://doi.org/10.22044/jadm.2019.7400.1876

Abstract

Text clustering and classification are two main tasks of text mining. Feature selection plays the key role in the quality of the clustering and classification results. Although word-based features such as term frequency-inverse document frequency (TF-IDF) vectors have been widely used in different applications, their shortcoming in capturing semantic concepts of text motivated researches to use semantic models for document vector representations. Latent Dirichlet allocation (LDA) topic modeling and doc2vec neural document embedding are two well-known techniques for this purpose.
In this paper, we first study the conceptual difference between the two models and show that they have different behavior and capture semantic features of texts from different perspectives. We then proposed a hybrid approach for document vector representation to benefit from the advantages of both models. The experimental results on 20newsgroup show the superiority of the proposed model compared to each of the baselines on both text clustering and classification tasks. We achieved 2.6% improvement in F-measure for text clustering and 2.1% improvement in F-measure in text classification compared to the best baseline model.

Keywords

Main Subjects

Document and Text Processing

Journal of AI and Data Mining

A Joint Semantic Vector Representation Model for Text Clustering and Classification

Volume 7, Issue 3
July 2019
Pages 443-450

A Joint Semantic Vector Representation Model for Text Clustering and Classification

Volume 7, Issue 3July 2019Pages 443-450

Volume 7, Issue 3
July 2019
Pages 443-450