Feature Selection using Variable Length Chromosome Genetic Algorithm for Sentiment Analysis

Closed

Tirana Noor Fatyanosa, Fitra A. Bachtiar, Mahendra Data

2018 3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018 - Proceedings Conference paper Cited by 11 Quartile

Abstract

Research in sentiment analysis often have very high features for the classification and it might affect the model's accuracy. In this paper, the Variable Length Chromosome Genetic Algorithm with Naïve Bayes (VLCGA-NB) is utilized to analyze the twitter sentiment. The tweets are preprocessed in several steps before using it in the algorithm. The preprocessing conducted to reduce the number of features. After the preprocessing performed, the features that produce a higher fitness value is selected. There are five classes: Very Positive, Positive, Neutral, Negative, and Very Negative to be classified. Comparison of NB and VLCGA-NB is conducted to show the models accuracy. The experiments show that VLCGA-NB produces the higher value for the Best F-Measure results of each sentiment, lower number of features (188 features), higher accuracy (75.2%), and higher fitness value (3.088953) than NB. © 2018 IEEE.

Affiliations

Faculty of Computer Science, Brawijaya University, Malang, Indonesia