Classification method comparison on Indonesian social media sentiment analysis

Closed

Tirana Noor Fatyanosa, Fitra A. Bachtiar

2017 Proceedings - 2017 International Conference on Sustainable Information Engineering and Technology, SIET 2017 Vol. 2018-January Conference paper Cited by 15 Quartile

Abstract

Sentiment analysis from social media has turned out to be essential since individuals are normally genuine with their sentiment on giving their perspective. However, in turning social media into a sentiment analysis possess challenges such as comments are usually ambiguous, language barrier problem, slang words, redundant comment, and sentiment classification. This study attempted to distinguish the issues of sentiment classification from Indonesian social media on Jakarta governor election. Several steps are taken to overcome those problems that include preprocessing. The preprocessing strategy used are removing the unrelated tweet, removing URL, deleting duplicate lines, deleting similar lines, removing the unrelated word, removing hashtag, removing Twitter username, removing number in the comment, removing punctuation, checking slang words, and converting the slang word into appropriate word. The preprocessed sentiment is then classified into positive, negative, and neutral. The classification method used in this study are Summation method, Average on Tweet, Average on Tweet with the threshold on objective score, Weighted Average, and Naïve Bayes method. The experimental results show that Naïve Bayes produce the highest precision, highest recall, and highest accuracy for neutral and positive sentiment. But, Naïve Bayes does not produce good results for negative sentiment. © 2017 IEEE.

Affiliations

Faculty of Computer Science, Brawijaya University, Malang, Indonesia