Classification of Protein Structure with A Variety of Code Features on the Same Index

Closed

Marji, Dian Eka Ratnawati, Edy Santoso

2018 3rd International Conference on Sustainable Information Engineering and Technology, SIET 2018 - Proceedings Conference paper Cited by 0 Quartile

Abstract

There have been several studies that use protein structure data using various methods, the highest accuracy of some of these studies is a maximum of 80%. In general, there are three ways to change the character code data of a protein structure to be numeric, that is by calculating the number of each code, using the PAM matrix and by using the concept of probability for each code in the same index. In this study, the data and methods are the same as those done by Tawang with a maximum accuracy of 79.17% [1], namely the classification of protein structure data using the Naïve Bayes Classifier method. The difference with this research is the features used. For research conducted by Tawang, the feature used is the length of one protein structure (393 characters). In this study, the features are taken from the diversity of codes in the same index because the results of the observation of many protein structures have the same code in the same index and the average number of diverse features as many as 250 indexes. This study produces the best accuracy (more than 90%) using test data smaller or equal to 50. With training data greater than 200 of 753 data will produce accuracy that tends to be low (less than 60%), then changes in value prior probability which have the smallest value among other priors. The test results show that changes in the prior value can increase the accuracy by up to 5%. © 2018 IEEE.

Affiliations

Departement of Informatics, Computer Science Faculty, Brawijaya University, Malang, Indonesia