Muhammad Ariz Putra Rianto, Novanto Yudistira, Agus Wahyu Widodo
This study aims to classify facial expressions of engagement using facial landmarks with a combined model between Graph Convolutional Network (GCN) and Convolutional Neural Network (CNN). With this research, researchers hope that this research can be useful for the world of education. In this research, the dataset used comes from the DAiSEE dataset about facial expressions in online learning. The dataset is a video to be extracted into each frame. All of the features in the frame were extracted through 68 landmark points and 52 types of action units. In this research, a comparison was made between GCN architecture and CNN for landmark extraction first to compare the results. This research also uses the LDAM loss function in the combined model to overcome imbalanced datasets. The final results of the model comparison between GCN and CNN using facial landmarks resulted in an accuracy of 50.56% and 49.61%, respectively. The combined GCN and CNN model produces an accuracy of 50.89%. Furthermore, the combined GCN and CNN model using the LDAM loss function produces an accuracy of 52.24%. Although the combined model showed modest performance, it runs significantly lower than temporal models like LSTM and TCN reported in similar studies. Future research could enhance classification accuracy by exploring temporal models more robustly suited to sequence data, such as LSTM or TCN, and by incorporating multimodal inputs to better capture complex engagement expressions in educational contexts. © 2025, Politeknik Negeri Padang. All rights reserved.
Informatics Engineering Department, Brawijaya University, Lowokwaru, Malang, Indonesia