Lazuardy Syahrul Darfiansa, Fitriyani, Sza Sza Amulya Larasati
Background: A major challenge in education is the dominance of exam questions that primarily assess basic thinking skills, such as remembering and understanding. Revised Bloom’s Taxonomy (BT), which classifies cognitive skills into six levels, offers a framework to promote higher-order thinking through better-designed assessments. Deep learning-based systems have shown promising results in automatically classifying questions by BT levels, supporting educators in creating more meaningful exams. Objective: This research aims to develop a classification system that can effectively classify Indonesian exam questions based on BT using IndoBERT pretrained models. These models were combined with Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) classifiers (referred to as IndoBERT-CNN and IndoBERT-LSTM) to determine the model with the highest performance. Methods: The dataset utilized was self-collected and underwent several stages of preparation, including expert labeling and splitting. Furthermore, preprocessing was conducted to ensure the dataset was consistent and free from irrelevant features related to case folding, tokenization, stopword removal, and stemming. Hyperparameter fine-tuning was subsequently carried out on IndoBERT, IndoBERT-CNN, and IndoBERT-LSTM. Model performance was evaluated using Accuracy, F-Measure, Precision, and Recall. Results: The fine-tuned IndoBERT model results showed that IndoBERT-LSTM outperformed IndoBERT-CNN. The optimal hyperparameter configuration, batch size of 64 and learning rate of 5e-5, showed the highest performance, achieving Accuracy of 88.75%, Precision of 85%, Recall of 88%, and F-Measure of 86%. Conclusion: IndoBERT, IndoBERT-CNN, and IndoBERT-LSTM reflected promising results, although the performance of the models was significantly affected by respective architectures and hyperparameter settings. IndoBERT-LSTM achieved the highest accuracy with larger batch sizes, while IndoBERT and IndoBERT-CNN performed best under different configurations. However, IndoBERT faced limitations due to its language-specific focus and the limited interpretability of predictions against expert-labeled data. © 2025 The Authors. Published by Universitas Airlangga.
Telkom University, Bandung, Indonesia; Universitas Brawijaya, Malang, Indonesia