Baghas Rizaluddin, Yuita Arum Sari, Sigit Adinugroho
Food leftovers in hospitals are a critical issue, directly impacting the quality of patient nutrition abd the efficiency of the hospital operating budget. This study proposes an automated framework to estimate the percentage of food leftovers from patient meals using deep learning-based image analysis. The proposed method utilizes the combination of Mask R-CNN model for object pixel-wise segmentation and Depth Anything V2 for monocular depth estimation, with the calculation of leftovers based on before and after object volumes. The dataset used in this study, referred to as Le-Food dataset, consists of manually captured food images in a controlled environment to maintain uniform conditions. Experimental results show that deeper backbones improve accuracy but require a higher amount of computational process, ResNet-50 + Depth Anything V2 VIT-L combination achieved the lowest Mean Absolute Error (MAE) of 10.516% with an inference time of 6.275s per pair. However, a lighter model, ResNet-50 + Depth Anything V2 VIT-S, offered a more practical compromise, producing a competitive MAE of 10.987% at 4.878s, the fastest combinations among all. These findings demonstrate that an objective, automated approach to monitoring food waste is viable. The proposed framework contributes to the improvement of hospital nutritional services and supports a data-driven solution in dietary planning for hospital patients. In addition, this approach can be generalized to other healthcare or food services where volume estimation is essential. © 2025 IEEE.
Faculty of Computer Science, Brawijaya University, Malang, Indonesia