An Integrated Deep Learning and Computer Vision Framework for Non- Destructive Assessment of Apple Nutritional Quality

Authors

  • Amisha Singh
  • Nishant Singh
  • Ankit Kumar
  • Uttam Sharma
  • Basant Kumar Das
  • Harsh P. Sharma
  • Anuj Yadav

DOI:

https://doi.org/10.63001/tbs.2026.v21.i02.S.I(2).pp21723-21756

Keywords:

Artificial intelligence; Deep learning;, Vision Transformer; EfficientNet;, Explainable artificial intelligence; Apple quality; Nutritional prediction.

Abstract

Rapid and non-destructive assessment of apple nutritional quality is essential for intelligent postharvest
management and automated grading. This study proposes an integrated framework combining RGB
computer vision, transfer learning, explainable artificial intelligence (XAI), and a decision support system
(DSS) for simultaneous prediction of multiple nutritional attributes, including soluble solids content (SSC),
moisture, firmness, vitamin C, total phenolic content (TPC), titratable acidity, total sugars, and dry matter.
A dataset comprising 720 Red Delicious apples represented by 4,800 multi-view RGB images was used to
develop and evaluate five deep learning architectures: CNN, ResNet50, DenseNet121, EfficientNet-B3, and
Vision Transformer (ViT). Among the evaluated models, ViT achieved the highest overall performance with
R² = 0.964, RMSE = 0.25, MAE = 0.18, MAPE = 2.10%, Pearson’s r = 0.982, and NSE = 0.960, while
attribute-specific prediction accuracies reached R² = 0.979 for SSC, 0.957 for moisture, 0.962 for firmness,
0.952 for vitamin C, and 0.957 for TPC. Bland–Altman analysis demonstrated excellent agreement with
laboratory measurements, with minimal mean bias (SSC: 0.03, moisture: −0.04, firmness: −0.21, and vitamin
C: 0.08). Grad-CAM and SHAP analyses confirmed that the models relied on physiologically meaningful
features, with peel colour and texture contributing most strongly to nutritional prediction. The proposed DSS
achieved an overall grading accuracy of 96.1% (692/720 samples). Although ViT provided the highest
prediction accuracy, EfficientNet-B3 achieved comparable performance (R² = 0.958) with substantially
lower computational requirements (48 MB model size; 21 ms/image; 48 FPS), making it more suitable for
real-time edge deployment. These findings demonstrate that the proposed framework provides an accurate,
rapid, and scalable solution for non-destructive nutritional assessment and automated apple quality grading.

Downloads

Published

2026-04-26 — Updated on 2026-08-08

Versions

How to Cite

Amisha Singh, Nishant Singh, Ankit Kumar, Uttam Sharma, Basant Kumar Das, Harsh P. Sharma, & Anuj Yadav. (2026). An Integrated Deep Learning and Computer Vision Framework for Non- Destructive Assessment of Apple Nutritional Quality. The Bioscan, 21(2), 21723–21756. https://doi.org/10.63001/tbs.2026.v21.i02.S.I(2).pp21723-21756 (Original work published April 26, 2026)