Submitted:
15 September 2026
Posted:
16 September 2026
You are already at the latest version
Abstract
Semantic segmentation of food images is an essential step for the automatic analysis of meal composition. This study compares two segmentation architectures derived from U-Net and ResNet34 on the FoodSeg103 dataset, progressively evaluating the contribution of skip connections and fine-tuning. During training, U-Net improves significantly, with accuracy increasing from 44.70% to 87.70% (+43 percentage points) over 50 epochs, while ResNet34 achieves high accuracy (97.27%) from the first epoch, which then improves only marginally. For the U-Net architecture, the F1 score improves in three stages: 57.33% for the simple model, 85.11% after the addition of skip connections, and then 94.40% after fine-tuning, representing a cumulative gain of 37.1 percentage points. For ResNet34, fine-tuning yields a significantly more modest improvement, with the F1 score increasing from 62.53% to only 63.05%. This could be explained by ResNet34’s already high baseline performance, leaving less room for improvement than for U-Net. A Grad-CAM analysis is applied to ResNet34, whose particularly rapid learning from the early epochs makes it a relevant case for visually examining the regions on which the model relies for its predictions, thus complementing the quantitative evaluation with a qualitative interpretation. This experimental framework thus provides an evidence-based comparison between U-Net and ResNet34 for food image segmentation, and highlights the differentiated contribution of fine-tuning depending on the architecture considered.
Keywords:
food semantic segmentation
; U-Net
; ResNet
; skip connections
; fine-tuning
; backpropagation
; cross-entropy loss
; deep learning
; Grad-CAM
; interpretability
1. Introduction
General Context
Global population growth, evolving consumption patterns, and increasing food losses are major challenges to the sustainability of food systems. According to the U.S. Environmental Protection Agency (EPA), food waste management must follow a hierarchy of value that prioritizes strategies with the greatest environmental and economic benefits. The EPA’s *Food Recovery Hierarchy* classifies the different ways of managing food surpluses, from preventing waste, considered the most sustainable approach, to landfill disposal, which is the least favorable solution due to its environmental impacts (EPA, 2023). Between these two extremes, redistribution for human consumption, food processing, animal feed, composting, and energy recovery are different strategies for preserving the value of food resources and promoting the principles of the circular economy (Papargyropoulou et al., 2014) [1].
However, the effective application of this hierarchy requires precise characterization of the food products involved. Indeed, before determining the best way to utilize surplus food, it is necessary to automatically identify its composition, condition, and the various components present. In real-world environments, a single food scene can simultaneously contain several food categories, such as rice, meat, vegetables, or various side dishes, making their automatic analysis particularly complex. Thus, beyond simple overall meal recognition, a nuanced understanding of the visual structure of food becomes essential for developing intelligent systems capable of assisting in the sorting, utilization, and sustainable management of food resources (Gao et al., 2026) [2].
From this perspective, computer vision and artificial intelligence technologies offer promising solutions for the automatic analysis of food scenes. Unlike traditional methods based on human inspection, approaches based on deep learning enable rapid, non-destructive analysis that can be adapted to real-world conditions. Convolutional neural networks (CNNs) have enabled major advances in automatic food recognition thanks to their ability to hierarchically extract visual features (LeCun et al., 2015) [3]. These advances have led to the development of numerous applications in the food sector, including dish recognition, nutritional analysis, and automatic monitoring of eating habits (Zhou et al., 2019) [4]. However, these approaches remain primarily focused on the overall classification of images and do not always allow for the precise identification of the different food components present in a complex scene.
Early approaches to food recognition relied primarily on image classification methods, aiming to assign an overall category to a given image. Among the most widely used databases, Food-101, proposed by Bossard et al. (2014) [5], is a major reference, providing over 100,000 images categorized into 101 food categories. However, this dataset is mainly used for classification purposes, while Food-Seg 103, introduced by Wu et al. in 2021 [6], is a dataset used for semantic segmentation of food. Published work using these datasets has led to significant advances in automatic food recognition. However, classification methods remain limited when an image contains multiple foods simultaneously, as they cannot accurately identify the location and shape of the different food components.
With the emergence of deep neural networks, convolutional neural networks (CNNs) have significantly improved the performance of food image analysis systems. Zhou et al. [4], presented a comprehensive study on the use of deep learning for automatic food understanding, including the recognition, classification, and analysis of visual features. However, a nuanced understanding of culinary scenes requires detailed spatial information, which has led to the development of semantic segmentation approaches.
Evolution of Food Recognition Towards Semantic Segmentation
Early research in the field of food vision focused primarily on image classification, the goal of which is to assign an overall category to a complete image. Among the most widely used databases, Food-101, introduced by Bossard et al. (2014) [5], is a major reference in the field of food recognition, with 101 annotated food categories. This database has enabled the evaluation of the performance of numerous deep learning models for automatic food recognition and has contributed to the rise of approaches based on convolutional neural networks. However, these global classification methods remain limited when an image simultaneously contains several food components, as they provide no information on the spatial location or proportion of each ingredient present in the scene. With recent advances in deep learning, convolutional neural networks (CNNs) have significantly improved the performance of visual food understanding systems. LeCun et al. (2015) [3], demonstrated the effectiveness of deep learning architectures for the automatic extraction of complex visual features, while Zhou et al. (2019) [4], presented a detailed analysis of deep learning applications in the food domain, including the automatic recognition, classification, and interpretation of food images. Despite these advances, classification remains insufficient for applications requiring a detailed understanding of a meal’s composition, especially when multiple foods appear in the same image.
Semantic segmentation thus represents a natural evolution that overcomes the limitations of classification by associating a category with each pixel of an image. Unlike classical approaches that only provide a global prediction, pixel-by-pixel segmentation allows for the precise identification of the position, shape, and extent of each food component. This capability is particularly important for complex food scenes, in which multiple ingredients may be mixed or partially overlapping.
In this context, several deep segmentation architectures have been proposed to improve object accuracy and localization. Ronneberger et al. (2015) [7], introduced U-Net, an encoder-decoder architecture widely used in biomedical segmentation and subsequently adapted to several computer vision domains thanks to its hop connections that preserve fine spatial information. Chen et al. (2018) [8], proposed DeepLab, which leverages atrous convolutions to improve multiscale context capture and the segmentation of objects of varying sizes. More recently, Xie et al. (2021) [9], introduced SegFormer, a visual transformer-based architecture that efficiently models global relationships in complex scenes while maintaining a lightweight architecture.
Food Segmentation and Specialized Databases
In the food industry, semantic segmentation is a crucial step for automatically analyzing meal composition and identifying the different ingredients present in a scene. Unlike traditional food recognition applications, food segmentation systems must handle significant visual variability related to differences in food presentation, lighting, texture, and shape. This complexity necessitates databases containing detailed, pixel-level annotations.
To address this need, Wu et al. (2021) [6], proposed FoodSeg103, a specialized database for semantic food segmentation comprising 103 food categories with pixel-by-pixel annotations. Unlike traditional classification databases such as Food-101, FoodSeg103 allows for more nuanced analysis of scenes containing multiple ingredients simultaneously. This characteristic makes it a suitable benchmark for evaluating food segmentation models under conditions closer to real-world applications.
Several recent studies have explored the use of artificial intelligence for the intelligent recognition, classification, and management of food waste. Mazloumian et al. [10] investigated deep learning-based approaches for automatic food waste classification using images acquired from smart waste bins, demonstrating the potential of computer vision techniques for food waste monitoring and characterization.
Rahman et al. [11] explored image-based waste analysis through the development of a dedicated waste image dataset and deep learning-based classification approaches, highlighting the importance of automated visual recognition for efficient waste segregation.
More recent studies, notably by Rahman et al. [11] and other data-driven approaches, have further demonstrated the increasing interest in artificial intelligence techniques for automated food waste segmentation, classification, sorting, and sustainable management processes.
Current Limitations and Motivation for the Study
Despite significant progress in food semantic segmentation, several challenges remain. First, the limited availability of pixel-level annotated images represents a major constraint for training deep learning models. Second, the high intra-class variability of foods, differences in presentation, and real-world acquisition conditions make model generalization difficult. Finally, modern deep learning architectures are often still considered difficult to interpret, limiting their adoption in applications requiring an understanding of the network’s decisions.
In this context, the explainability of deep learning models becomes essential. Interpretation methods such as Grad-CAM, proposed by Selvaraju et al. (2017), [12], allow visualization of the image regions that contributed to the model’s decision. This analysis provides a better understanding of network behavior and allows us to assess whether the extracted features truly correspond to relevant food regions.
Finally, to improve the transparency of the developed models, an interpretability analysis based on Grad-CAM is applied to the best-performing model. This analysis allows visualization of the food regions activated by the network and verification that the model’s decisions are based on relevant visual features. Thus, this work provides a comprehensive evaluation combining segmentation performance, architectural analysis, and explainability of deep models for automatic understanding of food images.
2. Materials and Methods
This research offers a comparative evaluation of two deep learning architectures for the semantic segmentation of food images. The experimental framework compares a standard U-Net model with modified versions incorporating skip connections and fine-tuning, as well as a ResNet34-based segmentation model analyzed with and without fine-tuning. The primary objective is to conduct an ablation study measuring the impact of architectural modifications and transfer learning on segmentation performance, while keeping the training and evaluation pipeline consistent across all models. Model evaluation relies on metrics such as pixel accuracy, Dice/F1 score, mean intersection over union (mIoU), precision, recall, and loss. Following quantitative analysis, the best-performing model is selected for qualitative evaluation using Grad-CAM, which provides visual insights into the regions influencing the model’s predictions, as illustrated in Figure 1.
2.1. Data Preparation
The FoodSeg103 dataset [6], consists of a training set containing 4,983 RGB food images and a validation set including 2,135 images. Each sample includes the original image, its corresponding pixel-level semantic segmentation mask, the food categories present in the image, and a unique identifier.
The segmentation masks provide pixel-wise annotations for 103 food categories, enabling supervised training and evaluation of semantic segmentation models. Owing to the large number of classes, complex meal compositions, intra-class variability, and frequent occlusions between food items, FoodSeg103 constitutes a challenging benchmark for semantic food segmentation. Before training, the FoodSeg103 images and their corresponding annotations were preprocessed according to the requirements of the segmentation and multilabel classification tasks.
2.1.1. U-Net-Based Segmentation Models
The basic U-Net, U-Net optimized with skip connections (U-Net Skip), and the fine-tuned U-Net (U-Net FT) used the same preprocessing procedure. All RGB images were resized to a fixed resolution of pixels and normalized to the interval according to:
where denotes the original RGB image and the normalized image provided as input to the network. The corresponding segmentation masks were resized using nearest neighbor interpolation to preserve the original class indices and avoid any modification of the pixel-level annotations. Thus, the same image and mask preprocessing procedure was consistently applied to the three U-Net-based models.
2.1.2. ResNet34-Based Semantic Segmentation
ResNet34 was employed as an encoder-based semantic segmentation model for pixel-level food image segmentation. The ResNet34-based segmentation model predicts a semantic class for each pixel of the input image, allowing the food regions to be explicitly delineated. The RGB input images were resized to pixels and normalized using the ImageNet channel-wise mean and standard deviation :
The corresponding segmentation masks were resized to the same spatial resolution using nearest-neighbor interpolation in order to preserve the original class indices and avoid the creation of invalid intermediate labels. Each pixel of the ground-truth mask therefore retains its original semantic class. Let denote the ground-truth segmentation mask, where C represents the number of semantic classes. The ResNet34-based segmentation model produces a pixel-wise prediction
where I is the input image and denotes the ResNet34-based encoder–decoder network. For each pixel , the predicted semantic class is obtained from the maximum class probability:
The ResNet34 encoder extracts hierarchical visual features from the input image, while the segmentation decoder progressively restores the spatial resolution to generate a dense pixel-level prediction. The model is trained using the corresponding ground-truth segmentation masks. Two configurations were evaluated in the comparative study: a ResNet34-based segmentation model without fine-tuning and the same architecture with fine-tuning (ResNet34 + FT). This comparison was designed to assess the contribution of fine-tuning to semantic segmentation performance while keeping the training and evaluation pipeline consistent across both configurations. The resulting segmentation predictions were evaluated using the underluying performance metrics. The best-performing model was subsequently selected for qualitative analysis and Grad-CAM-based visual interpretation of the learned feature responses.
2.2. Deep Learning Architectures
2.2.1. Construction of the UNet Model
SimpleUNet Baseline
The proposed SimpleUNet model is a lightweight convolutional architecture designed for pixel-wise semantic segmentation of FoodSeg103 images. The network takes an RGB image tensor of dimension , corresponding to a batch of eight images, and preserves the spatial resolution throughout the network. It produces a dense segmentation logit tensor of dimension . Here, 104 corresponds to the 103 food categories and the background class. The architecture consists of three convolutional layers with ReLU activation functions. The first convolutional layer transforms the three RGB channels into 64 feature maps using kernels, with 1,792 trainable parameters. The second convolutional layer increases the feature depth from 64 to 128 channels using kernels and contains 73,856 trainable parameters. Finally, a convolutional layer performs pixel-wise semantic classification by projecting the 128 feature maps at each spatial location onto the 104 semantic classes. This final layer contains 13,416 trainable parameters.
The resulting output is not an image-level classification vector but a spatially dense prediction map. For each pixel location , the network produces 104 class logits, corresponding to the 103 food categories and the background class. The predicted semantic label for each pixel is obtained as
where denotes the logit associated with class c at pixel location .
The SimpleUNet model, Table 1, contains 89,064 trainable parameters and no non-trainable parameters. It was used as a lightweight baseline for semantic segmentation to investigate the contribution of subsequent architectural improvements, particularly the introduction of skip connections and fine-tuning.
Forward Propagation
the successive convolutional layers of the SimpleUNet model. The RGB image, with a spatial resolution of pixels and three input channels, is processed by a sequence of convolutional layers that extract increasingly rich feature representations while preserving the spatial resolution.
The first convolutional layer transforms the three input channels into 64 feature maps, followed by a ReLU activation. The second convolution increases the feature representation to 128 channels while maintaining the same spatial resolution. Finally, a convolution projects the 128-dimensional feature representation at each pixel onto the 104 semantic classes. Consequently, the network produces a dense pixel-wise segmentation output rather than an image-level classification vector.
The forward propagation process of SimpleUNet is summarized in Table 2, showing the evolution of tensor dimensions from the input image to the final segmentation output.
2.2.2. Optimisation of U-Net with Skip Connections
A U-Net architecture with skip connections was implemented for semantic segmentation on the FoodSeg103 dataset. The model was designed with an encoder–decoder structure in which feature maps extracted by the encoder are transferred to the corresponding decoder layers through skip connections, [7]. The encoder consists of convolutional blocks with progressively increasing numbers of filters, from 64 to 128 and 256 filters, followed by a bottleneck with 512 filters. The decoder progressively reconstructs the spatial resolution using transposed convolution layers with 512, 256, 128, and 64 filters. The final layer is a convolution producing 104 output channels corresponding to the segmentation classes. Figure 2 shows the configurations of the two segmentation architectures. The SimpleUNet configuration is presented on the left, while the U-Net architecture with skip connections is shown on the right. The figure highlights the main differences in the layer organization and the feature connections between the two architectures.
Mathematical Formulation of U-Net and Skip Connections
The U-Net architecture proposed by Ronneberger et al. relies on a contracting path for extracting contextual features and an expansive path for reconstructing spatial resolution, enabling precise localization of segmented regions [7]. In an encoder block, a convolution followed by a ReLU activation is expressed as:
where is the feature map from the previous layer, the convolution filters, the bias, and [29]. Each encoder block uses two successive convolutions before downsampling [7], performed via max-pooling:
where is the pooling window (typically ) and c the channel.
The skip connections, central to U-Net, directly transmit high-resolution encoder features to the corresponding decoder levels, preserving fine spatial details lost during downsampling. With the encoder features and the upsampled decoder features, their fusion is given by:
combining high-level semantic information with fine spatial detail [7]. In the decoder, resolution is increased via transposed convolution:
after which features are refined by convolution:
The final layer projects features onto K semantic classes via a convolution followed by Softmax:
where is the logit for pixel i and class k. The predicted class is then:
Training relies on backpropagation, adjusting network parameters according to the error between predicted and reference segmentation. The pixel-wise cross-entropy loss is used:
where N is the total number of pixels, K the number of classes, and the ground-truth value for pixel i and class k. Gradients are propagated from the output layer to preceding layers via the chain rule, and parameters are updated by gradient descent:
where is the learning rate and t the iteration index. This process is repeated across batches to progressively reduce the segmentation error. Among the mechanisms formalized above, skip connections play a particularly important role, as they preserve and reintroduce fine spatial information into the decoder that would otherwise be lost during successive downsampling operations [7].
2.2.3. Optimization of the U-Net Model with Residual Connections by Fine-Tuning
The U-Net model used in this study is based on an encoder-decoder architecture incorporating jump connections to preserve the spatial information necessary for accurate segmentation [7]. To improve the network’s learning capacity, residual connections were also integrated, following the principle of residual blocks introduced in ResNet, to facilitate gradient propagation and the extraction of more discriminating features [16]. During fine-tuning, the encoder layers were initially frozen, while the decoder and task-specific layers were optimized on the target dataset. The deeper encoder layers were subsequently unfrozen to allow further adaptation, Figure 3.
Mathematical Foundations of Fine-Tuning for U-Net with Skip Connections
In a second step, a fine-tuning is applied to adapt the parameters learned by the model to the specific characteristics of the FoodSeg103 dataset. Unlike training performed from randomly initialized weights, fine-tuning starts from parameters already learned and continues their optimization on the target segmentation task. The loss function used to guide this adaptation can be expressed as the pixel-to-pixel cross-entropy loss [13], as:
where N represents the number of pixels considered, C the number of classes, the actual value of pixel i for class c, and the corresponding predicted probability. The network parameters are then adjusted from the pre-trained weights according to:
where denotes the learning rate used during fine-tuning, generally chosen to be lower than that used during initial training, in order to preserve relevant representations already acquired while allowing their adaptation to the target domain [14]. In the case of U-Net with skip connections, this adaptation jointly concerns the encoder, the decoder, and the parameters associated with the skip connections. The features extracted by the encoder can thus be specialized for food images, while the decoder progressively refines the reconstruction of segmented regions. Fine-tuning therefore makes it possible to exploit previously learned representations while improving their fit to the specific distribution of the segmentation data considered.
2.2.4. Construction of the ResNet34 Model Before Fine-Tuning
The ResNet34 model used for multi-label classification in this study is initially based on an architecture pre-trained on ImageNet. It comprises a succession of four groups of residual blocks, composed respectively of 3, 4, 6, and 3 BasicBlocks, with an increasing number of filters from 64 to 512. After the residual blocks, an Adaptive Average Pooling operation produces a representation of dimension 512, which is then flattened and passed to a fully connected layer transforming the 512 features into 103 outputs corresponding to the FoodSeg103 categories. A sigmoid function is used to obtain the probabilities associated with the different food categories. These same encoder weights, learned during multi-label classification training, are subsequently reused as the encoder of the segmentation model during the testing phase and fine-tuned on the target segmentation task during the training phase Figure 4.
Table 3.
Architecture of the ResNet34 model trained on the FoodSeg103 dataset for multi-label classification.
Table 3.
Architecture of the ResNet34 model trained on the FoodSeg103 dataset for multi-label classification.
| Block | Output Size | Description |
|---|---|---|
| Input Image | RGB image from the FoodSeg103 dataset | |
| Initial Convolution | convolution, stride 2 | |
| Max Pooling | Spatial downsampling | |
| Layer 1 | 3 residual blocks (BasicBlocks) | |
| Layer 2 | 4 residual blocks (BasicBlocks) | |
| Layer 3 | 6 residual blocks (BasicBlocks) | |
| Layer 4 | 3 residual blocks (BasicBlocks) | |
| Adaptive Average Pooling | Global feature extraction | |
| Flatten | 512 | Feature vector |
| Fully Connected Layer | 103 | Multi-label classification |
| Sigmoid Activation | 103 probabilities | Presence probability for each food category |
Mathematical Foundations for ResNet34 as an Encoder and Backpropagation
ResNet34 is used as the encoder to extract hierarchical representations from food images. ResNet introduces residual learning, where blocks learn a residual function relative to their input, easing the optimization of deep networks and improving gradient propagation [16].
For an input , a residual block outputs:
where is the learned residual transformation and is passed directly via the shortcut connection [16]. When dimensions differ, a linear projection ensures compatibility:
with the projection parameters [16].
Parameters are optimized via the segmentation loss , updated at iteration t as:
where is the learning rate, following the backpropagation principle of Rumelhart et al. [15].
2.2.5. Construction of the ResNet34 Model After Fine-Tuning
After fine-tuning, the overall architecture of the ResNet34 model remains unchanged, with the same residual blocks, the same representation of 512 features, and the same classification layer . The main modification concerns the configuration of trainable parameters: the convolutional backbone is frozen to preserve the learned visual representations, while only the classification layer is made trainable. The model thus retains its total of 21.3 million parameters, of which 52,839 parameters are trainable after the backbone is frozen.
Mathematical Formulation for Fine-Tuning of ResNet34
After initial training, fine-tuning is applied to adapt the pre-trained ResNet34 representations to the specific characteristics of the FoodSeg103 dataset, continuing optimization from previously learned parameters rather than a random initialization. Selected network parameters are again adjusted by backpropagation according to:
where is the fine-tuning learning rate and the loss computed on the target data. A lower learning rate than that of initial training is typically used to allow gradual adaptation while preserving previously learned visual characteristics.
For ResNet34 as an encoder, fine-tuning enables the gradual adaptation of representations across residual blocks, from low-level features such as contours and textures to higher-level semantic structures, which is particularly relevant when the target data distribution differs from the pre-training domain. The experimental analysis thus distinguishes ResNet34 before and after fine-tuning, allowing a direct evaluation of its contribution to segmentation performance.
2.3. Ablation Study
An ablation study was conducted to separately analyze the contribution of the main architectural components and the adaptation strategies used in the studied models. [20] For the U-Net architecture, three configurations were considered. The first corresponds to the reference U-Net, without hop connections and without fine-tuning. The second corresponds to the U-Net incorporating hop connections, but without fine-tuning. The third corresponds to the U-Net with hop connections and after fine-tuning.[7]. This organization allows us to separately evaluate the contribution of hop connections by comparing the first two configurations, and then the effect of fine-tuning by comparing the last two.
For ResNet34, two configurations were considered, corresponding respectively to the initial model and the model after fine-tuning. Comparing these two configurations allows us to evaluate the effect of adapting the model to the task of classifying food categories. In the fine-tuned configuration, the convolutional backbone is frozen while the classification layer is updated. All configurations are evaluated with the same experimental protocol and metrics to ensure a consistent comparison between the different variants.
2.4. Grad-CAM for Model Interpretability
To analyze the interpretability of the ResNet34 model’s predictions,
the Grad-CAM (Gradient-weighted Class Activation Mapping) method was used. Grad-CAM produces activation maps that locate image regions that contribute positively to a given prediction, by leveraging the gradients of the target class relative to the feature maps of a convolutional layer, [12]. This approach allows us to visualize the image areas primarily considered by the model during classification.
In this study, Grad-CAM is applied to the ResNet34 model to generate heat maps superimposed on the input images. These visualizations are used to provide a qualitative interpretation of the model’s decisions and to examine whether the highlighted regions correspond to the relevant visual characteristics of food categories.
2.4.1. Software and Computational Environment
The models were implemented in Python using the PyTorch framework. This framework was used for building the architectures, managing tensors, training, fine-tuning, and evaluating the models. The various image preprocessing and transformation operations were performed using libraries from the Python and PyTorch ecosystem.
The experiments were run on GPU-accelerated computing resources. Two NVIDIA Tesla T4 GPUs were used specifically for the different experimental phases. Additional GPU resources were also used when computational needs required it. The GPU environment was leveraged via CUDA to accelerate the computational operations necessary for deep learning.
For the analysis of the interpretability of the ResNet34 model, Grad-CAM (Gradient-weighted Class Activation Mapping) was integrated into the experimental environment to generate activation maps from the model’s gradients. These maps are used in the qualitative analysis of the predictions presented in the interpretability section.
2.5. Evaluation Metrics
2.5.1. Training Metrics
During the training phase, the loss function (Loss) and accuracy (Accuracy) were used to track the model’s learning progress. The loss function allows the evaluation of the gap between the model’s predictions and the reference annotations and to track the convergence of the optimization. Accuracy allows the tracking of the evolution of prediction performance over the different training epochs.
2.5.2. Evaluation Metrics on the Test Set
After training, the final models were evaluated on the independent test set. For segmentation models, performance was evaluated using pixel accuracy, mean intersection over union (mIoU), and Dice score. These metrics assess the proportion of correctly classified pixels, the overlap between predicted regions and reference annotations, and the similarity between the two segmentation maps. The loss function (Loss) and the accuracy were also calculated on the test data in order to evaluate the final performance of the models on data not used during parameter optimization.
2.6. Data Pipeline Orchestration for ResNet34 Fine-Tuning
For better understanding of the proposed fine-tuning procedure, the overall pipeline based on ResNet34 is illustrated in Figure 5. Three-stage experimental pipeline for ResNet34-based semantic segmentation, comprising (1) image–mask preprocessing and batch generation, (2) model training, fine-tuning, and best-model selection based on validation performance, and (3) evaluation of the selected model on the unseen test set through pixel-wise segmentation and quantitative performance metrics. To avoid replication, the same pipeline stages are applied to the U-Net with skip connections architecture, including its fine-tuning process, with the corresponding model-specific components adapted accordingly.
3. Results
This section presents the experimental results obtained with the improved U-Net model applied to the semantic segmentation of food categories in the FoodSeg103 dataset. The model’s performance is analyzed through learning evolution and quantitative evaluation across the test set.
3.1. Ablation Study: UNet Improvements
An ablation study was conducted to evaluate the contribution of each modification introduced into the baseline U-Net architecture, based on the comparison of training loss and accuracy over 50 epochs across configurations. Three configurations were compared: the standard U-Net, the Skip U-Net incorporating enhanced skip connections, and the fine-tuned Skip U-Net, the training results for the loss and the accuracy metrics through 50 epochs are reported in Figure 6.
The final performance of the optimized model was evaluated on the test set using commonly used semantic segmentation metrics, including accuracy, precision, mean intersection over union (mIoU), Dice score, and F1 score, Table 4. These metrics quantify the degree of overlap between the predicted masks and the reference annotations, [17].
The baseline U-Net achieved a pixel accuracy of 62.47%, with a mean IoU of 0.4116 and a Dice score of 0.4894. These results show that the conventional encoder-decoder architecture can learn semantic representations; however, the progressive downsampling operations may lead to the loss of fine spatial details, which affects the accuracy of object boundaries.
After introducing skip connections, the Skip U-Net significantly improved the segmentation performance, reaching 85.90% pixel accuracy, 0.7150 mean IoU, and 0.7899 Dice score. Compared with the baseline model, this corresponds to an increase of 23.43 percentage points in pixel accuracy, while the mean IoU and Dice score improved by 0.3034 and 0.3005, respectively. This improvement confirms the effectiveness of skip connections in recovering high-resolution features and preserving spatial details during the decoding process [7].
Following the fine-tuning stage, the optimized Skip U-Net achieved the best performance, with 94.45% pixel accuracy, 0.8626 mean IoU, and 0.9135 Dice score. Compared with the model before fine-tuning, the optimized model improved pixel accuracy by 8.55 percentage points, mean IoU by 0.1476, and Dice score by 0.1236. This improvement demonstrates that fine-tuning enables the adaptation of previously learned parameters to the target domain, [14], improving generalization and reducing segmentation errors.
The final results demonstrate that both architectural modification through skip connections and optimization through fine-tuning contribute to improving the segmentation quality. The evaluation using IoU and Dice metrics follows common practices in semantic segmentation evaluation, where overlap-based metrics are used to measure the agreement between predicted masks and ground-truth annotations [17]. The different scores are reported in Table 4.
3.1.1. Ablation Study for Adapting the ResNet34 Model to the FoodSeg103 Domain
To evaluate the impact of fine-tuning on the performance of the ResNet34 model, an ablation study was conducted on the FoodSeg103 dataset, comparing the base model with a fine-tuned version. The reported results in Figure 7, show how fine-tuning affects the training dynamics of the architecture. Without fine-tuning, the model exhibit a progressive improvement as the number of epochs increases toward 50, neither the tendency of the training loss nor that of the accuracy shows a noticeable slope. During this training phase, it can be observed that the fine-tuned Skip U-Net achieves training performance relatively close to that of ResNet34, despite the substantial difference in their initial training trajectories, suggesting that fine-tuning allows U-Net to progressively narrow the gap with the already high-performing ResNet34 model, Figure 8.
In the initial configuration, ResNet34 achieves an overall accuracy of 97.27%, with an F1 score of 62.53%, a mean IoU of 0.4549, and a Dice Score of 0.6253. After adjusting the classifier parameters on the FoodSeg103 data, the fine-tuned model showed a slight improvement in performance, with an Accuracy of 97.28%, a Recall increasing from 51.40% to 52.27%, an F1 score reaching 63.05%, and an increase in the Mean IoU and Dice Score to 0.4604 and 0.6305, respectively. These results demonstrate that fine-tuning allows for better adaptation of the model to the specific characteristics of the food domain, notably an increased ability to detect existing classes, while maintaining the robustness of the representations extracted by the ResNet34 backbone. The different scores are reported in Table 5.
3.2. Qualitative Visualization of ResNet34 Predictions
To complement the quantitative evaluation, qualitative visualizations were generated to examine the behavior of the ResNet34-based segmentation model before and after fine-tuning. Figure 9 presents ten representative food images together with their corresponding segmentation masks, obtained before fine-tuning, providing a visual representation of the regions identified by the model. Overall, the different food ingredients are well distinguished by their respective masks. However, some fine details are occasionally missed, such as the whipped cream inside a roll, which does not appear in the predicted mask. Conversely, in other cases, a partially filled glass is entirely colored by the model, as if it corresponded to a solid piece of meat. Such inconsistencies motivate the use of Grad-CAM to better understand the underlying reasoning behind these predictions.
After fine-tuning, Grad-CAM was additionally applied to visualize the spatial regions contributing most strongly to the model predictions. The resulting heatmaps highlight the areas receiving higher activation from the network, while their superposition onto the original images facilitates the visual interpretation of these activations in the context of the food images, Figure 10. The comparison shows how the selected model distributes its discriminative responses across the image and provides a complementary qualitative view of the model behavior beyond the quantitative segmentation metrics. The Grad-CAM analysis shows a color gradation from yellow to red over the food content of the plate, while blue tones appear outside the plate, indicating that the model correctly attends to food-relevant areas in the vast majority of cases.
4. Discussion
4.1. Impact of Architectural Improvements: Ablation Study
4.2. UNet-Model Architecture Analysis vs UNet with Skip Connections
To evaluate the impact of the various architectural improvements, an ablation study was conducted by comparing the reference convolutional model with the improved U-Net architectures.
Although initially named SimpleUNet, the first model used in this study does not implement the original U-Net encoder-decoder topology proposed by Ronneberger et al. (2015), [7]. Indeed, it contains neither spatial reduction operations by pooling, nor oversampling layers, nor jump connections between the encoding and decoding levels. Rather, it corresponds to a lightweight fully convolutional segmentation network used as a baseline model. In this study, it is therefore considered a Simple Fully Convolutional Network (Simple-FCN). Comparative analysis with U-Net architectures highlights the importance of skip connections in semantic segmentation tasks. Unlike a simple convolutional architecture, where detailed spatial information can be progressively lost during feature extraction, U-Net allows the high-resolution representations from the first layers of the encoder to be transferred directly to the corresponding layers of the decoder.
These connections facilitate the merging of the deep contextual information extracted by the encoder with the fine spatial details necessary for accurate contour reconstruction. This property is particularly important for segmenting food images, where objects may exhibit irregular shapes, similar textures, and low-contrast boundaries.
The results obtained after 50 training epochs confirm that the introduction of an encoder-decoder architecture with hop connections progressively improves segmentation performance compared to the Simple-FCN model. The observed improvement in segmentation metrics (Accuracy, Precision, Recall, Dice Score, and IoU) demonstrates the effectiveness of multi-scale feature fusion in preserving spatial information during the reconstruction process. This ablation study thus confirms that the performance gains do not stem solely from increased network depth, but primarily from the ability of U-Net architectures to efficiently combine global semantic features and the local details necessary for accurate pixel-by-pixel segmentation.
4.3. Optimization of the U-Net Model with Residual Connections by Fine-Tuning
For this U-Net model used, residual connections were also integrated, following the principle of residual blocks introduced in ResNet, to facilitate gradient propagation and the extraction of more discriminating features [16].
After an initial training phase, a fine-tuning step was applied using the already learned weights. This strategy allows the network parameters to be progressively adjusted to the specific characteristics of the food domain, without restarting the training from a random initialization. During this phase, the model continues to optimize the already extracted representations to improve the quality of the segmented masks. The gradual decrease in the loss function and the improvement in segmentation performance demonstrate that fine-tuning allows for a more precise adaptation of the network to the visual variations present in FoodSeg103.
4.4. Discussion Results Through Metrics
The improvements observed primarily concern better localization of regions belonging to different food categories, confirmed by the increase in the mIoU score, as well as a reduction in pixel-by-pixel classification errors, confirmed by the increase in the accuracy score, calculated from the pixel-by-pixel annotations available in FoodSeg103.
Since Dice and mIoU (mean intersection over union) scores are, according to recent literature, relatively insensitive to errors at the boundaries of segmented masks, particularly for large objects, [18], the improvement in the accuracy of segmented object contours cannot be validly inferred from the overall increase in these scores. A rigorous quantification of this improvement would require a dedicated boundary metric, such as Boundary IoU, [18], which was not calculated in this study.
Finally, the improvement attributed to better use of spatial information thanks to residual connections, [16], is based on the results of the ablation study: the configuration integrating these connections obtains higher scores than the equivalent configuration without residual connections, which allows us to attribute this improvement to the architectural mechanism considered, the latter constituting the only difference between the two configurations compared.
4.5. Discussion Resnet vs Unet
The results obtained show that segmentation performance depends not only on the depth or complexity of the network, but also on its ability to preserve and utilize fine spatial information. In our experiment on FoodSeg103, the U-Net model with skip connections and fine-tuning achieves a higher mIoU than that obtained with the ResNet34 encoder. This difference can be explained by the essential role of skip connections, which allow features extracted at shallow encoder levels to be transferred directly to the decoder. This information includes details about the contours, shapes, and precise location of food regions, which are particularly important for pixel-by-pixel segmentation. Conversely, although ResNet34 provides deeper and more discriminating semantic representations thanks to its residual structure, successive convolution and downsampling operations can lead to a loss of fine spatial information. Thus, a deeper architecture does not necessarily guarantee better segmentation. Fine-tuning also plays an important role, as it allows the learned representations to be adapted to the specific characteristics of FoodSeg103. The improvement in mIoU observed after fine-tuning therefore indicates that adapting the parameters to the domain data is crucial for improving segmentation accuracy. These results suggest that, for this type of data, preserving spatial detail and adapting the model to the domain may be more decisive than simply increasing the backbone depth. However, these results should be considered in light of the experimental protocol, particularly the fine-tuning parameters, the learning rate, the number of epochs, the pre-training weights, and the decoder architecture, to ensure a fair comparison between the different architectures.
Table 6.
Overview of representative food image analysis studies, including their tasks, methodologies, and contributions.
Table 6.
Overview of representative food image analysis studies, including their tasks, methodologies, and contributions.
![]() |
4.6. Discussion on the Trade-Off Between Performance, Complexity, and Interpretability
4.6.1. Analysis of the Computational Complexity of the Simple-FCN Model
The architectural simplicity of the SimpleUNet model, considered in this study as a lightweight, fully convolutional network (Simple-FCN), significantly reduces the number of parameters and the hardware resource requirements. The estimated computational cost is 46.70 GMAC, with an approximate memory consumption of 1.25 GB during training and inference.
This low complexity represents a significant advantage for applications requiring fast execution or deployment on resource-constrained platforms. However, this reduced complexity also limits the model’s ability to learn complex representations, particularly in food scenes containing multiple categories with similar textures and significant spatial interactions.
Thus, the Simple-FCN model provides a relevant baseline for evaluating the contribution of more advanced architectures incorporating multi-scale fusion mechanisms, learning transfer, or pre-trained encoders.
4.6.2. Impact of Architectural Improvements on the Quality of Segmentation: Ablation Study
U-Net architectures rely on an efficient encoder-decoder structure with jump connections that combine deep semantic features with fine spatial information. This strategy, initially proposed by Ronneberger et al. [7], Pixel-level segmentation has become a fundamental task in computer vision, enabling precise localization and delineation of objects within images [27].
In this study, several U-Net-based variants were evaluated, including classic U-Net, U-Net with optimized jump connections, and U-Net with fine tuning. Jump connections allow the preservation of spatial details lost during encoding by directly transferring high-resolution features to the decoding layers.
This property is particularly important for food images, where edge accuracy and separation between visually similar regions directly influence segmentation quality.
The ablation study results thus allow us to identify the actual contribution of each architectural improvement and confirm the value of hierarchical feature blending strategies.
4.6.3. Comparison Between CNN Classification and Segmentation Architectures
The experimental results highlight the complementary roles of convolutional neural network architectures for food image analysis. ResNet34 and U-Net are both based on convolutional operations; however, they address different learning objectives.
ResNet34 is primarily designed for image classification tasks. Through its deep residual blocks, it learns hierarchical visual representations by progressively extracting low- and high-level features from input images. The residual learning strategy introduced by He et al. enables efficient training of deep networks by improving gradient propagation and feature reuse [16].
In contrast, U-Net is specifically designed for pixel-level semantic segmentation. Its encoder-decoder structure combined with skip connections allows the network to preserve spatial information while recovering detailed object boundaries during the decoding process [7].
Compared with classification-oriented CNN models, U-Net provides a more precise localization capability because it generates pixel-wise predictions rather than a single class label. This characteristic is particularly important for food segmentation scenarios, where different ingredients may exhibit similar visual patterns but require accurate separation at the pixel level.
Recent food segmentation methods have explored alternative strategies to address this challenge. For instance, MVEANet [34] relies on a Transformer-based backbone (STViT) combined with a SAM-based decoder to achieve state-of-the-art segmentation accuracy, particularly in delineating fine-grained food boundaries. However, this comes at the cost of increased model complexity and inference time. In contrast, Muñoz et al. [33] proposed a lightweight CNN-based alternative by replacing the DeepLabv3+ backbone with EfficientNet-B1, achieving competitive segmentation performance while significantly reducing computational cost. These contrasting strategies illustrate the broader trade-off between Transformer-based and CNN-based architectures in food segmentation: the former favors global attention-driven accuracy, while the latter favors efficiency and deployability.
Building on this CNN-based direction, ResNet34 is commonly adopted as a classification backbone due to its residual connections, which mitigate the vanishing gradient problem and enable effective extraction of discriminative, high-level image features. Therefore, ResNet34 is effective for learning discriminative image-level features and category recognition, whereas U-Net-based architectures are more suitable for detailed food region delineation. The combination of deep feature extraction and spatial localization represents an effective strategy for intelligent food image analysis.
4.6.4. Interpretability Analysis by Grad-CAM
To improve the interpretation of the model’s decisions, a complementary analysis based on ResNet34 and Grad-CAM was performed.
ResNet34, proposed by He et al. [16], introduces residual connections, enabling efficient learning of deep representations. The Grad-CAM maps proposed by Selvaraju et al. [12], allow visualization of image regions that strongly contribute to the network’s predictions.
This analysis provides a better understanding of the behavior of deep learning models. Interpretability has become essential for promoting the adoption of artificial intelligence systems, particularly in applications where decisions must be analyzed by human experts [26].
Discussion of Grad-CAM Visualizations
The Grad-CAM visualizations provide additional qualitative evidence regarding the effect of fine-tuning on the ResNet34-based segmentation model. The superimposed activation maps indicate that, after fine-tuning, the model focuses more strongly on image regions associated with the food content. This behavior suggests that the adaptation of the pretrained representations enables the encoder to capture features that are more relevant to the target food-segmentation task. In particular, the concentration of high-activation regions within or around the segmented food areas indicates a better alignment between the discriminative features learned by the network and the spatial regions of interest. This qualitative behavior is consistent with the improvement observed in the quantitative segmentation metrics after fine-tuning, suggesting that the adaptation of ResNet34 contributes not only to higher segmentation performance but also to more task-relevant visual representations. Nevertheless, Grad-CAM should be considered as a complementary interpretability analysis rather than direct evidence of segmentation accuracy, since the activation maps indicate regions contributing to the model prediction and do not constitute segmentation masks themselves. Overall, the combined quantitative and qualitative analyses support the benefit of fine-tuning ResNet34 for adapting the pretrained encoder to the FoodSeg103 domain.
4.6.5. Synthesis of the Trade-Off Between Performance, Complexity, and Interpretability
The overall study shows that no architecture offers an absolute advantage for all the criteria considered. Convolutional models based on U-Net offer an effective trade-off between segmentation accuracy, architectural complexity, and deployment capacity. Thanks to their encoder-decoder structure and their jump connections, these architectures make it possible to preserve fine spatial information while improving the reconstruction of segmented regions [7].
The ablation study showed that integrating jump connections significantly improves segmentation performance by facilitating the transmission of spatial features between the encoder and the decoder. Furthermore, the fine-tuning phase allows for the gradual adaptation of the learned parameters to the specific characteristics of the food domain, thus improving the generalizability of the model [14].
In parallel, ResNet34 provides an efficient representation of visual features thanks to its residual architecture, enabling robust extraction of discriminating information for food category classification. The use of residual connections facilitates the learning of deep learning networks by improving gradient propagation.
Grad-CAM analysis complements this evaluation by adding a dimension of interpretability. This method allows visualization of the image regions that contributed to the network’s decisions, thus improving the understanding of the behavior of deep learning models.
Thus, the results obtained show that the combination of a suitable segmentation architecture (U-Net), fine-tuning optimization, and explanatory analysis provides a relevant compromise between performance, computational cost, and interpretability for intelligent systems of automatic food recognition and segmentation.
5. Conclusion
In this work, we proposed and evaluated a comparative approach to two Deep Learning architectures for the semantic segmentation of food images, with particular emphasis on the contribution of skip connections, fine-tuning and interpretability by Grad-CAM. Overall, the experimental results demonstrate that the proposed architectural improvements provide measurable benefits for semantic food segmentation. The ablation study confirms the contribution of skip connections to the preservation of spatial information, while fine-tuning further improves the adaptation of the models to the FoodSeg103 domain. More specifically, the results reveal a differentiated effect of this fine-tuning process between the two models: fine-tuning substantially improves the U-Net with skip connections, whereas ResNet34, which converges more rapidly during training, benefits only marginally from fine-tuning. This contrast suggests that the effectiveness of fine-tuning is closely tied to each architecture’s initial learning dynamics, with U-Net’s slower initial convergence leaving considerably more room for improvement compared with ResNet34’s rapid early convergence.
The complementary Grad-CAM analysis further provides visual evidence of the regions contributing to the model predictions, supporting the interpretability of the learned representations. These findings highlight the added value of combining architectural design, fine-tuning, quantitative evaluation, and explainability rather than relying exclusively on model complexity or final segmentation scores.
Data Availability Statement
The dataset used in this study is publicly available. The experiments were conducted on the FoodSeg103 dataset introduced by Wu et al. [6], which provides pixel-level annotations for food semantic segmentation. The dataset is publicly available through the repository provided by the authors: https://github.com/LARC-CMU-SMU/FoodSeg103-Benchmark-v1.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| CNN | Convolutional Neural Network |
| ResNet | Residual Network |
| U-Net | U-shaped Network (encoder-decoder segmentation architecture) |
| ReLU | Rectified Linear Unit |
| IoU | Intersection over Union |
| mIoU | mean Intersection over Union |
| CE | Cross-Entropy (loss) |
| FT | Fine-Tuning |
| Grad-CAM | Gradient-weighted Class Activation Mapping |
| CAM | Class Activation Mapping |
| GAP | Global Average Pooling |
| BN | Batch Normalization |
| SGD | Stochastic Gradient Descent |
| Adam | Adaptive Moment Estimation |
| LR | Learning Rate |
| GPU | Graphics Processing Unit |
| DSC | Dice Similarity Coefficient |
| DL | Dice Loss |
| FN | False Negative |
| FP | False Positive |
| TP | True Positive |
| TN | True Negative |
| AUC | Area Under the Curve |
| ROC | Receiver Operating Characteristic |
| FCN | Fully Convolutional Network |
| BCE | Binary Cross-Entropy |
| STViT | Super Token Vision Transformer |
| SAM | Segment Anything Model |
| ASPP | Atrous Spatial Pyramid Pooling |
References
- Papargyropoulou, E.; Lozano, R.; Steinberger, J.K.; Wright, N.; Ujang, Z. The Food Waste Hierarchy as a Framework for the Management of Food Surplus and Food Waste. J. Clean. Prod. 2014, 76, 106–115. [Google Scholar] [CrossRef]
- Gao, D.-M.; Song, J.-Q.; Fu, Z.-Q.; Liu, Z.; Li, G. Utilizing Multimodal Logic Fusion to Identify the Types of Food Waste Sources. Sensors 2026, 26, 851. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep Learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [PubMed]
- Zhou, L.; Zhang, C.; Liu, F.; Qiu, Z.; He, Y. Application of Deep Learning in Food: A Review. Compr. Rev. Food Sci. Food Saf. 2019, 18, 1793–1811. [Google Scholar] [CrossRef] [PubMed]
- Bossard, L.; Guillaumin, M.; Van Gool, L. Food-101 – Mining Discriminative Components with Random Forests. In European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2014; pp. 446–461. [Google Scholar] [CrossRef]
- Wu, X.; Fu, X.; Liu, Y.; Lim, E.-P.; Hoi, S.C.H.; Sun, Q. A Large-Scale Benchmark for Food Image Segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2021.
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. (NeurIPS) 2021, 34, 12077–12090. [Google Scholar]
- Mazloumian, A.; Rosenthal, M.; Gelke, H. Deep Learning for Classifying Food Waste. arXiv 2020. [Google Scholar] [CrossRef]
- Rahman, M.M.; Islam, M.S.; Hossain, M.A.; et al. BDWaste: A Comprehensive Image Dataset of Digestible and Indigestible Waste in Bangladesh. Data Brief. 2024, 53, 110153. [Google Scholar] [CrossRef] [PubMed]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017; pp. 618–626. [Google Scholar] [CrossRef]
- PyTorch, torch.nn.CrossEntropyLoss, PyTorch Documentation, 2026. [Online]. Available online: https://docs.pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html.
- Yosinski, J.; Clune, J.; Bengio, Y.; Lipson, H. How Transferable Are Features in Deep Neural Networks? Adv. Neural Inf. Process. Syst. (NeurIPS) 2014, 27, 3320–3328. Available online: https://arxiv.org/abs/1411.1792.
- Rumelhart, D. E.; Hinton, G. E.; Williams, R. J. Learning Representations by Back-Propagating Errors. Nature 1986, vol. 323, 533–536. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016; pp. 770–778. [Google Scholar] [CrossRef]
- Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef]
- Cheng, B.; Girshick, R.; Dollár, P.; Berg, A. C.; Kirillov, A. Boundary IoU: Improving Object-Centric Image Segmentation Evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021; pp. 15334–15342. [Google Scholar]
- Okamoto, K.; Yanai, K. UEC-FoodPIX Complete: A Large-scale Food Image Segmentation Dataset. International Conference on Pattern Recognition Workshops (ICPRW), Multimedia Assisted Dietary Management Workshop (MADiMa). 2021. [Google Scholar]
- Meyes, R.; Lu, M.; de Puiseau, C.W.; Meisen, T. Ablation Studies in Artificial Neural Networks. arXiv 2019, arXiv:1901.08644. [Google Scholar]
- Zhang, X.; Wang, Y.; Li, J.; et al. Multi-View Edge Attention Network for Fine-Grained Food Image Segmentation. Foods 2025, 14, 3016. [Google Scholar] [CrossRef]
- Wang, S.; Sun, G. Hybrid Decoding with Co-Occurrence Awareness for Fine-Grained Food Image Segmentation. Foods 2026, 15, 534. [Google Scholar] [CrossRef]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In International Conference on Learning Representations (ICLR); 2019.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR); 2021.
- Pan, S.J.; Yang, Q. A Survey on Transfer Learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef]
- Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; García, S.; Gil-López, S.; Molina, D.; Benjamins, R.; et al. Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef]
- Minaee, S.; Boykov, Y.; Porikli, F.; Plaza, A.; Kehtarnavaz, N.; Terzopoulos, D. Image Segmentation Using Deep Learning: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 3523–3542. [Google Scholar] [CrossRef] [PubMed]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. International Conference on Learning Representations (ICLR), 2015. [Google Scholar] [CrossRef]
- Nair, V.; Hinton, G.E. Rectified Linear Units Improve Restricted Boltzmann Machines. In Proceedings of the 27th International Conference on Machine Learning (ICML); Omnipress: Madison, WI, USA, 2010; pp. 807–814. [Google Scholar]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-Based Learning Applied to Document Recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- Hussain, M.A.; Khan, M.I.H.; Karim, A. Machine Learning in Transforming the Food Industry. Foods 2026, 15, 90. [Google Scholar] [CrossRef] [PubMed]
- Wang, S.; Sun, G. Hybrid Decoding with Co-Occurrence Awareness for Fine-Grained Food Image Segmentation. Foods 2026, 15, 534. [Google Scholar] [CrossRef]
- Muñoz, B.; Martínez-Arroyo, A.; Acevedo, C.; Aguilar, E. Lightweight DeepLabv3+ for Semantic Food Segmentation. Foods 2025, 14, 1306. [Google Scholar] [CrossRef] [PubMed]
- Liu, C.; Sheng, G.; Min, W.; Wu, X.; Jiang, S. Multi-View Edge Attention Network for Fine-Grained Food Image Segmentation. Foods 2025, 14, 3016. [Google Scholar] [CrossRef] [PubMed]
- Ben Ali, K.; Bouazizi, S.; Ksibi, M.; Hamdi, M. Integrating Deep Learning for Food Waste Classification and Intelligent Sorting in a Circular Economy Framework. Preprints.org 2026. [Google Scholar]
- Ben Ali, K.; Bouazizi, S.; Ksibi, M.; Hamdi, M. Integrating Deep Learning for Food Waste Classification and Intelligent Sorting in a Circular Economy Framework. Preprints.org. 2026. Available online: https://www.preprints.org/manuscript/202608.1216.
Figure 1.
Overall workflow of the proposed deep learning framework for food image semantic segmentation. The methodology starts from FoodSeg103 preprocessing, followed by parallel U-Net and ResNet34-U-Net branches, including skip connections and model adaptation, respectively. The two branches are then compared quantitatively, and the best-performing model (ResNet34-U-Net) is further analyzed through Grad-CAM-based explainability.
Figure 1.
Overall workflow of the proposed deep learning framework for food image semantic segmentation. The methodology starts from FoodSeg103 preprocessing, followed by parallel U-Net and ResNet34-U-Net branches, including skip connections and model adaptation, respectively. The two branches are then compared quantitatively, and the best-performing model (ResNet34-U-Net) is further analyzed through Grad-CAM-based explainability.

Figure 2.
Comparison between the SimpleUNet baseline and the U-Net architecture with skip connections. The skip connections directly transfer encoder feature maps to the corresponding decoder stages to preserve spatial information during segmentation.
Figure 2.
Comparison between the SimpleUNet baseline and the U-Net architecture with skip connections. The skip connections directly transfer encoder feature maps to the corresponding decoder stages to preserve spatial information during segmentation.

Figure 3.
Fine-tuning procedure of the U-Net with Skip Connections. The model is first trained for 50 epochs using a learning rate of . The learned weights are then used to initialize a second training phase of 50 additional epochs with a reduced learning rate of .
Figure 3.
Fine-tuning procedure of the U-Net with Skip Connections. The model is first trained for 50 epochs using a learning rate of . The learned weights are then used to initialize a second training phase of 50 additional epochs with a reduced learning rate of .

Figure 4.
ResNet34 architecture used for multi-label classification on FoodSeg103. The architecture remains identical before and after fine-tuning; only the trainable parameter configuration changes, as summarized in the comparison box.
Figure 4.
ResNet34 architecture used for multi-label classification on FoodSeg103. The architecture remains identical before and after fine-tuning; only the trainable parameter configuration changes, as summarized in the comparison box.

Figure 5.
Data pipeline orchestration and ResNet34 fine-tuning forsemantic segmentation on FoodSeg103.
Figure 5.
Data pipeline orchestration and ResNet34 fine-tuning forsemantic segmentation on FoodSeg103.

Figure 6.
Comparison of the training evolution of the Simple U-Net, U-Net with skip connections, and fine-tuned U-Net in terms of loss and accuracy. The fine-tuning process was performed for 50 epochs and therefore its curves terminate at epoch 50.
Figure 6.
Comparison of the training evolution of the Simple U-Net, U-Net with skip connections, and fine-tuned U-Net in terms of loss and accuracy. The fine-tuning process was performed for 50 epochs and therefore its curves terminate at epoch 50.

Figure 7.
Training loss and training accuracy of ResNet-34 before and after fine-tuning over 50 epochs. The left panel reports the training loss, while the right panel reports the training accuracy.
Figure 7.
Training loss and training accuracy of ResNet-34 before and after fine-tuning over 50 epochs. The left panel reports the training loss, while the right panel reports the training accuracy.

Figure 8.
Improvement model with Fine-tuning and Comparison U-Net vs ResNet34.

Figure 9.
Random food images and their corresponding predicted segmentation masks. First set of five random images and their masks. Second set of five random images and their masks.
Figure 9.
Random food images and their corresponding predicted segmentation masks. First set of five random images and their masks. Second set of five random images and their masks.

Figure 10.
Grad-CAM-based explainability analysis of the fine-tuned ResNet34 model on FoodSeg103 samples. For each subcaption, the first row presents the original food images, the second row shows the corresponding Grad-CAM attention maps highlighting the discriminative regions used by the model, illustrating the prediction overlays with the predicted classes.
Figure 10.
Grad-CAM-based explainability analysis of the fine-tuned ResNet34 model on FoodSeg103 samples. For each subcaption, the first row presents the original food images, the second row shows the corresponding Grad-CAM attention maps highlighting the discriminative regions used by the model, illustrating the prediction overlays with the predicted classes.

Table 1.
Detailed architecture of the SimpleUNet model used as a baseline for semantic segmentation on FoodSeg103.
Table 1.
Detailed architecture of the SimpleUNet model used as a baseline for semantic segmentation on FoodSeg103.
| Layer | Operation | Input Size | Output Size | Parameters |
|---|---|---|---|---|
| 1 | Conv2D (3→64, kernel=3, padding=1) | 1,792 | ||
| 2 | ReLU Activation | 0 | ||
| 3 | Conv2D (64→128, kernel=3, padding=1) | 73,856 | ||
| 4 | ReLU Activation | 0 | ||
| 5 | Conv2D (128→104, kernel=1) | 13,416 | ||
| Total Trainable Parameters | 89,064 | |||
Table 2.
Forward propagation flow of SimpleUNet.
| Stage | Tensor Representation |
|---|---|
| Input image | |
| Conv1 + ReLU | |
| Conv2 + ReLU | |
| Classification convolution | |
| Final segmentation map | Pixel-wise prediction over 104 classes |
Table 4.
Performance comparison of U-Net based models on FoodSeg103 dataset.
| Model | Pixel Acc. | Precision | Recall | F1-score | Mean IoU | Dice |
|---|---|---|---|---|---|---|
| U-Net | 62.47% | 58.54% | 62.47% | 57.33% | 0.4116 | 0.4894 |
| Skip U-Net Before Fine-tuning | 85.90% | 86.31% | 85.90% | 85.11% | 0.7150 | 0.7899 |
| Skip U-Net After Fine-tuning | 94.45% | 94.75% | 94.45% | 94.40% | 0.8626 | 0.9135 |
Table 5.
Ablation study of fine-tuning strategy on ResNet34 for FoodSeg103 semantic segmentation.
| Model | Accuracy | Precision | Recall | F1-score | Mean IoU | Dice |
|---|---|---|---|---|---|---|
| ResNet34 Baseline | 97.27% | 79.81% | 51.40% | 62.53% | 0.4549 | 0.6253 |
| ResNet34 + Fine-tuning | 97.28% | 79.42% | 52.27% | 63.05% | 0.4604 | 0.6305 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
