Preprint
Review

This version is not peer-reviewed.

Light-Assisted Computer Vision for Non-Destructive Assessment of Meat Quality: Intramuscular Fat and Freshness Estimation

Submitted:

05 August 2026

Posted:

07 August 2026

You are already at the latest version

Abstract
Meat quality at the point of purchase is largely assessed through subjective visual inspection, whereas accurate measurement of fat and freshness relies on destructive, time-consuming laboratory methods. Visible-range red–green–blue (RGB) computer vision offers a non-destructive, low-cost alternative, but evidence is fragmented across separate research strands. This review systematically synthesises peer-reviewed studies on light-assisted, visible-spectrum RGB methods for estimating fat and freshness in meat, published between 2017 and 2025 and retrieved from seven databases. Thirty-two studies are selected and analysed across three themes: sixteen focus on intramuscular fat (including marbling) and intermuscular fat; ten on colour-based freshness; and six on light-assisted capture. For fat, feature-based pipelines using CIELAB colour and visible fat area achieve R² ≈ 0.72–0.90, with errors near 0.2–1.3 percentage points under standardised loin or rib imaging. Deep segmentation achieves 96–99% accuracy, while smartphone setups yield R² ≈ 0.90 under strict protocols and 0.54–0.62 otherwise. For freshness, colour-feature models predict pH, TVB-N and microbial counts at R² up to ≈ 0.99 and classify spoilage at 90–95%. Reliable accuracy is achieved across the themes under controlled illumination, fixed geometry and calibration.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

A healthy, balanced diet that includes safe, high-quality animal-source foods remains important for human nutrition and productivity. Global demand for meat and other animal proteins is projected to continue rising as populations and incomes grow [1,2]. At the same time, livestock production and meat processing contribute substantially to land use and greenhouse-gas emissions, increasing pressure to improve efficiency and reduce negative environmental effects in the meat sector [3,4]. This contradiction hinders progress towards achieving the United Nations Sustainable Development Goals (SDGs). Indeed, SDG 2 (Zero Hunger) and SDG 3 (Good Health and Wellbeing) emphasise improved nutrition, while SDG 12 (Responsible Consumption and Production) calls for more sustainable food systems and more efficient use of resources [5]. In this context, the evaluation of meat product quality is becoming increasingly important, as it helps to balance production and consumption requirements and shapes how consumers perceive quality and value at retail [6,7].
In practice, buyers and consumers rely heavily on intrinsic cues such as colour, visible fat distribution and surface appearance to assess freshness and likely eating quality, especially when formal grading information is limited or absent [8,9,10]. Two cues are particularly influential. The first is visible fat, which occurs in two forms: (a) intramuscular fat (IMF), including so-called “marbling”, formed by fine fat flecks dispersed within the muscle, and (b) intermuscular fat, often identified as intermuscular adipose tissue (IMAT), which is the lighter, opaque seam fat deposited between adjacent muscles along fascial boundaries on a cut face. This review focuses on non-destructive, visible-range estimation of both. Intramuscular fat is associated with flavour, tenderness, and juiciness. Visible fat, in turn, is a prominent purchase cue that influences consumer acceptance and willingness to buy. [8,11]. The second cue is the surface colour. Bright red or pink meat is generally associated with freshness and acceptability, whereas brownish or discoloured surfaces are linked to lower perceived quality and higher perceived risk [10,12,13,14].
Established methods for quantifying these traits are poorly suited to routine point-of-sale use. Reference techniques for fat composition, such as solvent extraction and gas chromatography, are destructive, require specialised instruments and skilled personnel, and are time-consuming and costly, making them ill-suited to routine use outside the laboratory [15]. Marbling grading, meanwhile, still depends in many industries on visual assessment by trained graders, which can be inconsistent and vary between facilities and assessors, undermining standardisation [16]. There is thus a persistent gap between the objective quality information that would support fairer pricing and more informed purchasing and the tools available outside large, instrumented facilities.
Computer vision (CV) and machine learning (ML) have emerged as non-destructive, lower-cost alternatives. Recent studies show that consumer-grade images captured with smartphones or compact cameras can be used to classify tenderness and estimate intramuscular fat in beef and pork using convolutional models trained on the red–green–blue (RGB) colour domain [17]. These approaches can also detect fat-injection adulteration in beef using colour and texture features [18]. Because uncontrolled illumination is a primary source of error in RGB imaging, many of these approaches pair the camera with light-assisted capture, enclosures, diffusers, and calibrated references to stabilise lighting and colour. Reviews of artificial intelligence (AI) applications in meat processing [19] and across food systems more broadly [20] argue that such systems could support more objective grading and safer purchasing, while cautioning that many proposed solutions remain at the laboratory or algorithm-prototype stage and require validation in realistic production and retail environments.
Research on RGB-based meat assessment is spread across three largely separate strands: (a) estimation of intramuscular and intermuscular fat, (b) colour-based freshness or spoilage classification, and (c) light-assisted capture hardware, each studied using heterogeneous rigs, datasets, species and reporting conventions. As a result, it is difficult to determine what RGB-only light-assisted approaches can reliably achieve, how their accuracy compares with more expensive near-infrared (NIR) or hyperspectral (HS) systems, and which design choices transfer to low-cost, consumer-grade deployment. A systematic synthesis that consolidates imaging hardware, analysis techniques, and reported performance across these strands is therefore needed to consolidate the fragmented evidence base.
This review systematically surveys peer-reviewed literature published between 2017 and 2025 on light-assisted, visible-spectrum RGB computer-vision methods for the non-destructive assessment of intramuscular and intermuscular fat and freshness in meat. It addresses three research questions:
RQ1.
What imaging hardware and capture configurations, including light-assisted enclosures and calibration approaches, have been used for RGB-based assessment of meat quality?
RQ2.
What image-processing and machine-learning techniques are applied to estimate intramuscular and intermuscular fat and colour-based freshness, and what accuracy do they report?
RQ3.
What are the principal gaps and challenges that limit the low-cost, consumer-grade deployment of these methods?
The scope is limited to visible-spectrum RGB approaches, while NIR, HS, X-ray, and sensor-based modalities are considered only when studies use them as comparison baselines.
The remainder of the review is organised as follows. Section 2 outlines the review methodology, including the search protocol, inclusion and exclusion criteria, study selection, and data extraction. Section 3 presents the results, synthesised across the three thematic strands of intramuscular and intermuscular fat, freshness, and relevant light-assisted computer vision solutions. Section 4 discusses cross-cutting findings, research gaps, directions for future work, and the review's limitations. Section 5 concludes and summarises the study.

2. Methodology

This study follows a protocol-driven approach adapted from established guidelines for systematic literature reviews in computer science and software engineering [21,22]. The research questions, data sources, search strategy, selection criteria and data-extraction fields were defined in advance to support transparency and replicability.

2.1. Review Protocol and Research Questions

The review scope was structured using the PICOC framework (Population, Intervention, Comparison, Outcome, Context), from which the search keywords and research questions were derived [21,23]:
  • Population: meat and closely related animal products imaged on the cut or fillet surface; this includes red meat (beef, pork, lamb) and fish fillets, which are retained because they present the same visible-RGB fat- and colour-estimation problem on an exposed cut face.
  • Intervention: visible-spectrum RGB computer-vision imaging, including light-assisted capture (controlled-lighting enclosures, ring or strip lights, diffusers, cross-polarised setups).
  • Comparison: where reported, near-infrared or hyperspectral imaging, trained-grader assessment, or chemical and colourimeter reference measurements.
  • Outcome: estimation of intramuscular and intermuscular fat, and colour-based freshness or spoilage indicators.
  • Context: controlled laboratory rigs and smartphone or commodity-camera capture.
The review addresses the three research questions posed in Section 1 (RQ1–RQ3).

2.2. Data Sources and Search Period

The literature base was derived from a predetermined keyword-based search of major scientific databases, limited to peer-reviewed journals and conference proceedings published between 2017 and 2025, a window chosen to capture recent developments in light-assisted imaging and artificial intelligence for meat-quality assessment. The databases searched were:
  • IEEE Xplore
  • ScienceDirect (Elsevier)
  • ACM Digital Library
  • Wiley Online Library
  • Taylor & Francis Online
  • SPIE Digital Library
  • PubMed

2.3. Search Strategy

Search terms were grouped into three functional blocks mirroring the three themes. The fat and marbling block combined content terms (intermuscular fat, intramuscular fat, marbling, fat content, meat quality, grading), species and product terms (meat, pork, beef, lamb, steak), and imaging terms (RGB, visible imaging, camera, scanner, smartphone), along with controlled-lighting or light-assisted-capture phrases. The freshness block combined freshness-related terms (meat freshness, spoilage detection, shelf life, colour change) with RGB colour measures. The light-assisted and device block combined lighting and rig terms (controlled lighting, light-assisted, cross-polarised, diffuse dome, ring light, light box, dark box) and device terms (smartphone imaging, commodity camera, bench setup, industrial line), along with meat or food-quality terms.
Within each block, synonyms were combined using the Boolean operator OR, and the blocks were combined using AND to form the search string for each theme. Exclusionary terms were applied where needed to remove non-relevant uses of abbreviations, for example, to prevent “IMF” from retrieving “International Monetary Fund”.
A representative string for the fat and marbling theme is: (“intermuscular fat” OR “intramuscular fat” OR marbling OR “fat content”) AND (meat OR beef OR pork OR lamb) AND (RGB OR “visible imaging” OR camera OR smartphone) AND (“controlled lighting” OR “light-assisted”).

2.4. Inclusion Criteria

Following preliminary collection, the following criteria were used during title and abstract screening:
  • Studies must apply image-based analysis of meat or closely related animal products; studies using only visual inspection without imaging are excluded.
  • At least one target variable must relate to the three themes: intramuscular fat percentage, marbling scores, intermuscular fat, fat content, meat-quality grades (e.g., USDA Prime, Choice, Select), surface-colour measures, image-based freshness, or spoilage indicators.
  • Studies must include visible RGB capture; those using only NIR or HS sensors are excluded; multimodal studies are retained only where the RGB model is clearly documented.
  • Lighting conditions and capture geometry should be documented; studies using controlled or light-assisted setups and reporting ground-truth or reference measurements, chemical fat analysis, or standard grading scales for fat and marbling; colourimeter readings; defined storage time or laboratory freshness metrics for freshness were preferred [12,15].

2.5. Study Selection

Selection proceeded in stages: identification of records from the databases, removal of duplicates, title and abstract screening against the criteria above, and full-text assessment of the remaining records. Papers were then grouped by theme: RGB-based estimation of intramuscular and intermuscular fat (marbling); RGB-based freshness assessment from colour; and light-assisted RGB rigs and smartphone or commodity-camera workflows. This is because studies within each group address the same underlying problem and can be compared in terms of capture setup and modelling choices.

2.6. Data Extraction

For each included study, data were extracted into a structured form that captured the reference; the methodology (imaging hardware, lighting, region-of-interest selection, and feature or model choices); the results and findings (reported accuracy or error metrics); and the limitations. These fields structure the comparison tables presented in Section 3.

3. Results

3.1. Included Studies

The included studies are organised into the three themes defined in the protocol: RGB-based estimation of IMF (including marbling) and IMAT; RGB-based freshness assessment from colour cues; and light-assisted RGB capture rigs and smartphone or commodity-camera workflows. For each study, the imaging setup, modelling approach, reported performance and limitations were extracted into a comparison table, and the studies within each theme were then synthesised. In total, 32 studies met the inclusion criteria: 16 addressing intramuscular and intermuscular fat, 10 addressing freshness, and 6 addressing light-assisted capture.
Predictive performance across the reviewed studies is reported using two metrics. The coefficient of determination (R²) is the proportion of variance in the reference values explained by the model. It is bounded above by 1 and can be negative when a model performs worse than predicting the mean. Thus, higher values indicate greater explanatory power [24,25]. The root mean square error (RMSE) is the average magnitude of the prediction error, expressed in the same units as the target variable and giving greater weight to large errors. Lower values indicate a better fit, with RMSE = 0 denoting perfect prediction [26,27]. The two metrics are interpreted together, alongside the IMF error considered acceptable for the intended use.

3.2. Intramuscular and Intermuscular Fat

Studies in this theme produced fairly accurate estimates of IMF and marbling, as well as IMAT, using hardware with highly controlled illumination and imaging geometry, typically in laboratory rigs or with industrial cameras on production lines. There, lighting, distance and camera angles relative to the sample are predetermined, and the samples are carefully prepared before imaging. In this theme, “marbling” refers to intramuscular (within-muscle) fat, which is sometimes reported alongside IMF in the cited studies. Table 1 summarises the studies; the synthesis follows.

3.2.1. Equipment and Capture Techniques

Three hardware patterns recur. The first is bench-top RGB or colorimetric rigs around pork loins. In the selected and reviewed study [28], a Bisaro pork loin (longissimus thoracis et lumborum) was imaged using a digital RGB camera, and CIELAB colour measurements were taken in a controlled slaughterhouse environment, with the camera at a fixed distance and angle under uniform illumination to obtain colour features and visible IMF area from standardised photographs (CIELAB is a device-independent color space standardised by the International Commission on Illumination - CIE, where L* is lightness, a* the red–green axis, and b* the yellow–blue axis). Related industrial systems image meat on a conveyor under fixed illumination, capturing one image per steak as it passes a fixed camera with an LED line light, which limits shadow and white-balance variation [34,35].
The second pattern comprises specialised luminance and grading cameras. Wachta et al. [38] built a luminance-measurement system around a Canon camera in a shadowless tent with diffuse LED illumination, using an LMK matrix luminance meter to convert RAW images into luminance maps for IMF estimation from the spatial luminance distribution. The Q-FOM™ Beef computer-vision grading scanner combines a 3D camera and a high-resolution 2D camera, with two diffuse LED panels, within a shielded housing to capture ribeye images under standardised lighting at specified rib locations. It is factory-calibrated for colour and geometry [36].
The third pattern is controlled capture for segmentation-centric pipelines. Uttaro et al. [32] built a controlled station for intact pork loins, capturing RAW and JPEG images of the mid-loin chop and its external faces, with a consistent background, distance, camera settings, and standardised illumination and colour balance, so that marbling could be isolated regardless of lean colour variation. Steak images captured with a ruler for scale were used to drive a two-stage U-Net (a type of convolutional neural network) that first segmented the steak and then the fat region, supporting area-to-mass estimation [29].
A fourth pattern is handheld smartphone capture under a strict protocol. Meunier et al. [33] imaged the longissimus thoracis at the sixth rib in a slaughterhouse using a Samsung Galaxy S8 fitted with orthogonal cross-polarising filters to suppress surface glare, placed a 5×5 cm coloured reference card in frame for colour and scale normalisation, and extracted seventeen colour and texture features with an ImageJ macro. Unlike the rig-based systems, the camera was handheld under otherwise uncontrolled conditions, making this the closest approximation in the fat theme to a consumer capture workflow.
Across these systems, the common techniques are:
  • Fixed camera-to-sample distance and perpendicular viewing of the cut face to avoid perspective distortion [32].
  • High colour-rendering index (CRI) LED illumination with a stable spectrum and intensity, sometimes combined with enclosures or shrouds to control ambient light [28,38].
  • Device-independent colour spaces such as CIELAB, and image-derived features such as fat-area percentages, texture measures and luminance statistics, as predictors of IMF [28,38].
  • Predefined regions of interest at anatomically standardised locations, such as mid-loin chops or specific rib sections, to reduce sampling variability [32,40].

3.2.2. Findings

For feature-based regression in controlled pork rigs, Teixeira et al. [28] modelled Bisaro pork IMF with a polynomial support vector machine (SVM) using carcass weight, CIELAB coordinates and visible fat area, achieving R² = 0.88 (RMSE ≈ 0.18 pp) in cross-validation and a near-perfect best fit (R² ≈ 1.0, residual SE ≈ 0.04 pp); mixture discriminant analysis (MDA) classified IMF-defined juiciness groups with 100% accuracy on training and 92% in cross-validation. Chen et al. [31] combined CV scores with conventional traits across 200 pigs imaged in a standardised cabinet, with CV-derived IMF correlating at 0.68 with chemical IMF and about 88% (stepwise regression) and 89% (gradient boosting machine or GBM) of predictions falling within ±0.5 percentage points of chemical IMF, showing that combining RGB information with standard measurements improves on image features alone. Uttaro et al. [32] obtained R² of 0.72 (RAW) and 0.74 (JPEG) at the mid-loin, with external-face correlations of 0.56–0.72, indicating that external views can reasonably reflect internal IMF when imaging is standardised. Wachta et al. [38] reported near-identical mean IMF for luminance mapping and Soxhlet extraction (1.68% vs 1.67%) with a correlation of 0.79 across 68 pork samples, indicating low bias despite a lower correlation than other studies.
For segmentation-oriented pipelines, Liu et al. [37] reported fat and muscle segmentation accuracies of about 98–99% across 2,000 retail pork-cut images, using an improved U2-Net (a two-level nested U-shaped architecture designed for salient object detection, denoted as SOD) with attention, yielding about 93% IMF accuracy. Symeonidis et al. [29] reported steak segmentation accuracy of about 92–97% with a two-stage U-Net, showing that, with controlled or curated imaging, deep networks can produce highly accurate fat masks.
For industrial grading cameras, the Q-FOM™ Beef predicted chemical IMF in the longissimus thoracis at the fifth–sixth thoracic vertebra of crossbred calves with R² ≈ 0.91, and the root mean square error of cross-validation RMSECV ≈ 1.33 pp, though accuracy dropped to R² ≈ 0.48 on the narrow-range veal subsample [36]. Handheld MIJ cameras imaging the rib-eye in the abattoir reported substantially lower accuracy, with cross-validated IMF R² ≈ 0.4–0.5, errors of about 1.5–1.6 pp, and weaker marbling estimation [40]. On a Marel conveyor vision scanner, Pannier et al. [35] predicted chemical IMF from a single fresh-cut steak image under an LED line light with R² ≈ 0.87 and root mean square error of prediction RMSEP ≈ 1.16 pp, also reproducing the Meat Standards Australia (MSA, R² ≈ 0.82) and AUS-MEAT (R² ≈ 0.79) marbling scores, while a companion study found grading-site marbling correlated strongly with laboratory IMF (r ≈ 0.93) yet varied widely along the cube roll (about 316 MSA units). Thus, a single grading-site reading does not represent marbling across the whole primal [34]. Taken together, controlled RGB rigs produce R² of approximately 0.72–0.90 versus chemical IMF or reference grading, with absolute errors of about 0.2–1.3 percentage points IMF depending on product type and dataset size [28,32,33,35,36].
Two studies report genuine intermuscular (seam) fat rather than within-muscle marbling. Huang et al. [39] estimated IMAT between muscle blocks in rainbow trout fillets. The samples were imaged in a shadowless box using a macro lens, and fat areas were characterised by sixteen colour features. The random forest estimate achieved R² ≈ 0.91, with about 79% of predictions within the band and a mean absolute error of roughly 1.5–2.0 percentage points, comparable to the strongest intramuscular results above. Meunier et al. [33], in addition to their intramuscular marbling model, separately predicted intermuscular fat at the sixth-rib site, with R² ≈ 0.84–0.86. These are the only IMAT results reported in the reviewed studies; both still depend on controlled or strictly standardised capture, and the evidence base for visible intermuscular fat remains considerably thinner than that for intramuscular marbling.
Two further studies in this theme classify marbling categorically rather than estimating IMF percentage directly. Liao et al. [41] trained an EfficientNet model on 38,528 images of 602 beef striploin steaks captured in a studio light box, predicting marbling scores with about 96% accuracy within one grade and about 99.6% within two grades, while also identifying breed and diet and supporting image-based traceability. Zhang et al. [42] combined a ResNet50 (Microsoft’s 50-layer convolutional neural network) backbone with grey-level co-occurrence matrix (GLCM) texture features and a multi-head graph attention mechanism to classify beef cuts and marbling grades with accuracies of about 93.5% and 92.3%, respectively. These results indicate that RGB imaging supports grade-level marbling classification and continuous IMF estimation, although both rely on controlled capture. Zhang et al. [42] used local A1–A5 grades without a chemical reference, whereas Liao et al. [41] derived IMF% via proximate analysis and used the IMF-derived marble score as the ground truth, reporting that visual marble scores can diverge from chemically determined IMF by up to about 20%.

3.2.3. Advantages and Limitations

Controlled rigs offer clear advantages. Image precision and reproducibility are high because lighting, camera geometry and background are predetermined; the Q-FOM™ Beef camera is factory-calibrated to support comparability across plants, although transfer to new sites was not directly evaluated. The controlled environment also enables non-linear models, such as SVM and GBM, to capture subtle marbling variation [28,31,36]. Feature extraction is also simpler and more interpretable: visible fat percentages, CIELAB coordinates, and luminance statistics can be computed consistently because shadow, glare, and background clutter are minimised, thereby supporting model comparison and linking features to physical characteristics [32,38]. Finally, these systems integrate into established meat-industry processes at fixed-line positions, providing consistent data for payment systems [34].
For individual consumers or small-scale vendors, the disadvantages of controlled rig-type systems are substantial. They tend to be expensive and heavy, requiring specialised cameras, lenses, luminance meters, LED panels, and sometimes proprietary software. Examples include the LMK luminance meter [38] and the proprietary Q-FOM™ Beef scanner, both of which are used in large processing environments. They also require precise, reproducible sample presentation and access to specified anatomical areas. In addition, many models are tuned to a particular cut face and do not generalise to random retail cuts, off-angle views, or partially obscured surfaces. Their performance can deteriorate at low IMF or when fat-distribution patterns deviate from the calibration data [32,36,40]. The need for darkened enclosures or fixed light boxes further limits portability. Such rigs cannot be used in wet markets or home kitchens, where ambient lighting and handling vary greatly, and it is unrealistic to expect a consumer to replicate the industrial capture geometry. Lastly, some related work uses NIR or hyperspectral imaging to achieve higher R² and lower RMSE. These modalities fall outside the visible-only scope of this review and highlight that RGB systems operate under tighter information constraints than wider spectral methods.

3.2.4. Summary

Several solutions have consistently been linked to improved IMF prediction in the controlled-environment studies:
  • Stable, high-CRI LED lighting with minimal stray light, often housed within an enclosure or under shrouds, to ensure colour and luminance remain comparable over time and across locations [28,38].
  • Fixed capture geometry: camera distance, angle and focal settings are locked, with the cut surface perpendicular to the optical axis to minimise perspective distortion [32].
  • Anatomically standardised regions of interest, such as mid-loin chops or specific rib sections, to reduce sampling error and maintain valid comparisons [28,32].
  • Non-linear models such as SVM and GBM, cross-validated against chemical IMF to avoid overfitting [28,31]. ).
  • For segmentation pipelines, carefully labelled ground-truth masks and controlled or uniform backgrounds, as in U-Net and U2-Net segmentation of beef steak and pork-cut images [29,37].
These controlled RGB systems set an upper bound on what a light-assisted, visible-spectrum computer vision system can achieve when hardware and the environment are fully standardised. The literature indicates that chemical-IMF accuracy from RGB images is attainable, but only with fixed geometry, stable lighting and an anatomically correct cut-face representation [28,31,32,38,40].

3.3. Meat Freshness

This theme covers research that infers meat freshness from colour cues in RGB images. Some studies use small light boxes and relatively simple models to estimate laboratory-based freshness measures, while others apply deep learning to RGB images of supermarket-style products to classify them as “fresh” or “spoiled”. Table 2 summarises the studies.

3.3.1. Equipment and Capture Techniques

Four patterns recur. The first involves controlled light boxes and calibrated cameras that map colour values to laboratory freshness measures. Jin et al. [49] imaged lamb with a Vivo S7 phone in a dark imaging box over ten days of refrigerated modified-atmosphere packaging (MAP) or air storage, extracting RGB, CIELAB and HSV values to compare an SVR model, a genetic-algorithm backpropagation network and a small CNN for predicting pH and Total Volatile Basic Nitrogen (TVB-N), with SVR performing best. Cheng et al. [46] imaged four pork cuts stored at 4 °C in a standardised environment with a smartphone, computing an image-analysis colour difference (IA-ΔE). The ΔE, derived from smartphone images, was related to total viable count (TVC), TVB-N, and pH, with tree-based models performing best. Pereira et al. [50] built a dark box containing a Samsung S5 phone and a USB webcam, which was colour-calibrated using a ColorChecker chart. The distance of RGB values from a fresh reference was used to model shelf life and bacterial counts. These studies used stabilised illumination and geometry and generally preferred CIELAB or HSV values over raw RGB to separate lightness and chroma.
The second pattern is histogram- and colour-distribution-based approaches. Meza et al. [44] imaged five beef cuts at three time points over nine days, segmented the meat, and compared average colour difference (ΔE) with the Kullback–Leibler (KL) divergence between CIELAB histograms, finding that histogram divergence was more sensitive to loss of redness and browning. You et al. [51] placed a 12-patch printed colour card next to chicken meat, localised the card using thresholding and contour detection, trained a multivariate colour correction model, and grouped meat pixels into three freshness categories using k-means clustering on mean-corrected RGB values; their contribution was explicit colour-card calibration rather than distribution-based features.
The third pattern is webcams and compact HSV-feature pipelines. Putra et al. [43] imaged tuna flesh with a low-cost webcam, converted RGB to HSV, extracted Symlet wavelet features, and ranked four quality levels using k-nearest neighbours. Bahri et al. [45] imaged tuna eyes, converted them to HSV, and trained an unspecified model, reported through epoch-by-epoch loss minimisation rather than a named architecture, on 50 images to make fresh-versus-not-fresh decisions. Unfortunately, the paper does not specify its classifier and is internally inconsistent about the method employed (RGB red-channel isolation versus HSV quantisation).
The fourth pattern is the application of deep learning to RGB images for grading and spoilage classification. Tan et al. [52] trained an Inception-V4 classifier on the Meat Standards Australia (MSA) colour grades of 400 beef rib-eye steaks captured with a single Logitech webcam under controlled LED lighting. A U-Net/DenseNet pipeline (combining DenseNet’s strong, parameter-efficient feature extraction with U-Net’s symmetric encoder-decoder structure and spatial skip connections for high-accuracy localisation) is used primarily for image segmentation. It was employed to provide classification (Fresh, Half-fresh or Spoiled) on the more heterogeneous Kaggle Meat Freshness dataset [48]. Another multimodal system, combining an ESP32-CAM development board (an ESP32 Wi-Fi/Bluetooth SoC with an OV2640 camera module and a microSD card slot) with gas sensors, trained a custom CNN on images of beef and mutton to classify species and freshness [47]. Capture settings in this group range from fully controlled [47,52] to mixed-source conditions [48]. Overall, RGB-based and derived colour features can be linked to laboratory freshness indicators and categorical freshness labels, with performance reported using R² and RMSE.

3.3.2. Findings

For light boxes with laboratory freshness proxies, colour characteristics often align closely with laboratory indicators. Jin et al. [49] reported SVR R² values up to 0.99, with low error and clear superiority of SVR over the genetic algorithm-optimised back-propagation neural network (GA-BPNN) and CNN for pH and TVB-N. Cheng et al. [46] found that random-forest models typically achieved above 80% accuracy and decision-tree models above 90% for TVC, TVB-N, and pH using IA-ΔE, with smartphone image models superior to colorimetric measurements. Pereira et al. [50] reported an average colour difference of about 5% relative to a colour meter and R² values of about 0.68–0.87 for shelf life and for mesophilic and psychotropic counts, with a response time of roughly 3 seconds.
For histograms, HSV, and simple classifiers, Putra et al. [43] reported about 82% accuracy for held-out tuna grades, with the Symlet wavelet slightly ahead of Haar (81.8% versus 80.3%). Bahri et al. [45] reported 100% correct fresh versus not-fresh decisions on an 8-image test set (out of 50 images total) using an unspecified trained model, a result attributable to the very small, highly controlled set. Meza et al. [44] showed that the Kullback–Leibler divergence of CIELAB histograms separated early and late storage days more effectively than the mean colour difference alone. You et al. [51] found that colour-card correction plus three-cluster k-means produced the most distinct freshness groupings, as measured by the Calinski–Harabasz index.
For deep networks that classify fresh versus spoiled and inform grading, Tan et al. [52] achieved about 90% overall accuracy on unseen beef colour grades. However, the average F1 score (a key performance metric in image recognition that combines precision and recall into a single parameter) was only around 60%. This indicates weak per-class performance despite high aggregate accuracy. Sagiraju et al. [48] reported about 93% test accuracy with a ResNet-18 model but deployed a ResNet-50 model that minimised a custom misclassification cost. The multimodal CNN of Bhuiyan et al. [47] achieved about 99% for both species and freshness classification of beef and mutton, though only over two storage time points.

3.3.3. Advantages and Limitations

The application of RGB freshness systems delivered tangible benefits. Controlled light boxes achieved laboratory-like results for chemical freshness proxies, including pH and TVB-N [46,49], and for bacterial counts [46,50]. Low-cost webcams and HSV pipelines are inexpensive and lightweight, yet they achieved roughly 82–100% accuracy for species-specific tasks such as tuna grading and eye-based inspection [43,45]. Deep-learning models achieved roughly 90–95% accuracy for grade-level rating and fresh-versus-spoiled categorisation. Sagiraju et al. [48] reported results on a heterogeneous, publicly sourced Kaggle dataset, Tan et al. [52] reported results under controlled single-camera capture, and Bhuiyan et al. [47] reported results using a multimodal CNN technique, achieving about 99% accuracy on a controlled two-timepoint capture.
The restricted scope is a limitation of most datasets. For example, Meza et al. [44] analysed five beef cuts at three time points over nine days. Bahri et al. [45] analysed only fifty tuna eyes, thus making generalisation difficult. Most light-box studies do not evaluate varying lighting or multiple smartphone models [46,49]. Even when two capture devices are compared, they are tested only under identical controlled lighting conditions [50]. Because many systems rely solely on colour, they can misclassify samples when surface colour does not represent the underlying product, for example, under coloured films, in retail display packaging, or when heavily seasoned [46,50].

3.3.4. Summary

Three techniques are consistently associated with high performance in RGB-based freshness studies. First, standardising image capture so that colour changes reflect product condition rather than lighting, achieved through dark or shielded enclosures, fixed geometry and consistent controlled lighting [46,49,50]. Second, using features based on perceptually relevant colour distributions, CIELAB or HSV histograms and Euclidean distance from a fresh reference, which capture subtle browning better than raw RGB averages [44], together with explicit colour-card correction [51]. Third, matching the model to the capture environment: shallower models suit well-controlled, static single-camera settings [43,45], whereas deeper models become advantageous as image sources and conditions vary [48]. Depth alone is not sufficient, however: a deep network in a controlled single-camera setting reached high aggregate accuracy but weak per-grade performance [52].
Across the literature, controlling capture geometry, calibrating colour and establishing explicit links to laboratory standards are the steps needed for RGB images to provide reliable information about freshness trends in perishable foods. Within those constraints, a light-assisted enclosure with CIELAB or HSV features and a compact regression or classification model is a practical route to a colour-based freshness indicator, with outcomes typically reported using R² and RMSE.

3.4. Light-Assisted Computer Vision

This technique encompasses studies that use light-assisted, visible-range RGB rigs without NIR or hyperspectral imaging (HSI). The common approach places the product within a hooded or enclosed area, with a fixed camera position and stable lamps to maintain imaging geometry and colour stability [53,54,55]. It uses simple RGB imaging and optical elements, such as diffusers or polarising filters, to reduce glare and produce repeatable results [53,54]. At the same time, systems may also use fixed cameras and lamps without an enclosure [56]. Table 3 summarises the selected studies.

3.4.1. Equipment and Capture Techniques

A consistent pattern across the light-assisted computer vision approach is a controlled environment, in which the product is placed in a hooded or enclosed area with a fixed camera and stable lamps to maintain imaging geometry and colour stability. The first sub-pattern is the use of enclosed cabinets for the slices of the product under investigation. Cernadas et al. [55] used a Canon EOS 50D camera in a black cabinet with eight calibrated halogen lamps to image standardised ham slices on a tray. The arrangement allowed automatic extraction of a square region of interest per muscle and computation of colour- and texture-based features from RGB and CIELAB images (Haralick co-occurrence statistics, local binary patterns, wavelet transforms and Gabor filters), followed by comparison of linear models, SVR, model trees, GBM and random forest regressors.
The second sub-pattern is the use of a portable enclosure for marbling. Cardenas et al. [54] built a compact stainless-steel enclosure with LED lighting and diffusers, integrated with a Raspberry Pi 4, an 8-megapixel camera module, and a small touchscreen. A user marked a section of a ribeye image. The image was converted to HSV, fat was segmented using histogram-based thresholding and cleaned up, and fat-area descriptors were fed into a linear SVM, k-NN, or random forest to predict a USDA marbling grade.
The use of smartphone imaging tubes characterises the third sub-pattern. For instance, Yu et al. [53] placed a smartphone with a fisheye lens inside a cylindrical device that fits the loin gap, together with a ring light and crossed polarisers to reduce reflected light. They used DeepLab with a ResNet backbone to segment ribeye images taken on the slaughter floor.
The fourth sub-pattern concerns fish and beef rigs, as well as simple dark boxes. Sano et al. [56] combined an RGB camera with a time-of-flight depth sensor in a research conveyor rig, processing whole fish and head, body and tail crops with a VGG-based regressor to predict fat content, while lighting and conveyor geometry remained constant. Rahman et al. [57] used a simple dark box in which a single camera captured images of beef, which were analysed for CIELAB values using a laboratory colourimeter, as well as pH, drip loss, composition, and oxidation. Handayani and Masruriyah [58] studied a small number of beef cuts with expert-marked fat and a simple RGB threshold rule under two lighting scenarios. These cabinet and enclosure designs form the hardware foundation for the reported performance, as reflected in the R² and RMSE values.

3.4.2. Findings

Cernadas et al. [55] reported that expert-annotated muscle outlines achieved a correlation of about 0.95 and an average miss of 0.38 score units. Fully automatic region determination yielded a correlation of about 0.84–0.85 and an increased average miss rate of about 0.60, while the region finder kept more than 90% of pixels in the correct ROI. The process ran at roughly 30 ms per image, thus allowing in-plant grading. Cardenas et al. [54] achieved about 95% test accuracy for USDA marbling grade across approximately 4,900 ribeye regions using the portable stainless-steel enclosure. The chosen linear SVM agreed with the specialist grade in about 89% of cases on an expert subset, indicating that a small RGB box with HSV thresholding and simple fat-area descriptors generally aligns with experienced graders.
For the smartphone tube approach, Yu et al. [53] produced highly stable muscle masks with about 99% segmentation accuracy and an intersection over union (IoU) of about 98.5% for the ribeye. For fish grading, Sano et al. [56] achieved a mean absolute error of about 2.25% of fat content with the help of simple dark boxes. About 84% of samples were within ±4% of an NIR reference. The study used RGB processing together with data from a time-of-flight depth sensor. It provided substantial benefits in terms of cost savings, complexity and speed. Rahman et al. [57] reported a strong relationship between photometric and colourimeter lightness (R² ≈ 0.69) and a moderate relationship between redness and pH and drip loss, indicating that some freshness indicators can be captured from simple RGB images under good lighting. Handayani and Masruriyah [58] showed that their RGB threshold rule matched expert-marked fat areas and remained consistent when extra lamps were added. Unfortunately, no chemical IMF references were provided to further validate the results.
Overall, light-blocked enclosures (cabinets, tubes, and bench boxes) using RGB cameras allowed feature-based and deep-segmentation approaches to produce R² near 0.9 for quantitative targets and classification rates above 90% for grade-level assessment when species, cut, and lighting were well controlled [53,54,55], whereas the open-conveyor rig, e.g., [56], achieved a lower value R² ≈ 0.69 for the fat content.

3.4.3. Advantages and Limitations

Light-assisted RGB rigs shift the difficulty from the model to the optical and mechanical design. When the illumination spectrum and intensity are constant, and the camera position is constrained by an enclosure, glare can be suppressed with diffuse or polarised light, and many colour and texture descriptors, sometimes even manually defined thresholds, suffice for accurate prediction of marbling and fat content. This greatly reduces computational complexity and enables high-speed processing, as shown by the 30 ms per-image runtime [55] and real-time grading [54]. The enclosed design also shields the scene from external lights and shadows, improving repeatability over time and across users. The same hardware can support multiple tasks: colour, pH, and related quality proxies can be derived from the same rig by changing the analysis pipeline and retraining on new labels [57]. Smartphone-based tubes and mini-boxes show that the rigs can be built around commodity sensors, implying lower cost and easier maintenance than HSI or X-ray systems [50,53].
Most systems are optimised for specific products, plants, or laboratory environments. They require recalibration or model retraining for different meat cuts, lighting conditions, or camera models [54,55,56]. In addition, the headline accuracy reported by Sano et al. [56] relies on a time-of-flight depth channel rather than RGB alone, so it is not a purely RGB result. Moreover, specific enclosures, lamps, and the need for short capture protocols raise costs and reduce portability, particularly for small shops and markets. Some studies use expert scores as the primary ground truth, with limited chemical IMF measurements, so agreement between the system and a human grader does not necessarily imply exact accuracy of fat-content measurement [54,58]. Finally, optical cameras observe only the visible surface, so subsurface fat cannot be seen. It is only moderately predictable from colour and texture alone [57].

3.4.4. Summary

In the domain of operations assisted only by visible light (i.e., with no NIR or HSI application), several best practices recur:
  • The use of a shielded capture space with static light sources and a fixed camera position to minimise variability in colour and shape.
  • Utilisation of optical control and cross-polarised lenses [53] with diffuse and matte-coloured backgrounds [53,55] to reduce glare and colour casts, particularly for wet meat products.
  • Good region-of-interest handling via an automatic ROI finder or accurate deep segmentation ensuring that the IMF estimation is based on the correct muscle region [53,55].
  • Implementation and application of simple fat descriptors with classic ML models (e.g., linear SVMs and tree-type ensembles) for the rigidly fixed rigs, and broader colour and texture descriptors with architectures for less stable environments [54].
Within the stated constraint, light-assisted enclosures (e.g., benchtop cabinets, portable stainless boxes, and smartphone tubes) enable estimation of meat marbling and grading with accuracy approaching that of trained graders and reference devices, provided species, cut, and lighting are controlled; performance is reported using R² and RMSE.

4. Discussion

4.1. Principal Findings

Across the three themes, the literature indicates that visible-range RGB imaging, when combined with controlled capture and suitable models, can approximate chemical IMF and freshness proxies, but only under stringent conditions [28,49,55]. For IMF and IMAT (marbling), feature-based pipelines built on CIELAB colour and visible fat area achieve R² of about 0.72–0.90, with IMF errors of roughly 0.2–1.3 percentage points, when imaging is standardised at fixed loin or rib locations [28,32,38]. Segmentation of beef, pork and steak images reaches an accuracy of about 96–99% [29,37,53], and image-based marbling regression on cured ham reaches correlations of about 0.84–0.95 against expert scores [55]. The use of specialised cameras and stainless-steel boxes allows grade-level accuracy near 95% with IMF errors of about 1.3 percentage points [36,54]. Smartphones used under a strict capture protocol with cross-polarisers and a colour-calibration card reach R² near 0.90 with errors near 0.9 percentage points [33], while a standardised smartphone setup achieves more moderate accuracy, with R² of 0.54–0.62 and RMSE of about 1.2–2.6 percentage points [17].
For freshness, the use of boxes equipped with controlled smartphones and calibrated rigs enables high accuracy when compared to laboratory targets. For example, SVR on lamb colour features reaches R² near 0.99 for pH and TVB-N [49]; random-forest and decision-tree models on IA-ΔE indices exceed 80–90% for TVC, TVB-N and pH and outperform colourimetric models [46]; and the use of a simple dark box supports beef shelf life and mesophilic and psychrotrophic counts, with R² of about 0.68–0.87 in roughly three seconds [50]. Simpler HSV pipelines reach about 82–100% for species-specific tasks [43,45], and deep networks reach about 90–95% for grade-level rating under controlled single-camera capture [52] and for fresh-versus-spoiled classification on heterogeneous public images [48]. For the light-assisted theme, enclosed cabinets, tubes, and bench boxes stabilise scenes sufficiently that both feature-based and deep-segmentation approaches reach R² near 0.9 and grade-level classification above 90% when species, cut and lighting are controlled [53,54,55], although the open-conveyor rig of Sano et al. [56] reached a lower R² ≈ 0.69 for fat content.
The literature therefore converges on a common pipeline: stabilised, light-assisted capture (a hood or enclosure, fixed geometry, cross-polarisation and colour references, with a standardised region of interest for IMF), followed by shallow models such as SVM, GBM or random forest when the rig is rigid, or deeper networks when scenes vary, and anchored by laboratory chemical or colour references as ground truth. Where studies disagree, the disagreements concern trade-offs between enclosure rigidity and shallow-versus-deep modelling, rather than the overall need for controlled, light-assisted capture and laboratory-grounded evaluation.

4.2. Research Gaps

Despite encouraging overall results in light-assisted, computer vision-based assessment of meat quality, several gaps recur across the literature:
  • No like-for-like comparison of capture modalities. Bench and industrial rigs are assessed separately from smartphone-based systems, and no study captures the same object images under both a controlled rig and a smartphone setup using a common feature-extraction and modelling pipeline for both IMF and freshness. Therefore, the true accuracy cost of moving from laboratory to phone capture has not been measured directly.
  • IMF and freshness are rarely modelled jointly. Almost all studies treat the two as separate problems: IMF studies report fat, marbling and grade but seldom freshness, while freshness studies report pH, TVB-N, TVC or colour change but seldom IMF. No visible-spectrum RGB pipeline yields both an IMF estimate and a calibrated freshness band from a single image, despite the colour and texture features that could be shared between them.
  • Few strictly RGB-only approaches target consumer deployment. Several of the most accurate systems use NIR, hyperspectral, FoodScan, or gas-sensor instruments, which are impractical for a phone at a wet market or small shop; the visible-only systems that exist are largely industrial or research prototypes with specialised cameras and fixed enclosures, while less-controlled smartphone systems lose accuracy and lack calibration to chemical or industrial standards.
  • Confidence indicators are largely overlooked. Models report goodness of fit (R², RMSE or accuracy) but rarely translate these metrics into calibrated probabilities or risk bands that non-expert users can act on. Point predictions of pH or TVB-N can diverge from visual colour, creating a mismatch between “looks acceptable” and “is microbiologically safe” that is not surfaced at the interface.
  • Real-world capture variability is underexplored. Performance is sensitive to ribbing site, moisture, bone dust, blooming time and cut geometry, and it drops for very lean carcasses or out-of-calibration inputs. Smartphone freshness studies are typically conducted under simulated laboratory conditions rather than in the mixed lighting of wet markets, and little work quantifies how segmentation errors or user protocol mistakes propagate or evaluates error handling, such as image-quality checks and recapture prompts.
  • No end-to-end pipeline spans laboratory reference to consumer use. Prior works focus on individual components rather than complete pipelines with common data formats, feature representations and models spanning laboratory reference rigs and consumer-facing capture devices, leaving a gap between industrial and consumer environments.

4.3. Future Research Directions

The above-outlined gaps suggest several potential directions for future work:
  • Direct cross-modality comparison, capturing the same cuts while using a controlled rig and with a smartphone, and processing both through a single shared pipeline, so the accuracy difference between ideal and real-world capture can be quantified rather than assumed.
  • Joint modelling of IMF and freshness, using shared colour and texture features with separate prediction heads so that a single capture can return both an IMF estimate and a calibrated freshness band.
  • Strictly visible-only, low-cost pipelines designed for consumer or small-vendor use and explicitly validated against a controlled RGB laboratory reference using shared metrics such as R² and RMSE.
  • Calibrated uncertainty and interpretable confidence, conveyed through prediction intervals for regression and calibrated probabilities for classification, were surfaced as simple risk indicators that distinguish reliable from unreliable estimates.
  • Evaluation under real-world capture variability and domain shift, including wet-market lighting, handling effects and user error, with robustness strategies such as colour correction, region-of-interest refinement and robust segmentation.
  • End-to-end pipelines that are fully specified and bridge a laboratory reference and a consumer-facing device, using common formats, features and models.

5. Conclusions

RGB-based meat imaging has been shown in numerous studies to provide reliable and consistent estimates of intramuscular fat and freshness, with a smaller body of work addressing intermuscular fat estimation. However, these studies consistently rely on controlled visible illumination, carefully calibrated camera angles and distances, and models trained on tightly controlled datasets [33,35,46,49]. Under these conditions, the reported R² and RMSE values generally fall within the strong-performance bands. Low-cost, portable, light-assisted boxes and tubes further demonstrate that comparable estimates of marbling grade and muscle segmentation can be obtained even with low-end cameras, including smartphones, provided that lighting and camera-to-sample distance are kept constant and basic calibration is performed [53,54]. Taken together, light-assisted RGB imaging can reliably estimate both IMF and freshness indicators, provided imaging protocols are strictly followed, and performance is monitored with metrics such as R² and RMSE.
In addition to demonstrating the high potential of light-assisted RGB imaging for meat quality assessment, this review highlights significant technical and practical limitations to its widespread adoption that must be addressed. Most published reports on estimating IMF through RGB imaging, using smartphone or compact-camera images captured under fairly standardised conditions, report rather expected results that are less accurate and less precise than those from high-control laboratory systems, with noticeably lower R² and higher RMSE [17,40]. However, there is an almost complete absence of literature comparing a well-controlled laboratory RGB system with a standardised smartphone-style system using the same models on exactly the same samples. The primary barrier to adoption is therefore not the ability to achieve high R² and low RMSE under ideal conditions, but rather how to translate customer-sufficient accuracy into a practical, sufficiently accurate, chemistry-free tool usable by small vendors and consumers at the point of purchase.
Future research should develop an integrated, uncertainty-aware pipeline in which laboratory reference rigs, smartphone-based capture, and user-facing interpretation and output assembly are explicitly linked. On the image-processing front, studies will need to quantify how much of the laboratory-versus-real-world gap in R² and RMSE can be closed using a simple foldable phone hood and a colour card, as well as to assess whether deep segmentation can be combined with fat-area descriptors across muscles and carcasses, with joint training on laboratory and smartphone images to improve generalisation [29,37]. On the modelling front, multi-task architectures that share feature extraction for IMF regression and freshness classification, together with calibration techniques that yield prediction intervals or traffic-light-style confidence indicators, deserve evaluation. On the evaluation front, substantial work remains to quantify how accuracy and calibration degrade under realistic variability in lighting, handling, and device type. In addition, it is important to define protocols for rapid recalibration when a new phone model or market environment is introduced, with results continuing to be reported using R² and RMSE for direct comparison.
In summary, in controlled settings, light-assisted RGB imaging can assess IMF and freshness, with results that closely align with reference estimates obtained using chemical and microbiological techniques. Smartphone-based systems can approach this performance when appropriate capture protocols and light-assisted fixtures are used. To date, however, no purely visual pipeline delivers IMF percentage, a freshness band and confidence information in a single workflow designed for consumers purchasing meat. Closing that gap, the development of a low-cost vision system that is explicitly validated against a controlled reference and communicates its own uncertainty remains the central open challenge for the field.

Author Contributions

Conceptualisation, I.J.S, S.D. and V.K; methodology, I.J.S. and S.D.; formal analysis, I.J.S.; investigation, I.J.S.; writing—original draft preparation, I.J.S.; writing—review and editing, I.J.S. and S.D.; supervision, S.D. and V.K.; project administration, S.D.; laboratory research support, V.K. All authors have read and agreed to the published version of the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Acronyms and Abbreviations

The following abbreviations are used in this manuscript:
ACM Association for Computing Machinery
ANN Artificial Neural Network
AUS-MEAT Authority for Uniform Specification of Meat and Livestock
BPNN Backpropagation Neural Network
CBAM Convolutional Block Attention Module
CH Calinski–Harabasz – an index used to evaluate cluster separation
CIELAB Commission Internationale de l'Éclairage L*a*b* - a device-independent colour space mapping Lightness [L*], Red-Green [a*], and Blue-Yellow [b*] coordinates
CNN Convolutional Neural Network
CRI Colour Rendering Index
CV Computer Vision
ΔE / IA-ΔE Delta E / Image Analysis Delta E - mathematical metrics representing perceived colour differences)
DenseNet Densely Connected Convolutional Networks
DNN Deep Neural Network
DT Decision Tree
EfficientNet A highly efficient convolutional neural network architecture scale-optimised for performance
EMA Eye Muscle Area
EN Elastic Net - a regularised regression method
ESP32-CAM A low-cost microcontroller development board with a built-in camera module
FCM Fuzzy C-Means - a clustering algorithm where pixels can belong to multiple groups with varying degrees of membership
GA-BP Genetic Algorithm Backpropagation - a neural network optimized using genetic algorithms
GAT Graph Attention Network
GBM Gradient Boosting Machine
GLCM Gray-Level Co-occurrence Matrix - a statistical texture analysis method examining pixel spatial relationships
HS / HSI Hyperspectral / Hyperspectral Imaging
HSV Hue, Saturation, Value - an intuitive colour space representation separating chroma and luminance
IEEE The Institute of Electrical and Electronics Engineers
IMAT Intermuscular fat that is often identified as an Intermuscular Adipose Tissue
IMF Intramuscular Fat - fine marbling flecks dispersed within muscle tissue, reflecting meat tenderness, juiciness, and quality
IoU Intersection over Union - an evaluation metric measuring the accuracy of an automated image segmentation mask
JPEG Joint Photographic Experts Group’s compressed standard digital image format
KL Kullback–Leibler - a mathematical divergence metric used to measure differences between two distributions
KNN K-Nearest Neighbours - a proximity-based instance classification algorithm
LBP Local Binary Patterns - a visual texture descriptor comparing local pixel values
LDA Linear Discriminant Analysis
LMK Lichtmesstechnik - German for 'Light Measurement Technology,' referring to matrix-based luminance meters
LR Linear Regression
LT Longissimus thoracis - the thoracic portion of the ribeye/loin muscle
LTL Longissimus thoracis et lumborum - the full anatomical name of the primary loin muscle in pork and beef
MAE Mean Absolute Error
MAP Modified Atmosphere Packaging - a preservation method altering gas ratios inside packaging to slow microbial spoilage
MDA Mixture Discriminant Analysis
MIJ Meat Image Japan - a commercial brand of high-precision carcass cameras used for visual grading
ML Machine Learning
MLR Multiple Linear Regression
MP Megapixel
MSA Meat Standards Australia - an industry-accepted visual grading scale for red meat quality
NIR Near-Infrared - a wider-spectrum optical range commonly used as a reference benchmark
P / R / F1 Precision / Recall / F1-score - standard validation metrics used to assess categorical classifiers
PICOC Population, Intervention, Comparison, Outcome, Context - a structured framework defining systematic review boundaries
PLS Partial Least Squares - a regression technique optimised for highly correlated variable sets
pp Percentage points - the arithmetic difference between two percentages, e.g., 2% to 3% is a 1 pp increase
PSO-K-Means Particle Swarm Optimisation K-Means - a hybrid clustering algorithm combining swarm intelligence with distance metrics
/ R2 Coefficient of Determination - a statistical metric representing the proportion of variance explained by a model
RAW An uncompressed, unprocessed image format containing direct digital sensor data
re-ID / ID Re-identification / Identification - the physical tracking of carcasses or cuts across different images
ResNet Residual Network - a deep convolutional neural network using skip connections to combat vanishing gradients
RF Random Forest - an ensemble machine learning model built from numerous discrete decision trees
RGB Red, Green, Blue - the standard primary colour model utilised in conventional digital imaging
RMSE Root Mean Square Error – a standard metric measuring average predictive error
RMSECV Root Mean Square Error of Cross-Validation
RMSEP Root Mean Square Error of Prediction
ROI Region of Interest
SDG / SDGs Sustainable Development Goal(s) - United Nations global objectives for human development and sustainability
SPIE International Society for Optics and Photonics
SPLS Sparse Partial Least Squares
SR Stepwise Regression
SVM / SVR Support Vector Machine / Support Vector Regression
TVB-N Total Volatile Basic Nitrogen - a chemical indicator showing the degradation of proteins as an objective marker of freshness
TVC Total Viable Count - a microbiological metric showing the total number of living microorganisms present on a food sample
U-Net / U2-Net Highly recognised convolutional neural network architectures optimised for image segmentation Tasks
USDA United States Department of Agriculture (among many regulations defining the official beef quality grading scales
VGG Visual Geometry Group - a classic, deep convolutional neural network architecture developed by Oxford's VGG group

References

  1. Van Eenennaam, A. L. Addressing the 2050 demand for terrestrial animal source food. Proc. Natl. Acad. Sci. 2024, 121(50), e2319001121. [Google Scholar] [CrossRef] [PubMed]
  2. Gil, M.; Rudy, M.; Duma-Kocan, P.; Stanisławczyk, R.; Krajewska, A.; Dziki, D.; Hassoon, W. H. Sustainability of alternatives to animal protein sources: A comprehensive review. Sustainability 2024, 16(17), 7701. [Google Scholar] [CrossRef]
  3. Hegarty, R. S.; Tee, T. P.; Liang, J. B.; Abu Hassim, H.; Zainudin, M. H. M.; Azizi, A. A.; Widiawati, Y.; Pok, S.; Candyrine, S. C. L.; Rusli, N. D. Balancing future food security and greenhouse-gas emissions from animal-sourced protein foods in Southeast Asia. Anim. Prod. Sci. 2024, 64(18), AN24183. [Google Scholar] [CrossRef]
  4. Spiro, A.; Hill, Z.; Stanner, S. Meat and the future of sustainable diets—Challenges and opportunities. Nutr. Bull. 2024, 49(4), 572–598. [Google Scholar] [CrossRef] [PubMed]
  5. Nobre, F. S. Cultured meat and the sustainable development goals. Trends Food Sci. Technol. 2022, 124, 140–153. [Google Scholar] [CrossRef]
  6. Magqupu, S.; Chikwanha, O. C.; Katiyatiya, C. L. F.; Strydom, P. E.; Mapiye, C. Consumers’ purchasing behaviour and quality preferences for pork sold in the informal street markets of the Cape Metropole, South Africa. Agrekon 2024, 63(3), 113–132. [Google Scholar] [CrossRef]
  7. Jordaan, D.; Mielmann, A.; Brits, C. Consumers’ perceived value of pork meat: A segmentation on intrinsic and extrinsic cues. Foods 2025, 14(13), 2324. [Google Scholar] [CrossRef] [PubMed]
  8. Santos, D.; Monteiro, M. J.; Voss, H.-P.; Komora, N.; Teixeira, P.; Pintado, M. The most important attributes of beef sensory quality and production variables that can affect it: A review. Livest. Sci. 2021, 250, 104573. [Google Scholar] [CrossRef]
  9. Benli, H.; Gecgel Yildiz, D. Consumer perception of marbling and beef quality during purchase and consumer preferences for degree of doneness. Anim. Biosci. 2023, 36(8), 1274–1284. [Google Scholar] [CrossRef] [PubMed]
  10. Thies, A. J.; Altmann, B. A.; Countryman, A. M.; Smith, C.; Nair, M. N. Consumer willingness to pay (WTP) for beef based on color and price discounts. Meat Sci. 2024, 217, 109597. [Google Scholar] [CrossRef] [PubMed]
  11. Pandey, S.; Kim, S.; Kim, E. S.; Keum, G. B.; Doo, H.; Kwak, J.; Ryu, S.; Choi, Y.; Kang, J.; Kim, H.; Chae, Y.; Seol, K.-H.; Kang, S. M.; Kim, Y.; Seong, P. N.; Bae, I.-S.; Cho, S.-H.; Jung, S.; Kim, H. B. Exploring the multifaceted factors affecting pork meat quality. J. Anim. Sci. Technol. 2024, 66(5), 863–875. [Google Scholar] [CrossRef] [PubMed]
  12. Tomašević, I.; Đekić, I.; Font-i-Furnols, M.; Terjung, N.; Lorenzo, J. M. Recent advances in meat color research. Curr. Opin. Food Sci. 2021, 41, 81–87. [Google Scholar] [CrossRef]
  13. Liu, J.; Ellies-Oury, M.-P.; Stoyanchev, T.; Hocquette, J.-F. Consumer perception of beef quality and how to control, improve and predict it? Focus on eating quality. Foods 2022, 11(12), 1732. [Google Scholar] [CrossRef] [PubMed]
  14. Altmann, B. A.; Trinks, A.; Mörlein, D. Consumer preferences for the color of unprocessed animal foods. J. Food Sci. 2023, 88(3), 909–925. [Google Scholar] [CrossRef]
  15. Hewavitharana, G. G.; Perera, D. N.; Navaratne, S. B.; Wickramasinghe, I. Extraction methods of fat from food samples and preparation of fatty acid methyl esters for gas chromatography: A review. Arab. J. Chem. 2020, 13(8), 6865–6875. [Google Scholar] [CrossRef]
  16. Suzuki, Y.; Yue, B. Grading evaluation of marbling in Wagyu beef using fractal analysis. Eng 2024, 5(3), 2157–2169. [Google Scholar] [CrossRef]
  17. Menezes, G. L.; Valente Junior, D. T.; Ferreira, R. E. P.; Oliveira, D. A. B.; Araujo, J. A.; Duarte, M.; Dorea, J. R. R. Empowering informed choices: How computer vision can assist consumers in making decisions about meat quality. Meat Sci. 2025, 219, 109675. [Google Scholar] [CrossRef] [PubMed]
  18. Lin, H.-D.; Hsieh, Y.-T.; Lin, C.-H. Smartphone-based sensing system for identifying artificially marbled beef using texture and color analysis to enhance food safety. Sensors 2025, 25(14), 4440. [Google Scholar] [CrossRef] [PubMed]
  19. Jeong, K.; Jo, G.; Lee, J. H.; Kim, Y. H. B.; Choi, J.; Oh, H.; Jeong, J.-H.; Lee, E. Artificial intelligence in meat processing: A comprehensive review of data-driven applications and future directions. Meat Muscle Biol. 2025, 9(1), 20157. [Google Scholar] [CrossRef]
  20. Dhal, S. B.; Kar, D. Leveraging artificial intelligence and advanced food processing techniques for enhanced food safety, quality, and security: A comprehensive review. Discov. Appl. Sci. 2025, 7, 75. [Google Scholar] [CrossRef]
  21. Kitchenham, B.; Charters, S. Guidelines for performing systematic literature reviews in software engineering (EBSE Technical Report EBSE-2007-01). Keele University and Durham University. 2007.
  22. Carrera-Rivera, A.; Ochoa, W.; Larrinaga, F.; Lasa, G. How-to conduct a systematic literature review: A quick guide for computer science research. MethodsX 2022, 9, 101895. [Google Scholar] [CrossRef] [PubMed]
  23. Petersen, K.; Vakkalanka, S.; Kuzniarz, L. Guidelines for conducting systematic mapping studies in software engineering: An update. Inf. Softw. Technol. 2015, 64, 1–18. [Google Scholar] [CrossRef]
  24. Chicco, D.; Warrens, M. J.; Jurman, G. The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. PeerJ Comput. Sci. 2021, 7, e623. [Google Scholar] [CrossRef] [PubMed]
  25. James, G.; Witten, D.; Hastie, T.; Tibshirani, R.; Taylor, J. An introduction to statistical learning: With applications in Python; Springer, 2023. [Google Scholar] [CrossRef]
  26. Hodson, T. O. Root-mean-square error (RMSE) or mean absolute error (MAE): When to use them or not. Geosci. Model Dev. 2022, 15(14), 5481–5487. [Google Scholar] [CrossRef]
  27. Miller, C.; Portlock, T.; Nyaga, D. M.; O’Sullivan, J. M. A review of model evaluation metrics for machine learning in genetics and genomics. Front. Bioinform. 2024, 4, 1457619. [Google Scholar] [CrossRef] [PubMed]
  28. Teixeira, A.; Silva, S. R.; Hasse, M.; Almeida, J. M. H.; Dias, L. Intramuscular fat prediction using color and image analysis of Bísaro pork breed. Foods 2021, 10(1), 143. [Google Scholar] [CrossRef] [PubMed]
  29. Symeonidis, G.; Kiourt, C.; Kazakis, N. A.; Nerantzis, E.; Nestor, T. Fat calculation from raw-beef-steak images through machine learning approaches: An end-to-end pipeline. In Proceedings of the 26th Pan-Hellenic Conference on Informatics (PCI ’22); ACM, 2022; pp. 110–115. [Google Scholar] [CrossRef]
  30. Handayani, H. H.; Madenda, S.; Wibowo, E. P.; Kusuma, T. M.; Widiyanto, S.; Masruriyah, A. F. N. The best classification algorithm for identification beef quality based on marbling. In 2020 Fifth International Conference on Informatics and Computing (ICIC); IEEE, 2020a; pp. 63–67. [Google Scholar] [CrossRef]
  31. Chen, D.; Wu, P.; Wang, K.; Wang, S.; Ji, X.; Shen, Q.; Yu, Y.; Qiu, X.; Xu, X.; Liu, Y.; Tang, G. Combining computer vision score and conventional meat quality traits to estimate the intramuscular fat content using machine learning in pigs. Meat Sci. 2022, 185, 108727. [Google Scholar] [CrossRef] [PubMed]
  32. Uttaro, B.; Zawadski, S.; Larsen, I.; Juárez, M. An image analysis approach to identification and measurement of marbling in the intact pork loin. Meat Sci. 2021, 179, 108549. [Google Scholar] [CrossRef] [PubMed]
  33. Meunier, B.; Normand, J.; Albouy-Kissi, B.; Micol, D.; El Jabri, M.; Bonnet, M. An open-access computer image analysis (CIA) method to predict meat and fat content from an android smartphone-derived picture of the bovine 5th–6th rib. Methods 2021, 186, 79–89. [Google Scholar] [CrossRef] [PubMed]
  34. Pannier, L.; van de Weijer, T. M.; van der Steen, F. T. H. J.; Kranenbarg, R.; Gardner, G. E. Adding value to beef portion steaks through measuring individual marbling. Meat Sci. 2023a, 204, 109279. [Google Scholar] [CrossRef] [PubMed]
  35. Pannier, L.; van de Weijer, T. M.; van der Steen, F. T. H. J.; Kranenbarg, R.; Gardner, G. E. Prediction of chemical intramuscular fat and visual marbling scores with a conveyor vision scanner system on beef portion steaks. Meat Sci. 2023b, 199, 109141. [Google Scholar] [CrossRef] [PubMed]
  36. Drachmann, F. F.; Christensen, M.; Esberg, J.; Lauridsen, T.; Fogh, A.; Young, J. F.; Therkildsen, M. Beef-on-dairy: Meat quality of veal and prediction of intramuscular fat using the Q-FOM™ Beef camera at the 5th–6th thoracic vertebra. Meat Sci. 2024, 213, 109503. [Google Scholar] [CrossRef] [PubMed]
  37. Liu, H.; Zhan, W.; Du, Z.; Xiong, M.; Han, T.; Wang, P.; Li, W.; Sun, Y. Prediction of the intramuscular fat content of pork cuts by improved U2-Net model and clustering algorithm. Food Biosci. 2023, 53, 102848. [Google Scholar] [CrossRef]
  38. Wachta, H.; Tereszkiewicz, K.; Kulig, Ł. Luminance surface distribution measurements applied to assessing intramuscular fat content in meat. Measurement 2022, 193, 110846. [Google Scholar] [CrossRef]
  39. Huang, M.; Wang, L.; Wang, B.; Jiang, W.; Yu, Y.; Tang, Q.; Gao, Q.; Tian, Y. Integrating computer vision and machine learning technologies for model building to quantify intermuscular fat content in salmonid fillets. Food Control 2025, 175, 111293. [Google Scholar] [CrossRef]
  40. Stewart, S. M.; Gardner, G. E.; Williams, A.; Pethick, D. W.; McGilchrist, P.; Kuchida, K. Association between visual marbling score and chemical intramuscular fat with camera marbling percentage in Australian beef carcasses. Meat Sci. 2021, 181, 108369. [Google Scholar] [CrossRef] [PubMed]
  41. Liao, Q.; Gardner, B.; Barlow, R.; McMillan, K.; Moore, S.; Fitzgerald, A.; Arzhaeva, Y.; Botwright, N.; Wang, D.; Nelis, J. L. D. Improving traceability and quality control in the red-meat industry through computer vision-driven physical meat feature tracking. Food Chem. 2025, 480, 143830. [Google Scholar] [CrossRef] [PubMed]
  42. Zhang, R.; Shi, W.; Pan, Y.; Zhao, Y.; Hua, Z.; Zhao, Y.; Pan, Q.; Han, Z. Beef quality dual-label classification incorporating texture and multihead map attention mechanisms. J. Food Compos. Anal. 2025, 141, 107360. [Google Scholar] [CrossRef]
  43. Putra, I. G. S. E.; Putra, I. K. G. D.; Sudarma, M.; Sudana, A. A. K. O. Classification of tuna meat grade quality based on color space using wavelet and k-nearest neighbor algorithm. In 2023 International Conference on Smart-Green Technology in Electrical and Information Systems (ICSGTEIS); IEEE, 2023; pp. 35–40. [Google Scholar] [CrossRef]
  44. Meza, G.; Sánchez, C. N.; Orvañanos-Guerrero, M. T.; Domínguez-Soberanes, J. Analysis of meat color change using computer vision. In 2020 IEEE International Autumn Meeting on Power, Electronics and Computing (ROPEC); IEEE, 2020. [Google Scholar] [CrossRef]
  45. Bahri, M. I.; Anantama, R.; Bachtiar, M. I. Mackerel tuna freshness detection system using RGB color space image processing. IOP Conf. Ser. Earth Environ. Sci. 2025, 1454(1), 012021. [Google Scholar] [CrossRef]
  46. Cheng, H.; Li, J.; Yang, Y.; Zhou, G.; Xu, B.; Yang, L. Identifying freshness of various chilled pork cuts using rapid imaging analysis. J. Sci. Food Agric. 2025, 105(2), 747–759. [Google Scholar] [CrossRef] [PubMed]
  47. Bhuiyan, Z. W.; Haider, S. A. R.; Haque, A.; Uddin, M. R.; Hasan, M. IoT-based meat freshness classification using deep learning. IEEE Access 2024, 12, 196047–196069. [Google Scholar] [CrossRef]
  48. Sagiraju, B.; Casanova, N.; Chun, L.; Lohia, M.; Yoshiyasu, T. Meat freshness prediction. arXiv 2023, arXiv:2305.00986. [Google Scholar] [CrossRef]
  49. Jin, P.; Li, Z.; Zhang, X. Freshness prediction of modified atmosphere packaging lamb meat based on digital images from mobile portable devices. J. Food Process Eng. 2023, 46(12), e14444. [Google Scholar] [CrossRef]
  50. Pereira, L. M.; Lins, R. G.; Gaspar, R. Camera-based system for quality assessment of fresh beef based on image analysis. Meas. Food 2022, 5, 100013. [Google Scholar] [CrossRef]
  51. You, M.; Liu, J.; Zhang, J.; Xv, M.; He, D. A novel chicken meat quality evaluation method based on color card localization and color correction. IEEE Access 2020, 8, 170093–170100. [Google Scholar] [CrossRef]
  52. Tan, W. K.; Husin, Z.; Ismail, M. A. H. Feasibility study of beef quality assessment using computer vision and deep neural network (DNN) algorithm. In 2020 8th International Conference on Information Technology and Multimedia (ICIMU); IEEE, 2020; pp. 243–246. [Google Scholar] [CrossRef]
  53. Yu, H.; Lim, J.; Seo, Y.; Lee, A. Compact imaging system and deep learning based segmentation for objective longissimus muscle area in Korean beef carcass. Meat Sci. 2023, 206, 109325. [Google Scholar] [CrossRef] [PubMed]
  54. Cardenas, E.; Tabory, E.; Sanchez, A.; Kemper, G. An electronic equipment for marbling meat grade detection based on digital image processing and support vector machine. J. Saudi Soc. Agric. Sci. 2024, 23(7), 459–473. [Google Scholar] [CrossRef]
  55. Cernadas, E.; Fernández-Delgado, M.; Fulladosa, E.; Muñoz, I. Automatic marbling prediction of sliced dry-cured ham using image segmentation, texture analysis and regression. Expert Syst. With Appl. 2022, 206, 117765. [Google Scholar] [CrossRef]
  56. Sano, S.; Miyazaki, T.; Sugaya, Y.; Sekiguchi, N.; Omachi, S. Mackerel fat content estimation using RGB and depth images. IEEE Access 2021, 9, 164060–164069. [Google Scholar] [CrossRef]
  57. Rahman, M. F.; Iqbal, A.; Hashem, M. A.; Adedeji, A. A. Quality assessment of beef using computer vision technology. Food Sci. Anim. Resour. 2020, 40(6), 896–907. [Google Scholar] [CrossRef] [PubMed]
  58. Handayani, H. H.; Masruriyah, A. F. N. Determination of beef marbling based on fat percentage for meat quality. Int. J. Psychosoc. Rehabil. 2020b, 24(7), 3496–3502. [Google Scholar]
  59. Page, M. J.; McKenzie, J. E.; Bossuyt, P. M.; Boutron, I.; Hoffmann, T. C.; Mulrow, C. D.; Shamseer, L.; Tetzlaff, J. M.; Akl, E. A.; Brennan, S. E.; Chou, R.; Glanville, J.; Grimshaw, J. M.; Hróbjartsson, A.; Lalu, M. M.; Li, T.; Loder, E. W.; Mayo-Wilson, E.; McDonald, S.; …; Moher, D. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
Table 1. Comparison of RGB-based studies estimating intramuscular fat (IMF), marbling and intermuscular fat.
Table 1. Comparison of RGB-based studies estimating intramuscular fat (IMF), marbling and intermuscular fat.
Reference Imaging setup Model(s) Key results Main limitations
[28] Olympus EM-5, twin macro flash, cross-polarised, black background; ImageJ marbling area% + CIELAB; Bisaro pork loin (longissimus thoracis et lumborum - LTL) Mixture Discriminant Analysis - MDA (3 IMF groups); polynomial support vector machine (SVM) regression ~92% mean classification (best 100%); IMF% R2 ~0.88, RMSE ~0.18 pp 20 carcasses, single breed; narrow IMF (0.6-2.1%); controlled lighting only; manual selection of the region of interest (ROI)
[29] Single RGB photo; ruler marker for scale; butcher/web images Two-stage U-Net (VGG-16 convolutional neural network developed for image segmentation): steak then fat segmentation Steak segmentation ~92-97%; small fat error; area->mass (~125 g example) Small custom dataset; assumes fixed thickness/fat density; lighting/phone-dependent; not validated at scale
[30] Digital camera (5184x3456), auto-crop; 7 marbling features; beef Decision tree algorithm (vs linear discriminant analysis - LDA) ~91% marbling classification (SVM ~71%, LDA ~61%); precision ~0.87 Very small dataset (9 samples, 4 cuts); hand-crafted features; no smartphone/lighting tests
[31] Black-lined compartment, high-speed scanner; CV-IMF score + conventional traits; chilled pork loin Stepwise regression vs gradient boosting machine - GBM (best) Image-only r ~0.68; GBM ~89% within +/-0.5 IMF, R2 ~0.66; traits improve fit Needs controlled box + instruments; human density values; 2D approximates 3D; breed/site-specific
[32] Canon EOS 550D, cross-polarised halogen, manual ROI; JPEG+RAW, ImageJ; pork loin Stepwise regression (multi-view) Mid-loin r ~0.86; multi-view R2 ~0.76, error ~0.39 pp; JPEG ~ RAW Controlled lighting + manual setup; surface cues only; less stable at high marbling; no smartphone test
[33] Galaxy S8, cross-polarisers, 5x5 cm calibration card; ImageJ macro (17 features); slaughterhouse Spars partial least squares - SPLS (best) vs RF, clustered-MLR Table 6th-rib IMF% R2 ~0.90, error ~0.9 pp; fat/lean R2 ~0.72-0.86 Strict protocol (light, clean cut, coloured table); operator input; slaughterhouse-only validation
[34] Marel RGB conveyor scanner; marbling score vs chemical IMF% at 3 loin positions; beef Marel vision algorithm Large within-roll variation (~316 MSA units); grading-site vs lab IMF r ~0.93 Conveyor/industrial; cut/position-specific; single-site grading misrepresents variation
[35] Marel conveyor (RGB + LED line light), one image/steak; fresh-cut 15 mm; beef In-house segmentation (Mb1-Mb6; Mb4 best) IMF R2 ~0.87, error ~1.16 pp; MSA R2 ~0.82; AUS-MEAT R2 ~0.79 Needs fresh-cut surface; fixed conveyor rig; grader noise at high marbling
[36] Handheld Q-FOM Beef (RGB+3D), in-chiller, 5th-6th thoracic vertebra; DL LT segmentation; multi-breed PLS Overall R2 ~0.91, error ~1.33%; veal subset R2 ~0.48; 3-band ~63% Proprietary hardware; site-specific protocol; weaker at narrow/low IMF; Danish plants
[37] 2,000 retail images, 5 pork cuts; U2-Net (a two-level nested U-structure architecture that is designed for salient object detection - SOD) + CBAM background removal; smartphone/retail PSO-K-Means (fat/muscle) Segmentation ~98-99%; IMF ~93% accurate; faster than K-Means/FCM Single local dataset; indirect clustering; fine marbling hard; needs multi-phone/lighting tests
[38] LMK luminance meter (Canon) in shadowless tent, 2 LEDs; RAW->luminance; pork longissimus Threshold luminance count Image vs lab mean: 1.68% vs 1.67%; r ~0.79 Dedicated tent + luminance meter; dried surfaces; narrow pork range; not phone-friendly
[39] Imaging box, shadowless lamps, Nikon Z5 + 50 mm macro; fat% + 16 colour features; rainbow trout RF (best) vs DNN, EN, SR RF R2 ~0.91, ~79% within band; DNN 0.90; MAE ~1.5-2.0 Consistent lighting needed; colour-dependent; trout-specific; needs varied images
[17] iPhone XR, one photo/steak, consistent room light, 30 min bloom; beef LTL, pork loin Xception (end-to-end); RF/SVR/ANN Beef tenderness ~76.6%; IMF R2 ~0.62, RMSE ~2.6%; pork IMF R2 ~0.54, RMSE ~1.22% Controlled indoor only; pork IMF hard in RGB; re-tuning needed across phones/breeds
[40] Handheld MIJ cameras (MIJ-mirror, MIJ-30), abattoir, rib-eye after bloom; Beef Analyser II Models across 11 datasets, leave-one-out CV IMF R2 ~0.4-0.5, error ~1.5-1.6 pp; EMA R2 ~0.6-0.7; marbling weaker Varies by plant/cut/surface; bright spots/moisture mislead; marbling% misses’ distribution
[41] Studio light box, 602 striploin steaks (38,528 images) + phone/GoPro; U-Net segmentation EfficientNet-B0 (best) + re-ID Marbling ~96% within +/-1, ~99.6% within +/-2; breed ~91%; diet ~91%; re-ID F1 ~0.994 Specific setup/curated data; 1-10 grade classes mask fine differences; Australian cuts; recalibration needed
[42] RGB brisket/ribeye/sirloin 224x224, augmented; cut + grade (A1-A5) MS’ 50-layer convolutional neural network ResNet50 + Grey-level co-occurrence matrix GLCM + multi-head graph attention network GAT (best) Cut ~93.5%; grade ~92.3% (ResNet50 alone ~85.7%/85.0%) Small augmented single-source; local grading; no IMF% ground truth; controlled setting; shop lighting unproven
Source: Authors’ compilation. Alt text: five-column table listing sixteen studies with their imaging setup, models, key accuracy results and main limitations.
Table 2. Comparison of RGB-based studies estimating the freshness of meat or fish from colour cues.
Table 2. Comparison of RGB-based studies estimating the freshness of meat or fish from colour cues.
Reference Imaging setup Model(s) Key results Main limitations
[43] Camtech CT50 webcam, centre-crop; RGB->HSV, level-5 Symlet wavelet; tuna K-nearest neighbours (KNN); Symlet vs Haar ~82% overall; Symlet 81.8% vs Haar 80.3% Small local webcam dataset; colour-only; KNN may not generalise; labels/lighting unaudited
[44] 5 beef cuts, 3 timepoints over 9 days (refrigerated); meat segmented -> CIELab Average Euclidean distance (ΔE) vs Kullback Leibler (KL) divergence on colour histograms KL divergence separates mid/ late storage better than ΔE Very small; uncontrolled lighting; beef only; no chemical gold standard; trends not a calibrated score
[45] Tuna-eye photos; RGB/HSV colour features (source inconsistent on method) Unspecified trained model (not named) 100% on 8-image test set (50 total) Very small; classifier unspecified; source internally inconsistent; species/eye-specific; likely overfit
[46] Smartphone, 4 pork cuts at 4 C 0-7 d, consistent light; Lab -> IA-ΔE vs TVC/TVB-N/pH; vs colorimeter LR/KNN/SVR/BPNN/RF/DT (DT best) RF often >80%, DT >90%; phone-image beat colorimeter Lab-style lighting; standardised minced blocks; one source; colour-only; needs multi-phone/store tests
[47] IoT ESP32-CAM + MQ gas sensors + Raspberry Pi; species + freshness; 9,928 images at 0/48 h Custom CNN (best) vs ResNet-50, SVM, KNN CNN ~99% both tasks; ResNet-50 ~98%; SVM ~96-97% Two timepoints, one setup; no chemical labels; beef/mutton; overfitting risk; gas-sensor drift
[48] Kaggle Meat Freshness dataset (~1,816/452; red meat); Fresh/Half-fresh/Spoiled ResNet-18/50; U-Net->DenseNet (ResNet-50 deployed) ResNet-18 ~93%; ResNet-50 lowest business cost; U-Net route ~35% Cost-model assumptions; small red-meat-only; neat photos; no chemical ground truth
[49] Smartphone (Vivo S7), lamb 0-10 d under MAP/air, dark box, fixed light; RGB/Lab/HSV features SVR (best) vs GA-BP, CNN SVR (12 features) R2 up to 0.99, low error for pH and TVB-N Controlled mini-studio, one phone; single-source lamb-only; colour-only proxies; real-world untested
[50] Dark box rig, Samsung S5 + USB cam, ColorChecker-calibrated; RGB distance to fresh reference Regression (fitted equations) Phone/USB ~5% error vs colour meter; R2 ~0.68-0.87; ~3 s/image Needs a box + card; beef longissimus 4 C, small counts; colour-only; predicts risk not exact time
[51] Smartphone (Xiaomi MIX 2S), 12-patch colour card by raw chicken; card localisation + colour correction k-means (3 clusters, CH index) Corrected colours matched card; 3-level grouping clearest; card-finder robust to tilt/glare Needs printed card + protocol; chicken only ~24 h; colour proxy; unsupervised, no labels
[52] Logitech C905 webcam + LED, white-balanced; 400 rib-eye, Meat Standards Australia (MSA) colour-card labels Inception-V4 ~90% overall; train/val ~99/96%; per-class P/R/F1 ~62.8/61.3/60.4% Single camera/controlled lighting; colour-only; overfitting signs; phones/ lights untested
Source: Authors’ compilation. Alt text: five-column table listing ten studies with their imaging setup, models, key accuracy results and main limitations.
Table 3. Comparison of light-assisted visible-range RGB rigs for marbling, fat and quality assessment.
Table 3. Comparison of light-assisted visible-range RGB rigs for marbling, fat and quality assessment.
Reference Imaging setup Model(s) Key results Main limitations
[55] Canon EOS 50D in a black cabinet, calibrated halogen; 180 hams (714 slice images); auto square-ROI per muscle; RGB/Lab + Haralick/LBP/wavelet/Gabor features Linear, SVR, M5/Cubist, GBM, RF (RF/SVR best) Expert ROI correlation ~0.95, miss ~0.38; auto ROI ~0.84-0.85, miss ~0.60; ~30 ms/image Purpose-built cabinet; dry-cured ham slices/specific muscles; 2D approximates 3D; one acquisition setup
[56] Research conveyor rig; RGB + Time-of-Flight depth; whole-fish + head/body/tail crops; NIR ground truth VGG16 features + regression network MAE ~2.25% fat, RMSE <3, r ~0.83 (R2 ~0.69) vs NIR; ~84% within +/-4% Research conveyor rig (RGB + depth); whole fish (mackerel) not cuts; errors up to ~12%; accuracy not RGB-only
[53] Portable smartphone imaging tube (fisheye, ring light, cross-polarisers, light-blocking housing); 856 carcasses; rib-eye DeepLab-ResNet50 (best) vs FCN-8s, UNet, SegNet ~99% segmentation accuracy; Intersection over Union (IoU) ~98.5%; clean edges Purpose-built device, inserted into carcass gap; Korean beef, one setting; segmentation only, no fat% yet
[54] Portable stainless enclosure, LED, Raspberry Pi 4 + 8 MP camera + touchscreen; manual rib-eye ROI; HSV thresholding; fat-area descriptors Linear SVM (chosen), KNN, RF ~95% USDA grade on ~4,900 ROIs; ~89% vs expert subset Needs enclosure + protocol; manual ROI; judged mainly vs experts (limited chemical IMF); rib-eye context
[57] Canon IXUS in a dark box, fixed lighting, black background; beef longissimus 24 h; MATLAB L*a*b* features; lab ground truth Calibration/regression (Unscrambler) Lightness calib R2 ~0.73, pred ~0.69; redness/pH/drip moderate ~0.34-0.60; yellowness weak Small single-site; one camera; surface colour only; many traits weak; in-store/phones untested
[58] Simple bench, digital camera, two lighting scenarios; expert-marked fat; rule-based RGB threshold; 16-block fat-area% Rule-based threshold (no ML) Rule matched expert marks; ~35% fat area example; held with extra lamp Very small; no chemical IMF ground truth; thresholds hand-tuned; limited lighting/cuts; no smartphone test
Source: Authors’ compilation. Alt text: five-column table listing six light-assisted RGB studies with their imaging setup, models, key results and main limitations.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.