Submitted:
31 August 2026
Posted:
31 August 2026
You are already at the latest version
Abstract
Object detection has evolved dramatically over three decades, transitioning from handcrafted features to deep learning, transformers, and now large language models. This paper provides a comprehensive systematic review of this evolution, covering traditional methods (SIFT, HOG, Viola-Jones), CNN-based single- and two-stage detectors (R-CNN family, YOLO v1–v12, SSD), transformer-based architectures (DETR, ViT, Swin), and emerging LLM-driven approaches (GPT-4V, Grounding DINO, YOLO-World). We benchmark and compare these methods using key metrics—mAP, IoU, FPS—and analyze over 50 datasets across generic, autonomous driving, medical, remote sensing, and agricultural domains. Our critical synthesis identifies persistent research gaps: lightweight edge deployment, sensor fusion for real-time systems, bias and fairness in AI, and multimodal vision-language integration. By consolidating three decades of progress and outlining future challenges, this survey serves as a valuable resource for researchers and practitioners developing accurate, efficient, and reliable object detection systems.
Keywords:
computer vision
; object detection
; image processing
; deep learning
; YOLO algorithm
; ResNet
1. Introduction
Computer Vision (CV) is a branch of artificial intelligence that involves building machines or systems capable of interpreting visual data. Technically, this means giving the ability to extract meaningful information from visual data, that is, images and videos, to computer systems. Image segmentation, image classification, object tracking, and object detection are some of the most frequently used methods in CV that enable computers to process and understand visual data [1,2].
Object Detection (OD) in CV is a supervised learning task that involves detecting and localizing instances of predefined object classes in digital images and videos [1,2]. Detecting individual instances refers to the ability to identify specific objects within an image, while locating objects is the ability to accurately indicate the exact location of the objects by drawing bounding boxes around them [3]. OD faces key challenges due to the semantic gap between how humans perceive images and how computers represent them. While humans perceive images in the visual area (sense the colors and objects in images), computers represent images with matrices of different numbers, thereby leading to problems such as intra-class variations, illumination changes, viewpoint variations, and occlusion, among others [2,4].
However, despite the challenges of OD, and while still being a subject of substantial research in academia and industry, it has found wide applications in modern areas such as security, transportation [5,6,7], medicine [8,9], and military [10,11]. Other applications include video analysis, face recognition [12], and surveillance [1,4,13,14]. In the industry, applications include identifying defects on textile product surfaces, monitoring crop pests and diseases in agriculture [15,16], processing images for autonomous driving [17,18], generating images, reconstructing events, and measuring the essence of objects [19]. These applications have increased sharply since the introduction of Deep Learning (DL) and various CV features to solve problems in daily life.
Figure 1 shows the object detection applications in different fields. Businesses use CV-based object detection to automate their business processes, such as industrial quality control, production monitoring, inventory management, and fault analysis [20]. Some companies use CV techniques to monitor workers’ safety and also the safety of their work environment [21]. Various e-commerce companies that use the OD approach for product delivery, customer support, compliance, and business development have witnessed business growth in a much more profitable way [22,23]. OD is also used in agriculture, such as in the manufacturing of intelligent agricultural equipment [15,24], plant disease identification [25], harmful pest classification and prevention [26], and crop monitoring [27,28], which can change the dimension of modern agriculture towards a new era and the world without hunger and poverty [29].
Computer vision plays a vital role in medical image processing, for example, in preventing harmful diseases in humanity [9,30]. DL has made it easier to identify brain tumors [31,32], internal bleeding [33], lung cancer [34], pneumonia, COVID-19 [35,36,37] and other deadly diseases from medical images. In safety and surveillance systems, computer vision-based detection modeling in different aspects, such as maintaining safety and locating suspected persons and weapons in secured places [13,14,38], and CV has helped the military of different countries [39] ensure national security by detecting spy drones, missiles, or unwanted persons near border areas [10,11,40]. The entertainment industry is also connected to the object detection application. For instance, in football and cricket, the international organizing authorities use DL-based computerized game analyzers and referees to avoid conflicts and human error, while some use the computer vision-based object detection technique to locate the players first and then go to the other processes [41,42].
Based on the development of the object detection approach over the last thirty years, the object detection techniques could be divided into two categories by the development of CNN and DL, and they are the traditional object detectors and the deep learning-based object detectors [43,44]. Traditional techniques, such as the Viola-Jones detector and histogram of Gradients, are manually designed feature extractors and are characterized by slow processing speeds, low accuracy, and inadequate performance when working with new or unfamiliar datasets [2,3]. DL-based techniques integrate neural networks with convolutional neural networks (CNNs) and a transformer-based approach to identify objects. The CNN-based models are also divided into two groups: the region proposal-based approach and the regression- and classification-based approach, as shown in Figure 2. The CNN-based object detection approaches are widely adopted in several object detection models, such as Region-based CNN (R-CNN), the versions of the You-Only-Look-Once (YOLO) models, and have played important roles in the rapid development of OD [1,2]. These DL-based techniques have benefited greatly from benchmark image datasets, such as Pascal Visual Object Classes (Pascal VOC 07/12) [45], OpenImage [46], Microsoft Common Objects in Context (MS COCO) [47], and ImageNet [48]. Also, the backbone architectures of CNNs, such as AlexNet [49], VGG-16, GoogLeNet, and ResNet-50, which extract image features, have contributed to the success of DL-based techniques [1,4].
Owing to rapid improvements in OD, such as the creation of language vision models [50,51], and to effectively work with OD tools and gain sufficient understanding of the topic, it is important to possess knowledge of how the field has developed and significant topics, such as backbone architectures, object detectors, and benchmark datasets. Therefore, this survey aims to provide an up-to-date and concise overview of DL-based OD methods, benchmark datasets, and object detectors.
Furthermore, based on the development of object detection algorithms over the last three decades, as shown in Figure 3, the development of OD in CV can be classified into two periods. The first period was traditional object detection, when deep learning on object detection was not introduced, and the second period used deep learning-based algorithms. We refer to these two detection methods as traditional object detection and DL-based object detection [52].
This paper provides a comprehensive review of computer vision-based object detection algorithms and their trends in research. This paper will provide researchers and engineers with a solid idea of the changes and challenges following the trend of object detection in the last 30 years. They discussed traditional object detection algorithms, single- and two-stage object detection methods, and transformer-based object detectors. This paper discusses the object detection evaluation functions that are used to understand the accuracy of object detection algorithms. The key contributions of this review research can be listed as,
- Proposed a structured taxonomy of object detection approaches, covering traditional algorithms, CNN-based architectures, single and two-stage detectors, transformer-based models, and large language model (LLM) driven methods.
- Classified benchmark datasets by their application domains (e.g., generic, autonomous driving, medical, agricultural) and critically analyzed commonly used evaluation metrics such as mAP, IoU, and FPS, highlighting their implications for both accuracy and deployment efficiency.
- Identified emerging challenges, including lightweight edge deployment, multimodal vision-language integration, and fairness in AI in Object detection, for future research to guide the next wave of object detection systems in different fields.
This paper presents recent object detection methods with their applications and gaps in existing trends. This paper gives the review methodology in Section 2, discusses the traditional OD methods in Section 3, and the Convolutional Neural Network(CNN) based methods are discussed in Section 4. In Section 5, the paper discusses DL-based methods for OD algorithms, while Section 6 discusses some prominent OD datasets. A discussion and suggestions for future work are provided in Section 9, and the paper is wrapped with the conclusion given in Section 10.
2. Review Methodology
This systematic review adheres to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [53] to ensure methodological rigor, transparency, and reproducibility in synthesizing object detection literature from 1999 to January 2026, as shown in Figure 4. The literature search was conducted across three major academic databases—Google Scholar, IEEE Xplore, and Web of Science—yielding 4,862 initial records. Following the removal of 965 duplicate records and 6 records excluded for other reasons, 3,891 records proceeded to title and abstract screening. At this stage, 2,214 records were excluded based on irrelevance to object detection, non-English language, or lack of quantitative evaluation, leaving 1,677 records sought for retrieval. Of these, 966 full-text reports could not be retrieved, resulting in 711 reports assessed for eligibility through comprehensive full-text review. Applying predefined inclusion criteria—specifically, (i) focus on object detection algorithms, (ii) quantitative performance reporting (mAP, IoU, FPS), (iii) novel architectural contributions, and (iv) publication in English—298 reports were excluded due to insufficient methodological quality, incomplete experimental setups, or non-computer vision applications. The remaining 413 studies were included in the final review, comprising 384 reports of included studies supplemented by additional records identified through snowballing and forward citation tracking. This rigorous, multi-stage screening process ensures comprehensive coverage of the object detection landscape while minimizing selection bias, establishing a robust foundation for the qualitative synthesis and critical analysis presented in this review.
3. Traditional Methods of Detection
The journey of the computer vision-based object detection algorithm started in 1997 with the Scale Invariant Feature Transform (SIFT) algorithm. In Figure 3, it is observed that from 1997 to 2010, the traditional object detection period was considered when Scale Invariant Feature Transform (SIFT), Histogram of Gradients (HoG), and Speeded-Up Robust Features (SURF), other prominent object detection approaches, were introduced. The traditional object detection work begins with the process shown in Figure 5, which begins with input image processing techniques, such as color conversion, resizing of the input image, and region selection. Then it goes to feature extraction, where the image features are extracted, and then the extracted data passes through the classification of the object class. At the end, the classes represent the output image by determining and localizing the object class in the input image. To minimize the computational costs, most of the traditional object detection algorithms used to convert the RGB image to a grayscale images, where the feature selection of the grayscale image is comparatively easier than the RGB image.
Scale Invariant Feature Transform (SIFT) is an algorithm that performs object detection by determining and interpreting the local characteristics, such as corners and blobs, in the image [54]. David G. Lowe [55] from the University of British Columbia first proposed this method in 1999. As the name implies, the features determined by this algorithm are usually independent of changes in image rotation and size but depend on the position of the object in the image [54]. The SIFT method consists of four phases, which start with the selection of the scale-space peak, which indicates the possible feature placement. The second phase is the key point selection, which involves precisely identifying the feature key points. The third stage is the orientation assignment, where an orientation is assigned to each keypoint to achieve invariance to image rotation, and a neighborhood is taken around the keypoint location depending on the scale, and the gradient magnitude and direction are calculated in that region. The keypoint description and keypoint matching are the final stage, where a keypoint descriptor is created, and key points between two images are matched by identifying their nearest neighbors [56,57]. The advantages of the SIFT object detector include the local features that are robust against clutter and occlusion, the ability to create numerous features for small objects, and the ability to match individual features to a large database, making it prominent at that period [58]. The author in [59] provides a comprehensive review of image retrieval methods, where they classified various methods into different categories. Among those, the SIFT-based category is included due to its advantages in dealing with image transformations.
Figure 6.
Example of Haar-like rectangular features used to detect the frontal face of a lion. [60]
Figure 6.
Example of Haar-like rectangular features used to detect the frontal face of a lion. [60]

The cascade classifier-based object detector, also known as the Viola-Jones Object detector, was proposed in 2001 by Paul Viola and Michael Jones [61]. It was the first object detection technique at that time that could give a competitive detection rate. It could detect objects in real time, but it was mainly used for face detection because of its high accuracy (true positives) and very low false-positive rate, making it more robust than other algorithms at the time [62]. Despite having lower accuracy than modern CNNs or the LLM-based face detection algorithm, the Viola-Jones face detection algorithm is an efficient solution for resource-constrained devices. The working of the Viola-Jones approach depends on four steps, where it starts with selecting Haar cascade features from the datasets. Then, in the second stage, it creates an integral image formed by computing the rectangles using the Haar features. From the integral image structure, it performs face feature classification using the Adaboost learning algorithm. At the end, the cascade classifier combined the object features and localized the object in the visual region. Figure 6 shows the frontal face detection of a lion by the Haar cascade object detector. Along with its advantages, the cascade object detector had some limitations too, such as decreasing the accuracy ratio in various lighting conditions and consuming more processing time than other algorithms [60,63].
Table 1.
Comparison of Traditional Object Detection Methods.
| Method | Advantages | Disadvantages | Complexity | Applications |
|---|---|---|---|---|
| SIFT [55] | Robust to scale/rotation changes; strong feature matching | Slow; sensitive to illumination shifts | High | Object recognition, robotics vision |
| Cascade [61] | Real-time detection; efficient with integral image; good for frontal faces | Poor with non-frontal views and occlusions; sensitive to lighting and pose variation | Low | Face detection and real-time surveillance |
| HoG [64] | Effective for edge-based objects; lighting-invariant | Poor for textured/cluttered scenes | Medium | Pedestrian detection, vehicle recognition |
| DPM [65] | Handles occlusion; shape-invariant | Needs large datasets; slow inference | Very High | Face detection, cluttered scenes |
| SURF [66] | Faster than SIFT; decent accuracy | Unstable under extreme lighting/rotation | Medium | Image retrieval, 3D reconstruction |
In 2005, Navneet Dalal and Bill Triggs proposed the Histograms of Oriented Gradients(HOG)-based object detection approach for human detection [67]. In this method, the image is split into tiny cells, and in each of these tiny cells, histograms representing directions of the greatest steepness, that is, the gradient, are constructed. These histograms were further concatenated to form a descriptor [54,64]. The core concept behind the HOG-based object detector is to capture the distribution of gradient orientations in an image, which can be used to describe the shape and appearance of objects. The HOG features work by dividing an image into small cells, computing the gradient orientations within each cell, and then aggregating these orientations into a histogram. This histogram represents the distribution of gradient orientations, which can serve as a feature vector for object detection [68]. The HOG features-based object detector has several benefits that make it attractive compared to other traditional object detectors at that time, such as robustness to lighting changes and occlusions in lighting conditions, making it suitable for object detection in real-world scenarios. The HOG feature-based object detection algorithm was computed efficiently, making it suitable for real-time object detection applications where various classification algorithms, such as Support Vector Machines (SVMs) and Random Forests, are used [68]. Despite the advantages, the HOG-based object detector shows poor accuracy for textured and clustered scenes in object detection. The HOG detector can detect several kinds in addition to pedestrians, where the detection window size remains constant while the detector resizes the input image repeatedly through HOG to detect item sizes [69].
The Deformable Part Model (DPM) has demonstrated success in conventional object identification and has been recognized as the winner of the VOC recognition competition in 2008 and 2009, proposed by researchers from the University of Chicago and Toyota Technological Institute [65,70]. The HOG-based approach has been enhanced with the DPM detector, and the core principle of the Deformable Part Model (DPM) follows a divide-and-conquer strategy, where the training phase is crucial for analyzing and decomposing objects; the inference phase is regarded as a series of distinct steps in the object detection process. A standard DPM technique consists of two primary components: a root filter and a multiple-component filter. Unlike traditional methods, which manually define the configuration of part filters, the DPM detector employs a weakly supervised learning approach to automatically determine the structure of component filtration systems based on dependent variables [71,72]. Deformable Part Model (DPM) allows the object parts to move independently, and, as a result, the model finds objects that change their pose, shape, and perspective effectively. Flexibility in the representation of complex shapes, partial occlusion, and detection of objects in partial occlusion. DPM is flexible, representing each part separately, so it can work with partial occlusion, finding an object even when it is partly hidden. It is resistant to cluttered data and simple labeling data, in which deep learning models tend to fail. Despite being outperformed by modern detectors such as Faster R-CNN and YOLO, the DPM part-based approach is still useful in dealing with deformation and occlusion. Researcher of also improved DPM-based detection by combining it with CNN architecture [72] and R-CNN [73], and a refinement of candidate selections via a dense subgraph discovery filter.
The Speeded-Up Robust Features(SURF) method is a fast and robust approach for local, similarity-invariant image comparison and representation, which was proposed by Herbert Bay and his research group in 2006 [66]. The main attraction of this technique is its speedy computation of operators through the use of box filters, which enable real-time applications such as object detection and tracking [74]. The working of the SURF object detector is composed of two steps: feature extraction, where it extracts the image features from the input image, and feature description, where it cross-matches the features of the image with the class. SURF is robust to changes in scale and rotation, meaning it can detect the same object or feature even if the object appears larger, smaller, or the image is rotating [75]. This method is commonly used in pedestrian tracking applications.
Table 1, gives the comparison of the traditional object detection algorithms, where the SIFT algorithm is robust and has more efficiency in feature matching for object recognition and robotic vision applications, and the Cascade Viola-Jones algorithm is more efficient for real-time frontal face detection and visual surveillance, where the HOG-based detector is more efficient to define objects in different lighting conditions and this approach is effective for the autonomous vehicles and the pedestrian detection on road. The DPM approach handles occlusion and shape invariants, which is good for face detection applications. However, SURF is faster than the SIFT approach, which is used for image retrieval or 3D image reconstructions. But in unstable lighting and rotational conditions, the SURF approach decreases the object detection accuracy, which is one of its limitations.
4. Convolutional Neural Network (CNN)
After traditional object detection, in 2010, researchers started using Convolutional Neural Network (CNN) architectures for the detection of objects in computer vision. The CNN architecture is also a core component in deep learning-based object detection, where a CNN has four crucial layers: the convolution layer, activation layer, pooling layer, and fully connected layer. Figure 7 shows the working process of the deep learning object detector, which starts working with the processing of the input images, where the distortions in the input images are filtered. Then the filtered image goes to the feature extraction and classification step, where the image is extracted by the CNN architecture and cross-matched by the object classes in the output image.
In CNN layers, the convolution layer is the first stage of the network, through which the input image passes. This layer extracts useful information from the image via filters or kernels, which slide through the structure of the input image. This sliding operation is called convolution, and each convolution operation in the CNN results in a feature map that contains edges, boundaries, and other object features in the image, as shown in Figure 8. Therefore, the number of feature maps depends on the number of kernels in a CNN [3,4]. CNNs are used in object detection tasks because they preserve the spatial information of input images compared to a regular deep neural network that performs image flattening. The activation layer is crucial for adding nonlinearity to neural networks. Without this, deep neural networks can become linear classifiers. The activation layer performs spatial distortion and determines the neurons that are activated or deactivated. A Rectified Linear Unit (ReLU) is a common activation function used in deep learning. The ReLU sets the negative values in the filtered image to zero and only becomes activated when the node input exceeds an initially determined threshold. If the input is less than zero, the output is zero [3,4].
The pooling layer acts as a link between the two convolutional layers and aims to decrease the size of the feature maps, resulting in reduced computation for the network. This is accomplished by determining the largest value within a local group of units in the feature maps. The three most commonly used pooling techniques are the maximum, average, and minimum pooling. In the maximum pooling technique, the highest value within the region captured by the filter is selected, whereas in the minimum pooling technique, the lowest value is selected from the captured region. Average pooling involves down-sampling the feature map region by computing the average value within the region [3,4].
The fully connected layer receives the feature map output from the previous layer and flattens it into a one-dimensional vector. The fully connected layer, a feed-forward neural network, performs the detection and classification tasks. [3,4]. The CNN architecture is the backbone in deep learning-based object detection models, as it extracts feature maps from the input image. In most object detectors, backbone networks are responsible for performing classification tasks [1,2,3]. AlexNet [49], GoogLeNet [76], VGG-16 [77], and ResNet [78], the prominent CNN architectures showed in Table 2 highlighting their layers and top accuracy.
AlexNet, a CNN-based architecture for image classification, was proposed by Krizhevsky et al. [49] and won the ImageNet Large-Scale Visual Recognition Challenge (ILSVRC) 2012. It achieved a much higher accuracy of over 26% compared with the other models at that time. The architecture consists of eight layers, of which three are fully connected and five are convolutional, all of which can be learned during training. To prevent overfitting and increase the convergence speed during training, AlexNet uses dropout and ReLU as regularization techniques [2].
Szegedy et al. [76] proposed a locally sparsed connected architecture rather than a fully connected architecture to solve the high computational costs and large number of parameters associated with applying classification networks to real-world applications. GoogLeNet, a 22-layer deep network, consists of several inception modules stacked together. These inception modules had filters of different sizes at the same level. The feature maps pass through these filters and are concatenated before passing through the next layer. GoogLeNet also has supporting classifiers in the intermediate layers that help regulate and propagate gradients. The network demonstrates the high efficiency of the computation blocks while still achieving the same performance as other parameter-heavy networks. GoogLeNet achieved 93.3% top-5 accuracy on the ImageNet dataset and was faster than the other contemporary models [2,3].
The focus of AlexNet and its successors was to improve the accuracy of classification networks using smaller receptive window sizes. However, Simonyan and Zisserman [77] explored the effect of network depth on accuracy. They introduced VGG-16, which uses smaller convolution filters to build networks of varying depths. The use of small convolution filters allowed the capture of a larger receptive field, but it also significantly reduced the network parameters and resulted in faster convergence. The VGG consists of convolutional layers, three fully connected layers, and a softmax layer. The number of convolutional layers in VGG can vary between 8 and 16. It demonstrated superior performance compared to GoogLeNet in the category of single network performance, which was the winner of the ILSVRC 2014 [2,3].
He et al. [78] proposed integrating skip connections between stacked convolutional layers to reduce the degradation of CNN performance as they become deeper. This skip connection is an element-wise addition between the inputs and outputs of the block, which does not increase the number of parameters or computational complexity of the network. The ResNet architecture [78], for example, consists of a large () convolutional filter followed by 16 bottleneck modules, each consisting of a pair of small filters with an identity shortcut across them, and ends with a fully connected layer. The researchers also showed that their deeper ResNet architectures, such as the 101 and 152 layers, had higher accuracy and lower complexity than the VGG-16 network. ResNet has inspired many variants of itself, as well as other networks.
5. Deep Learning Object Detection Methods
Deep Learning-based methods use CNNs and CNN-based architectures as their backbones. As shown in Figure 2, the deep learning based object detection approaches can be separated into three categories: the CNN-based methods, which are also separated into two categories; the Region proposal-based method, which is known as the Two Stage Detectors(TSD), and the Regression and classification-based method, which is known as the Single Stage Detectors(SSD); the transformer based approaches and the Large Language Model(LLM) based object detectors [2,79,80].
In two-stage object detectors, regions in the images were classified in the first step, while classification and localization are performed on features extracted from these regions in the second stage. Single-stage detectors achieve the OD task in a single attempt by using dense sampling. They do not require any region of interest, but instead use predefined boxes or key points with different scales and aspect ratios to determine the location of objects in images [1,2,81]. A Transformer-based object detector is a type of deep learning model that uses a transformer architecture instead of CNN architectures, originally developed for natural language processing to perform object detection in images. Instead of relying solely on convolutional neural networks (CNNs), these models leverage the self-attention mechanism of transformers to capture the global context and relationships across the entire image [82,83]. Two-stage detectors (TSDs) have a higher detection accuracy but are slower than single-stage detectors. Single-stage detectors are time-efficient and well-suited for real-time applications, although their accuracy might not be as good as that of TSDs.
5.1. Two Stage Detectors
Two-stage detectors(TSD) operate by decoupling the detection process into sequential steps. In the initial stage, the model proposes a sparse set of potential object locations, often referred to as Regions of Interest (RoIs), where it performs the proposal generation within the image. The second stage then performs fine-grained classification and localization refinement on features extracted specifically from these proposed regions, as shown in Figure 9. TSDs, such as those based on the R-CNN family, are typically characterized by higher detection accuracy due to this focused, two-step approach but exhibit slower processing speeds [1,2,81].
Here, this paper briefly discusses some of the common two-stage detectors (TSDs) algorithms used in object detection. As mentioned earlier, these two-stage detectors primarily work by defining a specific region in the image frame. For this reason, TSDs are often called region-proposal-based object detectors. There are a few region-based detectors such as Region-based CNN(R-CNN), Fast Region-based CNN(Fast R-CNN), Faster Region-based CNN(Faster R-CNN), and Spatial Pyramid Pooling (SPP-Net). However, some modifications have been made to these object detectors in combination with other algorithms [2].
Proposed by Girshick et al. [84] by the researchers from Microsoft in 2014, is the first variant of region-based classes of TSDs. Girshick et al. [84] used the AlexNet [49] architecture and demonstrated how the object detection performance can be increased using CNNs. In the R-CNN, the input image is fed into the network, and its mean value is subtracted from it. This modified image is then fed into the region proposal module, which generates 2000 regions of interest (ROI), which are the areas of the image that have a greater potential to contain an object, using the selective search method [85,86] as shown in Figure 10. These object candidates obtained from the ROI were then transformed and passed through a CNN network consisting of convolutional layers and fully connected layers (five convolutional layers and two fully connected layers), which computed a feature vector with 4096 dimensions for each proposal. The input size of the adopted image was 227×227 pixels [1,2,81].
After the feature vectors are obtained, they are fed into Support Vector Machines (SVMs) that have been trained to identify specific classes to acquire confidence scores. The regions scored by the SVMs are then subjected to non-maximum suppression (NMS) based on their class and Intersection over Union (IoU). Once the class is determined, a trained bounding-box regressor is utilized to predict the bounding box by estimating four parameters: the center coordinates of the box, as well as its width and height [2,81]. R-CNN was characterized by slow performance (47 seconds per image) [84,87] and was time- and space-intensive, which is particularly problematic when dealing with smaller datasets.
The R-CNN training process involves several stages, including fine-tuning, pre-training, bounding box regression, and SVM classification. On the other hand, Fast R-CNN involves using a multi-task loss to train the network simultaneously on each labeled RoI [88]. Proposed by [88] in 2015, the Fast R-CNN algorithm obtains features from the complete input image and then uses an RoI pooling layer to obtain fixed-size features. These features are then fed into fully connected layers that perform the classification and bounding box regression tasks [89]. The algorithm extracts features from an entire image and uses them for both classification and localization via a CNN. The Fast R-CNN enhanced the speed of R-CNN by a factor of 146, although its accuracy improvement was seen as secondary [1]. Experimental results indicated that Fast R-CNN achieved a mean average precision(mAP) of 66.9% on the PASCAL VOC 2007 dataset, whereas R-CNN achieved 66.0%. Additionally, Fast R-CNN required only 9.5 hours of training time, which is significantly faster than R-CNN’s training time of 84 hours, making it 9 times faster [45]
Even though Fast R-CNN made progress toward achieving real-time object detection [88], its generation of region proposals remained significantly slower by some magnitude [2]. Fast R-CNN utilizes a selective search for RoI proposals, which is time-consuming. Faster R-CNN was proposed in [90] three months after Fast-RCNN [88] was proposed. It introduces a new approach called the Region Proposal Network (RPN), which is a fully convolutional network that efficiently predicts region proposals at different scales and aspect ratios. By sharing full-image convolutional features and convolutional layers with the detection network, RPN accelerates the process of generating region proposals [1,2]. In addition, a new approach for detecting objects of various sizes using multi-scale anchors as a reference point was used in the Faster R-CNN. This technique allows the creation of proposals for regions of different sizes without requiring different scales of input images or features, making the process much simpler [1,2]. The results showed that Faster R-CNN significantly enhanced both the accuracy and efficiency in detecting objects. When compared to Fast R-CNN with shared convolutional computations, Faster R-CNN had a mAP of 69.9% on the PASCAL VOC 2007 test set, while Fast R-CNN only achieved a mAP of 66.9%. Furthermore, when using the same VGG [77] backbone architecture, the total processing time of Faster R-CNN was almost 10 times faster than Fast R-CNN, with 198ms compared to 1830 ms, respectively. The processing rate of Faster R-CNN was also much faster at 5fps compared to 0.5 fps for Fast R-CNN [1].
An extension of Faster R-CNN, Mask R-CNN [91] integrates an additional branch for instance segmentation at the pixel level. This branch utilizes a fully connected network, which is applied to the RoIs to categorize each pixel into segments while maintaining low computation costs. The basic structure for the object proposal in Mask R-CNN is similar to that of Faster R-CNN. However, it includes a mask head alongside the classification and bounding box regressor head [91]. A notable difference is the use of the RoI Align layer rather than the RoI pool layer, which prevents any misalignment at the pixel level that may occur owing to spatial quantization [1,54]. Mask R-CNN is easy to train, easily adaptable, and performs well in different applications, such as keypoint detection and human pose estimation [2]. Compared to other single-model architectures, it demonstrates superior performance and adds the capability of instance segmentation with a minimal increase in computational overhead. However, it still falls short of real-time performance, typically requiring a processing rate of over 30 frames per second (fps) [2].
Spatial Pyramid Pooling (SPP-Net) represents a significant innovation in convolutional neural network (CNN) architectures that addresses the challenge of variable-sized image inputs [92]. Traditional CNNs often require fixed-resolution images, leading to information loss or distortion during resizing. SPP-net circumvents this limitation by integrating spatial pyramid pooling, which is a technique inspired by the hierarchical structure of spatial pyramid matching (SPM) [76,93]. This approach aggregates multi-scale contextual features through a series of grouping operations at varying grid resolutions, generating fixed-length feature vectors, irrespective of the input dimensions [92]. In practical terms, the SPP-net operates by decoupling the feature extraction from the input size constraints. Within a standard CNN framework, for instance, a seven-layer model, the initial convolutional layers process inputs of arbitrary sizes using sliding filters, producing spatially organized feature maps [94]. These maps retain the input aspect ratio and encode activation patterns across different regions. The subsequent spatial pyramid pooling layer divides each feature map into sub-regions at multiple scales and applies max-pooling to each sub-region [92]. The pooled outputs are concatenated into a unified vector, which is then fed into the fully connected layers for classification or detection tasks [95]. This architecture eliminates the need for repeated convolutional computations when processing multiscale inputs, thereby enhancing the efficiency of applications such as object detection. Moreover, by preserving spatial hierarchies and avoiding artificial resizing, the SPP-net improves the robustness of feature representations, enabling CNNs to generalize better across diverse visual scenarios [96]. Its adaptability has made it a cornerstone for advancing tasks such as multi-scale object recognition and scene understanding [92].
In 2017, Facebook researchers proposed a feature pyramid network (FPN) that analyzed images of varying sizes to produce layered maps that highlight details at different scales [97]. It operates independently of the underlying system used to process the image, making it adaptable for tasks such as efficiently detecting objects[2]. The system operates in two stages, where the first stage is called Bottom-Up Analysis, in which the image is processed, generating progressively simplified representations (each stage reduces resolution by half). The second stage is called the Top-Down Refinement stage, in which high-context layers from deeper stages are upscaled to match the resolution of the earlier layers [97,98]. These were then combined with the original, finer-grained layers from the initial analysis. While the earlier layers lacked a broader context, they retained precise positional details. Merging these layers creates a balanced output rich in both detailed and contextual understanding [99]. By blending simplified and detailed layers across scales, the FPN avoids resizing images or repeating calculations, enabling consistent detection of objects regardless of size while maintaining computational efficiency. Researchers have connected the FPN architecture to other models to enhance the efficiency of object detection in different scenarios [100,101,102].
5.2. Single Stage Detector
In contrast with two-stage detectors(TSDs), single-stage detectors achieve object detection tasks in a single forward pass through the network via dense sampling. These models bypass the explicit RoI proposal step, instead utilizing a dense grid of predefined anchor boxes or key points with varying scales and aspect ratios to predict object location and classification simultaneously across the entire image, as shown in Figure 11. Single-stage detectors prioritize speed and simplicity, making them ideal for real-time applications, such as video analysis, robotics, and edge devices [103]. While historically their accuracy trailed that of TSDs, the emergence of advanced architectures, such as YOLOv7 [104], has enabled contemporary SSDs to achieve comparable high accuracy in various OD scenarios. Starting with the single-stage detectors, the Yolo, the Single Shot Detector(SSD), and RetinaNet are the prominent approaches.
Table 3.
Comparison of Single-Stage vs. Two-Stage Object Detectors.
| Aspect | Single-Stage Detectors | Two-Stage Detectors |
|---|---|---|
| Speed | High (30−150 FPS) | Low to moderate (5−15 FPS) |
| Accuracy | Moderate (50−60 mAP on COCO) | High (60−70 mAP on COCO) |
| Complexity | Simple (single forward pass) | Complex (region proposals+refinement) |
| Use Case | Real-time applications (video, edge devices) | High-precision offline tasks (medical imaging, research) |
| Small Objects | Struggles due to coarse features | Better handles small/occluded objects |
The YOLO (You Only Look Once) family has evolved from YOLOv1 in 2015 [105], which introduced real-time object detection as a single regression problem, to the latest YOLOv12 [106], with continuous improvements in speed, accuracy, and deployment. YOLOv2 to v3 [107,108] enhanced detection with anchor boxes, multi-scale features, and better robustness, while YOLOv4 and YOLOv5 optimized training strategies and usability for wider adoption. Later versions, YOLOv6 to YOLOv8, focused on scalability, lightweight design, and edge deployment, and the most recent YOLOv9-v12 integrates advanced concepts like programmable gradients, dynamic label assignment, and stronger backbone-neck designs. In Figure 12, this roadmap illustrates YOLO’s consistent progress in striking a balance between efficiency and accuracy, making it a cornerstone of modern object detection research and applications over the past decade.
The You Only Look Once (YOLO) family of algorithms has become one of the most influential developments in computer vision due to its ability to perform real-time object detection while maintaining high accuracy. Introduced by Redmon et al. in 2015 [105], the original YOLO represented a paradigm shift compared to earlier region-based approaches such as R-CNN [1,2,81], Fast R-CNN [45], and Faster R-CNN. Whereas these earlier models relied on generating region proposals before classification, YOLO framed object detection as a single regression problem [109]. The input image, resized to 448×448, was divided into an SxS grid, and each grid cell directly predicted bounding box coordinates, confidence scores, and class probabilities [110]. The backbone consisted of 24 convolutional layers interspersed with four max-pooling layers, followed by two fully connected layers for final predictions [110]. This design enabled real-time processing at over 45 FPS, far surpassing earlier detectors. Figure 13 illustrates the architecture of the basic YOLO algorithm, and Figure 14 shows the basic working process of the YOLO object detector. However, the first version struggled with detecting multiple small objects, handling irregular bounding boxes, and accurately localizing objects near the grid boundaries, which limited its scalability in real-world scenarios [111].
To address these shortcomings, YOLOv2 (YOLO9000) was introduced in 2016 [107,112]. It replaced the fully connected layers with anchor boxes, which improved the handling of multiple objects and irregular aspect ratios [105]. The backbone was upgraded to Darknet-19, a more efficient CNN architecture with batch normalization for improved convergence and generalization. High-resolution classification, multi-scale training, and fine-grained features from intermediate layers further improved detection accuracy. A remarkable innovation in YOLOv2 was its ability to perform joint training on ImageNet and COCO, enabling detection of over 9,000 object categories, thus significantly broadening its applicability [107]. YOLOv3 (2018) introduced even more powerful enhancements where the backbone was upgraded to Darknet-53, a deeper network with residual skip connections inspired by ResNet, which improved gradient flow and allowed for better representation learning [108]. Instead of softmax, YOLOv3 employed independent logistic classifiers for multi-label classification, which was particularly useful in detecting overlapping categories. Moreover, it introduced multi-scale predictions at three different feature map scales, significantly improving the detection of small, medium, and large objects simultaneously. This made YOLOv3 robust across diverse datasets, though the model was heavier and required more computational resources compared to earlier versions [108]. In 2020, YOLOv4 pushed YOLO toward production readiness by focusing on both speed and accuracy. It integrated CSPDarknet-53 as the backbone, which reduced computational cost while improving learning capability [113]. Enhancements such as Spatial Pyramid Pooling (SPP), Path Aggregation Network (PANet), and advanced data augmentation strategies like Mosaic Augmentation and DropBlock Regularization made YOLOv4 significantly stronger [113]. It also benefited from techniques (e.g., CIoU loss, cross mini-batch normalization) and a bag of special techniques (e.g., Mish activation, SPP block), achieving state-of-the-art results without requiring expensive hardware [113]. In parallel with YOLOv4, two alternative branches emerged. YOLOR (2021) introduced a hybrid representation that combined explicit knowledge (visible features from the data) and implicit knowledge (latent representations not directly visible) [114]. This made the model more effective in capturing contextual dependencies. Around the same time, YOLOX proposed a fully anchor-free design, decoupled heads for classification and regression, and an innovative SimOTA label assignment strategy [115]. These changes improved training dynamics, reduced dependency on handcrafted anchor designs, and yielded better real-world performance [116]. YOLOv5, released by Ultralytics in 2020, marked a turning point due to its PyTorch implementation [117]. Although not officially from the original YOLO authors, it became one of the most widely adopted versions because of its modular design, active community support, and availability in multiple model sizes (s, m, l, x). Key improvements included the Focus layer, efficient augmentation strategies (e.g., Mosaic, Auto Augment), and scalability across hardware setups [117]. YOLOv5 could be used in lightweight deployment, and continuous updates made it standard in real-world applications ranging from robotics to surveillance applications [118]. Building on this momentum, YOLOv6 (2022) was designed with a strong emphasis on industrial use cases [119]. It optimized computational efficiency by redesigning the backbone with hardware-friendly operations, introduced an efficient decoupled head, and applied advanced label assignment strategies. YOLOv6 demonstrated strong performance in edge AI scenarios, where balancing accuracy and low-latency inference is crucial [119]. The release of YOLOv7 (2022) marked another milestone. It introduced the Extended Efficient Layer Aggregation Network (E-ELAN), which allowed dynamic expansion of depth and width while preserving computational efficiency [120]. It also proposed the concept of a ’trainable bag of freebies’ incorporating training optimizations such as coarse-to-fine label assignment without increasing inference cost. YOLOv7 set new benchmarks by outperforming contemporary transformers and CNN-based detectors, all while maintaining real-time inference speeds [116]. YOLOv8 (2023) [121] further advanced the framework by adopting an anchor-free detection head, simplifying the training and inference pipeline while improving accuracy for small and occluded objects8 [122,123]. It refined Mosaic augmentation for better generalization and introduced flexible model scaling strategies, making it applicable across classification, detection, and segmentation tasks [124]. YOLOv8’s versatility positioned it as one of the most practical versions for both research and deployment [47]. In 2024, YOLOv9 was introduced with Programmable Gradient Information (PGI) and Generalized ELAN (GELAN) [125]. PGI improved gradient flow during backpropagation, enabling the model to learn more effectively with fewer parameters. GELAN enhanced feature reuse and parameter efficiency, making the model not only lighter but also more robust across datasets. YOLOv9 significantly reduced computational cost compared to previous versions [125]. However, it still struggled in scenarios with extreme occlusions, densely packed small objects, and highly imbalanced datasets. To address these issues, YOLOv10 (2024) made a critical shift by eliminating the reliance on Non-Maximum Suppression (NMS) through consistent dual assignment [126]. This innovation allowed the model to handle overlapping objects more effectively and reduced post-processing latency. YOLOv10 also optimized architectural components such as CSPNet backbones, PANet necks, and dual prediction heads, delivering higher efficiency and accuracy simultaneously [126]. The most recent advancements include YOLOv11 (2024) [127] and YOLOv12 (2025) [106]. YOLOv11 introduced the C3k2 block and the C2PSA module, combining cross-stage partial connections with self-attention mechanisms. This design improved the ability to capture global context, which is critical for detecting small or occluded objects, while maintaining speed and reducing parameter complexity. YOLOv12 represents a paradigm shift by fully embracing attention-centric mechanisms within the YOLO framework [128]. It incorporates area attention and Residual Efficient Layer Aggregation Networks (R-ELAN), achieving remarkable gains in small-object detection and occlusion robustness [129,130]. YOLOv12 also streamlines the detection head into an anchor-free design, leverages advanced augmentation techniques such as Mosaic, MixUp, and copy-paste, and accelerates training with FlashAttention [106]. Although it requires specialized hardware for optimal deployment, YOLOv12 demonstrates how attention-driven architectures can surpass conventional CNN-based YOLO models [106], setting new state-of-the-art benchmarks.
Taken together, the evolution of the YOLO family illustrates a continuous trajectory of innovation. From the original grid-based regression model to anchor-based frameworks, then to anchor-free detection, and finally to attention-driven architectures, YOLO has consistently expanded the boundaries of real-time object detection [131]. Each version addressed specific limitations of its predecessor while introducing novel mechanisms that balanced accuracy, speed, and efficiency. This progression highlights not only the adaptability of YOLO to real-world demands but also its central role in shaping the future of computer vision applications across autonomous driving, robotics, surveillance, and beyond [106,132].
The evolution of You Only Look Once (YOLO) object detection algorithms from YOLOv1 to YOLOv12, including variants such as YOLOR and YOLOX, represents a significant advancement in the field of real-time object detection. YOLOv1 introduced a unified architecture based on GoogLeNet that enabled real-time detection, but suffered from poor localization accuracy and lower mean Average Precision (mAP) [110]. Subsequent versions, YOLOv2 and YOLOv3, adopted anchor-based methods and improved backbone networks, notably Darknet-19 and Darknet-53, which significantly enhanced the accuracy while maintaining high inference speeds. YOLOv4 incorporated additional modules, such as CSP Darknet-53 and PANet, along with data augmentation strategies, such as mosaic and self-adversarial training, pushing the performance further in terms of both speed and accuracy. YOLOv5 marked a shift towards a PyTorch implementation with increased modularity and deployment flexibility, achieving high mAP scores with minimal latency [110].
With the emergence of YOLOv6 and beyond, the architecture has transitioned towards anchor-free detection and decoupled heads, emphasizing lightweight and efficient designs for edge deployment. Notably, YOLOv7 and YOLOv8 introduced enhancements such as task alignment learning, extended support for multitask outputs including instance segmentation and pose estimation, and further optimizations for speed [110,133]. YOLOv9 to YOLOv12 introduced increasingly complex modules, including C2f, PSA, and multipath attention mechanisms, achieving better feature representation, gradient propagation, and inference stability. YOLOv12, in particular, adopted a generalist vision model approach capable of performing detection, segmentation, and pose estimation using a unified decoupled head and no-NMS design, thereby pushing the performance frontier in real-time vision systems [133].
YOLOR, an intermediate innovation, introduced a hybrid representation by combining implicit and explicit learning, though its performance in terms of speed and generalization was eventually surpassed by newer YOLO versions [118]. YOLOX, on the other hand, was instrumental in establishing a decoupled head and anchor-free paradigm along with SimOTA label assignment, setting a strong foundation for subsequent YOLO versions. Overall, the trajectory of the YOLO series demonstrates a clear trend towards unified, efficient, and multitask-capable models that balance high accuracy with real-time processing, making them highly suitable for modern autonomous systems and edge AI applications [110,118,133]. Table 4 shows a comparison table of YOLO models and their features. For the future development of the YOLO object detection family, there are a few things that need to be considered and added over time in future developments,
- Using cutting-edge deep learning techniques, data augmentation, and training approaches, scientists and engineers will keep improving the YOLO architectures, YOLO which will increase the model’s performance, resilience, and efficiency for a continual innovation process.
- A more recent and difficult benchmark may ultimately replace COCO 2017, the current benchmark for assessing object detection models. This move from the VOC 2007 benchmark, which was utilized in the first two YOLO versions, illustrates the shift toward more rigorous benchmarks as the models become more complex and precise.
- As the YOLO framework progresses, it is expected to witness an increase in the number of YOLO models released each year, along with a corresponding expansion of applications. As the framework becomes more versatile and powerful, it is likely to be employed in even more varied domains, from home appliance devices to autonomous cars.
- YOLO models have the potential to extend their capabilities beyond object detection and segmentation, branching into areas such as object tracking in videos and 3D keypoint estimation. As these models evolve, they may become the foundation for new solutions that address a broader range of computer-vision tasks.
- YOLO models will further span hardware platforms, from IoT devices to high-performance computing clusters. This adaptability enables the deployment of YOLO models in various contexts, depending on the application requirements and constraints. In addition, by tailoring the models to suit different hardware specifications, YOLO can be made accessible and effective for more users and industries, which will also promote ideas for solving real-life computer vision problems.
The single-shot detector(SSD) model was published by Wei Liu et al. [134] in 2015, shortly after the YOLO model, and was later refined in a subsequent paper [135]. Later, in 2017, Wei Liu and his group proposed a deconvolutional single-shot detector (DSSR), modifying the SSD algorithm [136].
and Figure 15 shows the working of the SSD object detector, which doesn’t split the whole image into grids of arbitrary size like the YOLO algorithm but predicts the offset of predefined anchor boxes for every location of the feature map. Each box has a fixed size and position relative to its corresponding cell, and all anchor boxes tile the entire feature map in a convolutional manner. Feature maps at different levels had different receptive field sizes. The anchor boxes at different levels are rescaled such that one feature map is only responsible for objects at a particular scale [134].
In 2017, Facebook AI Research introduced RetinaNet [137], a single-stage object detection framework, and introduced a novel approach to mitigate the pervasive challenge of class imbalance inherent in dense detection architectures [138]. Central to its design is the focal loss function, which dynamically modulates the standard cross-entropy loss function to prioritize learning from hard-negative examples during training [137]. This loss function incorporates a tunable modulating factor that decays exponentially as the confidence in the correct class increases, effectively downweighting the contribution of easily classified samples and redirecting the model’s focus toward misclassified or ambiguous instances [51]. The RetinaNet architecture comprises a unified network with three core components: the first one is the backbone convolutional neural network (e.g., ResNet) for extracting multi-scale feature maps across the input image, the second component is the classification subnet that performs dense object category prediction over these features, and lastly, a regression subnet tasked with bounding box coordinate refinement [51]. Critically, both subnetworks employ simple yet efficient convolutional designs optimized for single-stage grid-based detection, eliminating the need for region-proposal modules. The motivation for focal loss arises from the inherent limitations of conventional single-stage detectors compared with their two-stage counterparts. In two-stage frameworks (e.g., Faster R-CNN), class imbalance is mitigated through a cascade of stages: a region proposal network (RPN) first reduces candidate object locations to a sparse set (e.g., 1- 2k proposals), filtering out the majority of background samples [51]. The subsequent stages employ sampling heuristics, such as fixed foreground-background ratios or online hard example mining (OHEM), to balance training [137]. By contrast, single-stage detectors process a densely sampled set of candidate locations (often exceeding 100k per image), leading to an extreme imbalance between foreground and background. RetinaNet circumvents this issue entirely through focal loss, which intrinsically recalibrates the loss function to prioritize challenging examples without explicit sampling strategies [139,140,141].
5.3. Transformer-Based Object Detection
as
Transformer-based object detectors represent a modern paradigm shift, leveraging the self-attention mechanism borrowed from the Transformer architecture, which was originally developed for Natural Language Processing(NLP) rather than solely relying on traditional Convolutional Neural Networks(CNNs) [82]. By employing self-attention, these models effectively capture global contextual information and the long-range relationships between objects and features across the entire image. This global awareness contributes to their inherent robustness in object detection tasks [82,83].
The block diagram shows the working process of a Transformer-based object detector (e.g., DETR). As shown in Figure 16, the method is built on direct set prediction. The CNN Backbone first finds a reduced-resolution and semantically rich Feature Map of the input image. This 2D map is positionally encoded and linearized and fed into the Transformer Encoder. Multi-Head Self-Attention by the Encoder Global contextualizes the image features as well as creates dependencies between the entire range of the scene [142]. The Transformer Decoder is a process of these contextualized features, motivated by predetermined, instantiated Object Queries. The queries focus on the image features, with queries attending to each other through Cross-Attention and, in doing so, refining the object representations. Lastly, bounding box coordinates and class labels are directly output from the resulting object embeddings by Prediction Heads (FFNs). The architecture does away with the old-fashioned heuristics such as Non-Maximum Suppression (NMS) and anchor design, which make the object detection pipeline straightforward [143].
The Vision Transformer (ViT) represents a foundational development in the application of transformer architectures to computer-vision tasks. Initially proposed by Dosovitskiy et al.[144], ViT adapts the standard transformer model, which has been widely successful in natural language processing, to operate on image data.
Figure 17.
The basic architecture of Vision Transformer (ViT) mentioned by the Google research team [145].
Figure 17.
The basic architecture of Vision Transformer (ViT) mentioned by the Google research team [145].

The approach begins by dividing an input image into fixed-size patches (e.g., 16×16 pixels), each of which is flattened and linearly projected into an embedding space. These patch embeddings are then treated as a sequence of tokens, akin to words in text, and processed using a standard transformer encoder, as shown in Figure 17. This allows the model to capture global dependencies across the entire image, in contrast to convolutional neural networks (CNNs), which primarily rely on local receptive fields and hierarchical feature aggregation [146,147,148]. One of the strengths of ViT is its ability to model long-range relationships and contextual information, which are crucial in dense scenes or when multiple objects are present [149]. However, because it lacks the inductive biases inherent in CNNs, such as locality and translation invariance, ViT requires very large datasets and extensive computational resources for effective training. Although ViT was originally designed for image classification tasks, it has been integrated into various object detection frameworks, typically serving as a backbone feature extractor when combined with detection heads or region-based architectures [150,151]. The use of ViT in object detection has laid a strong foundation for the development of more advanced transformer-based detectors [152,153,154,155].
The DEtection TRansformer (DETR), introduced by Carion et al. in 2020, marked a significant departure from traditional object detection architectures [156]. DETR proposes a fully end-to-end transformer-based framework that eliminates several hand-designed components commonly used in conventional methods, such as anchor generation, region proposal networks, and non-maximum suppression (NMS). In DETR, a convolutional neural network (typically ResNet) is first used to extract visual features from the input image. These features are then passed to a transformer encoder-decoder structure. The encoder processes spatial features globally using self-attention mechanisms, whereas the decoder predicts a fixed number of object instances using learned object queries. Each query attends to different parts of the image and outputs a bounding box and a class prediction.
Figure 18.
An overview of the Detection Transformer(DETR) and its modifications proposed by recent methods to improve performance and training convergence [157].
Figure 18.
An overview of the Detection Transformer(DETR) and its modifications proposed by recent methods to improve performance and training convergence [157].

A set-based global loss, inspired by bipartite matching using the Hungarian algorithm, ensures that predictions are matched uniquely to ground-truth objects, thus avoiding duplicate detections. Despite its elegance and simplicity, DETR suffers from certain limitations, most notably its slow convergence rate and difficulty in detecting small or overlapping objects. Training DETR to perform on par with traditional detectors requires a large number of epochs and a large dataset [158]. Nonetheless, DETR introduced a novel way of thinking about object detection as a direct set prediction problem, and inspired a wide array of improved models in subsequent research.
Swin Transformers represent a significant advancement in object detection architectures, integrating the global contextual modeling capabilities of transformers with hierarchical feature processing inspired by convolutional neural networks (CNNs) [159]. The architecture employs a shifted window-based self-attention mechanism that partitions input images into non-overlapping local windows and computes self-attention within and across adjacent windows through a cyclic shifting strategy. This approach reduces the computational complexity while preserving the ability to model long-range dependencies, addressing the quadratic scaling limitations of standard transformer architectures. A key innovation lies in its hierarchical design, which progressively merges image patches in deeper layers to generate multi-scale feature maps (e.g., 1/4, 1/8, and 1/16 resolutions), emulating the pyramidal feature extraction of CNN-based systems, such as Feature Pyramid Networks (FPNs) [97]. This hierarchical processing enables swine transformers to achieve superior performance in dense prediction tasks, where both fine-grained localization and high-level semantic reasoning are critical.
When deployed as a backbone network in detection frameworks such as Mask R-CNN [160] or Cascade Mask R-CNN [161], Swin Transformers demonstrate state-of-the-art results on benchmark datasets like COCO [47], outperforming conventional CNN-based counterparts by significant margins in both average precision (AP) and small-object detection metrics. For instance, Swin-Large achieves a 58.7 box AP on COCO test-dev, surpassing ResNeXt-101-FPN by 4.5 AP [159]. The efficiency of the architecture is further enhanced through variants such as Swin-Tiny (Swin-T) and Swin-Large (Swin-L), which provide scalable solutions for resource-constrained and high-accuracy applications. By eliminating the need for handcrafted components, such as anchor boxes and non-maximum suppression (NMS), Swin Transformers also simplify the detection pipeline while maintaining robustness to occlusion and scale variations.
Spatially Modulated Co-Attention (SMCA) is a variant of the original DETR framework that seeks to address its inefficiencies in convergence and its inability to focus spatially on specific object regions. In the standard DETR model, attention is uniformly distributed across the image, which can dilute the focus and result in slower and less precise object localization. SMCA introduces a spatial modulation mechanism that enhances attention maps using a Gaussian-based prior centered around the reference points of object queries [162]. This modification encourages each query to attend more heavily to the localized regions that are likely to contain an object. The co-attention module incorporates spatial information into the transformer decoder, allowing it to correlate object queries better with relevant image regions [150]. This not only accelerates training convergence significantly but also improves detection performance, particularly for small objects and those in cluttered or crowded scenes. SMCA retains the set-based prediction nature of DETR, maintaining its elegant end-to-end learning paradigm while overcoming one of its critical drawbacks-spatial inefficiency [157]. The use of spatial modulation enhances the interpretability of the model and allows it to generate more focused and contextually relevant detections. Consequently, SMCA represents a practical enhancement for scenarios in which training time and localization accuracy are crucial considerations.
Anchor DETR is an evolution of the original DETR framework that integrates the concept of anchor boxes, a component long used in traditional object detection models like Faster R-CNN and YOLO [163]. In standard DETR, object queries are learned from scratch and do not incorporate any prior information regarding object locations or scales, which can slow down training and lead to suboptimal initial predictions. Anchor DETR addresses this issue by initializing object queries using predefined anchor boxes and effectively combining the strengths of anchor-based and transformer-based detection [82]. These anchors provide a rough estimation of object positions and serve as informative priors during the early stages of training. The transformer decoder then refines these predictions through multiple layers of attention, resulting in the final bounding boxes and class scores. This hybrid approach improves the convergence speed of DETR and enhances its ability to detect objects at varying scales and aspect ratios. By leveraging the prior knowledge inherent in anchor boxes, Anchor DETR benefits from faster optimization and improved stability while preserving DETR’s end-to-end learning and set-based output design of the DETR. This represents a practical bridge between traditional and modern paradigms in object detection and is particularly effective in datasets where object scales vary widely [82,164,165].
Deformable DETR, also known as Deformable Set Transformer (DESTR), is a significant improvement over the original DETR model, aimed at enhancing both training efficiency and detection accuracy [166]. One of the main limitations of DETR is its reliance on dense global attention, which makes it computationally expensive and slow to converge. Deformable DETR addresses this by introducing a deformable attention module that allows each query to attend to a sparse set of key sampling points around a reference location rather than the entire feature map [167,168]. This sparsity reduces computational complexity and enables the model to focus better on relevant regions, especially when objects are small or partially occluded. Moreover, Deformable DETR supports multiscale feature fusion, drawing information from different levels of the backbone network to improve object detection across varying resolutions. This is particularly advantageous for detecting objects of different sizes in a single image [82,165]. The model preserves the core structure of DETR, including the use of a transformer decoder and set-based prediction with bipartite matching, but introduces greater flexibility and efficiency through deformable attention. Consequently, Deformable DETR achieves state-of-the-art performance with significantly faster training and inference times, making it more practical for real-world applications that require both accuracy and speed.
Table 5.
Comparative Analysis of Transformer-based Object Detection and YOLO object detector.
| Model | Architecture | Backbone | Detection Approach | Key Features | Attention Mechanism | Speed vs. Accuracy | Strengths | Weaknesses | Use Cases |
|---|---|---|---|---|---|---|---|---|---|
| ViT | Single-stage | Vision Transformer | Query-based | Uses ViT backbone; often paired with detection heads (e.g., ViTDet) | Global self-attention | Moderate speed, High accuracy | Strong global context modeling; scalable to large datasets | Computationally heavy; requires pretraining | High-accuracy detection tasks |
| DETR | Single-stage | CNN + Transformer | Query-based (set prediction) | End-to-end; anchor-free; bipartite matching loss | Encoder-decoder cross-attention | Slow training, Moderate inference | Simplifies pipeline; handles occlusions well | Slow convergence; high memory usage | General-purpose detection |
| SMCA DETR | Single-stage | CNN + Transformer | Query-based | Spatial-Modulated Co-Attention to focus on likely regions early | Multi-scale conditional attention | Faster convergence than DETR | Reduces training time; improves spatial reasoning | Still slower than YOLO/RetinaNet | Real-time with better accuracy |
| Anchor DETR | Single-stage | CNN + Transformer | Anchor-guided queries | Anchor-based object query initialization; improves convergence | Anchor-based cross-attention | Faster convergence than DETR | Better localization; efficient query design | Limited to predefined anchor layouts | Balanced speed/accuracy tasks |
| Swin Transformer | Single-stage | Swin Transformer | Hierarchical query-based | Shifted window attention; multi-scale feature representation | Local + shifted window attention | High accuracy, Moderate speed | Scales efficiently; fine-grained feature capture | Complex architecture; high compute | High-resolution detection |
| DESTR | Single-stage | Deformable CNN + Transformer | Deformable attention | Combines deformable convolutions with transformers; dynamic feature sampling | Deformable attention | Faster than DETR | Efficient attention; good for small objects | Less mature; under development | Scene-heavy or crowded environments |
| YOLO Family | Single-stage | CNN (DarkNet, etc.) | Anchor-based | Grid-based predictions; multi-scale feature fusion (e.g., FPN in YOLOv3+) | None (CNN-only) | Very fast, Lower accuracy | Real-time performance; lightweight variants (e.g., YOLOv8-nano) | Struggles with small objects; anchor bias | Real-time applications (video, drones) |
5.4. LLM Based Models
The rapid evolution of large vision-language models (LVLMs) has redefined the landscape of object detection by leveraging the reasoning, compositional understanding, and zero-shot generalization power of Large Language Models (LLMs) [169]. Unlike traditional detectors that rely on supervised training with closed-vocabulary datasets, LLM-based approaches enable open-vocabulary recognition, spatial grounding, and contextual reasoning by integrating large-scale multimodal pretraining. These models align textual and visual modalities at scale, often incorporating billions of parameters, advanced tokenization strategies, and architectures ranging from encoder-only to hybrid encoder-decoder systems. Figure 19 shows the working principle of the Large language-based object detection models, where the input image is usually analyzed with the LLM model, where the visual encoder and the contextual proposal generator help the model to detect the class of object, draw the bounding box(BBOX), and enhance the detection accuracy[169].
Table 7 highlights recent multi-modal LVLMs tailored for object detection, summarizing their key contributions, advantages, and limitations. In object detection, LLMs can interpret natural language queries to identify and localize objects in images, reason about object relationships, and provide contextual descriptions of scenes [170]. Their applications span autonomous driving, medical imaging, remote sensing, and human-AI interaction, where flexible reasoning and cross-modal understanding are critical. Thus, LLM-based object detection represents a paradigm shift from task-specific perception models to general-purpose reasoning systems that unify language, vision, and contextual understanding. Below is a detailed overview of a few LLM-based object detection models,
GPT-4, introduced by OpenAI, is the first multimodal member of the GPT family that can process both images and text, thereby extending the generative pre-trained transformer framework into vision-language reasoning and object detection. Unlike conventional object detectors such as YOLO [105,118], which employ convolutional or transformer-based backbones with specialized detection heads, or DETR [156], which formulates detection as a direct set prediction problem using transformer encoders and decoders, GPT-4 operates as a foundation model without task-specific supervision. Instead, it leverages large-scale pre-training on multimodal data and is fine-tuned via Reinforcement Learning from Human Feedback (RLHF), enabling it to perform zero-shot and few-shot object detection by interpreting natural language prompts. Technically, GPT-4 relies on a transformer architecture with masked self-attention, extended context windows, and multimodal alignment layers that connect visual embeddings with textual representations [171]. This allows GPT-4 not only to localize objects but also to reason about object relationships and scene context in a way that YOLO, Faster R-CNN [90], and DETR struggle with, as those models are typically constrained to bounding box regression and class prediction. While GPT-4 demonstrates superior generalization across unseen domains, it also faces challenges such as hallucinations, high computational costs, and a lack of fine-grained precision compared to specialized detectors. Nonetheless, GPT-4 represents a paradigm shift in object detection by unifying perception and reasoning within a single generative framework, bridging the gap between vision-centric models and language-driven reasoning systems [172].
DeepSeek Janus-Pro represents an advanced multimodal large language model (LLM) framework that builds upon its predecessor, Janus, by integrating optimized training strategies, larger-scale data, and expanded model size to enhance both multimodal understanding and visual generation [173]. Unlike traditional object detectors such as YOLO [105] and DETR [156], which rely on task-specific architectures for bounding box regression, Janus-Pro adopts a decoupled visual encoding mechanism that enables it to process multimodal tasks in a unified autoregressive transformer framework. Specifically, it employs the SigLIP encoder [174] to extract high-dimensional semantic features for visual understanding, and a VQ-based tokenizer [175] to discretize images into codebook embeddings for visual generation. These feature sequences are mapped into the LLM input space via specialized adaptors and concatenated for joint processing, enabling Janus-Pro to perform not only object localization but also reasoning-rich understanding of object relationships within scenes. With its optimized training pipeline, including longer pretraining on ImageNet for pixel dependence modeling and focused Stage II training on dense text-to-image data, Janus-Pro achieves superior training efficiency compared to earlier methods. Its advantages lie in its ability to combine object detection with natural language reasoning, maintain strong multimodal comprehension, and support instruction-following text-to-image generation, positioning it as a versatile alternative to domain-specific detectors [172]. While limitations in input resolution (384×384) and fine-grained visual detail remain, Janus-Pro’s unified architecture demonstrates significant progress toward general-purpose object detection powered by LLMs [172].
KOSMOS-2.5 is a multimodal literate model designed by Microsoft to advance machine reading of text-intensive images, making it a unique contribution to large vision-language models (LVLMs) applied to object detection and document understanding [176]. Unlike conventional detectors such as YOLO [?] or DETR [156], which focus primarily on bounding box prediction, KOSMOS-2.5 introduces spatially-aware text block generation, where each detected text segment is assigned spatial coordinates and structured text output in Markdown format that preserves both layout and style [176]. Architecturally, it leverages a decoder-only autoregressive Transformer with task-specific prompts, pre-trained on a large-scale corpus of 357.4 million annotated document pages across diverse domains. For vision-text alignment, it automatically generates bounding box tokens that serve as anchors for localizing text-rich regions, enabling robust zero-shot detection and grounding in document-level tasks. A fine-tuned variant, KOSMOS-2.5-CHAT, achieves performance on par with much larger models (1.3B vs. 7B parameters) across nine text-rich VQA benchmarks, highlighting its parameter efficiency and scalable performance [172,176].
The key advantages of KOSMOS-2.5 in object detection lie in its ability to handle document-level visual grounding, zero-shot OCR-style detection, and structured output generation, making it particularly effective for tasks such as document parsing, infographic understanding, and scene-text localization [176]. Compared to GPT-4V, which excels in general multimodal reasoning but struggles with precise layout comprehension, and DeepSeek-VL2, which emphasizes high-resolution tiling for visual grounding, KOSMOS-2.5 stands out as a specialized model for literate multimodal intelligence [177]. However, challenges remain, particularly in handling multi-page documents and long-context reasoning, which limit its holistic comprehension in large-scale text-rich datasets. Despite these limitations, KOSMOS-2.5 provides a strong foundation for the next generation of LLM-based object detectors that combine OCR, layout analysis, and semantic reasoning into a unified multimodal framework [172].
DeepSeek-VL2 represents a major advancement in Mixture-of-Experts (MoE) Vision Language Models (VLMs), specifically designed to enhance multimodal reasoning and object detection through efficient high-resolution image processing [177]. Its architecture integrates three key components: a SigLIP-SO400M-384 vision encoder, a vision-language adaptor, and a DeepSeekMoE language model with Multi-head Latent Attention, which compresses the Key-Value cache into latent vectors for efficient inference and high throughput. A central innovation is the dynamic tiling strategy, where high-resolution images with varying aspect ratios are split into 384×384 tiles and processed along with a global thumbnail, enabling robust detection in dense visual tasks such as OCR, infographic understanding, and fine-grained object localization [177].
When compared with other LLM-based object detectors, DeepSeek-VL2 demonstrates notable advantages. GPT-4V, while a powerful general-purpose multimodal LLM, primarily focuses on semantic comprehension and zero-shot reasoning but often struggles with fine-grained spatial grounding [171]. Kosmos-2.5 introduces bounding box tokenization for spatial localization, but its reliance on fixed input resolutions limits its scalability in handling high-resolution or irregular-aspect-ratio images. LLaVA-Next improves vision-language alignment with CLIP-based encoders and supports bounding box outputs for visual question answering, yet it lacks the computational efficiency of MoE-driven architectures. In contrast, DeepSeek-VL2 combines dynamic resolution handling with MoE scaling, activating only a fraction of parameters (1.0B, 2.8B, or 4.5B) while delivering competitive or state-of-the-art performance across multimodal benchmarks [177].
Its advantages include efficient high-resolution image processing, better adaptability across domains, and superior scalability on commodity GPUs, making it more practical for real-world deployments where both semantic reasoning and precise object localization are required. Nonetheless, limitations such as a relatively short multimodal context window and occasional performance drops on blurry or unseen objects point to future directions for improvement. By striking a balance between efficiency, scalability, and fine-grained detection, DeepSeek-VL2 positions itself as a next-generation LLM-based object detector, surpassing many existing models in handling complex multimodal tasks [177].
InstructBLIP represents a significant advancement in vision-language instruction tuning, extending the pretrained BLIP-2 framework into a more generalized and instruction-aware multimodal model [178]. Unlike conventional object detectors that rely solely on bounding box regression and classification pipelines, InstructBLIP integrates an instruction-aware Query Transformer (Q-Former) that extracts visual features tailored to specific task instructions, enabling dynamic adaptation to diverse object detection scenarios. By leveraging 26 publicly available datasets covering a wide range of multimodal tasks and transforming them into instruction-tuning formats, InstructBLIP learns to follow natural language prompts while aligning them with visual representations. This design allows it to outperform its predecessors, such as BLIP-2 [179], Flamingo [180], and MiniGPT-4 [181], particularly in zero-shot detection and unseen task generalization. One of its core advantages lies in its instruction-tuned capability, where the model can flexibly adapt to user-defined object detection tasks (e.g., ’detect all vehicles and pedestrians in the scene’ or ’identify objects relevant to safety hazards’) without extensive retraining. Moreover, its integration of large-scale instruction data-both template-based and LLM-generated-enhances semantic understanding and fine-grained reasoning, offering superior interpretability and adaptability compared to rigid CNN- or Transformer-based detectors. These advantages make InstructBLIP not only a state-of-the-art vision-language model but also a promising foundation for next-generation instruction-driven object detection systems in autonomous driving, robotics, and intelligent surveillance [172,177].
LLaVA-OneVision is an open, large multimodal model (LMM) that advances object detection by unifying single-image, multi-image, and video scenarios within a single framework [182]. Developed from the LLaVA-NeXT series, it scales training with billions of image-text pairs, high-quality curated datasets, and stronger large language models (LLMs), enabling robust cross-scenario transfer learning. Technically, the model integrates a vision encoder with a decoder-only transformer LLM through an efficient alignment strategy, allowing it to process interleaved text-image sequences, perform instruction-driven object detection, and support multi-turn reasoning with precise spatial grounding [182]. Unlike prior models such as InstructBLIP [178], which rely heavily on Q-Formers, or Flamingo [180], which focuses mainly on few-shot adaptation, LLaVA-OneVision achieves state-of-the-art performance across detection benchmarks by leveraging joint training over single-image, multi-image, and video data. A notable advantage of LLaVA is its scalable architecture, which enables training and inference efficiency while eliminating the need for multiple task-specific detectors [183]. By consolidating insights in data, architecture, and multimodal representations, LLaVA-OneVision establishes itself as the first open-source model capable of general-purpose vision-language object detection with competitive performance against larger closed-source systems [184].
Grounding DINO 1.5, developed by IDEA Research, represents a major advancement in open-set object detection, offering both high-performance and edge-optimized variants. The suite includes Grounding DINO 1.5 Pro, designed for maximum accuracy and generalization, and Grounding DINO 1.5 Edge, tailored for real-time, resource-efficient deployment [185]. The Pro model scales up the architecture with an enhanced vision backbone and is trained on an expansive dataset of 20 million grounding-annotated images, resulting in richer semantic understanding and superior transferability. Empirically, it achieves 54.3 AP on COCO and 55.7 AP on LVIS-minival zero-shot transfer, setting new records in open-set detection [186]. In contrast, the Edge model prioritizes efficiency by reducing feature scales while maintaining robust detection performance; when optimized with TensorRT, it delivers 75.2 FPS with a 36.2 AP zero-shot score on LVIS-minival, making it ideal for edge computing and real-time applications [177]. Compared to earlier versions and contemporary detectors, Grounding DINO 1.5 strikes a balance between accuracy, scalability, and deployment efficiency, positioning itself as a versatile solution for diverse scenarios ranging from large-scale cloud-based detection pipelines to lightweight edge devices [172,177].
Florence-2 is a vision foundation model that introduces a unified, prompt-based architecture capable of handling a wide spectrum of computer vision and vision-language tasks, including object detection, captioning, visual grounding, and segmentation [187]. Unlike conventional large vision models that excel mainly in transfer learning but struggle with task diversity, Florence-2 leverages a sequence-to-sequence framework where text prompts serve as task instructions, enabling flexible adaptation across spatial hierarchies and semantic granularities [188]. A key strength of Florence-2 lies in its pre-training on FLD-5B, a large-scale dataset comprising 126 million images paired with 5.4 billion annotations, constructed using an iterative pipeline of automated annotation and model refinement. This massive and diverse dataset equips Florence-2 with fine-grained visual understanding at both coarse and detailed semantic levels. Technically, the model integrates multi-task learning within a unified representation space, allowing it to perform regional proposal segmentation, phrase grounding, and object detection without requiring task-specific architectures [189]. Empirical evaluations demonstrate that Florence-2 achieves state-of-the-art zero-shot and fine-tuning performance across multiple benchmarks, outperforming existing models in tasks such as detailed captioning and multimodal grounding[190]. Its advantages lie in its scalability, versatility, and superior generalization, positioning Florence-2 as a robust foundation for advancing general-purpose vision-language intelligence [172,177].
The YOLO-World object detector represents a significant advancement in the YOLO family by extending its capabilities from closed-set object detection to open-vocabulary detection through integration with vision-language modeling [191]. Unlike conventional YOLO detectors that rely on predefined categories [105,118], YOLO-World leverages large-scale vision-language pre-training to detect a wide range of novel objects in a zero-shot manner. Its core innovation lies in the Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN), which enables efficient cross-modal fusion between image regions and textual embeddings, while maintaining YOLO’s hallmark real-time efficiency [111]. Additionally, a region-text contrastive loss is introduced to align visual and linguistic representations, significantly boosting performance in open-world settings. Pre-trained with large-scale detection, grounding, and image-text datasets, YOLO-World achieves 35.4 AP at 52 FPS on the LVIS benchmark using a single V100 GPU, surpassing many state-of-the-art detectors in both accuracy and speed [191]. Beyond detection, fine-tuned YOLO-World also excels in downstream tasks such as open-vocabulary instance segmentation and visual grounding, making it a powerful and versatile real-time solution for real-world applications where object categories are dynamic and unpredictable [118,191].
Flamingo represents a pioneering family of Visual Language Models (VLMs) designed to rapidly adapt to novel multimodal tasks with only a few annotated examples, making it particularly suitable for object detection in low-data regimes [180]. Unlike conventional object detectors that rely on large-scale supervised training, Flamingo introduces architectural innovations to bridge pretrained vision-only and language-only models, enabling seamless integration of visual and textual modalities. Its design allows the model to process arbitrarily interleaved sequences of images, videos, and text, giving it robust flexibility in multimodal reasoning [180]. For object detection, Flamingo leverages in-context few-shot learning, where task-specific examples can be provided directly at inference without gradient-based retraining, allowing the system to generalize rapidly to unseen object categories [192]. Compared to contrastive models such as CLIP, Flamingo demonstrates superior adaptability across open-ended tasks (e.g., detecting objects described in natural language prompts) and structured tasks (e.g., multi-choice visual question answering) [193]. A major advantage of Flamingo lies in its general-purpose applicability: it can perform object localization, classification, and scene understanding in a unified framework without requiring separate models. Furthermore, its training on large-scale multimodal web corpora endows it with strong contextual grounding and transferability [180,180]. Although challenges such as hallucinations, sensitivity to in-context examples, and computational inefficiency remain, Flamingo sets a new paradigm in instruction-driven and few-shot object detection, offering significant potential for applications in autonomous systems, medical imaging, and intelligent surveillance [192,193].
The YOLOE object detector marks a major step forward in bridging the gap between efficiency and adaptability in open-world computer vision [194]. While conventional YOLO models achieve high accuracy and speed, they remain constrained by predefined categories, limiting their use in open and dynamic scenarios. YOLOE overcomes this limitation by unifying object detection and segmentation within a single framework that supports text prompts, visual prompts, and a prompt-free paradigm. Its core innovations include the Re-parameterizable Region-Text Alignment (RepRTA) strategy, which refines pretrained textual embeddings via a lightweight auxiliary network to enhance region-text alignment without additional inference cost; the Semantic-Activated Visual Prompt Encoder (SAVPE), which leverages decoupled semantic and activation branches for more accurate and efficient visual embedding; and the Lazy Region-Prompt Contrast (LRPC), which employs a built-in large vocabulary with specialized embeddings to identify objects in a prompt-free setting, eliminating the dependency on costly external language models [194]. Extensive experiments demonstrate YOLOE’s superior performance: on the LVIS dataset [195], YOLOE-v8-s outperforms YOLO-Worldv [191] by 3.5 AP, achieving 1.4 times faster inference and three times lower training cost, while on COCO, YOLOE-v8-L surpasses closed-set YOLOv8-L with a 0.6 APb and 0.4 APm gain, requiring nearly 4 times less training time. These results highlight YOLOE’s ability to deliver real-time, low-cost, open-vocabulary detection and segmentation, establishing it as a powerful and efficient benchmark for ’seeing anything’ in real-world applications [172].
Table 6 provides a comparative overview of four representative object detection models: YOLOv8, YOLO-World, Florence-2 [187], and YOLOE. The comparison highlights their architectural design, prompting mechanisms, key innovations, training datasets, performance benchmarks, and practical advantages. While YOLOv8 [121] remains a strong closed-set baseline for real-time detection, YOLO-World [191] extends YOLO with open-vocabulary capability through vision-language integration. Florence-2 represents a large-scale foundation model with unified prompt-based learning across diverse tasks, and YOLOE advances efficiency by integrating text, visual, and prompt-free mechanisms into a single framework, enabling real-time capabilities [194].
CogVLM is an open-source visual-language foundation model that redefines multimodal learning by shifting from shallow vision-language alignment to a deep fusion paradigm [196]. Unlike earlier approaches such as BLIP-2, which rely on mapping image features into the input space of a frozen language model, CogVLM introduces a trainable Visual Expert module integrated within the attention and feed-forward layers of the LLM, enabling seamless fusion of visual and textual representations without degrading natural language performance. The model, CogVLM-17B, achieves state-of-the-art results across 17 multimodal benchmarks, including image captioning (NoCaps, Flickr30k), visual question answering (OKVQA, TextVQA, ScienceQA), and visual grounding tasks (RefCOCO, RefCOCOg) [196]. Its successor, CogVLM2, extends this architecture with enhanced training recipes in 2024, a high-resolution input capability (up to 1344×1344 pixels), and specialized variants such as CogVLM2-Video, which supports temporal grounding through multi-frame input with timestamps [197]. These innovations allow the CogVLM2 family to achieve new state-of-the-art results on benchmarks such as MMBench, MM-Vet, MVBench, and VCG-Bench, while broadening applications to domains like document analysis, GUI comprehension, and video understanding. Despite these advancements, limitations remain, including sensitivity to hallucinations, the need for better supervised fine-tuning (SFT) [198] and reinforcement learning with human feedback (RLHF) [199], and trade-offs between computational overhead and deployment efficiency. Nevertheless, CogVLM and CogVLM2 establish a new benchmark for open-source visual language models, offering a scalable foundation for both academic research and real-world multimodal applications.
The OWL-ViT ( Open-World Localization with Vision Transformers ) [200] marks an important advance toward finding objects with open vocabulary using a conventional image detection Vision Transformer (ViT) [145] architecture based on pretraining on large datasets of images and language. In contrast to more traditional detectors based on closed-set categories, OWL-ViT is trained to denote a wide range of image-text pairs via contrastive learning (a common technique in CLIP), which allows detecting objects in zero-shot text-conditioned mode without such an object having explicit bounding box markings, regardless of the perceived category. The model is not only much more flexible towards open-world situations such as text queries (e.g., a red car) but it is also compatible with image queries in one-shot-detection scenarios. Although it involves very few architectural modifications (other than the ViT backbone) [145], OWL-ViT demonstrates high one-shot prediction on benchmarks, including LVIS [195], signifying the portability and transferability of image-level pretraining. But this depended on box-level annotations that were limited and middle-scale training scales, thus limiting performance, particularly on rare or long-tailed categories.
To this end, OWLv2, as an extension of OWL-ViT, was suggested, and an additional recipe, named OWL-ST (self-training), was suggested, combining an existing detector with pseudo-box annotations generated on web-scale image-text data [201]. This method is an effective way of moving the divide between large-scale weak supervision and object-level (during training) detection of images and text in a scale ranging between 10M and >1B pairs of images and text. OWLv2 supports the ViT backbone technically, eliminating the need to change it to address web-scale data, using efficient pseudo-label filtering, optimized label space choice, and optimized training plans. The outcome is that the performance of OWLv2 with L/14 architecture on LVIS rare categories is 31.2% to 44.6 points, which is significant (43), and is a percentage improvement in zero-shot detection over LWLv2 with L/14 architecture [195]. The key features of OWLv2 include scale to new objects and its strong generalization, in addition to high computational costs of training on a billion images, the possibility of reproducing pseudo-label noise, as well as fine-grained localization, which pose problems when using it with task-specific detectors[201].
Table 7.
Comparison of LLM-based Object Detection Models with their Key Contributions, Efficiency, Limitations, and Technical Features.
Table 7.
Comparison of LLM-based Object Detection Models with their Key Contributions, Efficiency, Limitations, and Technical Features.
| Model | Key Contributions | Efficiency | Limitations | Technical Features |
|---|---|---|---|---|
| GPT-4V [171] | Open-vocabulary detection, bounding box descriptions with reasoning | Scalable, real-time | Coarse localization, needs external grounding | 1.8T params, CLIP-ViT-L encoder, decoder-only |
| DeepSeek JanusPro [173] | High-resolution small-object localization, pretrained from scratch | Efficient at 7B | Training data undisclosed, unclear generalization | SigLIP-Large-Patch16-384, decoder-only |
| DeepSeek-VL2 [177] | Multi-task detection + captioning via MoE | Scalable with 74 experts | High computational cost | 4.5B×74 MoE, SigLIP + SAM-B, decoder-only |
| Kosmos-2.5 [176] | Zero-shot detection with spatial tokens | Lightweight (1.3B) | Limited fine-grained detection | ViT-L encoder, unigram tokenizer, enc-dec |
| InstructBLIP [178] | Instruction-tuned for versatile VQA + detection | Strong generalization | Not specialized for detection | 13B, ViT + Flan-T5/Vicuna, enc-dec |
| LLaVA-OneVision [182] | Bounding box outputs in VQA | Mid-scale (7B) | Limited dataset scale | CLIP-ViT-L + LLaMA-2, decoder |
| Grounding DINO1.5 [185] | LLM-integrated zero-shot detection (47.7 mAP) | Compact (110M) | Focused on detection, less reasoning | Swin-B + BERT, encoder |
| Florence-2 [187] | Unified captioning + detection | Moderate (5B) | Requires FLD-900M scale | ViT-g + T5, enc-dec |
| YOLO-World [191] | Real-time open-vocab detection (60+ FPS) | Highly efficient (42M) | Limited by YOLO backbone | YOLO-CSP + CLIP, encoder |
| YOLOE [194] | Prompt-free zero-shot detection | Real-time, efficient | Weaker reasoning than LVLMs | 1.3B, RepRTA + SAVPE, enc-dec |
| Flamingo [180] | Advanced open-vocab reasoning | High-capacity (80B) | Computationally heavy | ALIGN/M3W, Chinchilla backbone, decoder |
| CogVLM [71] | Contextual multimodal reasoning | Strong at 18B | Resource intensive | CLIP-ViT-L/14 + Vicuna, enc-dec |
| OWL-ViT v2 [200] | Efficient open-vocab detection (47.0 mAP) | Lightweight (630M) | Limited LLM reasoning | ViT-B/16 + BPE, encoder |
| DINO-GPT4-V [171] | Two-stage refinement with GPT-4 | Mid-scale (1.2B) | Requires GPT-4 integration | DINOv2 + GPT-4, hybrid |
| VisionLLM [169] | Unified prompt-based detection | Moderate (13B) | Not specialized for small objects | ViT-L + LLaMA, decoder |
DINO-GPT4-V is an advanced object detection framework [171] that builds on the DINO (DETR with Improved deNoising anchOr boxes) architecture [185]. Like DETR [156], DINO is an end-to-end Transformer-based detector composed of a backbone for multi-scale feature extraction, a multi-layer Transformer encoder-decoder, and multiple prediction heads. Its key innovations include a contrastive denoising training scheme that uses hard negative samples, a mixed query selection strategy for better anchor initialization, and a look-forward twice mechanism that improves gradient flow between decoder layers [202]. These upgrades lead to significant improvements in training efficiency and detection accuracy, achieving 49.4 AP in 12 epochs and 51.3 AP in 24 epochs on COCO with a ResNet-50 backbone, surpassing DN-DETR by +6.0 AP and +2.7 AP, respectively. Additionally, when scaled with stronger backbones such as Swin-L and pre-trained on large datasets like Objects365, DINO reaches state-of-the-art performance at 63.3 AP on COCO test-dev, outperforming competing models with less model size and training data [203]. By combining this high-performing DINO foundation with GPT4-V’s multimodal reasoning ability, the DINO-GPT4-V detector not only inherits DINO’s robust end-to-end optimization and high accuracy but also benefits from enhanced contextual understanding and vision-language alignment, making it highly effective for complex real-world object detection tasks [204].
Large language models (LLMs) have demonstrated remarkable progress toward artificial general intelligence by exhibiting strong zero-shot generalization and user-adaptive task capabilities [205,206]. However, in computer vision, existing vision foundation models (VFMs) remain largely confined to pre-defined tasks, limiting their flexibility compared to LLMs [169]. To address this gap, authors present VisionLLM, a novel LLM-based framework that treats images as a foreign language and unifies vision and language tasks through natural language instructions [207]. The architecture comprises three key components: a unified language instruction interface for defining and customizing vision-centric tasks, a language-guided image tokenizer for encoding visual content aligned with prompts, and an LLM-based open-ended task decoder that generates predictions tailored to the given instructions. This design enables VisionLLM to seamlessly bridge vision and language, supporting fine-grained object-level reasoning as well as coarse-grained task-level customization [207]. Extensive experiments demonstrate that VisionLLM achieves strong generalization across diverse tasks, attaining over 60% mAP on COCO, comparable to task-specific detectors while offering far greater flexibility. More recently, VisionLLM-v2 extends this paradigm into a comprehensive multimodal large language model (MLLM) that integrates perception [208], understanding, and generation within a single framework. With its novel super link mechanism for information transfer between the MLLM and task-specific decoders, VisionLLM-v2 effectively mitigates multi-task training conflicts and broadens applicability to tasks such as object localization, pose estimation, and image generation. Collectively, VisionLLM and VisionLLM v2 set new baselines for generalist multimodal models, highlighting the potential of LLM-driven approaches to unify and advance open-ended vision-centric tasks [169,207].
Table 8 compares four major deep learning methods of object detection (OD) with the primary focus on the underlying performance trade-offs on the choice of architecture. The Two-Stage Detectors (TSD) are typified by the Faster R-CNN family to focus on detection accuracy (mAP) through the use of a serial processing pipeline. The first stage produces sparse Region Proposals, and a second stage, with exclusive refinement and classification, comes after the first one. The success or failure is the most accurate mechanism, especially with smaller and occluded objects, but there is a cost of inference speed since it is bound to a multi-step latency cost. Single-Stage Detectors (SSD), or similar to the YOLO family, on the contrary, are fast and perform in real time [209]. They perform the OD task by a single network forward pass with dense sampling, which makes the pipeline simpler and computationally less consuming. Although historically poorer than TSDs in accuracy, the current versions of SSD have attained near-parity and are thus the standard when time is critical. Transformer-based object detectors are another recent evolution, which takes advantage of the mechanism of self-attention to express the relationship between global contexts across the entire image. This method facilitates an end-to-end set prediction, which is often heuristic-free, such as Non-Maximum Suppression (NMS). Transformer-based detectors are highly accurate and robust, with recent models having real-time speeds, and therefore they are very competitive in general-purpose detection where the global context is paramount [143,209].
Lastly, LLM-Based Detectors bring the concept of cross-modality, where the application combines vision with a Language Model [169]. Although they have demonstrated the greatest architectural complexity and longer latency, their main application is in open-vocabulary or zero-shot detection. They use linguistic cues to infer object presence and position, and allow seeing new object classes that have never been observed during training, which is especially useful with contextually relevant object identification. It is ultimately a matter of a trade-off of the limitations of latency, precision, and application sphere that decides which of these paradigms to use [169].
6. Datasets
Along with the object detection algorithms, the dataset plays a pivotal role in object detection by training and testing the object detection algorithms for various applications. The availability of large-scale image datasets has been one of the most critical drivers of progress in object detection research. Robust datasets not only provide the foundation for training deep learning models but also serve as standardized benchmarks for fair performance evaluation across competing algorithms [210]. Well-curated datasets such as Pascal VOC, MS COCO, and ImageNet have enabled the development of increasingly accurate and generalizable detectors, while domain-specific datasets (e.g., BDD100K for autonomous driving or CrowdHuman for surveillance) support specialized applications. However, despite their significance, existing datasets face persistent challenges, including limited diversity, label noise, class imbalance, dataset bias, and the high cost of annotation. These problems often lead to poor generalization when models are deployed in real-world, unconstrained environments [109]. Emerging solutions to these issues include semi-supervised and self-supervised learning that reduce dependence on exhaustive labeling, synthetic dataset generation using generative models and simulators to increase diversity, and active learning techniques to prioritize the most informative samples for annotation. Addressing these challenges is essential to bridge the gap between benchmark-driven research and practical, real-world object detection systems [83]. Based on the uses of the object detection approaches, the paper divided the benchmark datasets into a few groups, such as Generic object detection datasets, autonomous driving datasets, pedestrian detection, remote sensing, face detection, medical imaging, satellite imaging, and agricultural image datasets [83].
6.1. Generic Object Detection Datasets
Generic object detection datasets serve as the fundamental bedrock for advancement in computer vision research [83]. These datasets, such as the pioneering PASCAL VOC, the context-rich MS COCO, and the massively scaled Open Images, provide the essential annotated imagery required to train, validate, and benchmark algorithms. Their importance cannot be overstated; they establish standardized evaluation metrics like mean Average Precision (mAP), enabling fair comparison of state-of-the-art models and driving the field forward [2]. The evolution of these datasets-marked by increasing volume, diversity, and annotation complexity-has been directly correlated with performance breakthroughs, particularly with the advent of deep learning. The applications enabled by models trained on these datasets are vast, spanning autonomous vehicle navigation, robotic vision, image retrieval systems, automated video surveillance, and content-based image moderation, forming a critical component of modern intelligent systems. Table 9 shows a comparison of the generic object detection image datasets, highlighting their advantages, disadvantages, and key contributions.
The Caltech-101 [211] dataset, introduced in 2003 by the California Institute of Technology, is a widely used benchmark for computer vision research, particularly in image recognition, classification, and categorization. It contains 9,146 images distributed across 101 object categories (such as faces, ants, pianos, and watches) along with an additional background category. Each image is accompanied by annotations outlining object boundaries, and the dataset also includes a MATLAB script for visualization. The dataset was originally constructed by selecting object categories, retrieving candidate images from Google Images, and manually filtering them to ensure category consistency. Building on this, the Caltech-256 [212] dataset was released in 2006 to address some of the limitations of Caltech-101 [211]. It comprises 30,607 images spanning 256 object categories, thereby more than doubling the number of categories. Improvements over Caltech-101 include: (a) increasing the minimum number of images per category from 31 to 80, (b) avoiding artifacts caused by rotated images, and (c) introducing a larger and more diverse clutter category to enhance background rejection. The creators proposed multiple testing paradigms for evaluating classification performance and benchmarked the dataset using both simple evaluation metrics and advanced methods such as the spatial pyramid matching algorithm [44].
A significant amount of time, spanning from 2005 to 2012, was dedicated to developing and managing a set of benchmark datasets, which are generally used to detect basic object categories. These benchmark datasets are known as the PASCAL VOC datasets [45] and consist of 20 categories of objects (such as birds, bottles, persons, bicycles, and dogs) that were present in approximately 11,000 images [1,2]. These 20 categories were grouped into four main classes: vehicles, animals, household objects, and people. The dataset includes over 27,000 labelled object instance bounding boxes, almost 7,000 of which have detailed segmentations. Among the categories in the VOC2007 dataset, the "person" class is the largest, with nearly 20 times more instances in the training set compared to the smallest class, "sheep" [1,2].
Introduced in 2015, the MS-COCO dataset [47] has gained popularity as one of the most challenging benchmark datasets for object detection. It contains 91 common objects that a 4-year-old can easily recognize. These objects were captured in their natural context, and 82 of them had more than 5000 instances. It contains more than two million instances and an average of 3.5 categories per image [1,2,47]. Additionally, it has an average of 7.7 object instances per image, which is higher than in the other popular datasets. The images in MS-COCO are taken from various viewpoints and contain a lot of contextual information, as they also focus on the scene-understanding task. The most frequent category in this dataset is "person" with almost 800,000 instances, while the least frequent category is "hair dryer," with only about 600 instances in the entire dataset [1,2].
The Visual Genome dataset [213], introduced in 2016, is a large-scale benchmark designed to advance research in object detection, scene understanding, and relationship modeling. It contains 108,077 images with dense annotations that go beyond simple object bounding boxes. In total, the dataset provides over 3.8 million object instances across 1,600 object categories, annotated with 2.8 million attributes and 2.3 million pairwise relationships. Each image is further enriched with region descriptions, scene graphs, and question-answer pairs, making it one of the most comprehensive resources for studying visual understanding [44,213].
The Open Images dataset [46], created by Google in 2017, is a vast collection of 9.2 million images, each labeled with various annotations including image-level labels, object bounding boxes, and segmentation masks. Since its release, it has undergone six updates and is the largest dataset available for object-detection tasks. It contains 16 million bounding boxes for 600 categories found in 1.9 million images. The creators of Open Images carefully selected interesting, complex, and various images, resulting in an average of 8.3 object categories per image. It also provides annotations for visual relationships between pairs of objects in specific relations, such as ’beer on the table’, as well as segmentation masks for 2.8 million instances of objects across 350 different classes [1,2].
The ImageNet Large-Scale Visual Recognition Challenge (ILSVRC) [214] was a yearly contest held annually from 2010 to 2017 and has been widely used as a standard for assessing algorithm performance. It consists of over one million images with 1,000 object classification categories. For the OD task, 200 of these categories were selected manually, comprising more than 500,000 images. This dataset was constructed using ImageNet and Flickr [1,2].
6.2. Autonomous Driving Datasets
The key foundation on which a safe and robust vehicle perception system can be built is the object detection datasets used in autonomous driving, mainly those that are captured through the street-view camera perspective, with a focus on pedestrian detection. These datasets, gathered through advanced sensor packages, offer large-scale, annotated real-world data needed to teach deep learning models to recognize vehicles, infrastructure, and, most importantly, vulnerable road users like pedestrians [168]. The significance of these benchmarks cannot be underscored, given that they normalize evaluation and are the direct force behind the use of algorithms, which is increasingly important in terms of avoiding collisions as well as navigation [215]. Nonetheless, this area is faced with major challenges. The highly expensive nature of multi-sensor data capture and the complexity of balancing data to cover the wide range of rare long-tail situations, e.g., adverse weather or unexpected objects, exist in general street-view datasets. These questions are especially intensified in pedestrian-specific detection, where their task is thwarted by extremely high levels of occlusions, darkness, and even the large variety of human physiques and actions [215]. In addition, bias in training data that is based on geographic or demographic factors may severely impede generalization in the real world. To that end, the field is shifting towards more comprehensive datasets that focus on more time-stable data, more stringent attribute annotations (e.g., the degree of occlusion), and, most poignantly, a broader range of global environments [2]. The further development of these datasets is critical to create models able to reliably move in the complex and unpredictable real world and, ultimately, guarantee their safety to all road users.
One of the most well-known datasets for computer vision and autonomous driving is the KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) dataset [216]. It was the first large-scale dataset with a complete sensor suite for autonomous driving, which was also properly annotated for use as a benchmark, and was recorded in Karlsruhe in 2012 by the joint experiment of the Karlsruhe Institute of Technology and Toyota Technological Institute. Two camera streams (high-resolution RGB and grayscale stereo), a lidar with 100k points per frame, GPS and IMU measurements, object tracklets, and calibration data were all included in the collection. It can be applied to various autonomous driving tasks. The 7,500 training and 7,500 test photos for the 2D and 3D object detection benchmarks can be downloaded and viewed in several formats [216]. Table 10 shows the comparative analysis of autonomous driving datasets with their key features and use cases in object detection.
Berkeley Deep Drive, often known as BDD100K [217] released in 2017 by the University of California, Berkeley, and is the largest and most diverse autonomous driving dataset. A total of 100000 videos were divided into the training, validation, and test sets. Additionally, there were variants with similarly divided picture subsets from the films, each including 100,000 or 10,000 photos. The dataset comprises lane markers, pixel-wise semantic segmentation, instance segmentation, panoptic segmentation, pose-estimation labels, and ground truth annotations for all common road objects, with a total of 10 categories of vehicles, buses, pedestrians, bicycles, trucks, and motorcycles in the JSON format. The collection includes metadata such as the time of day and weather, in addition to ground truth labels. It is divided into a training set, test set, and validation set at a ratio of 7:2:1 [218].
The Cityscapes Dataset is a large-scale benchmark specifically designed for the semantic understanding of urban street scenes, serving as a critical resource for autonomous driving research [219]. Its primary strength lies in its high-quality, dense pixel-level annotations, which include not only semantic segmentation but also instance-level segmentation for vehicles and people, providing rich granularity for model training. Comprising 5,000 finely annotated and 20,000 coarsely annotated images, the dataset captures a diverse range of visual conditions across 50 European cities, including variations in season, weather, and urban layout. Beyond static images, Cityscapes is enriched with valuable metadata such as preceding and trailing video frames, stereo views, and ego-vehicle data, enabling research into temporal consistency and multi-task learning. This combination of precise polygonal annotations, scene diversity, and supplemental data has established Cityscapes as a premier benchmark for tasks like semantic, instance, and panoptic segmentation, driving advancements in the robustness of visual perception systems for autonomous vehicles [220].
The nuScenes dataset is a pivotal large-scale benchmark for autonomous driving research, renowned for its rich, multi-modal data and detailed 3D annotations [221]. Its core strength lies in 1,000 carefully curated scenes, each a 20-second snippet, which capture a wide array of challenging driving situations across diverse weather, lighting, and geographic conditions. While the original nuScenes focused on 3D object detection via LiDAR and radar, the nuImages extension significantly bolsters its 2D perception capabilities with 93,000 high-quality images annotated with instance masks, bounding boxes, and a comprehensive set of attributes for over 800,000 objects. This combination provides an unparalleled ecosystem that includes not only precise 3D cuboids but also 2D boxes, instance masks, and semantic segmentation labels [222]. The dataset’s meticulous design, which utilizes active learning to target challenging scenarios and rare classes, ensures a balanced and representative benchmark. Furthermore, the inclusion of temporal context, encompassing past and future frames, enables critical research into dynamic scene understanding, thereby solidifying nuScenes’ role as an essential resource for developing and evaluating robust perception algorithms for autonomous vehicles [223,224,225].
The Waymo Open Dataset has emerged as a cornerstone benchmark in the development of perception systems for autonomous driving, offering a comprehensive and large-scale resource for 3D object detection research [226]. Its principal advantages lie in the unparalleled richness of its multi-sensor data, which includes high-resolution, synchronized camera images and dense LiDAR point clouds across diverse and challenging urban driving scenarios. This extensive collection, annotated with precise 3D bounding boxes, provides the scale and variety necessary to train and evaluate complex deep learning models effectively. Furthermore, the dataset is supported by a robust metrics framework and accessible formats, including both a detailed TFRecord structure and an efficient column-based Parquet format, facilitating focused research. However, these advantages are tempered by significant challenges; the dataset’s immense size and complexity demand substantial computational resources for processing and model training, while the sparsity of annotations for rare objects and long-tail scenarios can limit generalization [227]. Additionally, the geographic and environmental focus of the data collection may introduce biases, potentially affecting model performance in underrepresented conditions. Despite these limitations, the Waymo Open Dataset remains an invaluable tool for advancing the state-of-the-art in automotive computer vision [228].
Table 10.
Comparative analysis of key autonomous driving datasets.
| Dataset (Year) | Sensor Modality | Size | Primary Annotations | Key Features / Use Case |
|---|---|---|---|---|
| KITTI (2012) [216] | Lidar, Stereo Cam, GPS/IMU | 15k frames | 3D/2D BBox, Tracklets | Pioneer dataset; complete sensor suite; General 3D perception |
| BDD100K (2017) [217] | Mono Camera | 100k videos | 2D BBox, Seg., Lanes | Time of day/weather metadata; 2D perception, scene parsing |
| Cityscapes (2016) [219] | Stereo Cam | 5k fine + 20k coarse imgs | Dense Semantic & Instance Seg. | 50 European cities; Pixel-wise urban scene understanding |
| nuScenes (2019) [221] | Lidar, Radar, 6x Cams | 1k scenes | 3D Cuboids, 2D BBox | Diverse weather; Temporal context; Multi-modal fusion |
| Waymo Open (2019) [226] | Lidar, 5x Cams | Large-scale | Precise 3D BBox | Diverse urban scenarios; Large-scale 3D detection |
| Caltech Ped. (2009) [229] | Mono Camera | 250k frames | 2D BBox (Peds) | Occluded, low-res pedestrians; Pedestrian detection benchmark |
| EuroCity Persons (2018) [230] | Mono Camera | 47.3k images | 2D BBox, Orientation | 31 EU cities; Pedestrian/cyclist detection and orientation |
| iROADS (2014) [231] | Mono Camera | 4,656 frames | Vehicle BBox | 7 adverse conditions; Vehicle detection in challenging weather |
| ONCE (2021) [232] | Lidar, Cams | 1M scenes, 7M imgs | 3D BBox | Massive scale; Self-supervised 3D perception |
| TU-DAT Dataset(2025) [233] | Camera and CCTV | 12,351 frames | 3D BBox | analyzing road accidents |
The Caltech Pedestrian Dataset [229], introduced as a benchmark of unprecedented scale, represents a significant milestone in pedestrian detection research, addressing critical limitations of earlier datasets by being two orders of magnitude larger and capturing the complexities of real-world environments. Its primary advantage lies in the collection of richly annotated video sequences from a moving vehicle, providing a realistic and challenging testbed with frequent occlusions, low-resolution subjects, and dynamic urban scenes that are crucial for developing robust automotive safety systems. This extensive data volume enabled more reliable benchmarking and revealed flaws in prevailing per-window evaluation metrics, leading to the proposal of more accurate per-image measures [234]. However, the dataset also highlighted key disadvantages, most notably the persistent failure of even state-of-the-art detectors-including the then-dominant HOG-based method-under these challenging conditions, particularly with small-scale or partially occluded pedestrians, which constituted the majority of real-world scenarios. Thus, while the Caltech dataset dramatically accelerated progress by providing a more realistic and scaled evaluation platform, it simultaneously exposed fundamental limitations in existing approaches, clearly delineating the need for future research into multi-frame analysis, contextual reasoning, and discriminative part-based models to achieve practical robustness for real-world applications [235].
The EuroCity Persons dataset [230] represents a significant advancement in the field of urban scene understanding, specifically designed to address the critical need for large-scale, diverse, and highly detailed data for pedestrian, cyclist, and rider detection. Collected from a moving vehicle across 31 cities in 12 European countries, it offers an unprecedented scale with over 238,200 meticulously annotated person instances in more than 47,300 images-nearly an order of magnitude larger than previous benchmarks. This extensive collection is further enriched with over 211,200 orientation annotations, providing valuable contextual information for developing more sophisticated detection models [234]. The dataset’s geographic and temporal diversity, encompassing both day and night conditions across varied European environments, enables robust evaluation of detector generalization and performance under real-world challenges. Benchmarking results with state-of-the-art deep learning architectures like Faster R-CNN, R-FCN, SSD, and YOLOv3 demonstrate that data volume and diversity remain pivotal factors in achieving superior detection accuracy, particularly for challenging cases involving small or occluded persons [236]. The EuroCity Persons 2.0 dataset was the updated version of a novel image dataset for person detection, tracking, and prediction in traffic [237]. The dataset was collected onboard a vehicle driving through 29 cities in 11 European countries. By facilitating a comprehensive analysis of factors such as training set size, annotation quality, and cross-dataset generalization, EuroCity Persons not only sets a new standard for evaluation, but also highlights future research directions toward developing life-saving perception systems for intelligent vehicles [238].
The iROADS (Intercity Roads and Adverse Driving Scenarios) dataset serves as a critical benchmark for developing robust vehicle detection and lane perception systems under challenging real-world conditions [231]. Comprising 4,656 color image frames captured from moving vehicles, its primary application lies in enhancing the resilience of Advanced Driver Assistance Systems (ADAS) by providing a diverse and comprehensive set of scenarios [18]. Unlike many contemporary datasets, iROADS is specifically curated to include seven distinct adverse weather and lighting categories, such as rainy day/night, snowy conditions, sun strokes, and tunnels, which are essential for testing and training classifiers to handle the complex visual noise and occlusions that severely impact lane marking visibility and road object detection. Although it lacks supplementary data like ego motion or stereo rectification, its focused variety in illumination and weather makes it a unique resource for benchmarking perception algorithms against the exact conditions where they are most likely to fail, ultimately aiming to improve safety and reliability in autonomous driving and driver-assistance technologies [239].
The ONCE (One millioN sCenEs) dataset represents a monumental leap in scale for 3D perception research in autonomous driving, addressing a critical bottleneck in the field [232], the reliance on massive, unlabeled real-world data to develop next-generation self-supervised and semi-supervised models. Its primary advantage lies in its unprecedented volume, comprising 1 million LiDAR scenes and 7 million corresponding camera images collected over 144 driving hours-a twenty-fold increase over existing benchmarks like nuScenes and Waymo. This vast scale, captured across diverse areas, periods, and weather conditions, is specifically designed to overcome the long-tail problem and facilitate robust exploration of label-efficient learning paradigms, offering an essential resource for training powerful industry-level perception models. However, the dataset’s formidable size also presents a significant disadvantage, as the computational resources and storage required for processing and experimenting with such a massive corpus may be prohibitive for many research institutions, potentially limiting its accessibility [240,241]. Furthermore, while it provides a benchmark for self-supervised 3D detection, its initial focus leaves benchmarks for other critical tasks like 2D detection or 3D segmentation as future work. Despite these challenges, the ONCE dataset’s colossal scale and diversity make it an indispensable catalyst for advancing the field beyond fully-supervised learning and toward more scalable and generalizable autonomous driving systems [241].
Introduced in May 2025, the TU-DAT dataset presents a significant advancement as a publicly available benchmark specifically designed for the computer vision analysis of traffic accidents from a roadside camera perspective [233]. Its primary effectiveness lies in addressing a critical gap in the field by providing a hybrid collection of approximately 280 real-world and simulated videos, complete with detailed spatiotemporal annotations for vehicle trajectories and collision types. The key advantages of TU-DAT include its unique focus on aggressive pre-crash behaviors like weaving and tailgating, its diversity in environmental conditions (daylight, night, fog, rain), and its proven utility in enhancing the performance of hybrid deep learning and logic-based reasoning frameworks for anomaly detection. However, the dataset is not without its disadvantages; the inclusion of simulated data, while necessary for scalability, may introduce a reality gap that could limit model generalization in purely real-world deployments. Furthermore, the accompanying text highlights practical challenges that also apply to models trained on TU-DAT, such as the legal and privacy constraints of deploying real-time roadside monitoring systems and the computational latency issues for edge device implementation. Despite these limitations, TU-DAT serves as an invaluable resource for developing proactive intelligent transportation systems, offering a foundational tool for research in predictive safety analytics and autonomous vehicle training [233].
The evolution of autonomous driving research has been significantly accelerated by the creation of large-scale, publicly available datasets. These benchmarks vary in sensor modalities, annotation types, environmental conditions, and intended tasks. The following table provides a comparative overview of the key datasets discussed in Table 10.
6.3. Remote Sensing (Aerial View) Datasets
Remote sensing datasets, comprised of imagery captured from satellites and unmanned aerial vehicles (UAVs), have become a cornerstone of geospatial analysis, providing a unique top-down perspective for large-scale object detection and monitoring [242,243]. Their importance in object detection is paramount, as they enable the identification, classification, and tracking of objects such as vehicles, aircraft, ships, and infrastructure across vast and often inaccessible geographical areas, facilitating applications that are impossible with ground-level imagery. The advantages of these datasets are demonstrated across a multitude of fields: in urban planning for mapping development and traffic flow, in precision agriculture for monitoring crop health, in disaster response for assessing damage and coordinating relief efforts, and in national security for surveillance and change detection, Figure 20 shows the images of different spots captured by drones and satellite in different weather condition [244]. However, these datasets also present significant limitations, including the high cost of capturing and storing high-resolution imagery, the immense computational resources required for processing large-scale geospatial data, and the technical challenges posed by object occlusion, extreme variations in scale and orientation, and atmospheric conditions like cloud cover that can obscure the scene, necessitating sophisticated algorithms to achieve robust performance [242,245]. The remote sensing datasets are discussed below with details, and the comparative analysis is shown in Table 11.
The Dataset for Object Detection in Aerial images (DOTAv2.0) [246], which was compiled from a range of sources, sensors, and platforms, is very well suited for detecting objects in aerial images. The images comprised items of various types, forms, and sizes, with resolutions ranging from 800x800 to 200,000x200,000 pixels. The dataset is constantly updated and is anticipated to expand. Expert aerial imaging annotators divided the ground truth annotations into 18 different groups with a total of 1.8M item instances shown in Figure 20. Each annotation additionally includes a difficulty score, which indicates how challenging it would be to discover the item in question using the oriented bounding box format (OBB). The dataset, which is divided into training, validation, and testing, is frequently utilized in computer vision competitions, such as LUAI2021 [247].
Introduced as a pivotal benchmark for overhead imagery analysis, the xView dataset represents a significant leap in large-scale object detection from satellite platforms, specifically captured by WorldView-3 satellites at an exceptionally high resolution of 0.3 meters per pixel [248]. Its foremost advantages lie in its unprecedented scale and diversity, featuring over 1 million meticulously annotated objects across 60 classes-encompassing everything from vehicles and buildings to complex mini-scenes-within 1,400 km2 of imagery, making it one of the most comprehensive public resources of its kind. This wide variety enables its application as a general-purpose benchmark for advancing critical research frontiers, including few-shot learning and domain adaptation for rapid deployment in disaster response and humanitarian efforts [249]. However, the dataset’s primary disadvantages and limitations stem from the inherent challenges of its domain: the "long-tail" distribution of its many classes likely leads to significant data imbalance, which can bias models towards more common objects, while the high resolution and vast geographical diversity, though beneficial, introduce substantial computational and algorithmic challenges in handling variations in scale, orientation, and lighting [250]. Furthermore, the complex, three-stage annotation process, while ensuring quality, underscores the immense cost and difficulty of curating such datasets, potentially limiting the frequency of similar future endeavors. Despite these limitations, xView serves as an indispensable unifying resource for bridging the gap between computer vision and geospatial analysis [248].
The VEDAI (Vehicle Detection in Aerial Imagery) dataset is a pioneering benchmark designed to advance the field of automatic target recognition by focusing on the complex challenge of detecting small vehicles in unconstrained aerial environments [251]. Its primary advantages lie in its meticulous design, where it provides over 3700 annotated vehicles across diverse conditions, including multiple orientations, lighting changes, occlusions, and varied backgrounds, and crucially offers imagery in both color and infrared spectral bands at multiple resolutions. This multi-modal approach allows for a more robust evaluation of detection algorithms. However, a significant limitation, as revealed by its own baseline studies, is the extreme difficulty of the task; state-of-the-art detectors at the time of its introduction performed poorly, achieving low recall rates and highlighting the dataset’s role as a challenging benchmark rather than a solved problem. This very limitation defines its future possibility: VEDAI serves as a critical tool for motivating the development of next-generation, more efficient algorithms capable of handling the intricacies of small object detection in aerial imagery, potentially inspiring advancements in model architecture, multi-spectral data fusion, and context-aware processing [252].
The DIOR (Dataset for Object Detection in Optical Remote Sensing Images) is a large-scale benchmark designed to advance research in deep learning-based object detection for optical remote sensing, containing 23,463 images and 192,472 instances across 20 diverse object categories, making it significantly larger and more comprehensive than many earlier datasets [253]. DIOR stands out for its rich variability in object sizes, both within and across classes, as well as across different spatial resolutions, imaging conditions, seasons, and weather scenarios. This diversity introduces challenging inter-class similarities and intra-class variations, enabling more robust evaluation of object detection models [244]. The dataset’s scale and diversity make it highly effective for developing and benchmarking data-driven methods in both computer vision and earth observation communities. However, despite these strengths, DIOR is not without limitations. While larger than previous datasets, it still does not fully capture the global diversity of remote sensing images, and some object classes may remain underrepresented. Nonetheless, DIOR provides a strong foundation and baseline for future research, helping to push the boundaries of object detection in remote sensing applications [244].
The INRIA Aerial Image Labeling Dataset was introduced as a pivotal benchmark to address a core challenge in remote sensing: the development of semantic segmentation models with strong generalization capabilities across geographically diverse urban landscapes [254]. Its primary advantage lies in its unique design, which rigorously tests a model’s ability to perform well on unseen data by strictly separating its training and testing sets by entire cities, using imagery from locations like Chicago for training and evaluating on entirely different cities like San Francisco or alpine towns in Austria. This structure, covering 810 km2 of high-resolution (0.3 m) imagery, forces models to learn fundamental "building" features that are invariant to changes in architectural style, urban density, illumination, and season, rather than simply memorizing the specifics of one location [255]. However, this strength is also the source of its main disadvantages and limitations. The dataset is restricted to a binary classification task ("building" vs. "not building"), which, while simplifying the problem, limits its utility for more complex scene understanding tasks that require multiple classes [256]. Furthermore, a significant limitation for the broader research community is that the ground truth labels for the test set are not publicly disclosed; performance can only be evaluated through an official online portal, which hinders independent error analysis and transparent benchmarking of new methods [245]. Despite these constraints, the dataset remains an invaluable resource for pioneering and evaluating methods designed for real-world applications, where models must perform reliably in novel and heterogeneous environments.
Table 11.
Comparative analysis of key remote sensing (Aerial view) datasets for object detection.
| Dataset | Scale | Resolution | Key Features & Advantages | Limitations |
|---|---|---|---|---|
| DOTA (v2.0) [246] | 1.8M instances | 800px to 200,000px | Extensive object categories (18). Oriented bounding boxes (OBB) for precise annotation. Includes instance difficulty scores. Widely adopted in international challenges (e.g., LUAI). | Extreme image size variation complicates processing. Complex and costly annotation process. High computational demands for full dataset utilization. |
| xView [248] | 1M+ instances, 1,400 km2 | 0.3 m/pixel | Exceptionally high resolution (WorldView-3). Unprecedented class diversity (60 object types). Supports few-shot learning and domain adaptation research. Designed for humanitarian and disaster response applications. | Pronounced class imbalance ("long-tail" distribution). Significant computational load due to data volume and resolution. Highly complex and expensive three-stage annotation pipeline. |
| VEDAI [251] | 3,700+ vehicles | Multiple Resolutions | Specialized focus on small vehicle detection. Unique multi-modal data (color and infrared spectra). Diverse conditions: occlusion, orientation, and lighting variations. Designed for challenging, unconstrained environments. | Limited scale compared to contemporary benchmarks. Extremely challenging; baseline performance remains low. Restricted to the vehicle category only. |
| DIOR [253] | 192,472 instances, 23,463 images | Varied Resolutions | Comprehensive object coverage (20 categories). Exceptional diversity in weather, season, and resolution. Significant intra-class variation and inter-class similarity. Highly effective for robust evaluation of data-driven methods. | Geographic diversity may not fully represent global variation. Potential underrepresentation in some object classes. Limited compared to newer datasets in total scale. |
| INRIA [255] | 810 km2 coverage | 0.3 m/pixel | Rigorous generalization testing (train/test on different cities). Focus on building/not-building segmentation. High-resolution orthorectified imagery. Covers diverse urban settlements across Europe and North America. | Limited to a binary classification task. Test set labels not publicly available (online evaluation only). Focuses primarily on urban environments. |
6.4. Face Detection Datasets
Face detection datasets have served as the cornerstone of advancement in facial analysis technologies, providing the critical foundation for training and benchmarking algorithms that locate and identify human faces in digital imagery [257]. The evolution of these datasets mirrors progress in the field itself, beginning with constrained laboratory collections that established initial baselines for verification and recognition tasks, then expanding to large-scale "in-the-wild" benchmarks that capture real-world challenges, including occlusion, extreme illumination, diverse poses, and varied demographic representations. Pioneering datasets such as Labeled Faces in the Wild (LFW) [258], WIDER Face dataset [259], and specialized collections like MAFA and the Niqab Dataset have driven innovation by offering structured complexity with precise annotations for various occlusion types and severity levels [260,261], enabling the development of robust deep learning models that approach human-level performance in unconstrained conditions. These datasets are indispensable for both academic research and practical applications in security, biometric authentication, and augmented reality [262]. However, they face significant limitations, including inherent specificity that often limits their utility across different tasks, pervasive biases related to age, ethnicity, and gender that lead to unequal performance across demographic groups, and concerns regarding privacy, consent, and potential misuse in surveillance systems. Furthermore, accessibility varies widely between public and restricted datasets, and the rapid evolution of real-world challenges can quickly render existing benchmarks outdated [263]. Consequently, future development is increasingly focused on creating more dynamic, diverse, and ethically governed datasets that prioritize fairness, transparency, and adaptability to support the next generation of universal and resilient facial analysis systems [262,263].
The WIDER FACE dataset represents a significant advancement in face detection research by addressing the critical gap between laboratory performance and real-world requirements through its massive scale and challenging content [259]. Its primary advantage lies in being an order of magnitude larger than previous datasets, containing rich annotations that encompass diverse facial variations, including extreme poses, substantial occlusions, and challenging lighting conditions across numerous event categories [259]. This comprehensive collection enables more robust training of detection algorithms and provides a realistic benchmark for evaluating model performance under conditions that mirror actual surveillance and public space scenarios. However, the dataset’s limitations include potential biases in its event-based categorization scheme and the inherent challenges of its extreme cases, particularly very small-scale faces and heavy occlusions, which can push beyond the capabilities of current detection architectures [43]. While exceptionally valuable for detection tasks, WIDER FACE’s lack of identity labels limits its direct applicability to face recognition research, and its emphasis on quantity over controlled variation makes it less suitable for studying specific isolated factors. Despite these limitations, it remains an essential resource for driving progress toward more robust real-world face detection systems [263].
The Labeled Faces in the Wild (LFW) dataset represents a landmark contribution to face recognition research by providing a large-scale, unconstrained benchmark designed to study real-world variability in facial imagery [264]. Its efficiency lies in its carefully curated collection of over 13,000 face images featuring 5,749 distinct individuals, captured under natural conditions that encompass vast diversity in pose, lighting, expression, background, and image quality [258]. The dataset’s structured organization-predefined training and testing splits with matched and mismatched pairs-enables streamlined and consistent evaluation of face verification algorithms, facilitating direct comparison between methods and accelerating research progress. However, LFW faces significant limitations: its focus on celebrity images introduces demographic biases toward well-known public figures, primarily of Western origin, while the manual annotation process may contain labeling inaccuracies or inconsistencies. Additionally, the dataset’s lack of controlled variables makes it less suitable for studying isolated factors like pose or occlusion, and its relatively small size by modern standards limits its utility for training deep learning models from scratch [265]. Despite these constraints, LFW remains a foundational resource for benchmarking generalization in unconstrained face recognition.
Table 12.
Comparative analysis of key face detection datasets.
| Dataset | Key Contributions & Importance | Advantages | Limitations |
|---|---|---|---|
| WIDER Face [259] | Bridged the critical gap between laboratory performance and real-world requirements. Set a new standard for scale and difficulty, becoming an essential benchmark for evaluating robust, real-world systems. | Massive scale; an order of magnitude larger than predecessors. Rich annotations for pose, occlusion, and lighting. Covers numerous real-world event categories. | Potential biases in event-based categorization. Lack of identity labels limits utility for recognition tasks. Extreme cases (small faces, heavy occlusion) remain highly challenging. |
| Labeled Faces in the Wild (LFW) [264] | A landmark benchmark that pioneered the study of unconstrained face verification. Provided standardized splits for direct method comparison, significantly accelerating research progress. | Carefully curated for real-world variability (pose, lighting, expression). Streamlined evaluation protocol with matched/mismatched pairs. | Demographic bias (celebrity-focused, Western origin). Relatively small scale for modern deep learning. Not suitable for studying isolated variables. |
| MAFA (MAsked FAces) [260] | Pioneered large-scale research on occluded face detection. Provides detailed annotations for precise occlusion analysis. Critical for applications in health and surveillance. | Large-scale and focused on a critical challenge. Comprehensive attribute annotations (orientation, severity, type). High-quality labels ensured via cross-validation. | Specialized focus on occlusion limits its general-purpose use. Occlusion types may not cover all real-world scenarios. |
The MAFA (MAsked FAces) dataset is a large-scale, specialized benchmark designed to address the critical challenge of detecting and analyzing occluded faces, particularly those obscured by masks and other accessories such as Hijabs or Niqabs in Muslim women [261,266,267]. Comprising over 30,000 images meticulously collected from internet sources, MAFA provides comprehensive annotations for more than 35,000 masked faces, capturing a wide spectrum of real-world occlusion scenarios [260]. Its richness lies in the detailed annotation of six key attributes per face: bounding box locations for faces, eyes, and masks; face orientation; occlusion severity (categorized as weak, medium, or heavy based on occluded facial regions); and mask type (e.g., simple, complex, human body, or hybrid) as shown in Figure 21(a). This structured approach enables precise evaluation of detectors under varying occlusion conditions, covering 60 unique combinations of orientation, occlusion degree, and mask type. Figure 21(b) shows the representative facial images with different attributes [260].
By focusing on challenges like mask occlusion-a problem of heightened importance in contexts like public health and surveillance-MAFA serves as an essential resource for developing robust face detection models capable of operating in complex, real-world environments [268]. Its rigorous annotation process, involving cross-validation by multiple annotators, ensures high-quality labels, making it a valuable benchmark for advancing occlusion-invariant computer vision systems [266].
6.5. Medical Imaging Datasets
Medical imaging datasets have emerged as a cornerstone of modern diagnostic and therapeutic innovation, providing the foundational data required to develop, validate, and deploy artificial intelligence models in healthcare [269]. These datasets comprise annotated images acquired from diverse modalities-such as X-ray, MRI, CT scans, and ultrasound-each tailored to specific clinical contexts, from detecting tumors and fractures to segmenting organs and monitoring disease progression [270]. One of their key advantages lies in enabling reproducible research and benchmarking, allowing for the comparison of algorithmic performance across institutions and pathologies. They also facilitate the training of robust models capable of recognizing subtle patterns indicative of disease, often surpassing human accuracy in tasks like early cancer detection [271]. However, these datasets face significant limitations, including stringent privacy concerns that restrict data sharing, high costs and expertise required for expert annotation, and inherent biases related to demographics, equipment variability, and uneven disease representation. Despite these challenges, the importance of medical imaging datasets cannot be overstated: they are indispensable for accelerating the translation of AI research into clinical practice, enhancing diagnostic precision, supporting telemedicine, and ultimately paving the way toward personalized and accessible healthcare solutions worldwide [272,273].
The Cancer Imaging Archive (TCIA) represents a pivotal, large-scale initiative in oncology research, serving as a centralized, open-access repository for multi-modal medical images (including CT, MRI, and PET) linked to rich clinical data across over 54 organ sites [274]. Its primary advantage lies in overcoming a significant barrier in the field by providing researchers with a vast, curated resource of high-quality, secondary data, which is indispensable for developing and validating artificial intelligence algorithms, conducting large-scale radiomics studies, and facilitating reproducible research without the prohibitive cost and time of new data collection [275]. However, its utility is inherently shaped by limitations such as data heterogeneity-stemming from varied acquisition protocols across contributing institutions-and the potential for incomplete clinical annotations, which can complicate analysis [276]. Despite these challenges, TCIA’s applications are profound, enabling advancements in cancer detection, diagnosis, prognosis, and the emerging field of radiogenomics, thereby accelerating the translation of imaging research into clinical practice [274,276].
The Alzheimer’s Disease Neuroimaging Initiative (ADNI) is a landmark, longitudinal multicenter study designed to identify and validate biomarkers for the early detection and tracking of Alzheimer’s disease (AD) [277]. Its principal advantage lies in its unprecedented scale and depth, providing a vast, standardized, and publicly available repository of longitudinal multimodal data-including MRI and PET scans, genetic profiles, and fluid biomarkers-from a well-characterized cohort of controls, individuals with Mild Cognitive Impairment (MCI), and AD patients [278]. This rich dataset has been instrumental in advancing the understanding of AD progression, enabling the development of predictive models for conversion from MCI to AD, and serving as a critical benchmark for validating diagnostic tools and therapeutic outcomes [279]. However, its limitations include a geographic focus on the US and Canada, which may limit the generalizability of findings to other populations, and the potential for selection bias, as participants in such intensive studies are often highly motivated and may not fully represent the broader population. Despite these constraints, ADNI’s applications are profound, forming the foundational data for thousands of studies in neuroimaging, genomics, and clinical trial design, thereby accelerating the development of biomarkers for early diagnosis and effective disease-modifying treatments [280].
The UK Biobank (UKB) stands as one of the most comprehensive and extensive prospective cohort studies globally, encompassing deep phenotyping data, including medical imaging, genetic sequences, and detailed health records, from over 500,000 participants [281]. Its unparalleled advantages for medical science include its massive sample size, the breadth of multi-modal data (genetic, imaging, lifestyle, biochemical), and its longitudinal design, which together provide an unprecedented resource for large-scale epidemiological studies, the development of multi-modal AI algorithms, and the investigation of complex disease mechanisms and risk factors [282]. However, these strengths are tempered by significant limitations, including the cohort’s healthy volunteer bias, which limits the generalizability of findings to the wider population, and the immense complexity of data wrangling, which requires specialized tools and computational resources for analysis [283]. Despite these challenges, the UKB’s applications are transformative, enabling research into disease prediction models, the identification of novel biomarkers for conditions ranging from cancer to mental health, and the exploration of genetic associations across a vast spectrum of human health, thereby accelerating the transition towards personalized and preventative medicine [284,285].
The Open Access Series of Imaging Studies (OASIS) is a foundational neuroimaging resource that provides freely available, high-quality brain MRI scans and, in its later versions, PET and clinical data [286], with a focus on normal aging and Alzheimer’s disease (AD) [287]. Its primary advantages include its open-access nature, which dramatically lowers the barrier to entry for research; its longitudinal design, allowing for the study of neurological changes over time; and its inclusion of both cognitively normal and impaired individuals, facilitating direct comparisons [288]. However, key limitations exist, such as demographic homogeneity (predominantly right-handed, older adults of Western origin) that can limit the generalizability of findings, and a potential selection bias as participants are volunteers [287,289], not fully representing the broader population. Despite these constraints, the OASIS datasets have become an indispensable tool in medical science, extensively applied to develop and validate automated brain segmentation algorithms, identify neuroanatomical biomarkers for early AD detection, and serve as a critical benchmark dataset for training and testing machine learning models in computational neuroscience [287,289].
The Medical Information Mart for Intensive Care Chest X-Ray (MIMIC-CXR) dataset is a large-scale, publicly available resource containing over 377,000 de-identified chest radiographs paired with their corresponding free-text radiology reports, sourced from the Beth Israel Deaconess Medical Center [290,291]. Its primary advantages are its immense size, the crucial pairing of image and text data, and its adherence to strict de-identification protocols (HIPAA Safe Harbor), which together make it an invaluable benchmark for developing and validating artificial intelligence models in medical imaging, particularly for tasks like automated pathology detection (e.g., pneumonia, pneumothorax) and natural language processing for report generation [291]. However, the dataset has significant limitations, including its retrospective, single-institution nature, which may introduce biases in patient population, imaging equipment, and clinical practices, limiting the generalizability of models trained on it; the potential for residual protected health information (PHI) in unstructured text despite de-identification efforts; and a lack of expert re-annotation, meaning labels extracted from reports can be noisy [292]. Despite these constraints, MIMIC-CXR has become a cornerstone for medical AI research, enabling groundbreaking applications in computer-aided diagnosis, the development of automated triage systems to alleviate radiologist workload, and the creation of assistive tools to extend radiological expertise to underserved and resource-poor regions globally [293].
The International Skin Imaging Collaboration (ISIC) datasets represent a cornerstone resource in dermatological AI, providing a large, publicly available collection of tens of thousands of dermoscopic images annotated with gold-standard diagnoses for skin lesion analysis, particularly melanoma and other skin cancer detection [294,295], and it is considered the standard benchmark for automated skin cancer detection. Their primary advantages include their scale, which is crucial for training data-hungry deep learning models, and their role in standardizing research through annual public challenges that have driven the field toward human-expert-level performance [296]. However, significant limitations have been identified, including the presence of numerous duplicate images within and across dataset versions-a flaw that can artificially inflate model performance metrics and introduce bias-as well as issues with class imbalance and label noise [297,298]. Despite these constraints, the curated ISIC datasets are extensively applied in medical science to develop and validate automated diagnostic assistants that can support clinicians in early skin cancer detection, serve as educational tools for training dermatologists, and have the potential to extend specialist-level diagnostic capability to primary care settings and underserved populations with limited access to healthcare [296].
The Breast Cancer Digital Repository (BCDR) is a benchmarking dataset designed to advance the development of computer-aided diagnosis (CADx) systems for breast cancer by providing a curated collection of mammograms, encompassing both low-resolution digitized film and high-resolution full-field digital images, paired with detailed clinical patient data [299]. Its key advantages include this dual-resolution design, which allows researchers to directly evaluate the impact of image quality on model performance, and the critical integration of image descriptors with clinical information, which has been shown to significantly boost diagnostic accuracy [300]. A primary limitation, common to many medical imaging repositories, is their relatively smaller scale compared to massive public datasets, which can constrain the training of highly complex deep learning models and potentially affect generalization [301]. Despite this, the BCDR’s meticulously annotated masses and structured format make it an invaluable resource for applications such as exploring and benchmarking machine learning classifiers, developing robust feature extraction techniques for mass characterization, and ultimately creating reliable CADx tools that provide a crucial second opinion to radiologists, aiming to improve early detection and diagnosis of breast cancer [301].
The Medical Segmentation Decathlon (MSD) is a landmark benchmarking dataset and challenge specifically designed to test the generalizability of semantic segmentation algorithms across ten vastly different clinical tasks, encompassing a wide range of anatomical regions (e.g., brain, liver, prostate, lungs) and imaging modalities (e.g., MRI, CT) [302]. Its primary advantage is this deliberate and unprecedented diversity, which moves beyond narrow, single-task challenges to provide a robust proving ground for developing AI models that can learn fundamental segmentation principles applicable to unseen problems, a hypothesis strongly validated by the success of the winning nnU-Net framework. A key limitation, however, is that while the dataset is diverse, the volume of data for each task is relatively modest compared to massive single-organ repositories, potentially constraining the performance ceiling for highly specific applications. Furthermore, the dataset does not fully address pervasive real-world challenges like domain shift between different hospital scanners [303]. Despite these constraints, the MSD’s major application and profound impact lie in catalyzing the commoditization of medical image segmentation; it demonstrated that highly accurate, state-of-the-art models can be trained automatically, thereby empowering clinical researchers without deep AI expertise to develop robust tools for tasks such as tumor volumetry, treatment planning, and quantitative organ assessment, ultimately accelerating the translation of AI from research into broader clinical practice[304].
The Retinal Fundus Multi-disease Image Dataset (RFMiD) addresses a critical gap in ophthalmological AI by providing a curated collection of fundus images annotated for a wide spectrum of 46 ocular conditions, including both common and rare, sight-threatening pathologies that are often absent in other repositories [305,306]. Its primary advantage is this unparalleled diversity, which moves beyond the standard focus on diabetic retinopathy, glaucoma, and AMD to enable the development of more generalizable and clinically comprehensive computer-aided diagnosis (CAD) tools [306]. A key limitation is the inherent challenge of its multi-source nature, as images were captured with three different fundus cameras, which can introduce variability that models must account for to avoid learning scanner-specific artifacts rather than pathological features. Despite this, RFMiD’s major application lies in pioneering robust AI systems capable of acting as a screening tool for a vast range of diseases, thereby assisting ophthalmologists in early detection, particularly for rare conditions that may be missed, and helping to alleviate the global burden of preventable vision impairment by extending expert-level diagnostic reach to underserved areas [307].
Medical imaging datasets are the fundamental pillars upon which the future of AI-driven healthcare is being built. As demonstrated by the diverse repositories analyzed-from the oncology-specific TCIA to the multi-disease ophthalmological RFMiD-these resources provide the essential fuel for developing, validating, and benchmarking algorithms that can detect, segment, and diagnose with increasing precision, as shown in the above sections. Their collective advantages, including scale, public accessibility, and multi-modal data integration, have democratized research and accelerated progress towards tools for automated screening and personalized medicine. However, significant challenges persist, such as inherent biases, annotation inconsistencies, data heterogeneity, and privacy concerns, which can limit the generalizability and real-world applicability of models trained on them. The path forward necessitates a concerted effort towards curating larger, more diverse, and harmonized datasets, developing robust federated learning frameworks to overcome data silos, and implementing rigorous validation protocols to ensure that these powerful tools translate equitably and effectively into clinical practice, ultimately fulfilling their promise to improve patient outcomes globally. Table 13 shows the comparative analysis of medical imaging datasets, highlighting their uses, annotation, key features, and limitations.
6.6. Agricultural Image Datasets
Agricultural image datasets have become a critical enabler of the digital transformation in farming, providing the foundational data for developing AI and computer vision models designed to optimize food production [308]. These datasets are typically composed of images captured by drones, satellites, and ground-based sensors, annotated to identify key agricultural elements such as crops, weeds, pests, diseases, and soil conditions. Their primary advantage lies in enabling precision agriculture-allowing for scalable monitoring of plant health, targeted resource application, and early detection of threats, which can significantly increase yields while reducing water, pesticide, and fertilizer use [309]. However, they also face significant limitations, including the high cost of precise annotation requiring agronomic expertise, a lack of standardization across datasets, and a sensitivity to extreme environmental variability in lighting, growth stages, and weather conditions that can hinder model generalization [310]. Despite these challenges, the importance of these datasets is profound; they are indispensable for advancing global food security, promoting sustainable farming practices by minimizing environmental impact, and building resilient agricultural systems capable of adapting to climate change [311].
Table 14.
Overview of agricultural object detection datasets (2025-2017).
| Dataset Name | Year | Plant Name (English) | Plant Part | Task | Annotation | Images |
|---|---|---|---|---|---|---|
| CornLeafInfection [312] | 2025 | Maize | Leaf | Disease Detection | Bounding Boxes and segmentation | 1,308 |
| PumkinDiseases [313,314] | 2024 | Pumpkin | Leaf | Disease Identification | Image Classes | 2,000 |
| MFWD [315] | 2024 | Sorghum, Maize | Whole plant | Weed Detection | Img. Cls., BBox, Seg. | 94,321 |
| SorghumAphids [316,317] | 2024 | Sorghum | Leaf | Pest Detection | Segmentation Masks | 54,742 |
| PhenoBench [318] | 2024 | Sugar beet | Leaf, Whole plant | Weed Detection | Segmentation Masks | 2,179 |
| SugarcanePlantCounting [319] | 2024 | Sugarcane | Cane, Fruit | Plant Counting | Bounding Boxes | 3,730 |
| SoybeanNet [320] | 2024 | Soybean | Panicle, Fruit | Bean Counting | Seg. Masks, Points | 196 |
| CottonWeedDet3 [321] | 2023 | Cotton | Leaf | Weed Detection | Bounding Boxes | 848 |
| CropAndWeed [322] | 2023 | Various (Maize, Beet, etc.) | Leaf, Whole plant | Weed/Crop ID | Seg. Masks, BBox | 8,034 |
| ImageWeeds [323] | 2023 | Various Weeds | Whole plant | Weed Identification | Bounding Boxes | 3,975 |
| LucasVision [324] | 2023 | Various Crops | Ground cover | Plant Identification | Image Classes | 16,946 |
| RiceLeafDiseases [325] | 2023 | Rice | Leaf | Disease Detection | Image Classes | 4,684 |
| RicePanicles [326] | 2023 | Rice | Panicle, Fruit | Panicle Detection | Bounding Boxes | 2,193 |
| RumexLeaves [327] | 2023 | Broad-leaved Dock | Leaf | Plant Detection | Seg. Mask, BBox | 809 |
| SoyNet [328] | 2023 | Soybean | Leaf | Disease Detection | Image Classes | 4,500 |
| TobaccoAerialDataset [329] | 2023 | Tobacco | Whole plant | Weed Detection | Segmentation Masks | 1,870 |
| WE3DS [330] | 2023 | Various (Bean, Pea, etc.) | Leaf, Whole plant | Weed/Crop ID | Segmentation Masks | 2,568 |
| YOLOWeeds [331] | 2023 | Various Weeds | Leaf | Weed Detection | Bounding Boxes | 5,648 |
| CottonWeedID15 [332] | 2022 | Cotton | Leaf | Weed Identification | Image Classes | 5,187 |
| ASDID [333] | 2022 | Soybean | Leaf | Disease Detection | Image Classes | 9,981 |
| RoboWeedMap [334] | 2022 | Barley | Whole plant | Weed Detection | Bounding Boxes | 1,147 |
| GlobalWheatHeadDetection [335,336] | 2021 | Wheat | Wheat Head | Plant Detection | Bounding Boxes | 6,422 |
| DPA [337] | 2020 | Maize, Sugar beet | Whole plant | Plant Classification | Segmentation Masks | 200 |
| EarlyCropWeed [338] | 2020 | Tomato, Cotton | Leaf | Weed Detection | Image Classes | 504 |
| OpenWeedPhenotype [339] | 2020 | Various Weeds | Leaf | Weed Detection | BBox, Img. Cls. | 7,590 |
| DeepWeeds [340] | 2019 | Various Weeds | Leaf | Weed Detection | Image Classes | 17,509 |
| DeepSeedling [341] | 2019 | Cotton | Seedlings | Seedling Detection | Bounding Boxes | 5,610 |
| WeedGrowthState [342] | 2018 | Various Weeds | Leaf | Growth Estimation | Image Classes | 12,165 |
| MaizeDiseaseSymptoms [343] | 2018 | Maize | Leaf | Disease Detection | Other (Polyline) | 18,222 |
| PlantSeedlingClassification [344] | 2017 | Various | Seedlings | Seedling Classification | Image Classes | 5,659 |
Based on the comprehensive collection of agricultural object detection datasets in Table 14, several key advantages and limitations have been shown with their contribution to AI-based development in agriculture. A significant strength lies in the remarkable diversity of tasks and species covered, as evidenced by datasets like CropAndWeed [322] and WE3DS [330], which encompass a wide variety of crops and weed species, and MFWD [315], which provides a massive scale of annotated imagery [311]. This diversity is crucial for developing robust, generalizable models. Furthermore, many datasets, such as DeepWeeds [340] and CottonWeedID15 [332], offer high-quality, expert-verified annotations including bounding boxes and segmentation masks, which are essential for supervised learning. The focus on real-world agricultural challenges, from early-stage Seedling Detection (DeepSeedling [341]) to in-season Weed Detection (PhenoBench [318]) and Disease Identification (ASDID [333], SoyNet [328]), demonstrates a strong applied relevance. The public availability of many of these resources significantly lowers the barrier to entry for research in agricultural computer vision [345].
However, these datasets are not without limitations. A primary concern is the issue of scale and imbalance; while some datasets are large (e.g., MFWD with 94k images) [315], many others are notably small, which can hinder the training of complex deep learning models and lead to overfitting [346]. There is also a pronounced geographic and species bias, with a heavy focus on major cash crops like wheat, maize, and soybean, potentially limiting model performance on less common crops or weeds prevalent in different regions. Annotation inconsistency is another challenge, as the type and quality of labels-ranging from image-level classes to precise segmentation masks-vary dramatically between datasets [346], complicating comparative analysis and multi-dataset learning. Finally, the highly controlled conditions in which many images are captured (e.g., CottonWeedID15 leaf images) often do not reflect the immense variability and complexity of a real-field environment [332], creating a significant sim2real gap that must be bridged for practical deployment. Collectively, these datasets provide a foundational but heterogeneous resource that drives innovation while simultaneously highlighting the critical need for larger, more balanced, and consistently annotated benchmarks that capture the true challenges of in-situ agricultural automation [346].
7. Evaluation Metrics of Object Detections
Object detection algorithms aim to replicate and surpass human visual reasoning, identifying and localizing objects with precision. To gauge their efficacy, researchers have relied on metrics that mirror human intuition. Average Precision (AP), the cornerstone of evaluation, reflects a model’s ability to balance correctness (Precision) and completeness (recall) across confidence thresholds [257]. At its core, the AP emerges from the precision-recall curve, which is a visualization of the trade-off between false positives and missed detections. However, detection is not merely classification: an accurate localization, measured by Intersection over Union (IoU), ensures that the boundaries align closely with ground-truth objects, with a common IoU threshold of 0.5 defining a successful detection, as shown in Figure 22 [116,347].
Modern challenges require nuanced metrics. For multi-class systems, the mean Average Precision (mAP) aggregates performance across categories, whereas COCO-style AP raises the bar by averaging results across stricter IoU thresholds (0.5-0.95) [103]. Class imbalance, a persistent hurdle, is addressed through F1-score optimization, harmonizing precision and recall, and adaptive thresholds. In addition to accuracy, real-world deployment hinges on efficiency-aware metrics such as inference speed (FPS) and computational cost (FLOPs), which prioritize practicality in edge devices or real-time systems, such as autonomous vehicles [82,168]. Crucially, confidence calibration-ensuring that a model’s self-reported certainty aligns with its actual performance emerged as a linchpin for trust in safety-critical domains such as medical imaging [257]. These metrics, rooted in mathematical rigor yet attuned to human-centric challenges, drive progress toward detectors that are not only accurate but also reliable, efficient, and equitable [118].
This paper outlined the evaluation matrices of the object detection approaches in four parts such as Fundamental Metrics that tell about the accuracy, precision, and the sensitivity results to measure how the object detection algorithm is working, IoU that measures the overlap between the predicted bounding box and the ground truth bounding box in the detected object by the algorithm, primary detection metrics give about average precision, and the mean average precision across all object classes and evaluate the performance over the entire dataset. This paper also discusses the metrics for more specific tasks derived from object detection, such as the Object Keypoint Similarity (OKS) and Panoptic Quality (PQ) for panoptic segmentation.
Precision is a metric used to measure the accuracy of positive predictions made by an object-detection model. It is defined as the ratio of true positive(TP) detections (correctly identified objects) to the total number of positive predictions (both true positives and false positives(FP)). Precision indicates the model’s ability to identify only relevant objects, without including irrelevant ones. Here represents the true positive, false positive as , false negative as , and the true negative is as in object detection. So we can say,
A high-precision value means that most of the objects identified by the model are relevant, with few false positives. Recall is a metric used to measure the completeness of an object detection model in identifying all the relevant objects within a dataset. This is defined as the ratio of true positive detections to the total number of actual objects (true positives and false negatives).
Recall indicates the ability of a model to find all relevant objects in the dataset. A high recall value indicates that the model successfully identifies most of the relevant objects, with few missed detections. Equ. 1 and Equ. 2 show the equations to calculate the precision and the recall of the object detection algorithms from their results. The ratio of all correctly predicted instances (both positive and negative) to all available instances. So we can say the accuracy of an object detection algorithm depends on the ratio of the sum of all detected objects and the true detections. Equ. 3 shows the process of calculating the overall accuracy of an object detection approach based on its true and false detections.
The F1 score serves as a pivotal evaluation metric in object detection, harmonizing precision and recall to provide a balanced assessment of an object detection model’s performance. From Equ. 4, we can see a mathematical formulation to calculate the F1 score, which is particularly valuable in scenarios requiring equilibrium between minimizing missed detections (critical in medical imaging or surveillance) and limiting false alarms (essential in autonomous driving).
While metrics such as Mean Average Precision (mAP) dominate benchmarks, F1 is favored in class-imbalanced or safety-critical contexts, where neither precision nor recall should be disproportionately prioritized. Researchers and engineers often calculate F1 at optimal confidence thresholds or derive it from precision-recall curves, with class-wise averaging (macro-F1) enhancing multiclass analysis. The simplicity and interpretability of this metric make it a practical tool for applications demanding a consolidated performance view, complementing nuanced metrics, such as mAP, in comprehensive detector assessments.
However, in dense object detection scenarios, where multiple objects are clustered within a single image, such as pedestrian detection in crowds or satellite imagery analysis, the traditional metrics, such as precision-recall (PR) curves, may inadequately capture performance because of the inherent challenges of high target density. Instead, the Miss Rate (MR) and False Positives Per Image (FPPI) metrics, along with their combined MR-FPPI curve, provide a more nuanced evaluation framework.
Equ. 5 and Equ. show the mathematical formulation to calculate the missing rate and the false positives per image for N images in the dataset and FP results. The FPPI quantifies the average number of incorrect detections per image across a dataset. MR measures the proportion of undetected ground-truth objects relative to the total number of actual positives.
Intersection over Union (IoU) is a fundamental metric in object detection that quantifies the spatial overlap between a predicted bounding box and its corresponding ground-truth annotation [210,347]. Mathematically, IoU is the ratio of the intersection area (the overlapping region) to the area of the union (the combined region) of the two bounding boxes shown in Figure 22, and Equ. 7 represents the mathematical formulation of IoU calculation,
A detection is typically classified as a True Positive (TP) if the IoU exceeds a predefined threshold (), commonly set to 0.5, in benchmarks such as PASCAL VOC, although stricter thresholds (e.g.,=0.75) are used to evaluate high-precision localization in datasets such as COCO, as shown in Figure 22. Although IoU effectively measures localization accuracy, it has limitations, such as it does not accounting for object orientation or shape (e.g., rotated or irregular objects) and being sensitive to minor positional deviations for small objects. Despite these limitations, IoU remains a cornerstone of detection evaluation because of its simplicity, interpretability, and widespread adoption in major benchmarks, serving as the foundation for deriving advanced metrics, such as Average Precision (AP) and Mean Average Precision (mAP).
As we know about the precision and recall in Equ. 2 and Equ. 1, the Average Precision (AP) is the primary metric for evaluating object detection performance on a single class. It summarizes the precision-recall curve into a single number, representing the model’s ability to make both accurate (high precision) and complete (high recall) detections. .
To get a complete picture across all classes, the mean Average Precision (mAP) is used, which is simply the average of the AP scores for every class in the dataset. This makes mAP the key benchmark metric for comparing overall object detection model performance. For N datasets, the mAP equation can be written as in Equ. 8 and Equ. 9, where is the Average Precision of the category. The key difference between the average precision(AP) and the mean average precision(mAP) is that AP evaluates the detection of objects based on the class, reflecting how well the detector balances the precision and recall while minimizing the false positives(FP) and false negatives(FN). However, mAP benchmarks the overall model performance that standardizes multiclass evaluation, enabling a fair comparison between the object detection algorithms. For example, if there is a dataset of four classes are humans, cats, dogs, and cars, and the average precision of each class was 0.85, was 0.87, was 0.89, and was 0.92. So the calculation of would be in Equ. 10,
While metrics like mean Average Precision (mAP) and IoU provide a general benchmark for bounding box detection, specialized computer vision tasks necessitate more nuanced evaluation metrics. These metrics are designed to align with the specific objectives and challenges of tasks such as keypoint detection and panoptic segmentation in different scenarios, offering a more precise measure of model performance.
The standard Intersection over Union (IoU) metric is ill-suited for evaluating the alignment of individual keypoints. For human pose estimation and similar keypoint detection tasks, the Object Keypoint Similarity (OKS) is employed as the canonical metric. The OKS between a set of predicted keypoints and their ground-truth counterparts is defined by Equ. 11,
where is the Euclidean distance (in pixels) between the i-th predicted keypoint and its corresponding ground-truth keypoint, s is a scale factor that normalizes the distance based on the object’s size. For a person, this is typically defined as , the square root of the area of the object’s bounding box. is the per-keypoint constant that controls the falloff, representing the acceptable error for that specific keypoint type (e.g., a small for eyes and ears, a larger for shoulders and hips). These values are dataset-dependent and are derived from the annotation uncertainty of each keypoint. is the visibility flag of the ground-truth keypoint. is an indicator function that is 1 if the keypoint is labeled and visible, and otherwise it is 0 (e.g., occluded or not labeled). This ensures unannotated keypoints do not contribute to the score.
The term can be interpreted as a univariate Gaussian function, assigning a value close to 1 for accurate predictions and penalizing predictions that deviate significantly from the truth, with the penalty scaled by the object’s size and the keypoints’ inherent uncertainty [47]. The final OKS score is the average of these values across all visible keypoints, yielding a number between 0 and 1. The overall performance is then evaluated by calculating the mean Average Precision (mAP) across a range of OKS thresholds (e.g., from 0.5 to 0.95 with a 0.05 increment), denoted as .
Panoptic segmentation, which requires the simultaneous segmentation of both "things" (countable object instances) and "stuff" (amorphous regions), is evaluated using the Panoptic Quality (PQ) metric. PQ elegantly combines segmentation quality with recognition quality into a single, unified score. The PQ is formally defined as:
A key insight is that PQ can be decomposed into the product of two independent components: Segmentation Quality(SQ) and Detection Quality(DQ) described in Equ. 13 and Equ. 14. Segmentation Quality (SQ) is simply the average IoU of all matched segments (TP), measuring how well the predicted masks align with the ground-truth masks.
Detection Quality (DQ) is identical to the F1 Score commonly used in detection and classification, which balances precision and recall in identifying segments. It measures the accuracy of the recognition process itself.
8. Applications of Object Detection
Earlier, this study started with the application of object detection algorithms in various fields with tremendous success owing to their efficiency. In the transportation sector, computer vision-based object detection is indispensable for autonomous vehicles [168,348]. By accurately detecting pedestrians, vehicles, and other obstacles or recognizing the number of vehicle plates, these systems can make real-time decisions to ensure safe navigation. In addition, traffic monitoring systems employ object detection to analyze traffic patterns, identify congestion points, and optimize traffic flow [349,350,351,352,353]. In research trends and development, collaboration of multi-agent systems has become very popular, where multiple unmanned vehicles collaborate while moving, and the detection of objects, group vehicles, and obstacles plays a vital role [354,355].
Another area where computer vision has had a significant impact is agriculture [167,356]. Object detection algorithms can be used to count crops, assess plant health, and detect pests and diseases [357,358]. This enables farmers to make informed decisions regarding irrigation, fertilization, and pest control, leading to increased yields and reduced costs. Using such technologies, companies are introducing new ones to expand agricultural development and increase crop production rates [359]. In the military, computer vision plays a crucial role in surveillance and reconnaissance. Object-detection systems can identify potential threats, track targets, and provide real-time intelligence [360]. Furthermore, drones equipped with computer vision capabilities can be used for search and rescue operations, disaster relief, and border patrols [361,362].
Industrial processes have also benefited from object detection based on computer-vision [363]. Automated inspection systems can detect defects in manufactured products, ensure quality control, and reduce wastage. In addition, object tracking systems can optimize production lines by monitoring the movement of materials and equipment [364,365,366].
The healthcare sector is another promising application area of computer vision. Object detection algorithms can assist in medical image analysis, such as identifying tumors in X-rays or detecting abnormalities in MRIs. Moreover, computer vision can be used for patient monitoring, gait analysis, and surgical assistance [367,368,369,370,371]. The application is not limited to healthcare or biological research, but computer vision has made major advances in molecular biology, particularly in better visualization approaches, such as Establishing tools for visualizing molecular structures or excessively simplified visuals of biomolecules [372,373]. However, computer vision allows for an agile and comprehensive representation of molecular structures such as proteins, nucleic acids, and significant macromolecular complexes, which become complex and time-consuming for humans [374]. Figure 23 shows the applications of computer vision in the research field of structural biology.
Furthermore, technological advances made possible by combining computer vision with structural biology have considerably increased the area’s integration and significance. Some researchers believe computer vision, together with open-source infrastructure and cloud computing, has enabled a global network of academics to do top-notch structural analysis. This data democratisation has resulted in more cross-disciplinary partnerships, allowing additional scientists to address complicated biological challenges, and these computational advancements are creating the framework for structural biology to grow into a key component in biomedical and pharmaceutical studies [375] in the future. Environmental monitoring is an important application of computer vision. Object detection can be used to track wildlife populations, monitor deforestation, and detect pollution levels. This information is crucial for conservation and environmental management [376,377]. Computer vision can enhance learning experiences in the field of education. For example, object detection can be used to create interactive learning environments in which students interact with virtual objects and explore concepts in a more engaging manner [378,379,380]. Recent research has focused on improving the accuracy and efficiency of object-detection algorithms. Deep learning techniques such as convolutional neural networks (CNNs) have achieved state-of-the-art performance in object detection tasks. In addition, advancements in hardware, including GPUs and TPUs, have enabled faster and more efficient object detection [381,382,383]. In conclusion, we can say that computer vision-based object detection has emerged as a powerful tool with applications across various domains. By accurately identifying and locating objects within images and videos, this technology is driving innovation and improving efficiency in transportation, agriculture, military, industrial processes, healthcare, the environment, and education. As research continues to advance, we can expect even more exciting developments and applications for object detection in the future.
9. Discussion and Future Works
The profound evolutionary trajectory of object detection, marked by the shift from complexity-laden two-stage models to highly efficient single-stage frameworks and the current ascent of vision transformers, confirms deep learning’s dominance in achieving real-time accuracy. Despite this success, several interconnected limitations persist, constraining generalization and deployment efficiency across diverse real-world applications. A foundational challenge lies in data consumption; deep learning models remain limited by a passive input process, learning only underlying patterns and trends frequently observed in the data without inherent contextual understanding. This dependency inherently amplifies the perceptions, biases, and intentional or unintentional data selection choices of the original dataset producers, potentially leading to critical errors in sensitive, decision-making applications [109]. Current public datasets, including high-density benchmarks like MS-COCO, often fail to comprehensively represent the long-tail distribution of real-world objects, occlusions, and intra-class variations, thereby yielding algorithms that struggle significantly when encountering novel or rare instances in unconstrained environments [81]. Future academic efforts must pivot towards developing mechanisms for automatic, objective data exploration and generation, complemented by label-efficient learning strategies (such as few-shot and weakly supervised methods) to mitigate the reliance on exhaustive, potentially biased annotations [384]. This focus is indispensable for addressing the vital, emerging challenge of "fairness in AI" and ensuring equitable performance across diverse populations and scenarios. The second critical barrier is the ongoing demand for substantial computational resources. While the introduction of lightweight architectures like the later YOLO variants (YOLOv8, YOLOE) has democratized real-time performance, many high-accuracy models, particularly complex transformer-based systems, still require costly, multi-GPU, high-speed configurations [384]. This remains a significant hurdle, limiting wider accessibility for researchers and delaying critical implementation in resource-constrained devices. Consequently, a core direction for future research is the aggressive pursuit of architectural re-engineering to design highly lightweight computational models that strictly balance accuracy (mean average precision) with high deployment efficiency (frames per second) for reliable lightweight edge deployment. Architecturally, the field is undergoing a paradigm shift towards generalist, open-vocabulary perception. The definitive future direction centers on robust multimodal vision-language integration. The recent success of LLM-driven architectures, including Grounding DINO, YOLO-World, Florence-2, and other algorithms, has demonstrated the unparalleled power of using natural language reasoning to achieve superior zero-shot generalization and contextual grounding, well beyond the limits of closed-set detection. Further academic research should focus on enhancing these LLM-based systems to develop finer-grained spatial localization, significantly speeding up generally slow training convergence, and ensuring stability in the presence of complex challenges related to depth ambiguity and dynamic human-object interaction. Successful navigation through these limitations will establish the next generation of generalist object detection systems that are suitable for safety-critical applications related to autonomous driving, industrial automation, and advanced medical diagnostics, with other new fields of engineering [2,384].
10. Conclusion
In conclusion, this review provides a comprehensive exploration of object detection algorithms over the last thirty years, from traditional algorithms to recent LLM models. This paper gives a comprehensive review and analysis of deep learning-based object detection models, such as the single-stage, two-stage, transformer-based detectors, and the LLM model-based object detectors. It also analysed the object detection datasets and benchmarks based on their usability, with an analysis of the evaluation metrics of object detection algorithms. This review paper has tracked the changes in research trends and future possible applications of object detection. The Analysis of the existing object detection approaches, identification of research gaps, changes in the research trends, and suggestions for future developments have illuminated the opportunities and challenges in object detection and its features through this review paper.
Author Contributions
The contributions of the authors are as follows: Md. Faishal Rahaman: Conceptualization, Methodology, Investigation, Writing – Original Draft, Writing – Review & Editing, Project administration; Xueyuan Li: Conceptualization, Supervision, Validation, Writing – Review & Editing; Md Nagib Mahfuz Sunny: Investigation, Visualization, Formal analysis, Data Curation, Writing – Original Draft; Muhammad Amjad: Validation, Writing – Review & Editing, Visualization; Xin Gao: Resources, Supervision, Writing – Review & Editing; Sakib Hasan: Investigation, Writing – Original Draft (Datasets & Evaluation Techniques sections).
Informed Consent Statement
Not Applicable.
Acknowledgments
The authors would like to express their sincere gratitude to the following individuals for their valuable assistance and contributions to this work K. M. Shihab Hossain from the College of Business Information Systems, Central Michigan University, Michigan, USA, for his assistance with the systematic review of deep learning-based object detection methods as well as his contributions to the analysis and description of medical imaging datasets. Zarqa Noor from the School of Life Science, Beijing Institute of Technology, Beijing, China, for her assistance with the writing and revision of the medical imaging and agricultural image datasets sections, and her valuable insights into the applications of computer vision in structural biological research. And S. M. Abul Bashar from the Department of Hydrogen Technology, Rosenheim Technical University of Applied Sciences, Rosenheim, Germany, for his assistance with the analysis of remote sensing datasets and evaluation techniques for sensor-based applications, as well as his contributions to the discussion on performance metrics for object detection in autonomous driving and environmental monitoring.
Conflicts of Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- Jiao, L.; Zhang, F.; Liu, F.; Yang, S.; Li, L.; Feng, Z.; Qu, R. A Survey of Deep Learning-Based Object Detection. IEEE Access 2019, 7, 128837–128868. [CrossRef]
- Zaidi, S.S.A.; Ansari, M.S.; Aslam, A.; Kanwal, N.; Asghar, M.; Lee, B. A survey of modern deep learning based object detection models. Digital Signal Processing 2022, 126, 103514. [CrossRef]
- Cherapanamjeri, J.; Rao, B.N.K. Neural Networks based Object Detection Techniques in Computer Vision. In Proceedings of the 2022 4th International Conference on Inventive Research in Computing Applications (ICIRCA), 2022, pp. 1092–1099. [CrossRef]
- Shetty, A.K.; Saha, I.; Sanghvi, R.M.; Save, S.A.; Patel, Y.J. A Review: Object Detection Models. In Proceedings of the 2021 6th International Conference for Convergence in Technology (I2CT), 2021, pp. 1–8. [CrossRef]
- Sarda, A.; Dixit, S.; Bhan, A. Object Detection for Autonomous Driving using YOLO algorithm. In Proceedings of the 2021 2nd International Conference on Intelligent Engineering and Management (ICIEM), 2021, pp. 447–451. [CrossRef]
- Masmoudi, M.; Ghazzai, H.; Frikha, M.; Massoud, Y. Object Detection Learning Techniques for Autonomous Vehicle Applications. In Proceedings of the 2019 IEEE International Conference on Vehicular Electronics and Safety (ICVES), 2019, pp. 1–5. [CrossRef]
- Wang, R.; Wang, Z.; Xu, Z.; Wang, C.C.; Li, Q.; Zhang, Y.; Li, H. A Real-Time object detector for autonomous vehicles based on YOLOV4. Computational Intelligence and Neuroscience 2021, 2021, 1–11. [CrossRef]
- Esteva, A.; Chou, K.J.; Yeung, S.; Naik, N.; Madani, A.; Mottaghi, A.; Liu, Y.; Topol, E.J.; Dean, J.; Socher, R. Deep learning-enabled medical computer vision. npj digital medicine 2021, 4. [CrossRef]
- Ganatra, N. A Comprehensive Study of Applying Object Detection Methods for Medical Image Analysis. In Proceedings of the 2021 8th International Conference on Computing for Sustainable Global Development (INDIACom), 2021, pp. 821–826.
- Ali, S.; Abdullah.; Athar, A.; Ali, M.; Hussain, A.; Kim, H.C. Computer Vision-Based Military Tank Recognition Using Object Detection Technique: An Application of the YOLO Framework. In Proceedings of the 2023 1st International Conference on Advanced Innovations in Smart Cities (ICAISC), 2023, pp. 1–6. [CrossRef]
- Bao, C.; Cao, J.; Hao, Q.; Cheng, Y.; Ning, Y.; Zhao, T. Dual-YOLO Architecture from Infrared and Visible Images for Object Detection. Sensors 2023, 23. [CrossRef]
- Rahaman, M.F.; Noman, M.A.A.; Ali, M.L.; Rahman, M. DESIGN AND IMPLEMENTATION OF A FACE RECOGNITION BASED DOOR ACCESS SECURITY SYSTEM USING RASPBERRY PI, 2021.
- Varma, S.; Sreeraj, M. Object detection and classification in surveillance system. In Proceedings of the 2013 IEEE Recent Advances in Intelligent Computational Systems (RAICS), 2013, pp. 299–303. [CrossRef]
- Xu, J. A deep learning approach to building an intelligent video surveillance system. Multimedia Tools and Applications 2020, 80, 5495–5515. [CrossRef]
- Wosner, O.; Farjon, G.; Bar-Hillel, A. Object detection in agricultural contexts: A multiple resolution benchmark and comparison to human. Computers and Electronics in Agriculture 2021, 189, 106404. [CrossRef]
- Kragh, M.; Jørgensen, R.N.; Pedersen, H. Object Detection and Terrain Classification in Agricultural Fields Using 3D Lidar Data. In Proceedings of the Computer Vision Systems; Nalpantidis, L.; Krüger, V.; Eklundh, J.O.; Gasteratos, A., Eds., Cham, 2015; pp. 188–197. [CrossRef]
- Noman, M.A.; LI, Z.; Almukhtar, F.; Rahaman, M.F.; Omarov, B.; Ray, S.; Miah, S.; Wang, C. A computer vision-based lane detection technique using gradient threshold and hue-lightness-saturation value for an autonomous vehicle. International Journal of Electrical and Computer Engineering (IJECE) 2023, 13, 347–357. [CrossRef]
- Al Noman, M.A.; Rahaman, M.F.; Li, Z.; Ray, S.; Wang, C. A Computer Vision-Based Lane Detection Approach for an Autonomous Vehicle Using the Image Hough Transformation andtheEdge Features. In Proceedings of the Advances in Information Communication Technology and Computing; Goar, V.; Kuri, M.; Kumar, R.; Senjyu, T., Eds., Singapore, 2023; pp. 53–66. [CrossRef]
- Yang, J.; Wang, C.; Jiang, B.; Song, H.; Meng, Q. Visual Perception Enabled Industry Intelligence: State of the Art, Challenges and Prospects. IEEE Transactions on Industrial Informatics 2021, 17, 2204–2219. [CrossRef]
- Villegas-Ch, W.; Navarro, A.M.; Sanchez-Viteri, S. Optimization of inventory management through computer vision and machine learning technologies. Intelligent Systems with Applications 2024, 24, 200438. [CrossRef]
- Yousif, I.; Samaha, J.; Ryu, J.; Harik, R. Safety 4.0: Harnessing computer vision for advanced industrial protection. Manufacturing Letters 2024, 41, 1342–1356. 52nd SME North American Manufacturing Research Conference (NAMRC 52), . [CrossRef]
- Bawack, R.E.; Wamba, S.F.; Carillo, K.; Akter, S. Artificial intelligence in E-Commerce: a bibliometric study and literature review. Electronic Markets 2022, 32, 297–338. [CrossRef]
- Liang, Z.; Tao, F. Research on the Application of Artificial Intelligence in E-commerce Design. In Proceedings of the 2020 International Conference on Innovation Design and Digital Technology (ICIDDT), 2020, pp. 455–458. [CrossRef]
- Ettalibi, A.; Elouadi, A.; Mansour, A. AI and Computer Vision-based Real-time Quality Control: A Review of Industrial Applications. Procedia Computer Science 2024, 231, 212–220. 14th International Conference on Emerging Ubiquitous Systems and Pervasive Networks / 13th International Conference on Current and Future Trends of Information and Communication Technologies in Healthcare (EUSPN/ICTH 2023), . [CrossRef]
- Hu, C.; Liu, Y.; Xiang, J.; Chen, Z. Detection of Crop Leaf Diseases Using Gl-Cgan Based Data Augmentation. In Proceedings of the 2021 International Conference on Machine Learning and Cybernetics (ICMLC), 2021, pp. 1–5. [CrossRef]
- Liu, J.; Wang, X. Plant diseases and pests detection based on deep learning: a review. Plant Methods 2021, 17. [CrossRef]
- Chari, S.; Parab, S.; Pawar, H.; Raut, D.; Chavan, G. Smart Agricultural Crop Monitoring System. In Proceedings of the 2022 5th International Conference on Advances in Science and Technology (ICAST), 2022, pp. 69–72. [CrossRef]
- Obu, U.; Sarkarkar, G.; Ambekar, Y. Computer Vision for Monitor and Control of Vertical Farms Using Machine Learning Methods. In Proceedings of the 2021 International Conference on Computational Intelligence and Computing Applications (ICCICA), 2021, pp. 1–6. [CrossRef]
- Raj, M.; Prahadeeswaran, M. Revolutionizing agriculture: a review of smart farming technologies for a sustainable future. Discover Applied Sciences 2025, 7. [CrossRef]
- Puttagunta, M.; Ravi, S. Medical image analysis based on deep learning approach. Multimedia Tools and Applications 2021. [CrossRef]
- Hossain, A.; Islam, M.T.; Almutairi, A.F. A deep learning model to classify and detect brain abnormalities in portable microwave based imaging system. Scientific Reports 2022, 12. [CrossRef]
- Abbasi, S.; Tajeripour, F. Detection of brain tumor in 3D MRI images using local binary patterns and histogram orientation gradient. Neurocomputing 2017, 219, 526–535. [CrossRef]
- Ito, K.; Sugano, S.; Iwata, H. Internal bleeding detection algorithm based on determination of organ boundary by low-brightness set analysis. In Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012, pp. 4131–4136. [CrossRef]
- Shimazaki, A.; Ueda, D.; Choppin, A.; Yamamoto, A.; Honjo, T.; Shimahara, Y.; Miki, Y. Deep learning-based algorithm for lung cancer detection on chest radiographs using the segmentation method. Scientific Reports 2022, 12. [CrossRef]
- Hammad, M.S.; Ghoneim, V.F.; Mabrouk, M.S.; Al-Atabany, W.I. A hybrid deep learning approach for COVID-19 detection based on genomic image processing techniques. Scientific Reports 2023, 13. [CrossRef]
- Liang, S.; Liu, H.; Gu, Y.; Guo, X.; Li, H.; Li, L.; Wu, Z.; Liu, M.; Tao, L. Fast automated detection of COVID-19 from medical images using convolutional neural networks. Communications biology 2021, 4. [CrossRef]
- Jafri, R.; Ali, S.A.; Arabnia, H.R.; Fatima, S. Computer vision-based object recognition for the visually impaired in an indoors environment: a survey. The Visual Computer 2013, 30, 1197–1222. [CrossRef]
- Jha, S.; Seo, C.; Yang, E.; Joshi, G.P. Real time object detection and trackingsystem for video surveillance system. Multimedia Tools and Applications 2020, 80, 3981–3996. [CrossRef]
- Janakiramaiah, B.; Kalyani, G.; Karuna, A.; Prasad, L.V.N.; Krishna, M. Military object detection in defense using multi-level capsule networks. Soft Computing 2021, 27, 1045–1059. [CrossRef]
- Pham, I.; Polasek, M. Algorithm for military object detection using image data. In Proceedings of the 2014 IEEE/AIAA 33rd Digital Avionics Systems Conference (DASC), 2014, pp. 3D3–1–3D3–15. [CrossRef]
- Archana, M.; Geetha, M.K. Object Detection and Tracking Based on Trajectory in Broadcast Tennis Video. Procedia Computer Science 2015, 58, 225–232. Second International Symposium on Computer Vision and the Internet (VisionNet’15), . [CrossRef]
- Zheng, Y.; Zhang, H. Video Analysis in Sports by Lightweight Object Detection Network under the Background of Sports Industry Development. Computational Intelligence and Neuroscience 2022, 2022, 1–10. [CrossRef]
- Dogra, A.K.; Sharma, V.; Sohal, H. A survey of deep learning techniques for detecting and recognizing objects in complex environments. Computer Science Review 2024, 54, 100686. [CrossRef]
- Salari, A.; Djavadifar, A.; Liu, X.; Najjaran, H. Object recognition datasets and challenges: A review. Neurocomputing 2022, 495, 129–152. [CrossRef]
- Everingham, M.; Van Gool, L.; Williams, C.K.; Winn, J.; Zisserman, A. The pascal visual object classes (voc) challenge. International Journal of Computer Vision 2009, 88, 303–308. [CrossRef]
- Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. International Journal of Computer Vision 2020, 128, 1956–1981. [CrossRef]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, 2014, pp. 740–755. [CrossRef]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. Commun. ACM 2017, 60, 84–90. [CrossRef]
- Liu, S.; Fan, L.; Johns, E.; Yu, Z.; Xiao, C.; Anandkumar, A. Prismer: A Vision-Language Model with An Ensemble of Experts. arXiv preprint arXiv:2303.02506 2023.
- Amjoud, A.B.; Amrouch, M. Object Detection Using Deep Learning, CNNs and Vision Transformers: A Review. IEEE Access 2023, 11, 35479–35516. [CrossRef]
- Zou, Z.; Chen, K.; Shi, Z.; Guo, Y.; Ye, J. Object detection in 20 years: A survey. Proceedings of the IEEE 2023, 111, 257–276.
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [CrossRef]
- Li, K.; Cao, L. A Review of Object Detection Techniques. In Proceedings of the 2020 5th International Conference on Electromechanical Control Technology and Transportation (ICECTT), 2020, pp. 385–390. [CrossRef]
- Lowe, D. Object recognition from local scale-invariant features. In Proceedings of the Proceedings of the Seventh IEEE International Conference on Computer Vision, 1999, Vol. 2, pp. 1150–1157 vol.2. [CrossRef]
- Burger, W.; Burge, M.J., Scale-Invariant Feature Transform (SIFT). In Digital Image Processing: An Algorithmic Introduction Using Java; Springer London: London, 2016; pp. 609–664. [CrossRef]
- Pagire, V.; Chavali, M.; Kale, A. A comprehensive review of object detection with traditional and deep learning methods. Signal Processing 2025, 237, 110075. [CrossRef]
- Tsourounis, D.; Kastaniotis, D.; Theoharatos, C.; Kazantzidis, A.; Economou, G. SIFT-CNN: When Convolutional Neural Networks Meet Dense SIFT Descriptors for Image and Sequence Classification. Journal of Imaging 2022, 8. [CrossRef]
- Zheng, L.; Yang, Y.; Tian, Q. SIFT Meets CNN: A Decade Survey of Instance Retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence 2018, 40, 1224–1244. [CrossRef]
- Burghardt, T.; Calic, J. Real-time Face Detection and Tracking of Animals. In Proceedings of the 2006 8th Seminar on Neural Network Applications in Electrical Engineering, 2006, pp. 27–32. [CrossRef]
- Viola, P.; Jones, M. Rapid object detection using a boosted cascade of simple features. In Proceedings of the Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, 2001, Vol. 1, pp. I–I. [CrossRef]
- Viola, P.; Jones, M. Robust Real-Time Face Detection. International Journal of Computer Vision 2004, 57, 137–154. [CrossRef]
- Dahirou, Z.; Zheng, M.; Yuxin, M. Face Detection with Viola Jones Algorithm. In Proceedings of the 2020 7th International Conference on Information Science and Control Engineering (ICISCE), 2020, pp. 602–606. [CrossRef]
- Zhang, G.; Gao, F.; Liu, C.; Liu, W.; Yuan, H. A pedestrian detection method based on SVM classifier and optimized Histograms of Oriented Gradients feature. In Proceedings of the 2010 Sixth International Conference on Natural Computation, 2010, Vol. 6, pp. 3257–3260. [CrossRef]
- Felzenszwalb, P.F.; Girshick, R.B.; McAllester, D. Cascade object detection with deformable part models. In Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2010, pp. 2241–2248. [CrossRef]
- Bay, H.; Tuytelaars, T.; Van Gool, L. Surf: Speeded up robust features. Lecture notes in computer science 2006, 3951, 404–417.
- Dalal, N.; Triggs, B. Histograms of oriented gradients for human detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), 2005, Vol. 1, pp. 886–893 vol. 1. [CrossRef]
- Lowe, D. Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision 2004, 60, 91–. [CrossRef]
- Edozie, E.; Nuhu, A.; K.J, U.; Sadiq, B. Comprehensive review of recent developments in visual object detection based on deep learning. Artificial Intelligence Review 2025, 58. [CrossRef]
- Felzenszwalb, P.F.; Girshick, R.B.; McAllester, D.; Ramanan, D. Object Detection with Discriminatively Trained Part-Based Models. IEEE Transactions on Pattern Analysis and Machine Intelligence 2010, 32, 1627–1645. [CrossRef]
- Yan, J.; Lei, Z.; Wen, L.; Li, S.Z. The Fastest Deformable Part Model for Object Detection. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 2497–2504. [CrossRef]
- Girshick, R.; Iandola, F.; Darrell, T.; Malik, J. Deformable part models are convolutional neural networks. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 437–446. [CrossRef]
- Li, J.; Wong, H.C.; Lo, S.L.; Xin, Y. Multiple Object Detection by a Deformable Part-Based Model and an R-CNN. IEEE Signal Processing Letters 2018, 25, 288–292. [CrossRef]
- Bay, H.; Ess, A.; Tuytelaars, T.; Van Gool, L. Speeded-Up Robust Features (SURF). Computer Vision and Image Understanding 2008, 110, 346–359. Similarity Matching in Computer Vision and Multimedia, . [CrossRef]
- Paul, M.; Karsh, R.K.; Ahmed Talukdar, F. Image Hashing based on Shape Context and Speeded Up Robust Features (SURF). In Proceedings of the 2019 International Conference on Automation, Computational and Technology Management (ICACTM), 2019, pp. 464–468. [CrossRef]
- Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 1–9. [CrossRef]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition, 2015, [arXiv:cs.CV/1409.1556].
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
- Yang, A.; Liu, M.; Wang, J.; He, Q.; Zhang, N.; Jia, J. Investigations of Object Detectors with Deep Learning Methods: A Review. In Proceedings of the 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI), 2022, pp. 351–358. [CrossRef]
- Zhou, Y. Research Advanced in Object Detection Based on Deep Learning. In Proceedings of the 2022 4th International Conference on Artificial Intelligence and Advanced Manufacturing (AIAM), 2022, pp. 694–698. [CrossRef]
- Liang, Y.j.; Cui, X.p.; Xu, X.h.; Jiang, F. A Review on Deep Learning Techniques Applied to Object Detection. In Proceedings of the 2020 7th International Conference on Information Science and Control Engineering (ICISCE), 2020, pp. 120–124. [CrossRef]
- Li, Y.; Miao, N.; Ma, L.; Shuang, F.; Huang, X. Transformer for object detection: Review and benchmark. Engineering Applications of Artificial Intelligence 2023, 126, 107021. [CrossRef]
- Ganga, B.; B.T., L.; K.R., V. Object detection and crowd analysis using deep learning techniques: Comprehensive review and future directions. Neurocomputing 2024, 597, 127932. [CrossRef]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
- Uijlings, J.R.R.; van de Sande, K.E.A.; Gevers, T.; Smeulders, A.W.M. Selective Search for Object Recognition. International Journal of Computer Vision 2013, 104, 154–171. [CrossRef]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Region-Based Convolutional Networks for Accurate Object Detection and Segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 2016, 38, 142–158. [CrossRef]
- Xie, X.; Cheng, G.; Wang, J.; Yao, X.; Han, J. Oriented R-CNN for Object Detection. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3500–3509. [CrossRef]
- Girshick, R. Fast R-CNN. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1440–1448. [CrossRef]
- Qiao, S.; Zhang, Q.; Wang, Z. A Review of Deep-Learning-Based SAR Image Ship Interpretation Technology: The Latest Advances. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2025, 18, 26152–26185. [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 2015, 28.
- He, H.; Xu, H.; Zhang, Y.; Gao, K.; Li, H.; Ma, L.; Li, J. Mask R-CNN based automated identification and extraction of oil well sites. International Journal of Applied Earth Observation and Geoinformation 2022, 112, 102875.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition. In Proceedings of the Computer Vision – ECCV 2014; Fleet, D.; Pajdla, T.; Schiele, B.; Tuytelaars, T., Eds., Cham, 2014; pp. 346–361.
- Razavian, A.S.; Azizpour, H.; Sullivan, J.; Carlsson, S. CNN Features Off-the-Shelf: An Astounding Baseline for Recognition. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2014, pp. 512–519. [CrossRef]
- Dewi, C.; Juli Christanto, H. Combination of Deep Cross-Stage Partial Network and Spatial Pyramid Pooling for Automatic Hand Detection. Big Data and Cognitive Computing 2022, 6. [CrossRef]
- Hozhabr, S.H.; Giorgi, R. A Survey on Real-Time Object Detection on FPGAs. IEEE Access 2025, 13, 38195–38238. [CrossRef]
- Dewi, C.; Chen, R.C.; Tai, S.K. Evaluation of Robust Spatial Pyramid Pooling Based on Convolutional Neural Network for Traffic Sign Recognition System. Electronics 2020, 9. [CrossRef]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 936–944. [CrossRef]
- Zhang, Y.; Han, J.H.; Kwon, Y.W.; Moon, Y.S. A New Architecture of Feature Pyramid Network for Object Detection. In Proceedings of the 2020 IEEE 6th International Conference on Computer and Communications (ICCC), 2020, pp. 1224–1228. [CrossRef]
- Ghiasi, G.; Lin, T.Y.; Le, Q.V. NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 7029–7038. [CrossRef]
- Lee, S.; Kim, E. Multiple Object Tracking via Feature Pyramid Siamese Networks. IEEE Access 2019, 7, 8181–8194. [CrossRef]
- Xing, H.; Wang, S.; Zheng, D.; Zhao, X. Dual attention based feature pyramid network. China Communications 2020, 17, 242–252. [CrossRef]
- Zhang, X.; Yan, L.; Gerada, C. A densely connected feature pyramid network for object detection. In Proceedings of the CSAA/IET International Conference on Aircraft Utility Systems (AUS 2020), 2020, Vol. 2020, pp. 699–703. [CrossRef]
- Zhang, Y.; Li, X.; Wang, F.; Wei, B.; Li, L. A Comprehensive Review of One-stage Networks for Object Detection. In Proceedings of the 2021 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), 2021, pp. 1–6. [CrossRef]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, 2022, [arXiv:cs.CV/2207.02696].
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection . In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Los Alamitos, CA, USA, 2016; pp. 779–788. [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors, 2025, [arXiv:cs.CV/2502.12524].
- Redmon, J.; Farhadi, A. YOLO9000: Better, Faster, Stronger, 2016, [arXiv:cs.CV/1612.08242].
- Redmon, J.; Farhadi, A. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 2018.
- Sanchez, S.A.; Romero, H.J.; Morales, A.D. A review: Comparison of performance metrics of pretrained models for object detection using the TensorFlow framework. IOP Conference Series: Materials Science and Engineering 2020, 844, 012024. [CrossRef]
- Vijayakumar, A.; Vairavasundaram, S. Yolo-based object detection models: A review and its applications. Multimedia Tools and Applications 2024, 83, 83535–83574.
- Hua, Z.; Aranganadin, K.; Yeh, C.C.; Hai, X.; Huang, C.Y.; Leung, T.C.; Hsu, H.Y.; Lan, Y.C.; Lin, M.C. A Benchmark Review of YOLO Algorithm Developments for Object Detection. IEEE Access 2025, 13, 123515–123545. [CrossRef]
- Keita, Z. YOLO object detection explained, 2022.
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 2020.
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. You Only Learn One Representation: Unified Network for Multiple Tasks, 2021, [arXiv:cs.CV/2105.04206].
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021, 2021, [arXiv:cs.CV/2107.08430].
- Bhumbla, S.; Gupta, D.K.; Nisha. A Review: Object Detection Algorithms. In Proceedings of the 2023 Third International Conference on Secure Cyber Computing and Communication (ICSCCC), 2023, pp. 827–832. [CrossRef]
- Jocher, G.; Stoken, A.; Borovec, J.; NanoCode012.; Chaurasia, A.; TaoXie.; Changyu, L.; V, A.; Laughing.; tkianai.; et al. ultralytics/yolov5: v5.0 - YOLOv5-P6 1280 models, AWS, Supervise.ly and YouTube integrations, 2021. [CrossRef]
- Ali, M.L.; Zhang, Z. The YOLO Framework: A Comprehensive Review of Evolution, Applications, and Benchmarks in Object Detection. Computers 2024, 13. [CrossRef]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications, 2022, [arXiv:cs.CV/2209.02976].
- Wang, C.; Bochkovskiy, A.; Liao, H.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. CoRR 2022, abs/2207.02696, [2207.02696]. [CrossRef]
- Varghese, R.; M., S. YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness. In Proceedings of the 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), 2024, pp. 1–6. [CrossRef]
- Ultralytics. Introducing ultralytics YOLOv8, 2023.
- Ultralytics. GitHub NEW - YOLOv8 in PyTorch, 2023.
- Lou, H.; Duan, X.; Guo, J.; Liu, H.; Gu, J.; Bi, L.; Chen, H. DC-YOLOv8: Small-Size Object Detection Algorithm Based on Camera Sensor. Electronics 2023, 12. [CrossRef]
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. Yolov9: Learning what you want to learn using programmable gradient information. arXiv preprint arXiv:2402.13616 2024.
- Ao Wang, Hui Chen, L.L. YOLOv10: Real-Time End-to-End Object Detection. arXiv preprint arXiv:2405.14458 2024.
- Jegham, N.; Koh, C.Y.; Abdelatti, M.; Hendawi, A. YOLO Evolution: A Comprehensive Benchmark and Architectural Review of YOLOv12, YOLO11, and Their Previous Versions, 2025, [arXiv:cs.CV/2411.00201].
- Khanam, R.; Hussain, M. YOLOv11: An Overview of the Key Architectural Enhancements, 2024, [arXiv:cs.CV/2410.17725].
- Sapkota, R.; Meng, Z.; Churuvija, M.; Du, X.; Ma, Z.; Karkee, M. Comprehensive Performance Evaluation of YOLOv12, YOLO11, YOLOv10, YOLOv9 and YOLOv8 on Detecting and Counting Fruitlet in Complex Orchard Environments, 2025, [arXiv:cs.CV/2407.12040].
- Hidayatullah, P.; Syakrani, N.; Sholahuddin, M.R.; Gelar, T.; Tubagus, R. YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review, 2025, [arXiv:cs.CV/2501.13400].
- Alif, M.A.R.; Hussain, M. YOLOv12: A Breakdown of the Key Architectural Features, 2025, [arXiv:cs.CV/2502.14740].
- Sapkota, R.; Qureshi, R.; Calero, M.F.; Badjugar, C.; Nepal, U.; Poulose, A.; Zeno, P.; Vaddevolu, U.B.P.; Khan, S.; Shoman, M.; et al. YOLOv12 to Its Genesis: A Decadal and Comprehensive Review of The You Only Look Once (YOLO) Series, 2025, [arXiv:cs.CV/2406.19407].
- Wei, J.; As’arry, A.; Anas Md Rezali, K.; Zuhri Mohamed Yusoff, M.; Ma, H.; Zhang, K. A Review of YOLO Algorithm and Its Applications in Autonomous Driving Object Detection. IEEE Access 2025, 13, 93688–93711. [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the Computer Vision – ECCV 2016; Leibe, B.; Matas, J.; Sebe, N.; Welling, M., Eds., Cham, 2016; pp. 21–37.
- Shi, X.; Liu, W.; He, L.; Jin, H.; Li, M.; Chen, Y. Optimizing the SSD Burst Buffer by Traffic Detection. ACM Trans. Archit. Code Optim. 2020, 17. [CrossRef]
- Fu, C.Y.; Liu, W.; Ranga, A.; Tyagi, A.; Berg, A.C. DSSD : Deconvolutional Single Shot Detector, 2017, [arXiv:cs.CV/1701.06659].
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2999–3007. [CrossRef]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 2020, 42, 318–327. [CrossRef]
- Nora, L.E.A.; Dias, R.L.; De Figueiredo, F.A.; Mafra, S.B.; Silva, H.S. Evaluating DETR, RetinaNet, and RTMDet Models for Object Detection in Rural Drone Imagery. In Proceedings of the 2024 International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA). IEEE, 2024, pp. 1326–1331.
- Nakada, A.; Niikura, R.; Otani, K.; Kurose, Y.; Hayashi, Y.; Kitamura, K.; Nakanishi, H.; Kawano, S.; Honda, T.; Hasatani, K.; et al. Improved object detection artificial intelligence using the revised RetinaNet model for the automatic detection of ulcerations, vascular lesions, and tumors in wireless capsule endoscopy. Biomedicines 2023, 11, 942.
- Shan, Y.; Jian, W.; Li, H.; Bo, L.; Hao, Z. Research on Occluded Object Detection by Improved RetinaNet. Journal of Computer Engineering & Applications 2022, 58.
- Lamichhane, B.R.; Srijuntongsiri, G.; Horanont, T. CNN based 2D object detection techniques: a review. Frontiers in Computer Science 2025, Volume 7 - 2025. [CrossRef]
- Yu, L.; Tang, L.; Mu, L. A Review of DEtection TRansformer: From Basic Architecture to Advanced Developments and Visual Perception Applications. Sensors 2025, 25. [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ArXiv 2020, abs/2010.11929.
- Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y.; et al. A Survey on Vision Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 2023, 45, 87–110. [CrossRef]
- Dubey, S.R.; Singh, S.K. Transformer-Based Generative Adversarial Networks in Computer Vision: A Comprehensive Survey. IEEE Transactions on Artificial Intelligence 2024, 5, 4851–4867. [CrossRef]
- Wang, J.; Ma, S.; An, Y.; Dong, R. A Comparative Study of Vision Transformer and Convolutional Neural Network Models in Geological Fault Detection. IEEE Access 2024, 12, 136148–136159. [CrossRef]
- Wang, Y.; Deng, Y.; Zheng, Y.; Chattopadhyay, P.; Wang, L. Vision Transformers for Image Classification: A Comparative Survey. Technologies 2025, 13. [CrossRef]
- Mia, M.S.; Arnob, A.B.H.; Naim, A.; Voban, A.A.B.; Islam, M.S. ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain. In Proceedings of the 2023 International Conference on the Cognitive Computing and Complex Data (ICCD), 2023, pp. 101–117. [CrossRef]
- V, N.; Sai, B.K.; Ishwariya, R. Deep Learning Based Binary Classification of Invasive Ductal Carcinoma: A Comparative Study on CNN and VIT Models. In Proceedings of the 2024 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE), 2024, pp. 398–403. [CrossRef]
- Bousaid, R.; El Hajji, M.; Es-Saady, Y. Facial Emotions Recognition Using Vit and Transfer Learning. In Proceedings of the 2022 5th International Conference on Advanced Communication Technologies and Networking (CommNet), 2022, pp. 1–6. [CrossRef]
- Ishfaq, M.; Saadia, A.; Alserhani, F.M.; Gul, A. Enhancing Security: Infused Hybrid Vision Transformer for Signature Verification. IEEE Access 2024, 12, 137504–137521. [CrossRef]
- Sadik, M.R.; Sony, R.I.; Prova, N.N.I.; Mahanandi, Y.; Maruf, A.A.; Fahim, S.H.; Islam, M. Computer Vision Based Bangla Sign Language Recognition Using Transfer Learning. In Proceedings of the 2024 Second International Conference on Data Science and Information System (ICDSIS), 2024, pp. 1–7. [CrossRef]
- Mohsan, M.M.; Akram, M.U.; Rasool, G.; Alghamdi, N.S.; Baqai, M.A.A.; Abbas, M. Vision Transformer and Language Model Based Radiology Report Generation. IEEE Access 2023, 11, 1814–1824. [CrossRef]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the Computer Vision – ECCV 2020; Vedaldi, A.; Bischof, H.; Brox, T.; Frahm, J.M., Eds., Cham, 2020; pp. 213–229.
- Shehzadi, T.; Hashmi, K.A.; Stricker, D.; Afzal, M.Z. Object Detection with Transformers: A Review, 2023, [arXiv:cs.CV/2306.04670].
- Zhu, M.; Gong, Y.; Tian, C.; Zhu, Z. A Systematic Survey of Transformer-Based 3D Object Detection for Autonomous Driving: Methods, Challenges and Trends. Drones 2024, 8. [CrossRef]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9992–10002. [CrossRef]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980–2988. [CrossRef]
- Cai, Z.; Vasconcelos, N. Cascade R-CNN: Delving Into High Quality Object Detection. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 6154–6162. [CrossRef]
- Gao, P.; Zheng, M.; Wang, X.; Dai, J.; Li, H. Fast convergence of detr with spatially modulated co-attention. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3621–3630.
- Wang, Y.; Zhang, X.; Yang, T.; Sun, J. Anchor DETR: Query Design for Transformer-Based Object Detection, 2022, [arXiv:cs.CV/2109.07107].
- Huang, Z.; Tao, X.; Liu, X. NAN-DETR: noising multi-anchor makes DETR better for object detection. Frontiers in Neurorobotics 2024, Volume 18 - 2024. [CrossRef]
- Muzammul, M.; Li, X. Comprehensive review of deep learning-based tiny object detection: challenges, strategies, and future directions. Knowledge and Information Systems 2025. [CrossRef]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable {DETR}: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations, 2021.
- Farjon, G.; Huijun, L.; Edan, Y. Deep-learning-based counting methods, datasets, and applications in agriculture: A review. Precision Agriculture 2023, 24, 1683–1711.
- Preeti.; Rana, C. Artificial intelligence based object detection and traffic prediction by autonomous vehicles – A review. Expert Systems with Applications 2024, 255, 124664. [CrossRef]
- Sapkota, R.; Karkee, M. Object detection with multimodal large vision-language models: An in-depth review. Information Fusion 2026, 126, 103575. [CrossRef]
- Liang, Z.; Xu, Y.; Hong, Y.; Shang, P.; Wang, Q.; Fu, Q.; Liu, K. A Survey of Multimodel Large Language Models. In Proceedings of the Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering, New York, NY, USA, 2024; CAICE ’24, p. 405–409. [CrossRef]
- Yang, Z.; Li, L.; Lin, K.; Wang, J.; Lin, C.C.; Liu, Z.; Wang, L. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision), 2023, [arXiv:cs.CV/2309.17421].
- Carolan, K.; Fennelly, L.; Smeaton, A.F. A Review of Multi-Modal Large Language and Vision Models, 2024, [arXiv:cs.CL/2404.01322].
- Chen, X.; Wu, Z.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. arXiv preprint arXiv:2501.17811 2025.
- Tschannen, M.; Gritsenko, A.; Wang, X.; Naeem, M.F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; et al. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features, 2025, [arXiv:cs.CV/2502.14786].
- Wang, Y.; Zhu, H.; Liu, M.; Yang, J.; Fang, H.S.; He, T. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers, 2025, [arXiv:cs.RO/2507.01016].
- Lv, T.; Huang, Y.; Chen, J.; Zhao, Y.; Jia, Y.; Cui, L.; Ma, S.; Chang, Y.; Huang, S.; Wang, W.; et al. KOSMOS-2.5: A Multimodal Literate Model, 2024, [arXiv:cs.CL/2309.11419].
- Wu, Z.; Chen, X.; Pan, Z.; Liu, X.; Liu, W.; Dai, D.; Gao, H.; Ma, Y.; Wu, C.; Wang, B.; et al. DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding, 2024, [arXiv:cs.CV/2412.10302].
- Dai, W.; Li, J.; Li, D.; Tiong, A.M.H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; Hoi, S. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning, 2023, [arXiv:cs.CV/2305.06500].
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, 2023, [arXiv:cs.CV/2301.12597].
- Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: a Visual Language Model for Few-Shot Learning. In Proceedings of the Advances in Neural Information Processing Systems; Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; Oh, A., Eds. Curran Associates, Inc., 2022, Vol. 35, pp. 23716–23736.
- Zhu, D.; Chen, J.; Shen, X.; Li, X.; Elhoseiny, M. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, 2023, [arXiv:cs.CV/2304.10592].
- Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. LLaVA-OneVision: Easy Visual Task Transfer, 2024, [arXiv:cs.CV/2408.03326].
- Li, F.; Zhang, R.; Zhang, H.; Zhang, Y.; Li, B.; Li, W.; Ma, Z.; Li, C. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models, 2024, [arXiv:cs.CV/2407.07895].
- Lu, J.; Srivastava, S.; Chen, J.; Shrestha, R.; Acharya, M.; Kafle, K.; Kanan, C. Revisiting Multi-Modal LLM Evaluation, 2024, [arXiv:cs.AI/2408.05334].
- Ren, T.; Jiang, Q.; Liu, S.; Zeng, Z.; Liu, W.; Gao, H.; Huang, H.; Ma, Z.; Jiang, X.; Chen, Y.; et al. Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection, 2024, [arXiv:cs.CV/2405.10300].
- Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H.; et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection, 2024, [arXiv:cs.CV/2303.05499].
- Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; Yuan, L. Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks, 2023, [arXiv:cs.CV/2311.06242].
- Abbas, M.H.; Xinyou, Z.; Manzoor, M.; Tabraiz, S.; Hassan, I. Visionxplore: Florence-2 for Robust Targeted Object Detection. In Proceedings of the 2024 21st International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), 2024, pp. 1–4. [CrossRef]
- Jindal, K.; Saxena, A.; Aggarwal, G.; Rajput, A.S.; Dutta, M.K. Image-Text to Text based Multimodal AI using FLORENCE-2 for Remote Sensing Application. In Proceedings of the 2025 International Conference on Engineering, Technology & Management (ICETM), 2025, pp. 1–6. [CrossRef]
- Ucar, A.; Ro, S.; Satwika, S.; Gayathri, P.Y.; Balsha, M.G. Fine-Tuning Florence2 for Enhanced Object Detection in Un-constructed Environments: Vision-Language Model Approach, 2025, [arXiv:cs.CV/2503.04918].
- Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; Shan, Y. YOLO-World: Real-Time Open-Vocabulary Object Detection. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 16901–16911. [CrossRef]
- Jing, Z.; Su, Y.; Han, Y. When Large Language Models Meet Vector Databases: A Survey. In Proceedings of the 2025 Conference on Artificial Intelligence x Multimedia (AIxMM), 2025, pp. 7–13. [CrossRef]
- Liu, D.; Yang, M.; Qu, X.; Zhou, P.; Cheng, Y.; Hu, W. A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends, 2024, [arXiv:cs.CV/2407.07403].
- Wang, A.; Liu, L.; Chen, H.; Lin, Z.; Han, J.; Ding, G. YOLOE: Real-Time Seeing Anything, 2025, [arXiv:cs.CV/2503.07465].
- Gupta, A.; Dollár, P.; Girshick, R. LVIS: A Dataset for Large Vocabulary Instance Segmentation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 5351–5359. [CrossRef]
- Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. CogVLM: Visual Expert for Pretrained Language Models, 2024, [arXiv:cs.CV/2311.03079].
- Hong, W.; Wang, W.; Ding, M.; Yu, W.; Lv, Q.; Wang, Y.; Cheng, Y.; Huang, S.; Ji, J.; Xue, Z.; et al. CogVLM2: Visual Language Models for Image and Video Understanding, 2024, [arXiv:cs.CV/2408.16500].
- P, M.; Velvizhy, P. A Comprehensive Review of Supervised Fine-Tuning for Large Language Models in Creative Applications and Content Moderation. In Proceedings of the 2025 International Conference on Inventive Computation Technologies (ICICT), 2025, pp. 1294–1299. [CrossRef]
- Cao, Y.; Ivanovic, B.; Xiao, C.; Pavone, M. Reinforcement Learning with Human Feedback for Realistic Traffic Simulation. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 14428–14434. [CrossRef]
- Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; et al. Simple Open-Vocabulary Object Detection with Vision Transformers, 2022, [arXiv:cs.CV/2205.06230].
- Minderer, M.; Gritsenko, A.; Houlsby, N. Scaling Open-Vocabulary Object Detection, 2024, [arXiv:cs.CV/2306.09683].
- Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; Chen, E. A survey on multimodal large language models. National Science Review 2024, 11, nwae403. [CrossRef]
- Zhang, X.; Lu, Y.; Wang, W.; Yan, A.; Yan, J.; Qin, L.; Wang, H.; Yan, X.; Wang, W.Y.; Petzold, L.R. GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks, 2023, [arXiv:cs.CV/2311.01361].
- Jiao, Q.; Chen, D.; Huang, Y.; Li, Y.; Shen, Y. From Training-Free to Adaptive: Empirical Insights into MLLMs’ Understanding of Detection Information, 2024, [arXiv:cs.CV/2401.17981].
- Husein, R.A.; Aburajouh, H.; Catal, C. Large language models for code completion: A systematic literature review. Computer Standards & Interfaces 2025, 92, 103917. [CrossRef]
- Shehab, M.A.; Wardat, M.; Omari, S.; Jararweh, Y. Evaluating Large Language Models for Code Generation: Assessing Accuracy, Quality, and Performance. In Proceedings of the 2024 2nd International Conference on Foundation and Large Language Models (FLLM), 2024, pp. 407–416. [CrossRef]
- Wang, W.; Chen, Z.; Chen, X.; Wu, J.; Zhu, X.; Zeng, G.; Luo, P.; Lu, T.; Zhou, J.; Qiao, Y.; et al. VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks, 2023, [arXiv:cs.CV/2305.11175].
- Wu, J.; Zhong, M.; Xing, S.; Lai, Z.; Liu, Z.; Chen, Z.; Wang, W.; Zhu, X.; Lu, L.; Lu, T.; et al. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks, 2024, [arXiv:cs.CV/2406.08394].
- Shehzadi, T.; Hashmi, K.A.; Liwicki, M.; Stricker, D.; Afzal, M.Z. Object Detection with Transformers: A Review. Sensors 2025, 25. [CrossRef]
- Tsirtsakis, P.; Zacharis, G.; Maraslidis, G.S.; Fragulis, G.F. Deep learning for object recognition: A comprehensive review of models and algorithms. International Journal of Cognitive Computing in Engineering 2025, 6, 298–312. [CrossRef]
- Li, F.F.; Andreeto, M.; Ranzato, M.; Perona, P. Caltech 101, 2003. [CrossRef]
- Griffin, G.; Holub, A.; Perona, P. Caltech 256, 2006. [CrossRef]
- Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.J.; Shamma, D.A.; et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 2017, 123, 32–73.
- Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision 2015, 115, 211–252.
- Liu, Q.; Li, Z.; Yuan, S.; Zhu, Y.; Li, X. Review on Vehicle Detection Technology for Unmanned Ground Vehicles. Sensors 2021, 21. [CrossRef]
- Geiger, A.; Lenz, P.; Urtasun, R. Are we ready for autonomous driving? The KITTI vision benchmark suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, 2012, pp. 3354–3361. [CrossRef]
- Yu, F.; Xian, W.; Chen, Y.; Liu, F.; Liao, M.; Madhavan, V.; Darrell, T. BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling. ArXiv 2018, abs/1805.04687.
- He, X. Vehicle target detection algorithm based on yolov5. Frontiers in Computing and Intelligent Systems 2023, 3, 56–59. [CrossRef]
- Cordts, M.; Omran, M.; Ramos, S.; Scharwächter, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes Dataset. In Proceedings of the CVPR Workshop on The Future of Datasets in Vision, 2015.
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Proceedings of the Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
- Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuScenes: A Multimodal Dataset for Autonomous Driving. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11618–11628. [CrossRef]
- Iqbal, H.; Sadia, H.; Al-Kaff, A.; Garcié, F. Novelty Detection in Autonomous Driving: A Generative Multi-Modal Sensor Fusion Approach. IEEE Open Journal of Intelligent Transportation Systems 2025, 6, 799–812. [CrossRef]
- Wang, N.; Shang, D.; Gong, Y.; Hu, X.; Song, Z.; Yang, L.; Huang, Y.; Wang, X.; Lu, J. Collaborative Perception Datasets for Autonomous Driving: A Review. IEEE Sensors Journal 2025, 25, 30255–30274. [CrossRef]
- Höhne, M.O.; Menke., M.; Bieshaar, M. Enhancing Data Efficiency for Training Object Detectors. In Proceedings of the 2025 IEEE Intelligent Vehicles Symposium (IV), 2025, pp. 285–292. [CrossRef]
- Ahmed Fime, A.; Mahmud, S.; Das, A.; Islam, M.S.; Kim, J.H. Automatic Scene Generation: State-of-the-Art Techniques, Models, Datasets, Challenges, and Future Prospects. IEEE Access 2025, 13, 95753–95796. [CrossRef]
- Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B.; et al. Scalability in Perception for Autonomous Driving: Waymo Open Dataset. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2443–2451. [CrossRef]
- Hu, X.; Zheng, Z.; Chen, D.; Zhang, X.; Sun, J. Processing, assessing, and enhancing the Waymo autonomous vehicle open dataset for driving behavior research. Transportation Research Part C: Emerging Technologies 2022, 134, 103490. [CrossRef]
- Mei, J.; Zhu, A.Z.; Yan, X.; Yan, H.; Qiao, S.; Chen, L.C.; Kretzschmar, H. Waymo Open Dataset: Panoramic Video Panoptic Segmentation. In Proceedings of the Computer Vision – ECCV 2022; Avidan, S.; Brostow, G.; Cissé, M.; Farinella, G.M.; Hassner, T., Eds., Cham, 2022; pp. 53–72.
- Dollar, P.; Wojek, C.; Schiele, B.; Perona, P. Pedestrian detection: A benchmark. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 304–311. [CrossRef]
- Braun, M.; Krebs, S.; Flohr, F.; Gavrila, D.M. EuroCity Persons: A Novel Benchmark for Person Detection in Traffic Scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence 2019, 41, 1844–1861. [CrossRef]
- Rezaei, M.; Terauchi, M. Vehicle Detection Based on Multi-feature Clues and Dempster-Shafer Fusion Theory. In Proceedings of the Image and Video Technology; Klette, R.; Rivera, M.; Satoh, S., Eds., Berlin, Heidelberg, 2014; pp. 60–72.
- Mao, J.; Niu, M.; Jiang, C.; Liang, H.; Chen, J.; Liang, X.; Li, Y.; Ye, C.; Zhang, W.; Li, Z.; et al. One Million Scenes for Autonomous Driving: ONCE Dataset, 2021, [arXiv:cs.CV/2106.11037].
- Pradeep Kumar, P.; Kant, K. TU-DAT: A Computer Vision Dataset on Road Traffic Anomalies. Sensors 2025, 25. [CrossRef]
- Zhou, Y.; Li, H. A Survey of Dense Object Detection Methods Based on Deep Learning. IEEE Access 2024, 12, 179944–179961. [CrossRef]
- Zhao, H.; Li, X.; Xu, C.; Xu, B.; Liu, H. A Survey of Automatic Driving Environment Perception. In Proceedings of the 2024 IEEE 24th International Conference on Software Quality, Reliability, and Security Companion (QRS-C), 2024, pp. 1038–1047. [CrossRef]
- Liu, M.; Yurtsever, E.; Fossaert, J.; Zhou, X.; Zimmer, W.; Cui, Y.; Zagar, B.L.; Knoll, A.C. A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook. IEEE Transactions on Intelligent Vehicles 2024, 9, 7138–7164. [CrossRef]
- Krebs, S.; Braun, M.; Gavrila, D.M. EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in Traffic. IEEE Transactions on Pattern Analysis and Machine Intelligence 2024, 46, 10929–10943. [CrossRef]
- van Andel, M.P.; Boekema, H.J.H.; Gavrila, D.M. SAM-Maps: Road Map Generation for Automated Vehicles in Urban Areas. In Proceedings of the 2025 IEEE Intelligent Vehicles Symposium (IV), 2025, pp. 221–228. [CrossRef]
- Rezaei, M.; Terauchi, M.; Klette, R. Robust Vehicle Detection and Distance Estimation Under Challenging Lighting Conditions. IEEE Transactions on Intelligent Transportation Systems 2015, 16, 2723–2743. [CrossRef]
- Yang, C.; Chen, Y.; Li, Z.; Wang, X.; Shi, K.; Yao, L.; Xu, G.; Guo, Z. Deep multimodal learning for time series analysis in social computing: a survey. International Journal of Multimedia Information Retrieval 2025, 14. [CrossRef]
- Gamerdinger, J.; Teufel, S.; Bringmann, O. Datasets for Lane Detection in Autonomous Driving: A Comprehensive Review, 2025, [arXiv:cs.CV/2504.08540].
- Jiang, C.; Ren, H.; Li, F.; Hong, Z.; Huo, H.; Zhang, J.; Xin, J. Object detection from aerial multi-angle thermal infrared remote sensing images: Dataset and method. ISPRS Journal of Photogrammetry and Remote Sensing 2025, 228, 438–452. [CrossRef]
- Cheng, G.; Han, J. A survey on object detection in optical remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing 2016, 117, 11–28. [CrossRef]
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing 2020, 159, 296–307. [CrossRef]
- Affek, M.; Szymański, J. A Survey on the Datasets and Algorithms for Satellite Data Applications. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2024, 17, 16078–16099. [CrossRef]
- Xia, G.S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; Zhang, L. DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3974–3983. [CrossRef]
- Xia, G.; Ding, J.; Qian, M.; Xue, N.; Han, J.; Bai, X.; Yang, M.Y.; Li, S.; Belongie, S.J.; Luo, J.; et al. LUAI Challenge 2021 on Learning to Understand Aerial Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, ICCVW 2021, Montreal, BC, Canada, October 11-17, 2021. IEEE, 2021, pp. 762–768. [CrossRef]
- Lam, D.; Kuzma, R.; McGee, K.; Dooley, S.; Laielli, M.; Klaric, M.; Bulatov, Y.; McCord, B. xView: Objects in Context in Overhead Imagery, 2018, [arXiv:cs.CV/1802.07856].
- Ninja, D. Visualization Tools for xView 2018 Dataset. https://datasetninja.com/xview , 2025. visited on 2025-08-29.
- Olamofe, J.; Dong, X.; Qian, L.; Shields, E. Performance Evaluation of Data Augmentation for Object Detection in XView Dataset. In Proceedings of the 2022 International Conference on Intelligent Data Science Technologies and Applications (IDSTA), 2022, pp. 26–33. [CrossRef]
- Razakarivony, S.; Jurie, F. Vehicle detection in aerial imagery : A small target detection benchmark. Journal of Visual Communication and Image Representation 2016, 34, 187–203. [CrossRef]
- Kumar, S.; Rajan, E.G.; Rani, S. A Study on Vehicle Detection through Aerial Images: Various Challenges, Issues and Applications. In Proceedings of the 2021 International Conference on Computing, Communication, and Intelligent Systems (ICCCIS), 2021, pp. 504–509. [CrossRef]
- Li, K.; Wan, G.; Cheng, G. DIOR, 2025. [CrossRef]
- huanran ye. Inria Aerial Image Labeling Dataset, 2022. [CrossRef]
- Maggiori, E.; Tarabalka, Y.; Charpiat, G.; Alliez, P. Can semantic labeling methods generalize to any city? the inria aerial image labeling benchmark. In Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2017, pp. 3226–3229. [CrossRef]
- Shoaib, H.A.; Nabil, H.R.; Rahman, M.A.; Kabir, M.M.; Mridha, M.; Shin, J. Advancements and challenges of deep learning architectures for aerial image analysis: A systematic review. Intelligent Systems with Applications 2025, 27, 200537. [CrossRef]
- Agrawal, S.C.; Sharma, V.; Bhardwaj, P. Face Recognition:A review of Datasets and Methods. In Proceedings of the 2021 5th International Conference on Information Systems and Computer Networks (ISCON), 2021, pp. 1–6. [CrossRef]
- Huang, G.B.; Mattar, M.A.; Berg, T.L.; Learned-Miller, E. Labeled Faces in the Wild: A Database forStudying Face Recognition in Unconstrained Environments, 2008.
- Yang, S.; Luo, P.; Loy, C.C.; Tang, X. WIDER FACE: A Face Detection Benchmark. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
- Ge, S.; Li, J.; Ye, Q.; Luo, Z. Detecting Masked Faces in the Wild With LLE-CNNs. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
- Umer, M.; Sadiq, S.; Alhebshi, R.M.; Alsubai, S.; Al Hejaili, A.; Eshmawi, A.A.; Nappi, M.; Ashraf, I. Face mask detection using deep convolutional neural network and multi-stage image processing. Image and Vision Computing 2023, 133, 104657. [CrossRef]
- Gao, G.; Gao, J.; Liu, Q.; Wang, Q.; Wang, Y. A survey of deep learning methods for density estimation and crowd counting. Vicinagearth 2025, 2, 1–37.
- Le, H.D.; Nguyen, Q.V.; Huy, N.T.V.; Jo, J. Knowledge-Assisted Small Object Detection. In Proceedings of the Databases Theory and Applications; Chen, T.; Cao, Y.; Nguyen, Q.V.H.; Nguyen, T.T., Eds., Singapore, 2025; pp. 403–418.
- Learned-Miller, E.; Huang, G.B.; RoyChowdhury, A.; Li, H.; Hua, G., Labeled Faces in the Wild: A Survey. In Advances in Face Detection and Facial Image Analysis; Kawulok, M.; Celebi, M.E.; Smolka, B., Eds.; Springer International Publishing: Cham, 2016; pp. 189–248. [CrossRef]
- Prihasto, B.; Choirunnisa, S.; Nurdiansyah, M.I.; Mathulaprangsan, S.; Chu, V.C.M.; Chen, S.H.; Wang, J.C. A survey of deep face recognition in the wild. In Proceedings of the 2016 International Conference on Orange Technologies (ICOT), 2016, pp. 76–79. [CrossRef]
- Alashbi, A.A.S.; Sunar, M.S. Occluded Face Detection, Face in Niqab Dataset. In Proceedings of the Emerging Trends in Intelligent Computing and Informatics; Saeed, F.; Mohammed, F.; Gazem, N., Eds., Cham, 2020; pp. 209–215.
- Alashbi, A.; Mohamed, A.H.H.; El-Saleh, A.A.; Shayea, I.; Sunar, M.S.; Alqahtani, Z.R.; Saeed, F.; Saoud, B. Human face localization and detection in highly occluded unconstrained environments. Engineering Science and Technology, an International Journal 2025, 61, 101893. [CrossRef]
- Alamarshadi, M.S.; Sunar, M.S.; Mandala, S.; Alashbi, A.; Alqathani, Z. Deep Learning Approaches for Facial Landmark Localization in Niqab-Occluded Face Recognition: A Survey. In Proceedings of the Emerging Science and Technology for Human Well-Being; Saidin, S.; Sunar, M.S.; Hau, Y.W.; Lee Ming, E.S.; Mualif, S.A.; Juhari, F.H.; Ibrahim, F., Eds., Cham, 2025; pp. 257–266.
- Jiménez-Sánchez, A.; Avlona, N.R.; de Boer, S.; Campello, V.M.; Feragen, A.; Ferrante, E.; Ganz, M.; Gichoya, J.W.; Gonzalez, C.; Groefsema, S.; et al. In the Picture: Medical Imaging Datasets, Artifacts, and their Living Review. In Proceedings of the Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, 2025; FAccT ’25, p. 511–531. [CrossRef]
- Li, J.; Zhu, G.; Hua, C.; Feng, M.; Bennamoun, B.; Li, P.; Lu, X.; Song, J.; Shen, P.; Xu, X.; et al. A Systematic Collection of Medical Image Datasets for Deep Learning. ACM Comput. Surv. 2023, 56. [CrossRef]
- Alabduljabbar, A.; Khan, S.; Alsuhaibani, A.; Almarshad, F.; Altherwy, Y. Medical imaging datasets, preparation, and availability for artificial intelligence in medical imaging. Journal of Alzheimer’s Disease Reports 2024, 8, 1471–1483. [CrossRef]
- Goel, N.; Yadav, A.; Singh, B.M. Medical image processing: A review. In Proceedings of the 2016 Second International Innovative Applications of Computational Intelligence on Power, Energy and Controls with their Impact on Humanity (CIPECH), 2016, pp. 57–62. [CrossRef]
- Yan, F.; Huang, H.; Pedrycz, W.; Hirota, K. Review of medical image processing using quantum-enabled algorithms. Artificial Intelligence Review 2024, 57, 300.
- Clark, K.; Vendt, B.; Smith, K.; Freymann, J.; Kirby, J.; Koppel, P.; Moore, S.; Phillips, S.; Maffitt, D.; Pringle, M.; et al. The Cancer Imaging Archive (TCIA): Maintaining and Operating a Public Information Repository. Journal of digital imaging 2013, 26. [CrossRef]
- Han, E.; Kwon, H.; Jung, I. A review on multi-omics integration for aiding study design of large scale TCGA cancer datasets. BMC Genomics 2025, 26. [CrossRef]
- Diana Albelda, C.; Garcia-Martin, A.; Bescós, J. A Review on Deep Learning Methods for Glioma Segmentation, Limitations, and Future Perspectives. Journal of Imaging 2025, 11, 269. [CrossRef]
- Aisen, P.S.; Donohue, M.C.; Raman, R.; Rafii, M.S.; Petersen, R.C.; for the Alzheimer’s Disease Neuroimaging Initiative. The Alzheimer’s Disease Neuroimaging Initiative Clinical Core. Alzheimer’s & Dementia 2024, 20, 7361–7368, [https://alz-journals.onlinelibrary.wiley.com/doi/pdf/10.1002/alz.14167]. [CrossRef]
- Weiner, M.W.; Veitch, D.P.; Aisen, P.S.; Beckett, L.A.; Cairns, N.J.; Green, R.C.; Harvey, D.; Jack, C.R.; Jagust, W.; Liu, E.; et al. The Alzheimer’s Disease Neuroimaging Initiative: A review of papers published since its inception. Alzheimer’s & Dementia 2013, 9, e111–e194, [https://alz-journals.onlinelibrary.wiley.com/doi/pdf/10.1016/j.jalz.2013.05.1769]. [CrossRef]
- Chu, N.N.; Gebre-Amlak, H. Navigating Neuroimaging Datasets ADNI for Alzheimer’s Disease. IEEE Consumer Electronics Magazine 2021, 10, 61–63. [CrossRef]
- Veitch, D.P.; Weiner, M.W.; Aisen, P.S.; Beckett, L.A.; DeCarli, C.; Green, R.C.; Harvey, D.; Jack Jr., C.R.; Jagust, W.; Landau, S.M.; et al. Using the Alzheimer’s Disease Neuroimaging Initiative to improve early detection, diagnosis, and treatment of Alzheimer’s disease. Alzheimer’s & Dementia 2022, 18, 824–857, [https://alz-journals.onlinelibrary.wiley.com/doi/pdf/10.1002/alz.12422]. [CrossRef]
- Dutt, R.K.; Hannon, K.; Easley, T.O.; Griffis, J.C.; Zhang, W.; Bijsterbosch, J.D. Mental health in the UK Biobank: A roadmap to self-report measures and neuroimaging correlates. Human Brain Mapping 2022, 43, 816–832, [https://onlinelibrary.wiley.com/doi/pdf/10.1002/hbm.25690]. [CrossRef]
- Schulz, M.A.; Yeo, B.; Vogelstein, J.; Mourao-Miranada, J.; Kather, J.; Kording, K.; Richards, B.; Bzdok, D. Different scaling of linear models and deep learning in UKBiobank brain images versus machine-learning datasets. Nature Communications 2020, 11, 4238. [CrossRef]
- Jiang, Y.; Zhao, B.; Wang, X.; Tang, B.; Peng, H.; Luo, Z.; Shen, Y.; Wang, Z.; Jiang, Z.; Wang, J.; et al. UKB-MDRMF: a multi-disease risk and multimorbidity framework based on UK biobank data. Nature Communications 2025, 16. [CrossRef]
- Littlejohns, T.; Holliday, J.; Gibson, L.; Garratt, S.; Oesingmann, N.; Alfaro-Almagro, F.; Bell, J.; Boultwood, C.; Collins, R.; Conroy, M.; et al. The UK Biobank imaging enhancement of 100,000 participants: rationale, data collection, management and future directions. Nature Communications 2020, 11. [CrossRef]
- Wang, M.; Wang, Z.; Wang, Y.; Zhou, Q.; Wang, J. Causal relationships involving brain imaging-derived phenotypes based on UKB imaging cohort: a review of Mendelian randomization studies. Frontiers in Neuroscience 2024, Volume 18 - 2024. [CrossRef]
- Salami, F.; Bozorgi-Amiri, A.; Hassan, G.M.; Tavakkoli-Moghaddam, R.; Datta, A. Designing a clinical decision support system for Alzheimer’s diagnosis on OASIS-3 data set. Biomedical Signal Processing and Control 2022, 74, 103527. [CrossRef]
- Marcus, D.S.; Wang, T.H.; Parker, J.; Csernansky, J.G.; Morris, J.C.; Buckner, R.L. Open Access Series of Imaging Studies (OASIS): Cross-sectional MRI Data in Young, Middle Aged, Nondemented, and Demented Older Adults. Journal of Cognitive Neuroscience 2007, 19, 1498–1507, [https://direct.mit.edu/jocn/article-pdf/19/9/1498/1936514/jocn.2007.19.9.1498.pdf]. [CrossRef]
- Marcus, D.S.; Fotenos, A.F.; Csernansky, J.G.; Morris, J.C.; Buckner, R.L. Open Access Series of Imaging Studies: Longitudinal MRI Data in Nondemented and Demented Older Adults. Journal of Cognitive Neuroscience 2010, 22, 2677–2684, [https://direct.mit.edu/jocn/article-pdf/22/12/2677/1940172/jocn.2009.21407.pdf]. [CrossRef]
- Koenig, L.N.; Day, G.S.; Salter, A.; Keefe, S.; Marple, L.M.; Long, J.; LaMontagne, P.; Massoumzadeh, P.; Snider, B.J.; Kanthamneni, M.; et al. Select Atrophied Regions in Alzheimer disease (SARA): An improved volumetric model for identifying Alzheimer disease dementia. NeuroImage: Clinical 2020, 26, 102248. [CrossRef]
- Rogers, P.; Wang, D.; Lu, Z. Medical Information Mart for Intensive Care: A Foundation for the Fusion of Artificial Intelligence and Real-World Data. Frontiers in Artificial Intelligence 2021, Volume 4 - 2021. [CrossRef]
- Johnson, A.E.W.; Pollard, T.J.; Greenbaum, N.R.; Lungren, M.P.; ying Deng, C.; Peng, Y.; Lu, Z.; Mark, R.G.; Berkowitz, S.J.; Horng, S. MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs, 2019, [arXiv:cs.CV/1901.07042].
- Bagalkot, L.; Kelapati, K. Advancements and challenges in deep learning techniques for lung disease diagnosis. Indonesian Journal of Electrical Engineering and Computer Science 2025, 39, 1053. [CrossRef]
- Johnson, A.; Pollard, T.; Berkowitz, S.; Greenbaum, N.; Lungren, M.; Deng, C.y.; Mark, R.; Horng, S. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 2019, 6, 317. [CrossRef]
- Cassidy, B.; Kendrick, C.; Brodzicki, A.; Jaworek-Korjakowska, J.; Yap, M.H. Analysis of the ISIC image datasets: Usage, benchmarks and recommendations. Medical Image Analysis 2022, 75, 102305. [CrossRef]
- DiSanto, N. ISIC Melanoma Dataset, 2023. [CrossRef]
- Kurtansky, N.; D’Alessandro, B.; Gillis, M.; Betz-Stablein, B.; Cerminara, S.; Garcia, R.; Girundi, M.; Gössinger, E.; Gottfrois, P.; Guitera, P.; et al. The SLICE-3D dataset: 400,000 skin lesion image crops extracted from 3D TBP for skin cancer detection. Scientific Data 2024, 11. [CrossRef]
- Jain, S.; jagtap, V.; Pise, N. Computer Aided Melanoma Skin Cancer Detection Using Image Processing. Procedia Computer Science 2015, 48, 735–740. International Conference on Computer, Communication and Convergence (ICCC 2015), . [CrossRef]
- Vibha, T.; Saravanan, C.; Divya, T. Melanoma Skin Cancer Detection and Classification Using Deep Learning and Image Processing. SN Computer Science 2025, 6. [CrossRef]
- Moura, D.; Guevara Lopez, M.A.; Cunha, P.; Posada, N.; Pollán, R.; Ramos, I.; Loureiro, J.; Moreira, I.; Araújo, B.; Fernandes, T. Benchmarking Datasets for Breast Cancer Computer-Aided Diagnosis (CADx), 2013. [CrossRef]
- Oza, P.; Sharma, P.; Patel, S.; Kumar, P. Computer-Aided Breast Cancer Diagnosis: A Study of Breast Imaging Modalities and Mammogram Repositories. Current Medical Imaging Formerly Current Medical Imaging Reviews 2022, 18. [CrossRef]
- Lawal, M.K.B.; Almousa, M.; Ibrahim, A.U.; Pwavodi, P.C.; Usman, A.G.; Aloraini, B. Artificial intelligent-powered detection of breast cancer. Journal of Radiation Research and Applied Sciences 2025, 18, 101422. [CrossRef]
- Antonelli, M.; Reinke, A.; Bakas, S.; Farahani, K.; Kopp-Schneider, A.; Landman, B.; Litjens, G.; Menze, B.; Ronneberger, O.; Summers, R.; et al. The Medical Segmentation Decathlon. Nature Communications 2022, 13, 4128. [CrossRef]
- Altini, N.; Prencipe, B.; Cascarano, G.D.; Brunetti, A.; Brunetti, G.; Triggiani, V.; Carnimeo, L.; Marino, F.; Guerriero, A.; Villani, L.; et al. Liver, kidney and spleen segmentation from CT scans and MRI with deep learning: A survey. Neurocomputing 2022, 490, 30–53. [CrossRef]
- Kuş, Z.; Aydin, M. MedSegBench: A comprehensive benchmark for medical image segmentation in diverse data modalities. Scientific Data 2024, 11. [CrossRef]
- Pachade, S.; Porwal, P.; Thulkar, D.; Kokare, M.; Deshmukh, G.; Sahasrabuddhe, V.; Giancardo, L.; Quellec, G.; Mériaudeau, F. Retinal Fundus Multi-disease Image Dataset (RFMiD), 2020. [CrossRef]
- Pachade, S.; Porwal, P.; Thulkar, D.; Kokare, M.; Deshmukh, G.; Sahasrabuddhe, V.; Giancardo, L.; Quellec, G.; Mériaudeau, F. Retinal Fundus Multi-Disease Image Dataset (RFMiD): A Dataset for Multi-Disease Detection Research. Data 2021, 6. [CrossRef]
- Wu, C.; Restrepo, D.; Nakayama, L.; Ribeiro, L.; Shuai, Z.; Barboza, N.; Sousa, M.; Fitterman, R.; Pereira, A.; Regatieri, C.; et al. A portable retina fundus photos dataset for clinical, demographic, and diabetic retinopathy prediction. Scientific Data 2025, 12. [CrossRef]
- Wang, R.; Jiao, L.; Liu, K. Large-Scale Agricultural Pest and Disease Datasets, 2023. [CrossRef]
- Yuan, Y.; Chen, L.; Wu, H.; Li, L. Advanced agricultural disease image recognition technologies: A review. Information Processing in Agriculture 2022, 9, 48–59. [CrossRef]
- Luo, Z.; Yang, W.; Yuan, Y.; Gou, R.; Li, X. Semantic segmentation of agricultural images: A survey. Information Processing in Agriculture 2024, 11, 172–186. [CrossRef]
- Lei, L.; Yang, Q.; Yang, L.; Shen, T.; Wang, R.; Fu, C. Deep learning implementation of image segmentation in agricultural applications: a comprehensive review. Artificial Intelligence Review 2024, 57. [CrossRef]
- Tchokogoué, T.; Noumsi, A.V.; Atemkeng, M.; Fono, L.A. Towards precision agriculture: A dataset for early detection of corn leaf pests. Data in Brief 2025, 59, 111394. [CrossRef]
- Biswas, J.; Ul Islam Rafim, A.R.; Paul, T.; Hossain, M.M.; Hasan, S.N.; Rashid, M.R.A.; Sathi, T.A. Automatic Detection of Pumpkin Leaf Diseases Using Transfer Learning and a Custom Dataset from Bangladesh. In Proceedings of the 2024 IEEE International Conference on Computing, Applications and Systems (COMPAS), 2024, pp. 1–6. [CrossRef]
- Rashid, M.R.A.; Biswas, J.; Hossain, M.M. Pumpkin Leaf Diseases Dataset From Bangladesh, 2024. Dataset, . [CrossRef]
- Genze, N.; Vahl, W.; Groth, J.; Wirth, M.; Grieb, M.; Grimm, D. Manually annotated and curated Dataset of diverse Weed Species in Maize and Sorghum for Computer Vision. Scientific Data 2024, 11. [CrossRef]
- Wang, R. Aphid Cluster Segmentation Dataset, 2023. [CrossRef]
- Rahman, R.; Indris, C.; Bramesfeld, G.; Zhang, T.; Li, K.; Chen, X.; Grijalva, I.; McCornack, B.; Flippo, D.; Sharda, A.; et al. A New Dataset and Comparative Study for Aphid Cluster Detection and Segmentation in Sorghum Fields, 2024, [arXiv:cs.CV/2405.04305].
- Weyler, J.; Magistri, F.; Marks, E.; Chong, Y.L.; Sodano, M.; Roggiolani, G.; Chebrolu, N.; Stachniss, C.; Behley, J. PhenoBench: A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain. IEEE Transactions on Pattern Analysis and Machine Intelligence 2024, 46, 9583–9594. [CrossRef]
- Ubaid, M.T.; Javaid, S. Precision Agriculture: Computer Vision-Enabled Sugarcane Plant Counting in the Tillering Phase. Journal of Imaging 2024, 10. [CrossRef]
- Li, J.; Magar, R.T.; Chen, D.; Lin, F.; Wang, D.; Yin, X.; Zhuang, W.; Li, Z. SoybeanNet: Transformer-based convolutional neural network for soybean pod counting from Unmanned Aerial Vehicle (UAV) images. Computers and Electronics in Agriculture 2024, 220, 108861. [CrossRef]
- Rahman, A.; Lu, Y.; Wang, H. Performance evaluation of deep learning object detectors for weed detection for cotton. Smart Agricultural Technology 2023, 3, 100126. [CrossRef]
- Steininger, D.; Trondl, A.; Croonen, G.; Simon, J.; Widhalm, V. The CropAndWeed Dataset: a Multi-Modal Learning Approach for Efficient Crop and Weed Manipulation. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 3718–3727. [CrossRef]
- Rai, N.; Mahecha, M.V.; Christensen, A.; Quanbeck, J.; Zhang, Y.; Howatt, K.; Ostlie, M.; Sun, X. Multi-format open-source weed image dataset for real-time weed identification in precision agriculture. Data in Brief 2023, 51, 109691. [CrossRef]
- Yordanov, M.; d’Andrimont, R.; Martinez-Sanchez, L.; Lemoine, G.; Fasbender, D.; van der Velde, M. Crop identification using deep learning on LUCAS crop cover photos, 2023, [arXiv:cs.CV/2305.04994].
- Lourdu Antony, L.P. Rice Leaf Diseases Dataset, 2023. Dataset, . [CrossRef]
- Rashid, M.R.A.; Hossain, M.S.; Fahim, M.; Islam, M.S.; Tahzib-E-Alindo.; Prito, R.H.; Sheikh, M.S.A.; Ali, M.S.; Hasan, M.; Islam, M. Comprehensive dataset of annotated rice panicle image from Bangladesh. Data in Brief 2023, 51, 109772. [CrossRef]
- Güldenring, R.; Andersen, R.E.; Nalpantidis, L. Zoom in on the Plant: Fine-grained Analysis of Leaf, Stem and Vein Instances, 2023, [arXiv:cs.RO/2312.08805].
- Rajput, A.S.; Shukla, S.; Thakur, S. SoyNet: A high-resolution Indian soybean image dataset for leaf disease classification. Data in Brief 2023, 49, 109447. [CrossRef]
- Moazzam, I. Tobacco Aerial Dataset, 2023. Dataset, . [CrossRef]
- Kitzler, F.; Barta, N.; Neugschwandtner, R.W.; Gronauer, A.; Motsch, V. WE3DS: An RGB-D Image Dataset for Semantic Segmentation in Agriculture. Sensors 2023, 23. [CrossRef]
- Dang, F.; Chen, D.; Lu, Y.; Li, Z. YOLOWeeds: A novel benchmark of YOLO object detectors for multi-class weed detection in cotton production systems. Computers and Electronics in Agriculture 2023, 205, 107655. [CrossRef]
- Chen, D.; Lu, Y.; Li, Z.; Young, S. Performance evaluation of deep transfer learning on multi-class identification of common weed species in cotton production systems. Computers and Electronics in Agriculture 2022, 198, 107091. [CrossRef]
- Bevers, N.; Sikora, E.J.; Hardy, N.B. Soybean disease identification using original field images and transfer learning with convolutional neural networks. Computers and Electronics in Agriculture 2022, 203, 107449. [CrossRef]
- Teimouri, N.; Jørgensen, R.N.; Green, O. Novel Assessment of Region-Based CNNs for Detecting Monocot/Dicot Weeds in Dense Field Environments. Agronomy 2022, 12. [CrossRef]
- David, E.; Madec, S.; Sadeghi-Tehran, P.; Aasen, H.; Zheng, B.; Liu, S.; Kirchgessner, N.; Ishikawa, G.; Nagasawa, K.; Badhon, M.A.; et al. Global Wheat Head Detection (GWHD) Dataset: A Large and Diverse Dataset of High-Resolution RGB-Labelled Images to Develop and Benchmark Wheat Head Detection Methods. Plant Phenomics 2020, 2020, [https://spj.science.org/doi/pdf/10.34133/2020/3521852]. [CrossRef]
- Global Wheat Head Detection 2021: An Improved Dataset for Benchmarking Wheat Head Detection Methods, 2021. [CrossRef]
- Riehle, D.; Reiser, D.; Griepentrog, H.W. Robust index-based semantic plant/background segmentation for RGB- images. Computers and Electronics in Agriculture 2020, 169, 105201. [CrossRef]
- Espejo-Garcia, B.; Mylonas, N.; Athanasakos, L.; Fountas, S.; Vasilakoglou, I. Towards weeds identification assistance through transfer learning. Computers and Electronics in Agriculture 2020, 171, 105306. [CrossRef]
- Leminen Madsen, S.; Mathiassen, S.K.; Dyrmann, M.; Laursen, M.S.; Paz, L.C.; Jørgensen, R.N. Open Plant Phenotype Database of Common Weeds in Denmark. Remote Sensing 2020, 12. [CrossRef]
- Olsen, A.; Konovalov, D.; Philippa, B.; Ridd, P.; Wood, J.; Johns, J.; Banks, W.; Girgenti, B.; Kenny, O.; Whinney, J.; et al. DeepWeeds: A Multiclass Weed Species Image Dataset for Deep Learning. Scientific Reports 2019, 9. [CrossRef]
- Jiang, Y.; Li, C.; Paterson, A.; Robertson, J. DeepSeedling: deep convolutional network and Kalman filter for plant seedling detection and counting in the field. Plant Methods 2019, 15. [CrossRef]
- Teimouri, N.; Dyrmann, M.; Nielsen, P.R.; Mathiassen, S.K.; Somerville, G.J.; Jørgensen, R.N. Weed Growth Stage Estimator Using Deep Convolutional Neural Networks. Sensors 2018, 18. [CrossRef]
- Wiesner-Hanks, T.; Stewart, E.; Kaczmar, N.; DeChant, C.; Wu, H.; Nelson, R.; Lipson, H.; Gore, M. Image set for deep learning: Field images of maize annotated with disease symptoms. BMC Research Notes 2018, 11. [CrossRef]
- Giselsson, T.M.; Jørgensen, R.N.; Jensen, P.K.; Dyrmann, M.; Midtiby, H.S. A Public Image Database for Benchmark of Plant Seedling Classification Algorithms, 2017, [arXiv:cs.CV/1711.05458].
- Prabavathy, K.; Bharath, M.; Sanjayratnam, K.; Reddy, N.S.S.R.; Reddy, M.S. Plant Leaf Disease Detection using Machine Learning. In Proceedings of the 2023 2nd International Conference on Applied Artificial Intelligence and Computing (ICAAIC), 2023, pp. 378–382. [CrossRef]
- Lu, Y.; Young, S. A survey of public datasets for computer vision tasks in precision agriculture. Computers and Electronics in Agriculture 2020, 178, 105760. [CrossRef]
- Padilla, R.; Netto, S.L.; da Silva, E.A.B. A Survey on Performance Metrics for Object-Detection Algorithms. In Proceedings of the 2020 International Conference on Systems, Signals and Image Processing (IWSSIP), 2020, pp. 237–242. [CrossRef]
- Zhao, R.; Wang, K.; Xiao, Y.; Gao, F.; Gao, Z. Leveraging Monte Carlo Dropout for Uncertainty Quantification in Real-Time Object Detection of Autonomous Vehicles. IEEE Access 2024, 12, 33384–33399. [CrossRef]
- Zhan, J.; Luo, Y.; Guo, C.; Wu, Y.; Meng, J.; Liu, J. YOLOPX: Anchor-free multi-task learning network for panoptic driving perception. Pattern Recognition 2024, 148, 110152. [CrossRef]
- Hong, X.; Huang, J.; Zhao, W.; Zou, H.; Lin, Z.; Chen, Y. Object detection for traffic management based on YOLO. In Proceedings of the International Conference on Smart Transportation and City Engineering (STCE 2023). SPIE, 2024, Vol. 13018, pp. 157–161.
- S, K.; P, K.; K, K.; S, R. Traffic Management through Cutting-Edge Vehicle Detection, Recognition, and Tracking Innovations. Procedia Computer Science 2024, 233, 793–800. 5th International Conference on Innovative Data Communication Technologies and Application (ICIDCA 2024), . [CrossRef]
- Hasan, S.; Sunny, M.N.M.; Al Nahian, A.; Yasin, M. Neural Network-Powered License Plate Recognition System Design. Engineering 2024, 16, 284–300.
- Ghahremannezhad, H.; Shi, H.; Liu, C. Object Detection in Traffic Videos: A Survey. IEEE Transactions on Intelligent Transportation Systems 2023, 24, 6780–6799. [CrossRef]
- Rahaman, M.F.; Li, X.; Zakaria, K.M.; Al, A.M.; Buse, K.; Bashar, S.A. Enhancing Multi-Robot Formation Control and Navigation Using Virtual Structures and Improved Path Planning Algorithms. In Proceedings of the 2024 4th International Conference on Innovative Research in Applied Science, Engineering and Technology (IRASET), 2024, pp. 1–8. [CrossRef]
- Amjad, M.; Sahin Ali, M.; Yao, S.; Faishal Rahaman, M.; Zheng, C.; Muhammad Kazim, R.; Zouaoui, B. Self and Target Locating With Cooperation of Heterogeneous Unmanned Vehicles in the Denial Environment. IEEE Access 2025, 13, 64699–64718. [CrossRef]
- Saranya, T.; Deisy, C.; Sridevi, S.; Anbananthen, K.S.M. A comparative study of deep learning and Internet of Things for precision agriculture. Engineering Applications of Artificial Intelligence 2023, 122, 106034.
- Khalid, S.; Oqaibi, H.M.; Aqib, M.; Hafeez, Y. Small pests detection in field crops using deep learning object detection. Sustainability 2023, 15, 6815.
- Attri, I.; Awasthi, L.K.; Sharma, T.P.; Rathee, P. A review of deep learning techniques used in agriculture. Ecological Informatics 2023, p. 102217.
- Albahar, M. A survey on deep learning and its impact on agriculture: Challenges and opportunities. Agriculture 2023, 13, 540.
- Zhuang, X.; Li, D.; Wang, Y.; Li, K. Military target detection method based on EfficientDet and Generative Adversarial Network. Engineering Applications of Artificial Intelligence 2024, 132, 107896.
- Evangelista, M. Innovation and the arms race: How the United States and the Soviet Union develop new military technologies; Cornell University Press, 2023.
- Skyruta, V.; Yurchuk, I. Military vehicles marking detection algorithm on a digital image. In Proceedings of the 2023 IEEE 4th KhPI Week on Advanced Technology (KhPIWeek). IEEE, 2023, pp. 1–6.
- Rane, N. YOLO and Faster R-CNN object detection for smart Industry 4.0 and Industry 5.0: applications, challenges, and opportunities. Available at SSRN 4624206 2023.
- Zhang, Z.; Zhou, M.; Wan, H.; Li, M.; Li, G.; Han, D. IDD-Net: industrial defect detection method based on deep-learning. Engineering Applications of Artificial Intelligence 2023, 123, 106390.
- Pham, D.L.; Chang, T.W.; et al. A YOLO-based real-time packaging defect detection system. Procedia Computer Science 2023, 217, 886–894.
- Saberironaghi, A.; Ren, J.; El-Gindy, M. Defect detection methods for industrial products using deep learning techniques: A review. Algorithms 2023, 16, 95.
- Gao, J.; Yang, Y.; Lin, P.; Park, D.S. Computer vision in healthcare applications. Journal of healthcare engineering 2018, 2018.
- Khang, A.; Abdullayev, V.; Litvinova, E.; Chumachenko, S.; Alyar, A.V.; Anh, P. Application of Computer Vision (CV) in the Healthcare Ecosystem. In Computer Vision and AI-Integrated IoT Technologies in the Medical Ecosystem; CRC Press, 2024; pp. 1–16.
- Javaid, M.; Haleem, A.; Singh, R.P.; Ahmed, M. Computer Vision to Enhance Healthcare Domain: An Overview of Features, Implementation, and Opportunities. Intelligent Pharmacy 2024.
- Ragab, M.G.; Abdulkader, S.J.; Muneer, A.; Alqushaibi, A.; Sumiea, E.H.; Qureshi, R.; Al-Selwi, S.M.; Alhussian, H. A Comprehensive Systematic Review of YOLO for Medical Object Detection (2018 to 2023). IEEE Access 2024.
- Pinto-Coelho, L. How artificial intelligence is shaping medical imaging technology: A survey of innovations and applications. Bioengineering 2023, 10, 1435.
- O’ Donoghue, S.; Goodsell, D.; Frangakis, A.; Jossinet, F.; Laskowski, R.; Nilges, M.; Saibil, H.; Schafferhans, A.; Wade, R.; Westhof, E.; et al. Visualization of Macromolecular Structures. Nature methods 2010, 7, S42–55. [CrossRef]
- Loewe, A.; Hunter, P.J.; Kohl, P. Computational modelling of biological systems now and then: revisiting tools and visions from the beginning of the century. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 2025, 383, 20230384, [https://royalsocietypublishing.org/doi/pdf/10.1098/rsta.2023.0384]. [CrossRef]
- Kozlikova, B.; Krone, M.; Lindow, N.; Falk, M.; Baaden, M.; Baum, D.; Viola, I.; Parulek, J.; Hege, H.C. Visualization of Biomolecular Structures: State of the Art. In Proceedings of the Eurographics Conference on Visualization (EuroVis) - STARs; Borgo, R.; Ganovelli, F.; Viola, I., Eds. The Eurographics Association, 2015, pp. 061–081.
- Medhat, B.; Shawish, A. A Novel Computer Vision Methodology for Intelligent Molecular Modeling and Simulation. In Proceedings of the Proceedings of the 11th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC 2018) - Volume 3: BIOINFORMATICS. INSTICC, SciTePress, 2018, pp. 97–104. [CrossRef]
- Sharma, H.; Kumar, H.; Mangla, S.K. Enablers to computer vision technology for sustainable E-waste management. Journal of Cleaner Production 2023, 412, 137396.
- González-Sabbagh, S.P.; Robles-Kelly, A. A survey on underwater computer vision. ACM Computing Surveys 2023, 55, 1–39.
- Ashley, M.I.K.D.; Chan, T.W. Intelligent tutoring systems; Springer, 1982.
- Dimitriadou, E.; Lanitis, A. A critical evaluation, challenges, and future perspectives of using artificial intelligence and emerging technologies in smart classrooms. Smart Learning Environments 2023, 10, 12.
- Alam, A. Harnessing the Power of AI to Create Intelligent Tutoring Systems for Enhanced Classroom Experience and Improved Learning Outcomes. In Intelligent Communication Technologies and Virtual Mobile Networks; Springer, 2023; pp. 571–591.
- Kaur, J.; Singh, W. A systematic review of object detection from images using deep learning. Multimedia Tools and Applications 2024, 83, 12253–12338.
- Amit, Y.; Felzenszwalb, P.; Girshick, R. Object detection. In Computer Vision: A Reference Guide; Springer, 2021; pp. 875–883.
- Zhou, Z.; Hirota, K.; Dai, Y.; Islam, M.S.; Mersha, B.W.; Dai, W.; Lin, Y. Fuzzy-Based Head Attitude Estimation for Improved Students’ Concentration Evaluation. In Proceedings of the Computational Intelligence and Industrial Applications; Xin, B.; Ma, H.; She, J.; Cao, W., Eds., Singapore, 2025; pp. 14–28.
- Chen, W.; Luo, J.; Zhang, F.; Tian, Z. A review of object detection: Datasets, performance evaluation, architecture, applications and current trends. Multimedia Tools and Applications 2024, 83, 65603–65661.
Figure 1.
Applications of Object Detection in different fields and domains.

Figure 2.
Classification of object detection algorithms.

Figure 3.
A road map of object detection and the prominent object detection algorithms.

Figure 4.
PRISMA(2020) flow diagram of the systematic review.

Figure 5.
Working principle of the traditional object detection approach.

Figure 7.
Working principle of Deep learning object detection.

Figure 8.
Architecture of the CNN in Object detection.

Figure 9.
Structure of a two-stage detector.

Figure 10.
Architecture of R-CNN Object detection [86].
Figure 10.
Architecture of R-CNN Object detection [86].

Figure 11.
Structure of Single Stage detector.

Figure 12.
The road map of the YOLO object detector family.

Figure 13.
The architecture of the YOLO object detector [105].
Figure 13.
The architecture of the YOLO object detector [105].

Figure 14.
The basic working process of YOLO object detector [105].
Figure 14.
The basic working process of YOLO object detector [105].

Figure 15.
Working of the SSD Object detector [134].
Figure 15.
Working of the SSD Object detector [134].

Figure 16.
Simplified architectural block diagram of a Transformer-based object detector (e.g., DETR), highlighting the sequential flow from feature extraction to direct set prediction.
Figure 16.
Simplified architectural block diagram of a Transformer-based object detector (e.g., DETR), highlighting the sequential flow from feature extraction to direct set prediction.

Figure 19.
Overview of the LLM-based object detection approach.

Figure 20.
Example remote sensing datasets captured by drones and satellites in different imaging conditions, weather, seasons, and image quality [244].
Figure 20.
Example remote sensing datasets captured by drones and satellites in different imaging conditions, weather, seasons, and image quality [244].

Figure 21.
Definition of facial attributes and representative facial images in MAFA [260].
Figure 21.
Definition of facial attributes and representative facial images in MAFA [260].

Figure 22.
Demonstration of intersection over union(IoU).

Figure 23.
Applications of Computer Vision in Structural Biological Research.

Table 2.
Comparison Between AlexNet, VGG-16, GoogLeNet, and ResNet-50 as presented in [2].
Table 2.
Comparison Between AlexNet, VGG-16, GoogLeNet, and ResNet-50 as presented in [2].
| Model | Year | Layers | Parameters (Millions) | Top Accuracy (%) |
|---|---|---|---|---|
| AlexNet | 2012 | 7 | 62.4 | 63.3 |
| VGG-16 | 2014 | 16 | 138.4 | 73 |
| GoogLeNet | 2014 | 22 | 6.7 | - |
| ResNet-50 | 2015 | 50 | 25.6 | 76 |
Table 4.
Comparison of YOLO Object Detection Algorithms
| Model | Year | Architecture | Key Improvements | Output Head | Anchor | Features |
|---|---|---|---|---|---|---|
| YOLOv1 | 2015 | GoogLeNet | Unified detection, real-time speed | Fully connected | No | Single CNN for detection |
| YOLOv2 | 2016 | Darknet-19 | BatchNorm, high-res classifier, anchor boxes | Convolutional | Yes | Dimension clustering, fine-grained features |
| YOLOv3 | 2018 | Darknet-53 | Residual blocks, multi-scale prediction | Convolutional | Yes | 3-scale prediction, better backbone |
| YOLOv4 | 2020 | CSPDarknet53 | Weighted residual connections, PANet | Convolutional | Yes | Mish activation, SPP, SAM, Mosaic augmentation |
| YOLOv5 | 2020 | Custom PyTorch | Light-weight, flexible deployment | Convolutional | Yes | Auto-learning bounding boxes, faster training |
| YOLOR | 2021 | Hybrid CNN | Implicit + explicit feature representation | Convolutional | Yes | Unified representation learning |
| YOLOX | 2021 | CSPDarknet + SimOTA | Decoupled head, anchor-free | Decoupled Head | No | Strong baseline, SimOTA label assignment |
| YOLOv6 | 2022 | EfficientRep | Optimized for edge devices | Decoupled Head | No | SimOTA, higher FPS |
| YOLOv7 | 2022 | E-ELAN | Extended efficient layer aggregation | Decoupled Head | Yes | Model scaling, multi-task support |
| YOLOv8 | 2023 | Custom (C2f) | Modular, unified task support | Decoupled Head | No | Instance segmentation, pose estimation |
| YOLOv9 | 2024 | GELAN+PGI | Programmable Gradient Information (PGI), Generalized ELAN (GELAN), information bottleneck | Decoupled Head | No | Sparse sampling, high accuracy |
| YOLOv10 | 2024 | Enhanced C2f | Re-parameterized ELAN | Decoupled Head | No | Multi-head attention, inference stability |
| YOLOv11 | 2024 | Improved C2f + Sparse Attention | Lightweight generalist model | Decoupled Head | No | Enhanced gradient flow, real-time optimization |
| YOLOv12 | 2025 | R-ELAN + Area Attention | Attention-centric design, Area Attention, R-ELAN, FlashAttention | Decoupled Head + NMS-free | No | Joint detection, segmentation, and pose estimation |
Table 6.
Comparison of YOLOE, YOLO-World, YOLOv8, and Florence-2 object detectors.
| Model | Architecture | Prompt Mechanism | Key Innovations | Datasets / Training | Performance Highlights | Advantages |
|---|---|---|---|---|---|---|
| YOLOv8 [121] | Convolutional YOLO framework with CSP-based backbone and PAN-FPN head | Closed-set (predefined categories only) | Anchor-free design, decoupled head, mosaic augmentation | COCO and related closed-set datasets | High accuracy and speed in closed-set detection | Real-time, efficient, lightweight, widely used baseline |
| YOLO-World [191] | YOLO + vision-language integration via RepVL-PAN | Text prompts (open vocabulary) | RepVL-PAN (vision-language fusion), region-text contrastive loss | Pre-trained with detection, grounding, and large image-text datasets | 35.4 AP @ 52 FPS on LVIS (V100) | Open-vocabulary, strong zero-shot, efficient real-time |
| Florence-2 [187] | Transformer-based unified vision foundation model | Text prompts (task instructions) | Sequence-to-sequence prompt-based multi-task learning | FLD-5B (126M images, 5.4B annotations) | Strong zero-shot detection, grounding, segmentation, captioning | Foundation-scale, versatile across CV and V&L tasks |
| YOLOE [194] | Unified detection + segmentation YOLO framework | Text, visual, and prompt-free | RepRTA (text alignment), SAVPE (visual prompts), LRPC (prompt-free vocab) | LVIS, COCO, large vocabulary datasets | +3.5 AP vs YOLO-Worldv2-S on LVIS, 1.4× faster, 3× less training cost | Multi-prompt support, real-time “see anything,” efficient training |
Table 8.
Comparison between single-stage, two-stage, Transformer-based, and the LLM-based object detectors.
Table 8.
Comparison between single-stage, two-stage, Transformer-based, and the LLM-based object detectors.
| Feature | Single-Stage Detectors | Two-Stage Detectors | Transformer-Based Detectors | LLM-Based Detectors |
|---|---|---|---|---|
| Architecture | Single CNN pass; Direct Regression to bounding boxes and classes. | Stage 1: Region Proposal Network (RPN); Stage 2: Classification & Refinement. | Encoder-Decoder Transformer; Uses Self-Attention to capture global context. | Vision Model (often Transformer) + Language Model; Uses text prompts for zero-shot detection. |
| Detection Speed | Fastest in Real-time. Inference is a single step. | Slowest (High latency). Two sequential steps are required. | Highly variable (initially slow, but modern variants are real-time). | Generally Slower than SSD/modern T-detectors due to language model overhead. |
| Accuracy (mAP) | High, but historically lower than TSDs. The gap has been significantly narrowed by recent YOLO variants. | Highest (Traditionally, the state-of-the-art for precision). Refinement boosts accuracy. | Very High (Competitive with TSDs). Excels with global context. | High in zero-shot/open-vocabulary tasks; Accuracy varies on specific prompts. |
| Complexity | Low to Moderate. Simpler pipeline (no RoI Pooling/NMS-free in new versions). | High. Complex multi-component pipeline (RPN, RoI Align, NMS). | High. Transformer architecture is parameter-heavy; Slow convergence in training. | Highest. Involves two complex, large models (Vision + Language). |
| Small Object Effectiveness | Moderate. Struggles due to dense sampling and limited resolution in the final detection layers. | High. The two-stage refinement process and use of multi-scale features (FPN) significantly boost small object accuracy. | Moderate to High. Early models struggled, but modern variants with multi-scale features and higher resolution handle small objects well. | High in Context. Can use language to infer the location of small, hard-to-see objects based on scene context. |
| Primary Use Case | Real-time systems, video surveillance, autonomous driving (for speed). | Offline analysis, medical imaging, quality inspection (where accuracy is paramount). | General-purpose OD, scenarios needing a strong global context (e.g., crowded scenes). | Open-Vocabulary detection, zero-shot learning, language-guided search. |
Table 9.
Comparison of Generic Object Datasets.
| Dataset | Classes | Images | Applications | Advantages | Disadvantages | Contributions |
|---|---|---|---|---|---|---|
| Caltech-101 (2003) [211] | 102 | 9144 | Generic Object Recognition | Early dataset moving beyond iconic views to natural environments | Limited to one object per image and Not designed for modern detection tasks | Early efforts to include object categories in natural environments, primarily for classification |
| PASCAL VOC (2005) [45] | 20 | 22591 | Generic Object Detection | Relatively small and convenient for benchmarking, and established standard evaluation metrics (mAP) | Small size by modern standards, Limited to 20 classes, superseded by larger datasets | The foundational dataset that set the benchmark for object detection and evaluation metrics. |
| Caltech-256 (2006) [212] | 257 | 30307 | Generic Object Recognition | Early dataset expanding object categories in natural environments | Limited to one object per image and not designed for modern detection tasks | Extended Caltech-101 with more categories and improved dataset construction |
| MS COCO (2014) [47] | 91 | 328000 | Generic Object Detection | High object density in natural, contextual environments with Rich annotations. | Basic object categories lack fine granularity, and the annotation process is complex and expensive | Established a benchmark emphasizing context and segmentation beyond simple bounding boxes |
| Visual Genome (2016) [213] | 76340 | 108000 | Scene Understanding | Dense annotations of objects, attributes, relationships, and QA pairs with a massive vocabulary of object classes | Annotations may be noisy/inconsistent due to crowdsourcing, where not all images are fully annotated | Introduced dense scene graphs linking visual data with structured semantics |
| Open Images (2017) [46] | 600 | 9178275 | Large-Scale Object Detection | -Extremely large scale (9M+ images), including relationship annotations that provide negative labels | Initially lacked segmentation masks (added later), where severe object classes are imbalanced due to web sourcing | One of the largest datasets, emphasizing real-world diversity and relationship annotations |
Table 13.
Comprehensive comparative analysis of medical imaging datasets.
| Dataset | Modality | Primary Focus | Scale | Annotations | Key Features / Advantages | Primary Limitation |
|---|---|---|---|---|---|---|
| TCIA [274] | CT, MRI, PET | Oncology | 54+ organs; Numerous collections | Tumor segmentation, clinical data | Public, diverse cancer types, linked clinical data | Dataset-specific; varying annotation depth |
| ADNI [277] | MRI, PET | Neurology (Alzheimer’s) | 3000+ subjects | Longitudinal segmentation, diagnosis | Longitudinal, multi-modal, well-established | Focused primarily on Alzheimer’s |
| UK Biobank [281] | MRI, CT, X-ray | Population Health | >100,000 participants | Organ segmentation, genetic, EHR | Unprecedented scale, genetic + imaging + EHR data | Access requires approval process |
| OASIS [287] | MRI | Neurology (Aging, Alzheimer’s) | 1500+ subjects | Brain segmentation, clinical diagnosis | Curated for brain studies, publicly accessible | Smaller scale than UK Biobank |
| MIMIC-CXR [290,291] | X-ray (Chest) | Pulmonary | 377,110 images | Free-text reports, labels from NLP | Large-scale, paired with EHR for NLP tasks | Labels automated from text, not pixel-wise |
| ISIC [294,295] | Dermoscopy | Dermatology | 70,000+ images | Lesion segmentation, diagnosis | High-quality, standard benchmark for skin cancer | Mostly 2D images of lesions |
| BCDR [299] | Mammography | Breast Oncology | Thousands of cases | Lesion type, density, BIRADS | Focused on mammography, film and digital | Smaller than some modern benchmarks |
| MSD [302] | CT, MRI | Multi-organ Segmentation | 10 challenges | Pixel-wise segmentation | Tests generalizability across 10 tasks | Designed as a challenge, not a unified db |
| RFMiD [306] | Fundus | Ophthalmology | 3,200 images | Unparalleled pathology diversity | Multi-source variability | small in demographic scale |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.