Submitted:
22 July 2026
Posted:
23 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Neural Networks
3. Convolution Neural Networks
4. Recurrent Neural Networks
5. Variational Autoencoder
6. Generative Adversarial Network
7. Transformers
8. Diffusion Models
9. Application Specific Architectures
10. Conclusion
Acknowledgments
References
- McCulloch, W.S.; Pitts, W. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics 1943, 5, 115–133.
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Advances in neural information processing systems 2020, 33, 1877–1901.
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. Llama: Open and efficient foundation language models. arXiv:2302.13971 2023.
- Esteva, A.; Robicquet, A.; Ramsundar, B.; Kuleshov, V.; DePristo, M.; Chou, K.; Cui, C.; Corrado, G.; Thrun, S.; Dean, J. A guide to deep learning in healthcare. Nature medicine 2019, 25, 24–29.
- Huang, J.; Chai, J.; Cho, S. Deep learning in finance and banking: A literature review and classification. Frontiers of Business Research in China 2020, 14, 13.
- Sreenu, G.; Durai, S. Intelligent video surveillance: a review through deep learning techniques for crowd analysis. Journal of Big Data 2019, 6, 1–27.
- Muhammad, K.; Ullah, A.; Lloret, J.; Del Ser, J.; De Albuquerque, V.H.C. Deep learning for safe autonomous driving: Current challenges and future directions. IEEE Transactions on Intelligent Transportation Systems 2020, 22, 4316–4336.
- Zhan, H. The potential of large language models to achieve artificial general intelligence. Topoi 2025, pp. 1–9.
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 2012, 25.
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 2014.
- Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9.
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 2016, 39, 1137–1149.
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440.
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2015, pp. 234–241.
- LeCun, Y.; Boser, B.; Denker, J.; Henderson, D.; Howard, R.; Hubbard, W.; Jackel, L. Handwritten digit recognition with a back-propagation network. Advances in neural information processing systems 1989, 2.
- Graves, A.; Mohamed, A.r.; Hinton, G. Speech recognition with deep recurrent neural networks. In Proceedings of the 2013 IEEE international conference on acoustics, speech and signal processing. Ieee, 2013, pp. 6645–6649.
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural computation 1997, 9, 1735–1780.
- Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv:1412.3555 2014.
- Kingma, D.P.; Welling, M. Auto-encoding variational bayes. arXiv:1312.6114 2013.
- Sohn, K.; Lee, H.; Yan, X. Learning structured output representation using deep conditional generative models. Advances in neural information processing systems 2015, 28.
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Advances in neural information processing systems 2014, 27.
- Radford, A.; Metz, L.; Chintala, S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv:1511.06434 2015.
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Advances in neural information processing systems 2017, 30.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929 2020.
- Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv:2010.02502 2020.
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems 2020, 33, 6840–6851.
- Schmidhuber, J. Deep learning in neural networks: An overview. Neural networks 2015, 61, 85–117.
- Khan, A.; Sohail, A.; Zahoora, U.; Qureshi, A.S. A survey of the recent architectures of deep convolutional neural networks. Artificial intelligence review 2020, 53, 5455–5516.
- Li, Z.; Liu, F.; Yang, W.; Peng, S.; Zhou, J. A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems 2021, 33, 6999–7019.
- Salehinejad, H.; Sankar, S.; Barfett, J.; Colak, E.; Valaee, S. Recent advances in recurrent neural networks. arXiv:1801.01078 2017.
- Lin, T.; Wang, Y.; Liu, X.; Qiu, X. A survey of transformers. AI open 2022, 3, 111–132.
- Chakraborty, T.; Reddy KS, U.; Naik, S.M.; Panja, M.; Manvitha, B. Ten years of generative adversarial nets (GANs): a survey of the state-of-the-art. Machine Learning: Science and Technology 2024, 5, 011001.
- Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; Yang, M.H. Diffusion models: A comprehensive survey of methods and applications. ACM computing surveys 2023, 56, 1–39.
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep learning; Vol. 1, MIT press Cambridge, 2016.
- Bishop, C.M.; Bishop, H. Deep learning: Foundations and concepts; Springer Nature, 2023.
- Koch, C. The Brain: Neurons, Synapses, and Neural Networks. In Encyclopedia of Religious Psychology and Behavior; Springer, 2025; pp. 1–4.
- Rosenblatt, F. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review 1958, 65, 386.
- Minsky, M.; Papert, S. An introduction to computational geometry. Cambridge tiass., HIT 1969, 479, 104.
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. nature 1986, 323, 533–536.
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE 1998, 86, 2278–2324.
- Fukushima, K. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological cybernetics 1980, 36, 193–202.
- Lowe, D.G. Distinctive image features from scale-invariant keypoints. International journal of computer vision 2004, 60, 91–110.
- Dalal, N.; Triggs, B. Histograms of oriented gradients for human detection. In Proceedings of the 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05). Ieee, 2005, Vol. 1, pp. 886–893.
- Csurka, G.; Dance, C.; Fan, L.; Willamowski, J.; Bray, C. Visual categorization with bags of keypoints. In Proceedings of the Workshop on statistical learning in computer vision, ECCV. Prague, 2004, Vol. 1, pp. 1–2.
- Zeiler, M.D.; Fergus, R. Visualizing and understanding convolutional networks. In Proceedings of the Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13. Springer, 2014, pp. 818–833.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
- Zagoruyko, S.; Komodakis, N. Wide residual networks. arXiv preprint arXiv:1605.07146 2016.
- Iandola, F.N.; Han, S.; Moskewicz, M.W.; Ashraf, K.; Dally, W.J.; Keutzer, K. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and<0.5 MB model size. arXiv preprint arXiv:1602.07360 2016.
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 2017.
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6848–6856.
- Zoph, B.; Vasudevan, V.; Shlens, J.; Le, Q.V. Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 8697–8710.
- Tan, M.; Chen, B.; Pang, R.; Vasudevan, V.; Sandler, M.; Howard, A.; Le, Q.V. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 2820–2828.
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141.
- Tan, M.; Le, Q.E.; et al. Rethinking model scaling for convolutional neural networks. In Proceedings of the International conference on machine learning, Long Beach, CA, USA, 2019, Vol. 15.
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), 2018, pp. 3–19.
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803.
- Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; Zhang, L. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 22–31.
- Guo, J.; Han, K.; Wu, H.; Tang, Y.; Chen, X.; Wang, Y.; Xu, C. Cmt: Convolutional neural networks meet vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 12175–12185.
- Dai, Z.; Liu, H.; Le, Q.V.; Tan, M. Coatnet: Marrying convolution and attention for all data sizes. Advances in neural information processing systems 2021, 34, 3965–3977.
- Liu, Z.; Mao, H.; Wu, C.Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11976–11986.
- Elman, J.L. Finding structure in time. Cognitive science 1990, 14, 179–211.
- Hopfield, J.J. Neural networks and physical systems with emergent collective computational abilities. national academy of sciences 1982, 79, 2554–2558.
- Werbos, P.J. Backpropagation through time: what it does and how to do it. Proceedings of the IEEE 1990, 78, 1550–1560.
- Jordan, M.I. Serial order: A parallel distributed processing approach. In Advances in psychology; Elsevier, 1997; Vol. 121, pp. 471–495.
- Hochreiter, S. Untersuchungen zu dynamischen neuronalen Netzen. Diploma, Technische Universität München 1991, 91, 31.
- Schuster, M.; Paliwal, K.K. Bidirectional recurrent neural networks. IEEE transactions on Signal Processing 1997, 45, 2673–2681.
- Gers, F.A.; Schraudolph, N.N.; Schmidhuber, J. Learning precise timing with LSTM recurrent networks. Journal of machine learning research 2002, 3, 115–143.
- Rabiner, L.; Juang, B. An introduction to hidden Markov models. ieee assp magazine 2003, 3, 4–16.
- Lafferty, J.; McCallum, A.; Pereira, F.C. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the International Conference on Machine Learning, 2001.
- Mikolov, T.; Chen, K.; Corrado, G.; Dean, J. Efficient estimation of word representations in vector space. arXiv:1301.3781 2013.
- Sutskever, I.; Vinyals, O.; Le, Q.V. Sequence to sequence learning with neural networks. Advances in neural information processing systems 2014, 27.
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 2014.
- Luong, M.T.; Pham, H.; Manning, C.D. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025 2015.
- Shi, X.; Chen, Z.; Wang, H.; Yeung, D.Y.; Wong, W.K.; Woo, W.c. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Advances in neural information processing systems 2015, 28.
- Vinyals, O.; Fortunato, M.; Jaitly, N. Pointer networks. Advances in neural information processing systems 2015, 28.
- Sukhbaatar, S.; Weston, J.; Fergus, R.; et al. End-to-end memory networks. Advances in neural information processing systems 2015, 28.
- Gulrajani, I.; Kumar, K.; Ahmed, F.; Taiga, A.A.; Visin, F.; Vazquez, D.; Courville, A. Pixelvae: A latent variable model for natural images. arXiv:1611.05013 2016.
- Higgins, I.; Matthey, L.; Pal, A.; Burgess, C.P.; Glorot, X.; Botvinick, M.M.; Mohamed, S.; Lerchner, A. beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. In Proceedings of the ICLR (Poster). OpenReview.net, 2017.
- Van Den Oord, A.; Vinyals, O.; et al. Neural discrete representation learning. Advances in neural information processing systems 2017, 30.
- Vahdat, A.; Kautz, J. NVAE: A deep hierarchical variational autoencoder. Advances in neural information processing systems 2020, 33, 19667–19679.
- Mirza, M.; Osindero, S. Conditional generative adversarial nets. arXiv:1411.1784 2014.
- Perarnau, G.; Van De Weijer, J.; Raducanu, B.; Álvarez, J.M. Invertible conditional gans for image editing. arXiv:1611.06355 2016.
- Odena, A.; Olah, C.; Shlens, J. Conditional image synthesis with auxiliary classifier gans. In Proceedings of the International conference on machine learning. PMLR, 2017, pp. 2642–2651.
- Denton, E.L.; Chintala, S.; Fergus, R.; et al. Deep generative image models using a laplacian pyramid of adversarial networks. Advances in neural information processing systems 2015, 28.
- Zhang, H.; Xu, T.; Li, H.; Zhang, S.; Wang, X.; Huang, X.; Metaxas, D.N. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, 2017, pp. 5907–5915.
- Huang, X.; Li, Y.; Poursaeed, O.; Hopcroft, J.; Belongie, S. Stacked generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5077–5086.
- Isola, P.; Zhu, J.Y.; Zhou, T.; Efros, A.A. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1125–1134.
- Chen, X.; Duan, Y.; Houthooft, R.; Schulman, J.; Sutskever, I.; Abbeel, P. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. Advances in neural information processing systems 2016, 29.
- Lee, H.Y.; Tseng, H.Y.; Huang, J.B.; Singh, M.; Yang, M.H. Diverse image-to-image translation via disentangled representations. In Proceedings of the European conference on computer vision (ECCV), 2018, pp. 35–51.
- Tran, L.; Yin, X.; Liu, X. Disentangled representation learning gan for pose-invariant face recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1415–1424.
- Nowozin, S.; Cseke, B.; Tomioka, R. f-gan: Training generative neural samplers using variational divergence minimization. Advances in neural information processing systems 2016, 29.
- Zhu, J.Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, 2017, pp. 2223–2232.
- Karras, T.; Aila, T.; Laine, S.; Lehtinen, J. Progressive growing of gans for improved quality, stability, and variation. arXiv:1710.10196 2017.
- Arjovsky, M.; Chintala, S.; Bottou, L. Wasserstein generative adversarial networks. In Proceedings of the International conference on machine learning. PMLR, 2017, pp. 214–223.
- Tulyakov, S.; Liu, M.Y.; Yang, X.; Kautz, J. Mocogan: Decomposing motion and content for video generation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1526–1535.
- Vondrick, C.; Pirsiavash, H.; Torralba, A. Generating videos with scene dynamics. Advances in neural information processing systems 2016, 29.
- Clark, A.; Donahue, J.; Simonyan, K. Adversarial video generation on complex datasets. arXiv:1907.06571 2019.
- Brock, A.; Donahue, J.; Simonyan, K. Large scale GAN training for high fidelity natural image synthesis. arXiv:1809.11096 2018.
- Karras, T.; Laine, S.; Aila, T. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4401–4410.
- Zhang, H.; Goodfellow, I.; Metaxas, D.; Odena, A. Self-attention generative adversarial networks. In Proceedings of the International conference on machine learning. PMLR, 2019, pp. 7354–7363.
- Wang, Z.; Zheng, H.; He, P.; Chen, W.; Zhou, M. Diffusion-gan: Training gans with diffusion. arXiv:2206.02262 2022.
- Kang, M.; Zhang, R.; Barnes, C.; Paris, S.; Kwak, S.; Park, J.; Shechtman, E.; Zhu, J.Y.; Park, T. Distilling diffusion models into conditional gans. In Proceedings of the European Conference on Computer Vision. Springer, 2024, pp. 428–447.
- Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. Improving language understanding by generative pre-training. OpenAI 2018.
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805 2018.
- Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q.V.; Salakhutdinov, R. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv:1901.02860 2019.
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 2020, 21, 5485–5551.
- Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T.B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; Amodei, D. Scaling laws for neural language models. arXiv:2001.08361 2020.
- Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; Jégou, H. Training data-efficient image transformers & distillation through attention. In Proceedings of the International conference on machine learning. PMLR, 2021, pp. 10347–10357.
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International conference on machine learning. PmLR, 2021, pp. 8748–8763.
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022.
- Fedus, W.; Zoph, B.; Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research 2022, 23, 1–39.
- Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv:2006.16668 2020.
- Wang, S.; Li, B.Z.; Khabsa, M.; Fang, H.; Ma, H. Linformer: Self-attention with linear complexity. arXiv:2006.04768 2020.
- Lou, C.; Jia, Z.; Zheng, Z.; Tu, K. Sparser is faster and less is more: Efficient sparse attention for long-range transformers. arXiv:2406.16747 2024.
- Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; Ré, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in neural information processing systems 2022, 35, 16344–16359.
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 2022, 35, 27730–27744.
- Simonyan, K.; Zisserman, A. Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems 2014, 27.
- Hannun, A.; Case, C.; Casper, J.; Catanzaro, B.; Diamos, G.; Elsen, E.; Prenger, R.; Satheesh, S.; Sengupta, S.; Coates, A.; et al. Deep speech: Scaling up end-to-end speech recognition. arXiv:1412.5567 2014.
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
- Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1440–1448.
- Schroff, F.; Kalenichenko, D.; Philbin, J. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 815–823.
- Bertinetto, L.; Valmadre, J.; Henriques, J.F.; Vedaldi, A.; Torr, P.H. Fully-convolutional siamese networks for object tracking. In Proceedings of the European conference on computer vision. Springer, 2016, pp. 850–865.
- Newell, A.; Yang, K.; Deng, J. Stacked hourglass networks for human pose estimation. In Proceedings of the European conference on computer vision. Springer, 2016, pp. 483–499.
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European conference on computer vision. Springer, 2016, pp. 21–37.
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788.
- Wojke, N.; Bewley, A.; Paulus, D. Simple online and realtime tracking with a deep association metric. In Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP). IEEE, 2017, pp. 3645–3649.
- Cao, Z.; Simon, T.; Wei, S.E.; Sheikh, Y. Realtime multi-person 2d pose estimation using part affinity fields. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7291–7299.
- Carreira, J.; Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6299–6308.
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988.
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence 2017, 39, 2481–2495.
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv:1412.7062 2014.
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969.
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), 2018, pp. 801–818.
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. Language models are unsupervised multitask learners. OpenAI blog 2019, 1, 9.
- Sun, K.; Xiao, B.; Liu, D.; Wang, J. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5693–5703.
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211.
- Deng, J.; Guo, J.; Xue, N.; Zafeiriou, S. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4690–4699.
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European conference on computer vision. Springer, 2020, pp. 213–229.
- Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; Schmid, C. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6836–6846.
- Bertasius, G.; Wang, H.; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? In Proceedings of the ICML. PMLR, 2021, Vol. 139, Proceedings of Machine Learning Research, pp. 813–824.
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, 2021, Vol. 35, pp. 11106–11115.
- Cheng, B.; Schwing, A.; Kirillov, A. Per-pixel classification is not all you need for semantic segmentation. Advances in neural information processing systems 2021, 34, 17864–17875.
- Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 2022, 35, 23716–23736.
- Li, J.; Li, D.; Xiong, C.; Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International conference on machine learning. PMLR, 2022, pp. 12888–12900.
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 2022, 35, 27730–27744.
- Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video diffusion models. Advances in neural information processing systems 2022, 35, 8633–8646.
- Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D.P.; Poole, B.; Norouzi, M.; Fleet, D.J.; et al. Imagen video: High definition video generation with diffusion models. arXiv:2210.02303 2022.
- Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1290–1299.
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026.
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak supervision. In Proceedings of the International conference on machine learning. PMLR, 2023, pp. 28492–28518.
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In Proceedings of the International Conference on Learning Representations, 2023.
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. Advances in neural information processing systems 2023, 36, 34892–34916.
- Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H.; et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In Proceedings of the European conference on computer vision. Springer, 2024, pp. 38–55.
- Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv:2311.15127 2023.
- Zheng, Z.; Peng, X.; Lou, Y.; Shen, C.; Young, T.; Guo, X.; Wang, B.; Xu, H.; Liu, H.; Jiang, M.; et al. Open-sora 2.0: Training a commercial-level video generation model in $200 k. arXiv:2503.09642 2025.
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025, 645, 633–638.
- Salimans, T.; Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv:2202.00512 2022.
- Song, Y.; Sohl-Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score-based generative modeling through stochastic differential equations. arXiv:2011.13456 2020.
- Vahdat, A.; Kreis, K.; Kautz, J. Score-based generative modeling in latent space. Advances in neural information processing systems 2021, 34, 11287–11302.
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10684–10695.
- Nichol, A.Q.; Dhariwal, P. Improved denoising diffusion probabilistic models. In Proceedings of the International conference on machine learning. PMLR, 2021, pp. 8162–8171.
- Song, Y.; Durkan, C.; Murray, I.; Ermon, S. Maximum likelihood training of score-based diffusion models. Advances in neural information processing systems 2021, 34, 1415–1428.
- Dhariwal, P.; Nichol, A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems 2021, 34, 8780–8794.
- Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 2022, 35, 36479–36494.
- Ho, J.; Salimans, T. Classifier-free diffusion guidance. arXiv:2207.12598 2022.
- Song, Y.; Dhariwal, P.; Chen, M.; Sutskever, I. Consistency models. In Proceedings of the International Conference on Machine Learning, 2023, pp. 32211–32252.
- Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video diffusion models. Advances in neural information processing systems 2022, 35, 8633–8646.
- Zhang, L.; Rao, A.; Agrawala, M. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 3836–3847.


| Method | Year | Pub. | App. | Base Arch. | Param. | Dataset | Perf. | Summary |
| Two Stream [118] | 2014 | NeurIPS | HAR | CNN | UCF-101 | 88.0 Acc | Two-stream architecture, one for spatial and the other for temporal, employing optical flow to feed motion information | |
| DeepSpeech [119] | 2014 | arXiv | ASR | RNN | SWB | 12.6 WER | Trains an end-to-end RNN model using speech spectrograms as input. It eliminates the need for a phoneme dictionary or handcrafted features | |
| RCNN [120] | 2014 | CVPR | Object Det. | CNN | VOC 2010 | 53.7 mAP | Performs selective search to extract 2000 regions, followed by CNN processing, SVM-based classification and bounding box regression | |
| FCN [14] | 2015 | CVPR | Image Seg. | CNN | VOC 2012 | 62.2 mIoU | Employs pre-trained image classification models for feature extraction, suggested skip connections with summation from the lower layers to the decoder | |
| U-Net [15] | 2015 | MICCAI | Image Seg. | CNN | Extends FCN with symmetric concatenated skip connections from the encoder to the decoder | |||
| FastRCNN [121] | 2015 | ICCV | Obj. Det | CNN | VOC 2012 | 68.4 mAP | Speeds up object detection by processing all regions through the CNN at once. Additionally, proposed region of interest pooling | |
| FasterRCNN [13] | 2015 | NeurIPS | Obj. Det | CNN | VOC 2012 | 75.9 mAP | Replaced selective searching with trainable region proposal network with end-to-end training | |
| FaceNet [122] | 2015 | CVPR | Face Recog. | CNN | 140M | LFW | 99.63 Acc. | Employs triplet loss to bring embeddings of the same faces closer in the Euclidean space and vice versa |
| SiamFC [123] | 2016 | ECCV | Tracking | CNN | VOT-15 | 53.3 Acc | A Siamese architecture to track an object in the larger image by finding similarities through a sliding window | |
| Stacked Hourglass [124] | 2016 | ECCV | Human Pose Est. | CNN | MPII | 90.9 PCKh | Stacks bottom-up and top-down modules with pose supervision at each block for iterative refinement | |
| SSD [125] | 2016 | ECCV | Obj. Det | CNN | VOC 2012 | 80.0 mAP | Uses pre-defined anchors with multi-scale feature maps for object detection, eliminating the need for region proposal network | |
| YOLO [126] | 2016 | CVPR | Obj. Det | CNN | 45M | VOC 2012 | 57.9 mAP | Repurposed object detection as a regression problem by unifying class and bounding box prediction into a single network |
| DeepSORT [127] | 2017 | ICIP | Tracking | CNN | 2.8M | MOT-16 | 61.4 MOTA | Tracks pedestrians by using Kalman filter and appearance descriptors even in occluded scenarios |
| OpenPose [128] | 2017 | CVPR | Human Pose Est. | CNN | 26.2M | MPII | 75.6 mAP | Cascaded two-branch architecture, one to predict keypoint heatmap and the other for part affinity fields |
| I3D [129] | 2017 | CVPR | HAR | CNN | 25M | UCF-101 | 98.0 Acc | Two-stream temporal 3D-CNN, one processes RGB and the other optical flow. It inflates pre-trained 2D CNNs into 3D CNN |
| YOLOv2 | 2017 | CVPR | Obj. Det | CNN | ∼51M | VOC 2012 | 73.4 mAP | It is a faster and more accurate variant with an efficient Darknet backbone and predicts boxes as an offset to the anchor boxes |
| RetinaNet [130] | 2017 | ICCV | Obj. Det | CNN | COCO | 40.8 mAP | Proposed focal loss to balance the class imbalance between foreground and background classes in object detection | |
| SegNet [131] | 2017 | PAMI | Image Seg. | CNN | ∼15M+ | SUNRGB-D | 31.84 mIoU | Proposed upsampling by using indices used for downsampling in max-pooling, contrary to deconvolution in FCN |
| DeepLab [132] | 2017 | PAMI | Image Seg. | CNN | VOC 2012 | 79.7 mIoU | Proposed atrous spatial pyramid pooling to capture information at multiple scales and uses fully connected conditional random fields at the last layer | |
| MaskRCNN [133] | 2017 | ICCV | Image Seg. | CNN | City-Scapes | 32.0 mAP | Extended FasterRCNN for instance segmentation by adding a parallel branch to the object detection path. Moreover, it proposed ROI Align that does not do quantization as opposed to ROI Pooling | |
| GPT [104] | 2018 | arXiv | TG | T | 117M | GLUE | 72.8 | Large-scale unsupervised pre-training on unlabelled data followed by task-specific fine-tuning consistently outperforms task-specific training |
| DeepLabv3+ [134] | 2018 | ECCV | Image Seg. | CNN | VOC 2012 | 89.0 mIoU | Extended DeepLabv3 with decoder modules and employing depthwise separable convolution in atrous spatial pyramid pooling and decoder | |
| GPT-2 [135] | 2019 | arXiv | TG | T | 1.5B | LAMBA-DA | 8.63 PPL | Large-scale unsupervised pre-training is enough for models to perform well on diverse tasks without supervision |
| HRNeT [136] | 2019 | CVPR | Human Pose Est. | CNN | 63.6M | COCO | 77.0 mAP | Parallely stacks multi-resolution sub-networks, each operating on a different scale with multi-scale fusion |
| SlowFast [137] | 2019 | ICCV | HAR | CNN | 32.9M | Kinetics-400 | 75.6 Acc | Proposes two video processing streams, one operating on low frame rate for spatial semantics and the other on high frame rate for motion learning |
| ArcFace [138] | 2019 | CVPR | Face Recog. | CNN | 65M | LFW | 99.83 Acc | Proposes angular margin loss which improves class separability in embedding space |
| DETR [139] | 2020 | ICCV | Obj. Det | T | 60M | COCO | 44.9 mAP | Transformer-based architecture that eliminates the need for non-maximal suppression and anchor boxes, instead directly predicts objects |
| T5 [107] | 2020 | JMLR | TG | T | 11B | GLUE | 90.3 | Pre-trains by treating all language-modeling tasks in a unified text-to-text approach followed by task-specific transfer learning |
| GPT-3 [2] | 2020 | NeurIPS | TG | T | 175B | Super-GLUE | 73.2 Few-Shot | Illustrates that increasing the pre-training scale further brings more gains in text generation |
| CLIP [110] | 2021 | PMLR | ITM | T | 428M | ImageNet | 76.2 Zero-Shot | Jointly learns image and text encoders to pair related images and texts via contrastive pre-training. CLIP shows excellent zero-shot performance |
| ViViT [140] | 2021 | ICCV | VU | T | Kinetics-400 | 78.9 Acc | Suggests a combination of parallel and sequential spatial and temporal attention video transformer variants | |
| TimeSformer [141] | 2021 | ICML | VU | T | ∼121M | Kinetics-400 | 80.7 Acc | Suggests temporal attention followed by spatial attention is better than other spatio-temporal attention variants |
| Informer [142] | 2021 | AAAI | Time-Series | CNN+T | Weather | 0.831 MSE | Suggests efficient attention mechanism by attending only for top queries in long sequence time-series forecasting | |
| MaskFormer [143] | 2021 | NeurIPS | Image Seg. | T | 212M | ADE20K | 55.6 mIoU | Improved semantic segmentation by fusing mask embeddings with pixel embeddings extracted from transformer and pixel decoder, respectively |
| Flamingo [144] | 2022 | NeurIPS | ITM | T | 80B | VQA | 45.3 Few-Shot | Trains on arbitrarily interleaved visual and text data on a large dataset for few-shot in-context learning. It can be used in open-ended VQA, image captioning, etc. |
| BLIP [145] | 2022 | PMLR | ITM | T | VQA | 78.25 | Pre-trains with image-text contrastive, image-text matching, and language modeling losses for VLM to be good at both understanding and generation tasks | |
| Instruct-GPT [146] | 2022 | NeurIPS | TG | T | 175B | SQuADv2 | 69.93 Few-Shot | Training pre-trained LLMs with human feedback steers models to generate human desired output |
| Video Diffusion [147] | 2022 | NeurIPS | Video Gen. | CNN+T | Kinetics-600 | 16.2 FVD | Extends 2D diffusion models to 3D for video generation with factorized space-time attention blocks | |
| Imagen Video [148] | 2022 | arXiv | Video Gen. | CNN+T | Trains cascaded text-conditioned video diffusion models to generate high fidelity videos, containing interleaved spatial and temporal super resolution models | |||
| Mask2Former [149] | 2022 | CVPR | Image Seg. | T | 216M | COCO | 50.1 mAP | Improved MaskFormer by proposing masked attention to attend to only foreground regions |
| SAM [150] | 2023 | ICCV | Image Seg. | T | 636M | COCO | 46.5 Zero-Shot | A model trained on 1 billion masks, capable of segmenting regions based on diverse prompts, such as masks, points, and bounding boxes. SAM achieves excellent zero-shot transferability |
| Whisper [151] | 2023 | ICML | ASR | T | 1.5B | SWB | 13.8 | Illustrated large-scale weakly supervised seq2seq pre-training on 680k hours labeled audio data, collected from the internet, achieves significant performance gains |
| LLaMA [3] | 2023 | arXiv | TG | T | 65B | MMLU | 68.9 Few-Shot | Shows open-source datasets can perform well or equivalent to the models trained on closed-source datasets like GPT-3 |
| TimesNet [152] | 2023 | ICLR | Time-Series | CNN | Weather | 0.259 MSE | Captures intra- and inter-period patterns in time series data by reshaping to a 2D format to process through 2D kernels | |
| LLaVA [153] | 2023 | NeurIPS | ITM | T | ScienceQA | 92.53 | Proposed vision-language instruction tuning using pre-trained LLMs and vision encoder | |
| Grounding DINO [154] | 2024 | ECCV | Obj. Det | T | 172M | COCO | 52.5 Zero-Shot | Achieves open set object detection by introducing language modeling to closed set object detection |
| Stable Video Diff [155] | 2023 | arXiv | Video Gen. | CNN+T | 1.5B | UCF-101 | 242.02 FVD (Zero-Shot) | Scales latent video diffusion model to large datasets |
| Open-Sora [156] | 2025 | arXiv | Video Gen. | CNN+T | 1.1B | VBench | 79.76 | An open-source video generation implementation, trying to replicate the OpenAI Sora model |
| DeepSeek-R1 [157] | 2025 | arXiv | TG | T | 671B | MMLU | 90.8 | Minimizes human-annotated data reliance by training LLMs via pure reinforcement learning |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).