Submitted:
01 June 2023
Posted:
02 June 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction

2. Related work
3. Proposed approach
- Scaling to 224x224 without preserving the aspect ratio (Resize(224x224)).
- Scaling to 256 with the aspect ratio preserved, then cropping out a 224x224 square from the center of the image (Resize(256) + CenterCrop(224x224)).
- Scaling to 224x224 with the aspect ratio preserved, black bars appear (Resize(224x224) with borders).

- Resize(256) + CenterCrop(224x224) is a standard transformation used in classification. Its disadvantage is that part of the image is cropped, and thus information is lost.
- Resize(224x224) loses aspect ratio information, which makes the image look very different visually.
- Resize(224x224) with borders preserves the aspect ratio information but reduces the effective resolution.

4. Experimental evaluation
| Model name | Pre-training Dataset |
|---|---|
| Resnet18_in1k | ResNet18 pre-trained on ImageNet-1K |
| Resnet50_in1k | ResNet50 pre-trained on ImageNet-1K |
| resnetv2_50x1_bitm_in21k | ResNet50V2 pre-trained on ImageNet-21K |
| Model name | Pre-training Dataset |
|---|---|
| ViT B/16 in 1k | ViT B/16 pre-trained on ImageNet-1K |
| ViT B/16 in 21k | ViT B/16 pre-trained on ImageNet-21K |
| BEiT ViT B/16 in 21k | BEiT ViT B/16 pre-trained on ImageNet-21K |


| Neural network | Dimensions | Resizing technique | mAP@100 |
|---|---|---|---|
| Resnet18_in1k | 512 | Resize(256) + CenterCrop(224x224) | 11.05 |
| Resize(224x224) | 12.23 | ||
| Resize(224x224) with borders | 11.25 | ||
| Resnet50_in1k | 2048 | Resize(256) + CenterCrop(224x224) | 11.55 |
| Resize(224x224) | 13.19 | ||
| Resize(224x224) with borders | 13.62 | ||
| resnetv2_50x1_bitm_in21k | 2048 | Resize(256) + CenterCrop(224x224) | 14.02 |
| Resize(224x224) | 14.93 | ||
| Resize(224x224) with borders | 15.17 | ||
| ViT B/16 in 1k | 768 | Resize(256) + CenterCrop(224x224) | 7.44 |
| Resize(224x224) | 8.30 | ||
| Resize(224x224) with borders | 7.35 | ||
| ViT B/16 in 21k | 768 | Resize(256) + CenterCrop(224x224) | 14.65 |
| Resize(224x224) | 15.97 | ||
| Resize(224x224) with borders | 16.83 | ||
| BEiT ViT B/16 in 21k | 768 | Resize(256) + CenterCrop(224x224) | 18.01 |
| Resize(224x224) | 19.63 | ||
| Resize(224x224) with borders | 20.19 |

| Neural network | mAP@100 before applying PCAw |
mAP@100 after applying PCAw |
|---|---|---|
| resnetv2_50x1_bitm_in21k | 15.17 | 18.50 |
| BEiT ViT B/16 in 21k | 20.19 | 25.12 |

| Method | mAP@100 |
|---|---|
| BEiT ViT B/16 in 21k | 20.19 |
| BEiT ViT B/16 in 21k +pcaW | 25.12 |
| BEiT ViT B/16 in 21k +pcaW+aQE | 28.46 |
| BEiT ViT B/16 in 21k +pcaW+aQE+reranking | 30.62 |
| BEiT ViT B/16 in 21k +pcaW+aQE+reranking + local_features | 31.23 |
5. Discussion

6. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| ViT | Vision transformer |
| CNN | Convolutional neural network |
| CBIR | Content-based image retrieval |
| mAP | Mean average precision |
| NAR | Normalized average rank |
| MR | Multi-resolution |
| SMAC | Sum and max activation of convolution |
| URA | Unsupervised regional attention |
| AQE | Average query expansion |
| QE | -weighted query expansion |
| K-NN | K nearest neighbors |
| SMNN | Second mutual nearest neighbors |
References
- Organization, W.I.P. World Intellectual Property Indicators 2021. https://www.wipo.int/edocs/pubdocs/en/wipo_pub_941_2021.pdf, 2021. Online: accessed 31.05.2023.
- Tursun, O.; Aker, C.; Kalkan, S. A large-scale dataset and benchmark for similar trademark retrieval. arXiv preprint 2017, arXiv:1701.05766 2017. [Google Scholar]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. Communications of the ACM 2017, 60, 84–90. [Google Scholar] [CrossRef]
- Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9.
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint 2014, arXiv:1409.1556 2014. [Google Scholar]
- others. Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; others. Imagenet large scale visual recognition challenge. International journal of computer vision 2015, 115, 211–252. [Google Scholar]
- Perez, C.A.; Estévez, P.A.; Galdames, F.J.; Schulz, D.A.; Perez, J.P.; Bastías, D.; Vilar, D.R. Trademark image retrieval using a combination of deep convolutional neural networks. 2018 International Joint Conference on Neural Networks (IJCNN). IEEE, 2018, pp. 1–7.
- Babenko, A.; Lempitsky, V. Aggregating local deep features for image retrieval. Proceedings of the IEEE international conference on computer vision, 2015, pp. 1269–127.
- Kalantidis, Y.; Mellina, C.; Osindero, S. Cross-dimensional weighting for aggregated deep convolutional features. Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I 14. Springer, 2016, pp. 685–701. 8 October.
- Tolias, G.; Sicre, R.; Jégou, H. Particular object retrieval with integral max-pooling of CNN activations. arXiv preprint 2015, arXiv:1511.05879 2015. [Google Scholar]
- Radenović, F.; Tolias, G.; Chum, O. Fine-tuning CNN image retrieval with no human annotation. IEEE transactions on pattern analysis and machine intelligence 2018, 41, 1655–1668. [Google Scholar] [CrossRef] [PubMed]
- Tursun, O.; Denman, S.; Sivapalan, S.; Sridharan, S.; Fookes, C.; Mau, S. Component-based attention for large-scale trademark retrieval. IEEE Transactions on Information Forensics and Security 2019, 17, 2350–2363. [Google Scholar] [CrossRef]
- Cao, J.; Huang, Y.; Dai, Q.; Ling, W.K. Unsupervised trademark retrieval method based on attention mechanism. Sensors 2021, 21, 1894. [Google Scholar] [CrossRef] [PubMed]
- Tursun, O.; Denman, S.; Sridharan, S.; Fookes, C. Learning test-time augmentation for content-based image retrieval. Computer Vision and Image Understanding 2022, 222, 103494. [Google Scholar] [CrossRef]
- Tursun, O.; Denman, S.; Sridharan, S.; Fookes, C. Learning regional attention over multi-resolution deep convolutional features for trademark retrieval. 2021 IEEE International Conference on Image Processing (ICIP). IEEE, 2021, pp. 2393–2397.
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; others. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint 2020, arXiv:2010.11929 2020.
- Sharif Razavian, A.; Azizpour, H.; Sullivan, J.; Carlsson, S. CNN features off-the-shelf: an astounding baseline for recognition. Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2014, pp. 806–813.
- Chum, O.; Philbin, J.; Sivic, J.; Isard, M.; Zisserman, A. Total recall: Automatic query expansion with a generative feature model for object retrieval. 2007 IEEE 11th International Conference on Computer Vision. IEEE, 2007, pp. 1–8.
- Jin, Y.; Mishkin, D.; Mishchuk, A.; Matas, J.; Fua, P.; Yi, K.M.; Trulls, E. Image matching across wide baselines: From paper to practice. International Journal of Computer Vision 2021, 129, 517–547. [Google Scholar]
- Barath, D.; Noskova, J.; Ivashechkin, M.; Matas, J. MAGSAC++, a fast, reliable and accurate robust estimator. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1304–1312.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
- Bao, H.; Dong, L.; Piao, S.; Wei, F. Beit: Bert pre-training of image transformers. arXiv preprint 2021, arXiv:2106.08254 2021. [Google Scholar]
- Kotenko, I.; Kalameyets, M.; Chechulin, A.; Chevalier, Y. A visual analytics approach for the cyber forensics based on different views of the network traffic. Journal of Wireless Mobile Networks, Ubiquitous Computing, and Dependable Applications 2018, 9, 57–73. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).