Submitted:
27 December 2024
Posted:
30 December 2024
Read the latest preprint version here
Abstract
Keywords:
I. Introduction
II. Background
i. Image Resizing and LLMs
ii. LLM-Related Work
iii. Image and Vision Generation Work
iv. Image Resizing and Seam Carving Research
v. Impact of SeamCarver and Future Directions
III. Functionality
- LLM-Augmented Region Prioritization: LLMs analyze semantics or textual inputs to prioritize key regions, ensuring critical areas (e.g., faces, text) are preserved.
- LLM-Augmented Bicubic Interpolation: LLMs optimize bicubic interpolation for high-quality enlargements, adjusting parameters based on context or user input.
- LLM-Augmented LC Algorithm: LLMs adapt the LC algorithm by adjusting weights, ensuring the preservation of important image features during resizing.
- LLM-Augmented Canny Edge Detection: LLMs guide Canny edge detection to refine boundaries, enhancing clarity and accuracy based on contextual analysis.
- LLM-Augmented Hough Transformation: LLMs strengthen the Hough transformation, detecting structural lines and ensuring the preservation of geometric features.
- LLM-Augmented Absolute Energy Function: LLMs dynamically adjust energy maps to improve seam selection for more precise resizing.
- LLM-Augmented Dual Energy Model: LLMs refine energy functions, enhancing flexibility and ensuring effective seam carving across various use cases.
- LLM-Augmented Performance Evaluation: CNN-based classification experiments on CIFAR-10 are enhanced with LLM feedback to fine-tune resizing results.
IV. LLM-Guided Region Prioritization
i. Method Overview
ii. Energy Map Adjustment
iii. Energy Map Adjustment
iv. Pseudocode
| Algorithm 1:LLM-Guided Region Prioritization for Seam Carving |
|
Initialize:
Compute the initial energy map for the image I;
Obtain semantic importance scores from LLM based on image content or user description;
Normalize the importance scores to a suitable range.
Adjustment:
1. For each pixel , compute the adjusted energy map:
2. Set to control the influence of semantic importance on the energy map.
3. Repeat for all pixels to generate the adjusted energy map .
Output:
The adjusted energy map for guiding seam carving.
|
V. LLM-Augmented Bicubic Interpolation
VI. LLM-Augmented LC (Loyalty-Clarity) Policy
i. Global Contrast Calculation with LLM Influence
ii. Frequency-Based Refinement with LLM Augmentation
iii. Application in Image Resizing

VII. LLM-Augmented Canny Line Detection
i. Algorithm Overview
ii. Gaussian Filter Application
iii. Gradient Calculation with LLM Augmentation
iv. Edge Enhancement
v. Significance in Image Resizing


VIII. LLM-Augmented Hough Transformation
i. Algorithm Overview
ii. Mathematical Formulation
iii. LLM-Augmented Hough Transformation Algorithm
| Algorithm 2:LLM-Augmented Hough Transformation for Line Detection |
|
iv. Significance of LLM-Augmented Hough Transformation in Image Resizing



IX. LLM-Augmented Absolute Energy Equation
i. Semantic Weighting
ii. Gradient Refinement
iii. Cumulative Energy Update
X. LLM-Augmented Dual Gradient Energy Equation
i. Numerical Differentiation with LLM Adjustments
ii. Gradient Approximation with Adaptive Refinements
iii. Energy Calculation with LLM Refinements
XI. Result Evaluation
i. Experimental Setup
ii. Methodology

iii. Results and Discussion







iv. Conclusion
References
- Li, K.; Liu, L.; Chen, J.; Yu, D.; Zhou, X.; Li, M.; Wang, C.; Li, Z. Research on reinforcement learning based warehouse robot navigation algorithm in complex warehouse layout. arXiv preprint, 2024; arXiv:2411.06128 2024. [Google Scholar]
- Hu, Z.; Lei, F.; Fan, Y.; Ke, Z.; Shi, G.; Li, Z. Research on Financial Multi-Asset Portfolio Risk Prediction Model Based on Convolutional Neural Networks and Image Processing. arXiv preprint, 2024; arXiv:2412.03618 2024. [Google Scholar]
- Ke, Z.; Yin, Y. Tail Risk Alert Based on Conditional Autoregressive VaR by Regression Quantiles and Machine Learning Algorithms. arXiv.org 2024. [Google Scholar]
- Xiang, A.; Qi, Z.; Wang, H.; Yang, Q.; Ma, D. A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product 2024. arXiv:cs.CV/2403.08511].
- Wu, C.; Yu, Z.; Song, D. Window views psychological effects on indoor thermal perception: A comparison experiment based on virtual reality environments. E3S Web of Conferences 2024, 546, 02003. [Google Scholar] [CrossRef]
- Ke, Z.; Xu, J.; Zhang, Z.; Cheng, Y.; Wu, W. A Consolidated Volatility Prediction with Back Propagation Neural Network and Genetic Algorithm. arXiv, 2024; arXiv:2412.07223 2024. [Google Scholar]
- Ke, Z.; Yin, Y. Tail Risk Alert Based on Conditional Autoregressive VaR by Regression Quantiles and Machine Learning Algorithms. arXiv, 2024; arXiv:2412.06193 2024. [Google Scholar]
- Guo, F.; Mo, H.; Wu, J.; Pan, L.; Zhou, H.; Zhang, Z.; Li, L.; Huang, F. A hybrid stacking model for enhanced short-term load forecasting. Electronics 2024, 13, 2719. [Google Scholar] [CrossRef]
- Zhang, Z.; Li, P.; Al Hammadi, A.Y.; Guo, F.; Damiani, E.; Yeun, C.Y. Reputation-based federated learning defense to mitigate threats in eeg signal classification. In Proceedings of the 2024 16th International Conference on Computer and Automation Engineering (ICCAE). IEEE; 2024; pp. 173–180. [Google Scholar]
- Bu, X.; Wu, Y.; Gao, Z.; Jia, Y. Deep convolutional network with locality and sparsity constraints for texture classification. Pattern Recognition 2019, 91, 34–46. [Google Scholar] [CrossRef]
- Dan, H.C.; Lu, B.; Li, M. Evaluation of asphalt pavement texture using multiview stereo reconstruction based on deep learning. Construction and Building Materials 2024, 412, 134837. [Google Scholar] [CrossRef]
- Xiang, J.; Chen, J.; Liu, Y. Hybrid Multiscale Search for Dynamic Planning of Multi-Agent Drone Traffic. Journal of Guidance, Control, and Dynamics 2023, 46, 1963–1974. [Google Scholar] [CrossRef]
- Liu, D.; Lai, Z.; Wang, Y.; Wu, J.; Yu, Y.; Wan, Z.; Lengerich, B.; Wu, Y.N. Efficient Large Foundation Model Inference: A Perspective From Model and System Co-Design 2024. arXiv:cs.DC/2409.01990].
- Liu, D.; Pister, K. LLMEasyQuant – An Easy to Use Toolkit for LLM Quantization 2024. arXiv:cs.LG/2406.19657].
- Cao, H.; Zhang, Z.; Li, X.; Wu, C.; Zhang, H.; Zhang, W. Mitigating Knowledge Conflicts in Language Model-Driven Question Answering 2024. arXiv:cs.CL/2411.11344].
- Lai, Z.; Wu, J.; Chen, S.; Zhou, Y.; Hovakimyan, N. Residual-based Language Models are Free Boosters for Biomedical Imaging 2024. arXiv:cs.CV/2403.17343].
- Xin, W.; Wang, K.; Fu, Z.; Zhou, L. Let Community Rules Be Reflected in Online Content Moderation 2024. arXiv:cs.SI/2408.12035].
- Li, K.; Wang, J.; Wu, X.; Peng, X.; Chang, R.; Deng, X.; Kang, Y.; Yang, Y.; Ni, F.; Hong, B. Optimizing automated picking systems in warehouse robots using machine learning. arXiv, 2024; arXiv:2408.16633 2024. [Google Scholar]
- Li, S.; Sun, K.; Lai, Z.; Wu, X.; Qiu, F.; Xie, H.; Miyata, K.; Li, H. ECNet: Effective Controllable Text-to-Image Diffusion Models 2024. arXiv:cs.CV/2403.18417].
- Guo, Y.; Chen, S.; Zhan, R.; Wang, W.; Zhang, J. LMSD-YOLO: A lightweight YOLO algorithm for multi-scale SAR ship detection. Remote Sensing 2022, 14, 4801. [Google Scholar] [CrossRef]
- Dan, H.C.; Huang, Z.; Lu, B.; Li, M. Image-driven prediction system: Automatic extraction of aggregate gradation of pavement core samples integrating deep learning and interactive image processing framework. Construction and Building Materials 2024, 453, 139056. [Google Scholar] [CrossRef]
- Peng, J.; Bu, X.; Sun, M.; Zhang, Z.; Tan, T.; Yan, J. Large-scale object detection in the wild from imbalanced multi-labels. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp.; pp. 9709–9718.
- Fang, X.; Si, S.; Sun, G.; Sheng, Q.Z.; Wu, W.; Wang, K.; Lv, H. Selecting workers wisely for crowdsourcing when copiers and domain experts co-exist. Future Internet 2022, 14, 37. [Google Scholar] [CrossRef]
- Wu, W. Alphanetv4: Alpha Mining Model. arXiv, 2024; arXiv:2411.04409 2024. [Google Scholar]
- Hu, Y.; Cao, H.; Yang, Z.; Huang, Y. Improving text-image matching with adversarial learning and circle loss for multi-modal steganography. In Proceedings of the International Workshop on Digital Watermarking. Springer; 2020; pp. 41–52. [Google Scholar]
- Cheng, Y.; Yang, Q.; Wang, L.; Xiang, A.; Zhang, J. Research on Credit Risk Early Warning Model of Commercial Banks Based on Neural Network Algorithm 2024. arXiv:q-fin.RM/2405.10762].
- Liu, D.; Waleffe, R.; Jiang, M.; Venkataraman, S. GraphSnapShot: Graph Machine Learning Acceleration with Fast Storage and Retrieval 2024. arXiv:cs.LG/2406.17918].
- Liu, D.; Yu, Y. MT2ST: Adaptive Multi-Task to Single-Task Learning 2024. arXiv:cs.LG/2406.18038].
- Kiess, H. Improved Edge Preservation in Seam Carving for Image Resizing. Computer Graphics Forum 2014, 33, 421–429. [Google Scholar] [CrossRef]
- Zhang, W.; Wu, C.; Li, X. Comparison of Image Resizing Techniques: A Case Study of Seam Carving vs. Traditional Resizing Methods. Journal of Visual Communication and Image Representation 2015, 29, 149–158. [Google Scholar] [CrossRef]
- Frankovich, R. Enhanced Seam Carving: Energy Gradient Functionals and Resizing Control. In Proceedings of the Proceedings of the IEEE International Conference on Image Processing (ICIP), 2011; pp. 2157–2160. [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems 2017, 30. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Proceedings of NAACL-HLT; 2019. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).