Submitted:
11 July 2026
Posted:
13 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We propose a modification of the dual encoder network by evaluating the best CNN encoder and Transformer encoder to replace the C-encoder and Transformer encoder of the standard DECTNet for binary semantic segmentation of focal liver segmentation of triphasic CT scan.
- We added the feature fusion strategy that will optimize the evaluation of the modified dual encoder network in comparison with the baseline models such as standard U-Net, DECTNet, and nnU-Net.
- We evaluated the models using 3-fold strategy cross-validation against the baseline models using Generalization Metrics: Dice and Intersection over Union, Image Segmentation Metrics: Accuracy, Sensitivity, Specificity, Precision, and F1-Score, Boundary and Shape Metrics: 95th Percentile Hausdorff Distance (HD95) and Normalized Surface Dice (NSD), and Qualitative Results.
- We performed statistical analysis for the validation of all the metrics for the model performance using two-tailed (dependent-samples) t-tests on the fold-wise scores on shared cross-validation folds. P-values within metrics were adjusted with the Benjamini-Hochberg procedure to control the false discovery rate.
2. Related Literature
2.1. Multi-Phase Frameworks Built on nnU-Net
2.2. Dual-Encoder Optimization Architectures
2.3. Critical Literature for Binary Task Significance
3. Methods
3.1. Preparation of Datasets
3.2. Data Pre-Processing
3.3. Network Architecture


3.4. Training Configuration and Optimization
| Settings | Value | Settings | Value | |
|---|---|---|---|---|
| Optimizer | AdamW | Dropout/MLP ratio/ Window Size | 0.1/ 0.4/ 8 | |
| Initial learning rate | CNN Branch init | ImageNet-pretrained (time) | ||
| Weight Decay | Transformer Branch init | Trained from scratch | ||
| LR Scheduler | ReduceLROnPlateau (Monitor val loss factor 0.5; Patience 3; min LR ) | Global Seed | 42 (per-fold, seed = 42 + k) | |
| Max Epochs | 100 | Cross-Validation | 3-fold, patient-grouped | |
| Early Stopping | Patience 10 epochs on validation Dice; Best-Dice weights restored | Mixed Precisions | Not used | |
| Batch Size | 3 | |||
| Input Resolution | 256 × 256 |
4. Results
4.1. Generalization Metrics
| Dice Per Fold | Intersection over Union Per Fold | |||||||
|---|---|---|---|---|---|---|---|---|
| Settings | Fold 1 | Fold 2 | Fold 3 | Average | Fold 1 | Fold 2 | Fold 3 | Average |
| UNet | 0.6769 | 0.6900 | 0.5777 | 0.6482+0.061 | 0.6564 | 0.6710 | 0.5599 | 0.6291+0.06 |
| DECTNet | 0.7520 | 0.7555 | 0.7419 | 0.7498+0.007 | 0.7352 | 0.7399 | 0.7266 | 0.7339+0.007 |
| nn-UNet | 0.7187 | 0.7131 | 0.7570 | 0.7296+0.024 | 0.6065 | 0.6014 | 0.6396 | 0.6158+0.021 |
| VGG-19+Swin | 0.8946 | 0.9028 | 0.9055 | 0.901+0.006 | 0.8766 | 0.8862 | 0.8889 | 0.8839+0.006 |
| ConvNeXT+CSwin | 0.909 | 0.9101 | 0.9083 | 0.9091+0.001 | 0.8928 | 0.8944 | 0.8924 | 0.8924+0.001 |
4.2. Segmentation Performance Metrics
| Precision Per Fold | F1 Score Per Fold | |||||||
|---|---|---|---|---|---|---|---|---|
| Settings | Fold 1 | Fold 2 | Fold 3 | Average | Fold 1 | Fold 2 | Fold 3 | Average |
| UNet | 0.6637 | 0.6787 | 0.56230 | 0.6349+0.063 | 0.6860 | 0.6978 | 0.5841 | 0.656+0.063 |
| DECTNet | 0.7659 | 0.7625 | 0.7460 | 0.7581+0.011 | 0.7694 | 0.7723 | 0.7612 | 0.7676+0.006 |
| nn-UNet | 0.7655 | 0.7365 | 0.8159 | 0.7726+004 | 0.7187 | 0.7131 | 0.7570 | 0.7296+0.024 |
| VGG-19+Swin | 0.9070 | 0.9188 | 0.9221 | 0.9163+007 | 0.8971 | 0.9066 | 0.9087 | 0.9041+0.006 |
| ConvNeXT+CSwin | 0.9377 | 0.9369 | 0.9311 | 0.9352+004 | 0.9451 | 0.9479 | 0.9523 | 0.9118+0.001 |
4.3. Boundary and Shape Metrics
4.4. Computational Efficiency
| Model | Parameters (M) | FLOPs (C) | Average Dice | Inference Time (sec) |
|---|---|---|---|---|
| Standard UNet | ~25 M | Scaled | ~0.65 | 0.052 sec |
| DECTNet | ~143 M | High | ~0.52 | 0.178 sec |
| nn-UNet | ~31.22 M | ~412.35 G Extremely High |
~0.73 | 0.165 sec |
| VGG-19+Swin | ~45 M | Scaled | ~0.90 | 0.222 sec |
| ConvNeXT Small+CSwin | ~29 M | Scaled | ~0.91 | 0.111 sec |
4.5. Qualitative Results

4. Discussion
5. Conclusions
Abbreviations
| FLLs | Focal Liver Lesions |
| HCC | Hepatocellular Carcinoma |
| ICC | Intrahepatic Cholangiocarcinoma |
| CT | Computed Tomography |
| VGG-19 | Visual Geometry Group—19 layers |
| MCT-LTDiag | Multi-phase CT Dataset for Liver Tumor Diagnosis |
| DECTNet | Dual Encoder Network |
| HD95 | 95th Percentile Hausdorff Distance |
| NSD | Normalized Surface Distance or Normalized Surface Dice |
| BH-FDR | Benjamini-Hochberg False Discovery Rate |
References
- Gul, S.; Khan, M.S.; Hossain, M.S.A.; Chowdhury, M.E.H.; and Sumon, M.S.I. A Comparative Study of Decoders for Liver and Tumor Segmentation Using a Self-ONN-Based Cascaded Framework. Diagnostics. MDPI. 2024. [CrossRef]
- Reyad, M.; Sarhan, A.M.; and Arafa, M. Architecture Optimization for Hybrid Deep Residual Networks in Liver Tumor Segmentation Using GA. International Journal of Computationa Intelligence System, Vol. 17 Art. 209. Springer Nature Link. 2024. [CrossRef]
- Debnath, R.K., Rahman, M.A., Azam, A, Jonkman, M. FSS-ULivR: A Clinically-inspired Few Shot Segmentation Framework for Liver Imaging Using Unified Representations and Attention Mechanism. Journal of Cancer Research and Clinical Oncology, Vol. 15 Issue 7, 215. National Library of Medicine. July 17, 2025.
- Ly, D.V.A.; Pham, T.T.H.; and Le, T.H. Comparative Study of UNet-based Architectures for Liver Tumor Segmentation in Multi-phase Contrast-Enhanced Computed Tomography. Computer Vision and Pattern Recognition. Computer Science. January 20, 2026.
- Gul, S.; Khan, M.S.; Bibi, A.; Khandakar, A.; Ayari, M.A.; and Chowdhury, M.E.H. Deep Learning Techniques for Liver and Liver Tumor Segmentation: A Review. Medicine. Vol. 147, 105620. Elsevier. August 2022. [CrossRef]
- Gao, F.; Hu, Z.; Xian, J.; and Lu, W. Mixed U-Net: Segmentation of Focal Liver Lesions Using a Hybrid 2D and 3D Model. Journal of Applied Clinical Medical Physics. Vol. 27 Issue 1. PubMed Central. December 28, 2025. [CrossRef]
- d’Albienzo, G.; Kamkova, Y.; Naseem, R.; Ullah, M.; Colonnese, S.; Cheik, F.A.; and Kumar, R.P. A Dual Encoder Concatenation Y-Shape Network for Precise Volumetric Liver and Lesion Segmentation. Computers in Biology and Medicine. Vol. 179, 108870. Elsevier. 2024. [CrossRef]
- Bilic, P.; Christ, P.; Li, H.B.; The Liver Tumor Segmentation Benchmark (LiTS) Medical Image Analysis. Vol. 84, 102680. Elsevier. 2023.
- Huang, W.; Liu, W.; Zhang, X.; Yin, X.; Han, X.; Li, C.; Gao, Y.; Shi, Y.; Lu, L.; Zhang, L.; Zhang, L.; and Yan, K. LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning. Lecture Notes in Computer Science. Springer Nature Link. 2024.
- Harini, G. and Karthika, R.; Liver and Liver Tumor Segmentation Using Modified Encoder Decoder Network and Find the Tumor Geometry. 2023 IEEE 4th Annual Flagship India Council International Subsections Conference (INDISCO). IEEE Explore. October 10, 2023.
- Lyu, P.; Liu, W.; Lin, T.; Zhang, J.; Liu, Y.; Wang, C.; and Zhu, J. Semi-Supervised Segmentation of Abdominal Organs and Liver Tumor: Uncertainty Rectified Curriculum Labeling Meets X-Fuse. Machine Learning, Science and Technology. May 23, 2024. [CrossRef]
- Li, B.; Xu, Y.; Wang, Y.; and Li, X. Accurate Semi-Supervised Medical Image Segmentation Using DECTNet Combined with DSST Framework. 2024 5th International Conference on Computer Vision, Image and Deep Learning (CVIDL). IEEE. 2024.
- Gul, S.; Khan, M.S.; Hossain, M.S.A.; Chowdurry, M.E.H.; and Sumon, M.S.I. A Comparative Study of Decoders for Liver and Tumor Segmentation Using a Self-ONN-Based Cascaded Framework. Diagnostics. MDPI. 2024.
- Almotairi, S.; Kareem, G.; Aouf, M.; Almutairi, B.; and Salem, M.A.-M. Liver Tumor Segmentation in CT Scans Using Modified SegNet, Sensors, 2020, 1516, MDPI. [CrossRef]
- Mrugesan, R.; Devaki, K. Liver Lesion Detection Using Semantic Segmentation and Chaotic Cuckoo Search Algorithm. Information Technology and Control Vol. 52 No. 3. ITC KTU. 2023. [CrossRef]
- Hussein, A.-J.; Disha, D.; Azhar, A.S.; Lamya, H.; and Malak, E.-A. A Review of Deep Learning Algorithms and Their Applications in Healthcare. Algorithms. Vol 15 2022.
- Lu, M.; Yaoyu, T.; and Sihang, B. Liver tumor segmentation based on 3D convolutional neural network with dual scale. PubMed 2020.
- Jiang, L.; R, Ou.; Y, Liu.; T, Zuo.; H, Xie, T.; Xiao, H.; and Bai, T. RMAU-Net: Residual Multi-Scale Attention U-Net For Liver and Tumor Segmentation in CT Images. Computers in Biology and Medicine, Elsevier, 2023.
- Xu, Y.; Cai, M.; Lin, L.; Zhang, Y.; Hu, H.; Peng, Z.; Zhang, Q.; Chen, Q.; Mao, X.; Iwamoto Y.; Han, X.-H.; Chen, Y.-W.; and Tong, R. PA-ResSeg: A Phase Attention Residual Network for Liver Tumor Segmentation from Multi-Phase CT Images. Medical Physics Vol. 48 Issue 7 pp. 3572-3766, PubMed. [CrossRef]
- Nabizadeh, N.; Dorodchi, M.; and Sihang, B. Automatic tumor lesion detection and segmentation using modified winnow algorithm. IEEE. 2015.
- Zhu, S.; Zou, M.; Wu, Q.; Go, Z.; Huang, Z.; Zou, Y.; Tan, T.; You, Y.; Dong, X.; and Lou, H. STD-Net: A Spatio-Temporal Decoupling Network for Multiphasic Liver Lesion Segmentation and Characterization. NPJ Digital Medicine Article Number 13 (2026), Nature Careers. [CrossRef]
- Zhang, C.; Wang, L.; Zhang, C.; Zhang, Y.; Li, J.; and Wang, P. Liver Tumor Segmentation Based on Multi-Scale Deformable Feature Fusion and Global Context Awareness. Biomimetics Vol. 10 Issue 9. MDPI. 2025. [CrossRef]
- Yang, Z.; and Li, S. “Dual-Path Network for Liver and Tumor Segmentation in CT Images Using Swin Transformer Encoding Approach,” Current Medical Imaging Vol. 19 Issue 10 pp. 1114-1123 (2023), PubMed. [CrossRef]
- Wang, D.; Sun, Y.; Chen, H.; and Zhao, X. Image Segmentation Network Based on Enhanced Dual Encoder. Scientific Data Article Number 35983. Nature. 2025. [CrossRef]
- Skorupko, G.; Avgoustidis, F.; Martin-Isla, C.; Garrucho, L.; Kessler, D.A.; Pujadas, E.r.; Diaz, O.; Bobowicz, M.; Gwozdziewicz, K.; Bargallo, X.; Jarusevoicius, P.; Osula, R.; Kushibar, K.; and Ledakir, K. Federated nnU-Net for privacy-preserving medical image segmentation. Scientific Reports. Vol 15 (38312). PubMed Central. November 2025. [CrossRef]
- Isensee, F.; Petersen, J.; Klein, A.; Zimmerer, D.; Jaeger, P.F.; Kohl .; Wassertal, J.; Koehler, G.; Narajitra, T.; Wirkert, S.; and Maiser-Hein, K.H. nnU-Net: Self-Adapting Framework for U-Net-Based Medical Image Segmentation. Bildverarbeitung fur die Medizine 2019, February 7, 2019, Springer Nature Link.
- Elbatel, M.; Ghonim, M.; Mao, J.; Lin, Z.; Ecstein, K.; Mora, A.M.; Deissler, J.; and Li, X. TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT. TriaALS 2026. May 2026.
- Li, B.; Xu, Y.; Wang, Y.; and Zhang, B. DECTNet: Dual Encoder Network Combined Convolution and Transformer Architecture for Medical Imaging Segmentation. PlosOne. April 4, 2024. [CrossRef]
- Wu, X.; Su, H.; Hua, Y.; Xu, Y.; Wang, L.; Wang, X.; Wang, S.; Jin, B.; Liu, X.; Wan, X.; Sun, Q.; Wang, X.; and Du, S. A Multi-phase CT Dataset for Automated Differential Diagnosis of Liver Tumors. Scientific Data, Data Descriptors. Vol. 13 Article 31. Springer Nature. December, 2025. [CrossRef]
- Li, H.; Hu, D.; Liu, H.; Wang, J.; and Oguz, I. CATS: Complementary CNN and Transformer Encoders for Segmentation. 2022 IEEE 19th InternationalSymposium on Biomendical Imaging (ISBI). IEEE Explore. March 2022.

| Accuracy Per Fold | Sensitivity Per Fold | |||||||
|---|---|---|---|---|---|---|---|---|
| Settings | Fold 1 | Fold 2 | Fold 3 | Average | Fold 1 | Fold 2 | Fold 3 | Average |
| UNet | 0.9958 | 0.9954 | 0.9900 | 0.9937+0.003 | 0.9762 | 0.9749 | 0.9836 | 0.9782+0.005 |
| DECTNet | 0.9978 | 0.9967 | 0.9958 | 0.9966+0.001 | 0.9365 | 0.9439 | 0.9406 | 0.9403+0.004 |
| nn-UNet | 0.9991 | 0.9991 | 0.9992 | 0.9991+0 | 0.7352 | 0.7413 | 0.7512 | 0.7426+0.008 |
| VGG-19+Swin | 0.9986 | 0.9985 | 0.9988 | 0.9986+0 | 0.9600 | 0.9563 | 0.9576 | 0.958+0.002 |
| ConvNeXT+CSwin | 0.9988 | 0.9987 | 0.9990 | 0.9988+0 | 0.9451 | 0.9479 | 0.9523 | 0.9484+0.004 |
| a) Dice | b) IoU | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet |
| UNet | 0.035 | 0.036 | 0.267 | 0.120 | 0.029 | 0.029 | 0.804 | 0.029 |
| DECTNet | 0.005 | 0.009 | 0.375 | 0.004 | 0.005 | 0.029 | ||
| nn-UNet | 0.016 | 0.016 | 0.005 | 0.005 | ||||
| VGG-19+Swin | 0.171 | 0.143 | ||||||
| a) Accuracy | b) Sensitivity | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet |
| UNet | 0.143 | 0.143 | 0.143 | 0.149 | 0.029 | 0.029 | 0.001 | 0.015 |
| DECTNet | 0.143 | 0.143 | 0.143 | 0.076 | 0.040 | 0.002 | ||
| nn-UNet | 0.117 | 0.065 | 0.001 | 0.002 | ||||
| VGG-19+Swin | 0.000 | 0.078 | ||||||
| a) Precision | b) F1 Score | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet |
| UNet | 0.038 | 0.038 | 0.162 | 0.081 | 0.038 | 0.038 | 0.278 | 0.109 |
| DECTNet | 0.005 | 0.019 | 0.663 | 0.004 | 0.009 | 0.184 | ||
| nn-UNet | 0.038 | 0.038 | 0.015 | 0.015 | ||||
| VGG-19+Swin | 0.110 | 0.184 | ||||||
| a) HD95 | b) NSD | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet | ConvNeXT Small | VGG-19 | nn-UNet | DECTNet |
| UNet | 0.072 | 0.077 | 0.709 | 0.267 | 0.029 | 0.038 | 0.038 | 0.081 |
| DECTNet | 0.060 | 0.070 | 0.442 | 0.056 | 0.093 | 0.056 | ||
| nn-UNet | 0.060 | 0.060 | 0.326 | 0.056 | ||||
| VGG-19+Swin | 0.155 | 0.056 | ||||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).