Submitted:
22 December 2025
Posted:
23 December 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction

- We propose GCRoCL, a novel Graph-Enhanced Cross-domain Robust Contrastive Learning framework, which integrates dynamic graph construction, hybrid graph-temporal encoding, and adversarial training to learn noise-robust and domain-invariant representations for medical time series data.
- We introduce dynamic graph construction and graph-level data augmentation techniques that effectively capture complex topological relationships within medical time series and enhance the model’s resilience to various noise sources and data incompleteness.
- We incorporate a Cross-domain Adversarial Contrastive Learner that explicitly mitigates domain shift by aligning feature distributions across different medical domains, thereby significantly improving the model’s cross-domain generalization ability, particularly in scenarios with limited labeled data.
2. Related Work
2.1. Self-Supervised Representation Learning for Medical Time Series
2.2. Graph Neural Networks and Domain Generalization in Medical AI
2.3. Graph Neural Networks and Graph Representation Learning
2.4. Domain Generalization and Adaptation
2.5. Applications and Challenges in Medical AI
3. Method

3.1. Graph Construction & Augmentation Module
3.1.1. Dynamic Graph Construction
3.1.2. Graph-Level Data Augmentation
3.2. Graph-Temporal Feature Encoder
3.2.1. Spatial Feature Extraction (GCN/GAT)
3.2.2. Temporal Feature Extraction (Transformer)
3.3. Cross-Domain Adversarial Contrastive Learner
3.3.1. Instance-Level Contrastive Loss
3.3.2. Domain Adversarial Alignment (DAA) Module
3.3.3. Overall Objective Function
3.4. Downstream Classifier Fine-Tuning
4. Experiments
4.1. Experimental Setup
4.1.1. Datasets
- AD (Alzheimer’s Disease) EEG Dataset: This dataset comprises electroencephalogram (EEG) recordings from subjects with Alzheimer’s disease and healthy controls. It is particularly challenging due to inherent EEG noise and the subtle nature of disease biomarkers. We used the same source as referenced in previous studies for consistency.
- TDBrain Dataset: This dataset includes brain-derived neuro-signal data, often used for brain-computer interface (BCI) applications or neurological disorder classification tasks. It exhibits complex spatio-temporal patterns and varying signal characteristics across subjects. We matched the data source with prior relevant works.
- PTB-XL / PTB (PhysioBank) ECG Datasets: These comprehensive electrocardiogram (ECG) datasets contain recordings from patients with various cardiac conditions. They are known for their diversity in signal morphology, recording equipment, and patient demographics, making them ideal for cross-domain generalization studies.
4.1.2. Data Preprocessing
- Standardization: Each channel’s signal was standardized to have zero mean and unit variance.
- Filtering: A band-pass filter (e.g., 0.5-30Hz for EEG, 0.5-100Hz for ECG) was applied to remove drift, power-line noise, and high-frequency muscle artifacts.
- Resampling: All signals were resampled to a unified sampling rate to ensure consistency across different recordings and datasets.
- Segmentation: Long continuous recordings were segmented into fixed-length time windows (e.g., 2-5 seconds) suitable for batch processing and feature extraction.
4.1.3. Training Phases
- GCRoCL Pre-training: In this self-supervised stage, the Graph-Temporal Feature Encoder and the Domain Classifier were trained on large volumes of unlabeled or partially labeled medical time series data. The objective functions, as defined in Equation (9), optimize for both instance-level discrimination through InfoNCE loss (Equation (4)) and domain invariance via the DAA module (Equations 7 and (8)). This stage aims to learn robust and domain-invariant representations.
- Fine-tuning: Following pre-training, the DAA module was removed. A lightweight linear classifier or a small multi-layer perceptron was appended to the pre-trained Graph-Temporal Feature Encoder. This composite model was then fine-tuned on target-specific, labeled data using a standard cross-entropy loss. We evaluated fine-tuning performance using varying proportions of labeled data (e.g., 100% for full fine-tuning, and 10% for low-data regimes) to simulate real-world scenarios where labeled data is scarce.
4.1.4. Baseline Models
- TS2vec: A universal framework for time series representation learning via contrastive learning.
- TF-C: A transformer-based contrastive learning approach for time series.
- Mixing-up: A method that leverages mixing strategies for data augmentation in self-supervised time series learning.
- TS-TCC: A temporal and channel-wise contrastive learning framework for physiological signals.
- SimCLR: A well-known contrastive learning framework adapted for time series data.
- DAAC: A discrepancy-aware adaptive contrastive learning method.
4.1.5. Evaluation Metrics
4.2. Quantitative Results
4.2.1. Performance on AD Dataset with Full Fine-Tuning
4.2.2. Cross-Domain Generalization Performance
4.2.3. Performance Under Limited Labeled Data
4.3. Ablation Studies
- GCRoCL w/o GE & DAA (Base CL): This configuration removes both the Graph Construction & Augmentation Module (GE) and the Domain Adversarial Alignment (DAA) module. The model then functions as a standard contrastive learning framework with only the temporal encoder (e.g., Transformer).
- GCRoCL w/o Graph Enhancement: In this setup, the DAA module is retained, but the graph construction and graph-level augmentations are removed. The encoder processes raw time series directly with a temporal model.
- GCRoCL w/o Domain Adversarial Alignment: This variant includes the Graph Construction & Augmentation Module but disables the DAA module. The encoder learns graph-enhanced features with contrastive loss but without explicit domain-invariant objectives.
- GCRoCL (Full): Our complete proposed framework, integrating all modules.
4.3.1. Impact of Graph Module Design Choices
4.4. Hyperparameter Sensitivity Analysis
4.4.1. Sensitivity to Temperature Parameter
4.4.2. Sensitivity to Gradient Reversal Strength
4.5. Human Evaluation Results
5. Conclusions
References
- Dou, Y.; Forbes, M.; Koncel-Kedziorski, R.; Smith, N.A.; Choi, Y. Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2022, pp. 7250–7274. [CrossRef]
- Zhang, W.; Stratos, K. Understanding Hard Negatives in Noise Contrastive Estimation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 1090–1101. [CrossRef]
- Wang, B.; Lapata, M.; Titov, I. Meta-Learning for Domain Generalization in Semantic Parsing. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 366–379. [CrossRef]
- Ye, Q.; Lin, B.Y.; Ren, X. CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021, pp. 7163–7189. [CrossRef]
- Tang, Y.; Gong, H.; Dong, N.; Wang, C.; Hsu, W.N.; Gu, J.; Baevski, A.; Li, X.; Mohamed, A.; Auli, M.; et al. Unified Speech-Text Pre-training for Speech Translation and Recognition. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2022, pp. 1488–1499. [CrossRef]
- Kim, T.; Yoo, K.M.; Lee, S.g. Self-Guided Contrastive Learning for BERT Sentence Representations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 2528–2540. [CrossRef]
- Wu, C.; Wu, F.; Huang, Y. DA-Transformer: Distance-aware Transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 2059–2068. [CrossRef]
- Yan, Y.; Li, R.; Wang, S.; Zhang, F.; Wu, W.; Xu, W. ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 5065–5075. [CrossRef]
- Zhou, Y.; Li, X.; Wang, Q.; Shen, J. Visual In-Context Learning for Large Vision-Language Models. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024. Association for Computational Linguistics, 2024, pp. 15890–15902.
- Zhou, Y.; Zhang, J.; Chen, G.; Shen, J.; Cheng, Y. Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models, 2024.
- Liu, Y.; Yu, R.; Yin, F.; Zhao, X.; Zhao, W.; Xia, W.; Yang, Y. Learning quality-aware dynamic memory for video object segmentation. In Proceedings of the European Conference on Computer Vision. Springer, 2022, pp. 468–486.
- Liu, Y.; Bai, S.; Li, G.; Wang, Y.; Tang, Y. Open-vocabulary segmentation with semantic-assisted calibration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3491–3500.
- Han, K.; Liu, Y.; Liew, J.H.; Ding, H.; Liu, J.; Wang, Y.; Tang, Y.; Yang, Y.; Feng, J.; Zhao, Y.; et al. Global knowledge calibration for fast open-vocabulary segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 797–807.
- Zhang, F.; Cheng, Z.; Deng, C.; Li, H.; Lian, Z.; Chen, Q.; Liu, H.; Wang, W.; Zhang, Y.F.; Zhang, R.; et al. Mme-emotion: A holistic evaluation benchmark for emotional intelligence in multimodal large language models. arXiv preprint arXiv:2508.09210 2025. [CrossRef]
- Zhang, F.; Li, H.; Qian, S.; Wang, X.; Lian, Z.; Wu, H.; Zhu, Z.; Gao, Y.; Li, Q.; Zheng, Y.; et al. Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond. arXiv preprint arXiv:2511.00389 2025. [CrossRef]
- Liu, F.; Shareghi, E.; Meng, Z.; Basaldella, M.; Collier, N. Self-Alignment Pretraining for Biomedical Entity Representations. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 4228–4238. [CrossRef]
- Yan, A.; He, Z.; Lu, X.; Du, J.; Chang, E.; Gentili, A.; McAuley, J.; Hsu, C.N. Weakly Supervised Contrastive Learning for Chest X-Ray Report Generation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics, 2021, pp. 4009–4015. [CrossRef]
- Xiong, G.; Jin, Q.; Lu, Z.; Zhang, A. Benchmarking Retrieval-Augmented Generation for Medicine. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics, 2024, pp. 6233–6251. [CrossRef]
- Yang, J.; Yu, Y.; Niu, D.; Guo, W.; Xu, Y. ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2023, pp. 7617–7630. [CrossRef]
- Chao, L.; He, J.; Wang, T.; Chu, W. PairRE: Knowledge Graph Embeddings via Paired Relation Vectors. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, 2021, pp. 4360–4369. [CrossRef]
- Saxena, A.; Kochsiek, A.; Gemulla, R. Sequence-to-Sequence Knowledge Graph Completion and Question Answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2022, pp. 2814–2828. [CrossRef]
- Hardalov, M.; Arora, A.; Nakov, P.; Augenstein, I. Cross-Domain Label-Adaptive Stance Detection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021, pp. 9011–9028. [CrossRef]
- Han, C.; Fan, Z.; Zhang, D.; Qiu, M.; Gao, M.; Zhou, A. Meta-Learning Adversarial Domain Adaptation Network for Few-Shot Text Classification. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, 2021, pp. 1664–1673. [CrossRef]
- Zhou, Y.; Geng, X.; Shen, T.; Zhang, W.; Jiang, D. Improving Zero-Shot Cross-lingual Transfer for Multilingual Question Answering over Knowledge Graph. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 5822–5834. [CrossRef]
- Zheng, L.; Tian, Z.; He, Y.; Liu, S.; Chen, H.; Yuan, F.; Peng, Y. Enhanced mean field game for interactive decision-making with varied stylish multi-vehicles. arXiv preprint arXiv:2509.00981 2025. [CrossRef]
- Lin, Z.; Tian, Z.; Lan, J.; Zhao, D.; Wei, C. Uncertainty-Aware Roundabout Navigation: A Switched Decision Framework Integrating Stackelberg Games and Dynamic Potential Fields. IEEE Transactions on Vehicular Technology 2025, pp. 1–13. [CrossRef]
- Tian, Z.; Lin, Z.; Zhao, D.; Zhao, W.; Flynn, D.; Ansari, S.; Wei, C. Evaluating scenario-based decision-making for interactive autonomous driving using rational criteria: A survey. arXiv preprint arXiv:2501.01886 2025. [CrossRef]
- Zhang, H.; Lu, J.; Jiang, S.; Zhu, C.; Xie, L.; Zhong, C.; Chen, H.; Zhu, Y.; Du, Y.; Gao, Y.; et al. Co-sight: Enhancing llm-based agents via conflict-aware meta-verification and trustworthy reasoning with structured facts. arXiv preprint arXiv:2510.21557 2025.
- Wu, Y.; Zhan, P.; Zhang, Y.; Wang, L.; Xu, Z. Multimodal Fusion with Co-Attention Networks for Fake News Detection. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, 2021, pp. 2560–2569. [CrossRef]
- Devaraj, A.; Marshall, I.; Wallace, B.; Li, J.J. Paragraph-level Simplification of Medical Texts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 2021, pp. 4972–4984. [CrossRef]
- Zhou, Y.; Zheng, H.; Chen, D.; Yang, H.; Han, W.; Shen, J. Reasoning as the Engine: The Evolution from Medical LLMs to Versatile Medical Agents. In Proceedings of the OpenReview, 2025.
- Zhang, F.; Cheng, Z.Q.; Zhao, J.; Peng, X.; Li, X. LEAF: unveiling two sides of the same coin in semi-supervised facial expression recognition. Computer Vision and Image Understanding 2025, p. 104451. [CrossRef]


| Model | Accuracy | Precision | Recall | F1 score | AUROC | AUPRC |
| TS2vec | 81.26 ± 2.08 | 81.21 ± 2.14 | 81.34 ± 2.04 | 81.12 ± 2.06 | 89.20 ± 1.76 | 88.94 ± 1.85 |
| TF-C | 75.31 ± 8.27 | 75.87 ± 8.73 | 74.83 ± 8.98 | 74.54 ± 8.85 | 79.45 ± 10.23 | 79.33 ± 10.57 |
| Mixing-up | 65.68 ± 7.89 | 72.61 ± 4.21 | 68.25 ± 6.97 | 63.98 ± 9.92 | 84.63 ± 5.04 | 83.46 ± 5.48 |
| TS-TCC | 73.55 ± 10.00 | 77.22 ± 6.13 | 73.83 ± 9.65 | 71.86 ± 11.59 | 86.17 ± 5.11 | 85.73 ± 5.11 |
| SimCLR | 54.77 ± 1.97 | 50.15 ± 7.02 | 50.58 ± 1.92 | 43.18 ± 4.27 | 50.15 ± 7.02 | 50.42 ± 1.06 |
| Ours (GCRoCL) | 82.55 ± 1.95 | 82.48 ± 2.01 | 82.60 ± 1.90 | 82.40 ± 1.98 | 90.30 ± 1.65 | 90.05 ± 1.70 |
| Model | Src. Dom. | Tgt. Dom. | AUROC |
| TS2vec | D1+D2 | D3 | 82.10 ± 1.85 |
| TS2vec | D1+D2 | D4 | 82.55 ± 1.90 |
| TF-C | D1+D2 | D3 | 78.40 ± 2.10 |
| TF-C | D1+D2 | D4 | 79.15 ± 2.05 |
| Mixing-up | D1+D2 | D3 | 75.60 ± 2.30 |
| Mixing-up | D1+D2 | D4 | 76.25 ± 2.25 |
| TS-TCC | D1+D2 | D3 | 80.90 ± 1.95 |
| TS-TCC | D1+D2 | D4 | 81.30 ± 2.00 |
| SimCLR | D1+D2 | D3 | 68.20 ± 2.50 |
| SimCLR | D1+D2 | D4 | 69.10 ± 2.45 |
| Ours (GCRoCL) | D1+D2 | D3 | 88.75 ± 1.50 |
| Ours (GCRoCL) | D1+D2 | D4 | 89.10 ± 1.40 |
| Model | 1% Labeled Data | 5% Labeled Data | 10% Labeled Data | 25% Labeled Data |
| TS2vec | 68.50 ± 2.50 | 75.10 ± 2.15 | 78.90 ± 1.90 | 80.80 ± 1.80 |
| TF-C | 62.30 ± 3.10 | 69.80 ± 2.80 | 73.50 ± 2.50 | 76.80 ± 2.20 |
| TS-TCC | 65.80 ± 2.60 | 72.40 ± 2.30 | 76.10 ± 2.00 | 79.20 ± 1.90 |
| SimCLR | 55.10 ± 3.50 | 60.50 ± 3.00 | 65.20 ± 2.80 | 70.10 ± 2.50 |
| Ours (GCRoCL) | 75.20 ± 2.10 | 80.50 ± 1.80 | 83.10 ± 1.70 | 85.50 ± 1.60 |
| Model Configuration | Accuracy | AUROC |
| GCRoCL w/o GE & DAA (Base CL) | 75.12 ± 2.50 | 82.35 ± 2.10 |
| GCRoCL w/o Graph Enhancement | 78.45 ± 2.20 | 85.10 ± 1.95 |
| GCRoCL w/o Domain Adversarial Alignment | 79.88 ± 2.15 | 87.20 ± 1.80 |
| GCRoCL (Full) | 82.55 ± 1.95 | 90.30 ± 1.65 |
| Graph Construction | Graph Augmentation | AUROC |
| Pearson Correlation | Full Suite | 88.50 ± 1.75 |
| DTW | Only NF_Pert. | 89.05 ± 1.75 |
| DTW | Full Suite | 89.95 ± 1.70 |
| Model | Avg. Clinician Confidence (1-5) | Agreement with Expert Diagnosis (%) |
| SVM (Hand-crafted features) | 2.8 ± 0.4 | 68.5 ± 3.2 |
| TS2vec | 3.5 ± 0.3 | 78.9 ± 2.8 |
| GCRoCL | 4.2 ± 0.2 | 85.1 ± 2.0 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).