Submitted:
16 June 2025
Posted:
17 June 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Focused Temporal Analysis of rPPG: Establishes foundational insights into the capabilities and limitations of using exclusively temporal physiological information, providing a rigorous benchmark.
- Multi-scale Temporal Dynamics Encoder (MTDE): Effectively captures physiologically meaningful ANS responses across multiple timescales, addressing complexity in subtle temporal emotional signals.
- Adaptive Sparse Attention: Precisely identifies transient, emotionally relevant physiological segments amidst noisy rPPG data, significantly enhancing robustness.
- Gated Temporal Pooling: Sophisticatedly aggregates emotional information across temporal chunks, effectively mitigating noise and irrelevant features.
- Curriculum Learning Strategy: Systematically addresses learning complexities associated with weak labels, noise, and temporal sparsity, ensuring robust, stable model learning.
2. Related Work
2.1. The Physiological Signals for Emotion Recognition
2.2. Remote PPG Signal Extraction and Denoising
2.3. Emotion Recognition from rPPG/PPG
- CNN-LSTM. Mellouk & Handouzi [23] combine 2-D convolutions with an LSTM on four-second segments and report the first subject-balanced results on MAHNOB-HCI. The network architecture and training protocol are fully specified, making it the most appropriate domain-equivalent baseline.
2.4. Comparison and Key Differences
- Temporal-only perspective. We analyse the BVP waveform in isolation, establishing a lower-bound benchmark that future multimodal extensions can build upon.
- Noise-robust aggregation. GatedPooling jointly weights time steps and feature channels, reducing artefacts propagated from the rPPG extractor.
- Rigorous protocol and efficiency. A dedicated hold-out test set, weighted-F1 scoring, and a 198 k-parameter model that runs at 0.66 s per two-minute video establish(Section 3.7) both statistical and practical credibility.
3. Methodology
3.1. Dataset and Experimental Protocol
3.1.1. MAHNOB-HCI Dataset
3.1.2. Data Partitioning and Protocol Rationale
3.1.3. A Two-Stage Process for Optimal Method Selection
3.1.4. Final Method and Implementation
3.2. Overall Framework
3.3. Training Modules
3.3.1. rPPG Extraction Front-End (PhysMamba)
3.3.2. Multi-Scale Temporal Dynamics Encoder (MTDE)
- Short-scale branch (RF0.2 s): This scale is sensitive to rapid, beat-to-beat changes in pulse morphology, reflecting high-frequency dynamics driven by parasympathetic (vagal) fluctuations, analogous to the HF-likeband of HRV [4].
- Medium-scale branch (RF2.2 s): This branch targets slower oscillations from the interplay between sympathetic and parasympathetic systems (e.g., baroreflex), operating in a window comparable to the LF-like band [4].
- Long-scale branch (RF4.3 s): This branch integrates very-slow modulations across the entire chunk, reflecting hormonal or thermoregulatory influences, which are thought to contribute to the VLF-like band [41]. Here, as the chunk is limited to 128 frames, the three RFs cover only the order-of-magnitude ranges rather than exact periods.
3.3.3. AttnScorer
3.3.5. GatedPooling
3.3.4. Auxiliary Components
3.3.6. Main Classifier
3.4. Training Curriculum
3.4.1. Phase 0 (Epochs 0–14): Exploration and Representation Learning.
- Cognitive/Computational Link: This initial phase corresponds to the “pervasive exploration” observed in human learning, where a system first captures a wide array of sensory inputs to build a general understanding of the feature space before focusing on specific tasks. This aligns with classic models of attention where an initial, broad orientation precedes focused analysis [49].
- Objective: The primary objective is to train the MTDE and related components (AttnScorer and ChunkProjection) to produce robust and diverse embedding representations for individual temporal chunks. During this phase, the GatedPooling module and the Main Classifier are not used for the primary loss computation.
- Primary Losses: The total loss in Phase 0 is a combination of the Supervised Contrastive Loss and an Entropy Regularization Loss. The Supervised Contrastive Loss () [47] is applied to the normalized embeddings from the ChunkProjection. This loss encourages embeddings from chunks originating from the same session (sharing the same label) to be closer in the representation space, while pushing embeddings from different sessions apart.
3.4.2. Phase 1 (Epochs 15–29): Chunk-Level Discrimination and Attentional Refinement.
- Cognitive/Computational Link: This second phase corresponds to the establishment of selective attention. After the initial exploration, learning resources are focused on task-relevant signals. By using the AttnScorer and Focal Loss on only the Top-K chunks, our model progressively narrows its attentional focus, mimicking how neural systems learn to prioritize information-rich stimuli over time [51].
- Objective: To significantly enhance the discriminative capacity of the individual chunk embeddings and to refine the AttnScorer’s ability to identify emotionally salient temporal segments. During this phase, the ChunkProjection module is frozen, while the MTDE, AttnScorer, and ChunkAuxClassifier are actively trained.
- Primary Losses: The main objective is driven by a chunk-level classification task using the ChunkAuxClassifier. To handle class imbalance and focus on challenging examples, we employ Focal Loss [52] with (Appendix C):
3.4.3. Phase 2 (Epochs ≥ 30): Session-Level Exploitation and Fine-Tuning
- Objective: To optimize the entire end-to-end pipeline for the final session-level emotion recognition task. The ChunkAuxClassifier is removed, and the Main Classifier is initialized with its weights. The MTDE, AttnScorer, GatedPooling, and Main Classifier are all fine-tuned.
- Primary Loss Function: The sole objective function in this phase is the Session-level Cross-Entropy Loss () applied to the output of the Main Classifier based on the GatedPooling session embedding.
3.5. Evaluation Metrics
- Accuracy: Defined as the proportion of correctly classified sessions out of the total number of sessions in the test set:
- Weighted F1-score: This metric is calculated based on the Precision (), Recall (), and F1-score () for each individual class . The formulas for these class-specific metrics are:
3.6. Baseline
3.7. Experimental Setup
4. Results
4.1. Main Results
- Accuracy: 66.04% vs. 61.31%
- Positive-class F1-score: 74.29% vs. 50.96%
- Weighted F1-score: 61.97% vs. 59.46%
- Physiological limitations: Arousal is closely associated with sympathetic nervous system (SNS) activity, which often manifests as acute, high-magnitude changes in cardiovascular patterns [56]. In contrast, valence is linked to more complex and subtle interactions involving the parasympathetic nervous system (PNS) and cortical patterns, which are not as easily captured in peripheral signals [4,57].
- Modality Constraints: By spatially averaging the facial region into a single 1D signal, our approach cannot access fine-grained spatial information. This is a notable limitation, as recent studies suggest that regional facial blood flow patterns, such as subtle temperature changes around key facial areas, may hold valuable cues related to emotional valence [58].
- Architectural Inductive Bias: Our framework’s event-driven design, particularly its reliance on sparse attention to identify transient, high-magnitude physiological events, introduces an inductive bias. This approach is highly effective for detecting the ‘spikes’ characteristic of SNS-driven arousal responses but is inherently less suited to integrating the subtle, sustained, and context-dependent patterns that often characterize valence.
- Data Imbalance: A notable class imbalance in valence labels may also contribute to biased predictions.
4.2. Ablation Studies
4.2.1. Impact of Temporal Aggregation Strategy
4.2.2. Validation of the Multi-Scale Architecture in MTDE
4.2.3. Ablation Study on Pooling and Attention Mechanisms (Arousal)
4.3. Qualitative Analysis of Sparse Attention
5. Discussion
5.1. Limitations
5.2. Future Work
6. Conclusions
7. Patents
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| Acc | Accuracy |
| ANS | Autonomic Nervous System |
| AttnScorer | Attention Scorer |
| BVP | Blood Volume Pulse |
| CE | Confusion Matrix |
| CM | Convolutional Neural Network |
| EDA | Electrodermal Activity |
| HRV | Heart Rate Variability |
| MIL | Multiple Instance Learning |
| MTDE | Multi-scale Temporal Dynamics Encoder |
| rPPG | remote Photoplethysmography |
| WF1 | Weighted F1-score |
Appendix A. Architecture Details of MTDE
A.1. Architecture Overview
- SlimStem: Two Conv1D layers (kernel=5, then 3 with stride=2) for initial noise reduction and temporal downsampling (T=128 → 64).
- MultiScaleTemporalBlock (MSTB): Three parallel branches with different kernel sizes and dilations to model short-, mid-, and long-range temporal dynamics.
A.2. Physiological Rationale & Receptive Fields
| Branch | Kernel | Dilation | Effective RF | Approx. Duration | Physiological Role |
|---|---|---|---|---|---|
| Short | 3 | 3 | 6 | ∼0.2 s | Parasympathetic (vagal) fluctuations, analogous to HF band of HRV. |
| Medium | 5 | 8 | 66 | ∼2.2 s | Sympathetic-parasympathetic interplay (e.g., baroreflex), analogous to LF band. |
| Long | 3 | 32 | 129 | ∼4.3 s | Very-slow modulations (e.g., hormonal, thermoregulatory), analogous to VLF band. |
A.3. Pooling Layer: SoftMax Pool
Appendix B. Attention Modules: AttnScorer and GatedPooling
B.1. AttnScorer: Phase-Aware Attention Scoring
B.1.1. Architecture
- 2-layer MLP: Linear (D, D/2) → GELU → Linear (D/2, 1)
- Zero-mean score normalization:
- scaling with EMA-based adjustment:
- Raw score scaling:
B.1.2. Phase-Dependent Attention Mechanism
| Phase | Epoch Range | Attention Type | Notes |
|---|---|---|---|
| 0 | 0–14 | Softmax | Entourage diversity |
| 1 | 15–29 | -Entmax () | Sharp, sparse, differentiable Top-K |
| 2 | ≥30 | Raw scores only | Passed to GatedPooling |
B.1.3. Entmax Scheduling During Phase 1
B.2. GatedPooling: Sparse Temporal Aggregation
Appendix C. Phase-Wise Training Schedule
| Phase | Epochs | Objective | Active Modules |
|---|---|---|---|
| 0 | 0–14 | Embedding diversity (SupCon) | MTDE, AttnScorer, ChunkProjection |
| 1 | 15–29 | Chunk-level discrimination | MTDE, AttnScorer, ChunkAuxClassifier, GatedPooling✓ (E≥25) |
| 2 | 30–49 | Session-level classification | GatedPooling, Classifier |
| Epoch | Top-K Ratio | (Temp) | (AttnScorer) | (Gated) | |||
|---|---|---|---|---|---|---|---|
| 0 | 0.0 | 1.00 | 0.0 | 0.1 | 1.2 | — | — |
| 14 | 0.0 | 0.44 | 0.0 | 0.1 | 0.7 | — | — |
| 15 | 0.6 | 0.0 | 0.5 | 0.0 | 1.0 | 1.0 | — |
| 25 | 0.3 | 0.0 | 0.5 | 0.0 | 1.0 | 1.7 | start = 1.7 |
| 30 | — | 0.0 | 0.7 | 0.0 | 1.0 | raw only | 1.7 → 2.0 |
| 50 | — | 0.0 | 1.0 | 0.0 | 1.0 | raw only | 2.0 |
Appendix D. Contusion Matrices
| Predicted Low | Predicted High | |
|---|---|---|
| Actual Low | 9 | 18 |
| Actual High | 0 | 26 |
- Accuracy: 66.04%
- Weighted F1-score: 61.97%
- The model shows strong sensitivity to high arousal states (recall: 100%), with most misclassifications occurring in the low-arousal category.
| Predicted Low | Predicted High | |
|---|---|---|
| Actual Low | 13 | 10 |
| Actual High | 10 | 20 |
- Accuracy: 62.26%
- Weighted F1-score: 62.26%
- The model demonstrates relatively balanced performance but reveals confusion between low and high valence categories, indicating the nuanced nature of valence detection from unimodal temporal signals.
Appendix E. Contusion Matrices
E.1. Description of Comparative Models
- Temporal Convolutional Network (TCN): A standard TCN with a stack of dilated causal convolutions¹ was implemented to serve as a strong, non-recurrent backbone baseline. The TCN consisted of two residual blocks, each with two layers of dilated causal convolutions (kernel size=7, 256 channels), and dilation factors following an exponential growth (1, 2, 4, 8).
- 1D-CNN + LSTM: A classic hybrid architecture, combining a simple 1D-CNN feature extractor with an LSTM layer to model sequential dependencies.
- Multi-scale + Bi-LSTM: This model replaces the final attention pooling layer of our MTDE with a Bi-directional LSTM to evaluate a standard recurrent aggregation method.
- Multi-scale + SE block: This model integrates Squeeze-and-Excitation (SE) blocks into our MTDE to assess the effect of a channel-wise attention mechanism as an alternative to our temporal attention.
E.2. Result and Analysis
| Comparison Focus | Backbone / Variant | Accuracy (%) | Weighted F1 (%) |
|---|---|---|---|
| Ours (3-branch MTDE)* | Ours (MTDE)* | 66.04* | 61.97* |
| Backbone Comparison | Standard TCN | 58.49 | 52.05 |
| 1D-CNN + LSTM | 58.49 | 52.69 | |
| Aggregation Comparison | Multi-scale + Bi-LSTM | 52.83 | 46.34 |
| Attention Comparison | Multi-scale + SE block | 60.38 | 52.50 |
- Backbone Analysis (MTDE vs. TCN, 1D-CNN+LSTM): The proposed MTDE outperforms both the standard TCN and the classic hybrid model. This suggests that for this specific task, the physiologically-aligned receptive fields of the MTDE are more effective at capturing the relevant, multi-scale dynamics of the BVP waveform than a generic temporal model like TCN. The relatively lower performance of the LSTM-based variant may be attributed to the nature of our input data. Recurrent models are designed for long-range dependencies, but in our framework which processes short, 4-second chunks, the most critical emotional cues are embedded in the local waveform morphology. In such cases, the added complexity of a recurrent layer may not provide a significant advantage.
- Aggregation & Attention Analysis (vs. Bi-LSTM, SE block): Our model also outperforms the variants with a Bi-LSTM aggregator or an SE block. This suggests that our combined temporal attention (AttnScorer) and feature-gating (GatedPooling) mechanism is more effective than a standard recurrent aggregator. The underperformance of the SE block variant further implies that for this task, identifying when an emotionally salient event occurs (temporal attention) is more critical than re-weighting which feature channels are active (channel-wise attention).
- Note on Transformer-based Models: While Transformer architectures are powerful, we deliberately did not include them as a primary baseline for two reasons. First, Transformers are notoriously data-hungry and are prone to overfitting on datasets of limited scale and session diversity like MAHNOB-HCI. Second, their core strength lies in modeling very long-range global dependencies, which, as observed with the LSTM variants, may be suboptimal for our event-driven, short-chunk analysis. As our ablation studies suggest, for this specific problem context, increasing architectural complexity does not necessarily guarantee better performance.
References
- Picard, R.W. Affective Computing; MIT Press, 2000; ISBN 978-0-262-66115-7. [Google Scholar]
- Saffaryazdi, N.; Wasim, S.T.; Dileep, K.; Nia, A.F.; Nanayakkara, S.; Broadbent, E.; Billinghurst, M. Using Facial Micro-Expressions in Combination With EEG and Physiological Signals for Emotion Recognition. Front. Psychol. 2022, 13. [Google Scholar] [CrossRef]
- Mauss, I.B.; Robinson, M.D. Measures of Emotion: A Review. Cogn Emot 2009, 23, 209–237. [Google Scholar] [CrossRef]
- Kreibig, S.D. Autonomic Nervous System Activity in Emotion: A Review. Biological Psychology 2010, 84, 394–421. [Google Scholar] [CrossRef]
- Shaffer, F.; Ginsberg, J.P. An Overview of Heart Rate Variability Metrics and Norms. Front Public Health 2017, 5, 258. [Google Scholar] [CrossRef]
- Yu, S.-N.; Lin, I.-M.; Wang, S.-Y.; Hou, Y.-C.; Yao, S.-P.; Lee, C.-Y.; Chang, C.-J.; Chu, C.-S.; Lin, T.-H. Affective Computing Based on Morphological Features of Photoplethysmography for Patients with Hypertension. Sensors (Basel) 2022, 22, 8771. [Google Scholar] [CrossRef]
- Sanchez, E.; Tellamekala, M.K.; Valstar, M.; Tzimiropoulos, G. Affective processes: Stochastic modelling of temporal context for emotion and facial expression recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 9070–9080. [Google Scholar]
- Wang, W.; den Brinker, A.C.; Stuijk, S.; de Haan, G. Algorithmic Principles of Remote PPG. IEEE Transactions on Biomedical Engineering 2017, 64, 1479–1491. [Google Scholar] [CrossRef]
- Pirzada, P.; Wilde, A.; Doherty, G.H.; Harris-Birtill, D. Remote photoplethysmography (rPPG): A state-of-the-art review. bioRxiv, 2329. [Google Scholar] [CrossRef]
- Dietterich, T.G.; Lathrop, R.H.; Lozano-Pérez, T. Solving the Multiple Instance Problem with Axis-Parallel Rectangles. Artificial Intelligence 1997, 89, 31–71. [Google Scholar] [CrossRef]
- Xie, Y.; Yu, Z.; Wu, B.; Xie, W.; Shen, L. SFDA-rPPG: Source-free domain adaptive remote physiological measurement with spatio-temporal consistency. arXiv arXiv:2401.01234, 2024.
- Russell, J.A. A Circumplex Model of Affect. Journal of Personality and Social Psychology 1980, 39, 1161–1178. [Google Scholar] [CrossRef]
- Jemioło, P.; Storman, D.; Mamica, M.; Szymkowski, M.; Żabicka, W.; Wojtaszek-Główka, M.; Ligęza, A. Datasets for Automated Affect and Emotion Recognition from Cardiovascular Signals Using Artificial Intelligence—A Systematic Review. Sensors 2022, 22, 2538. [Google Scholar] [CrossRef] [PubMed]
- A survey on physiological signal-based emotion recognition. Available online: https://www.mdpi.com/2306-5354/9/11/688 (accessed on 14 June 2025).
- Akselrod, S.; Gordon, D.; Ubel, F.A.; Shannon, D.C.; Berger, A.C.; Cohen, R.J. Power Spectrum Analysis of Heart Rate Fluctuation: A Quantitative Probe of Beat-to-Beat Cardiovascular Control. Science 1981, 213, 220–222. [Google Scholar] [CrossRef]
- Vu, T.; Huynh, V.T.; Kim, S.-H. Multi-scale transformer-based network for emotion recognition from multi-physiological signals. In Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2023; Volume 14408, pp. 113–122. [Google Scholar]
- Martins, A.F.T.; Farinhas, A.; Treviso, M.; Niculae, V.; Aguiar, P.M.Q.; Figueiredo, M.A.T. Sparse and Continuous Attention Mechanisms 2020. arXiv:2006.07214.
- Corbetta, M.; Shulman, G.L. Control of Goal-Directed and Stimulus-Driven Attention in the Brain. Nat Rev Neurosci 2002, 3, 201–215. [Google Scholar] [CrossRef] [PubMed]
- Briggs, F.; Mangun, G.R.; Usrey, W.M. Attention Enhances Synaptic Efficacy and the Signal-to-Noise Ratio in Neural Circuits. Nature 2013, 499, 476–480. [Google Scholar] [CrossRef] [PubMed]
- Lee, C.-Y.; Gallagher, P.; Tu, Z. Generalizing Pooling Functions in CNNs: Mixed, Gated, and Tree. IEEE Transactions on Pattern Analysis and Machine Intelligence 2018, 40, 863–875. [Google Scholar] [CrossRef]
- Bengio, Y.; Louradour, J.; Collobert, R.; Weston, J. Curriculum Learning. In Proceedings of the 26th Annual International Conference on Machine Learning; ACM: Montreal Quebec Canada, 14 June 2009; pp. 41–48. [Google Scholar]
- Elman, J.L. Learning and Development in Neural Networks: The Importance of Starting Small. Cognition 1993, 48, 71–99. [Google Scholar] [CrossRef] [PubMed]
- Mellouk, W.; Handouzi, W. CNN-LSTM for Automatic Emotion Recognition Using Contactless Photoplythesmographic Signals. Biomedical Signal Processing and Control 2023, 85, 104907. [Google Scholar] [CrossRef]
- (PDF) Lang, P.J.; Bradley, M.M. The affect system has parallel and integrative processing components. Available online: https://www.researchgate.net/publication/232564098_The_Affect_System_Has_Parallel_and_Integrative_Processing_Components (accessed on 14 June 2025).
- Zhou, K.; Schinle, M.; Stork, W. Dimensional Emotion Recognition from Camera-Based PRV Features. Methods 2023, 218, 224–232. [Google Scholar] [CrossRef] [PubMed]
- Talala, S.; Shvimmer, S.; Simhon, R.; Gilead, M.; Yitzhaky, Y. Emotion Classification Based on Pulsatile Images Extracted from Short Facial Videos via Deep Learning. Sensors 2024, 24, 2620. [Google Scholar] [CrossRef]
- Adisa, O.; Akinduyite, O.; Badeji-Ajisafe, B. A comprehensive review of remote photoplethysmography techniques for accurate and robust heart-rate monitoring. In Proceedings of the IEEE 5th International Conference on Electro-Computing Technologies for Humanity (NIGERCON), Minna, Nigeria, 23–25 Jan 2024; pp. 1–5. [Google Scholar]
- Soleymani, M.; Lichtenauer, J.; Pun, T.; Pantic, M. A Multimodal Database for Affect Recognition and Implicit Tagging. IEEE Transactions on Affective Computing 2012, 3, 42–55. [Google Scholar] [CrossRef]
- Unke, O.T.; Meuwly, M. PhysNet: A Neural Network for Predicting Energies, Forces, Dipole Moments and Partial Charges. J. Chem. Theory Comput. 2019, 15, 3678–3693. [Google Scholar] [CrossRef] [PubMed]
- Yu, Z.; Shen, Y.; Shi, J.; Zhao, H.; Torr, P.; Zhao, G. PhysFormer: Facial Video-Based Physiological Measurement with Temporal Difference Transformer. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New Orleans, LA, USA, June, 2022; pp. 4176–4186. [Google Scholar]
- Zou, B.; Guo, Z.; Chen, J.; Zhuo, J.; Huang, W.; Ma, H. RhythmFormer: Extracting Patterned rPPG Signals Based on Periodic Sparse Attention. 2025; arXiv:2402.12788. [Google Scholar]
- Luo, C.; Xie, Y.; Yu, Z. PhysMamba: Efficient Remote Physiological Measurement with SlowFast Temporal Difference Mamba. 2024; arXiv:2409.12031. [Google Scholar]
- A Preliminary Approach to Identify Arousal and Valence Using Remote Photoplethysmography | Request PDF. ResearchGate 2025. [CrossRef]
- A Real-Time QRS Detection Algorithm. IEEE Transactions on Biomedical Engineering 1985, 32, 230–236. [CrossRef]
- Bobbia, S.; Macwan, R.; Benezeth, Y.; Mansouri, A.; Dubois, J. Unsupervised Skin Tissue Segmentation for Remote Photoplethysmography. Pattern Recognition Letters 2019, 124, 82–90. [Google Scholar] [CrossRef]
- Rinella, S.; Massimino, S.; Fallica, P.G.; Giacobbe, A.; Donato, N.; Coco, M.; Neri, G.; Parenti, R.; Perciavalle, V.; Conoci, S. Emotion Recognition: Photoplethysmography and Electrocardiography in Comparison. Biosensors (Basel) 2022, 12, 811. [Google Scholar] [CrossRef] [PubMed]
- Zhu, Z.; Wang, X.; Xu, Y.; Chen, W.; Zheng, J.; Chen, S.; Chen, H. An Emotion Recognition Method Based on Frequency-Domain Features of PPG. Front. Physiol. 2025, 16. [Google Scholar] [CrossRef]
- Liu, X.; Narayanswamy, G.; Paruchuri, A.; … rPPG-Toolbox: Deep Remote PPG Toolbox; GitHub repository, 2023. Available online: https://github.com/XYZ/rPPG-Toolbox (accessed on 15 Jan 2025).
- La Fountaine, M.F.; Toda, M.; Testa, A.J.; Hill-Lombardi, V. Autonomic Nervous System Responses to Concussion: Arterial Pulse Contour Analysis. Front. Neurol. 2016, 7. [Google Scholar] [CrossRef]
- Berntson, G.G.; Cacioppo, J.T.; Quigley, K.S. Autonomic Determinism: The Modes of Autonomic Control, the Doctrine of Autonomic Space, and the Laws of Autonomic Constraint. Psychol Rev 1991, 98, 459–487. [Google Scholar] [CrossRef]
- Task Force of the European Society of Cardiology and the North American Society of Pacing and Electrophysiology. Heart-rate variability: Standards of measurement, physiological interpretation and clinical use. Circulation 1996, 93, 1043–1065. [Google Scholar] [CrossRef]
- Bradley, M.M.; Lang, P.J. Measuring Emotion: The Self-Assessment Manikin and the Semantic Differential. Journal of Behavior Therapy and Experimental Psychiatry 1994, 25, 49–59. [Google Scholar] [CrossRef]
- Reynolds, J.H.; Heeger, D.J. The Normalization Model of Attention. Neuron 2009, 61, 168–185. [Google Scholar] [CrossRef]
- Peters, B.; Niculae, V.; Martins, A.F.T. Sparse Sequence-to-Sequence Models Available online:. Available online: https://arxiv.org/abs/1905.05702v2 (accessed on 15 June 2025).
- O’Reilly, R.C.; Munakata, Y. Computational Explorations in Cognitive Neuroscience; MIT Press: Cambridge, MA, USA, 2000. [Google Scholar]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Computation 1997, 9, 1735–1780. [Google Scholar] [CrossRef]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised Contrastive Learning Available online:. Available online: https://arxiv.org/abs/2004.11362v5 (accessed on 15 June 2025).
- Anderson, J.R. Acquisition of Cognitive Skill. Psychological Review 1982, 89, 369–406. [Google Scholar] [CrossRef]
- Posner, M.I.; Petersen, S.E. The Attention System of the Human Brain. Annu Rev Neurosci 1990, 13, 25–42. [Google Scholar] [CrossRef]
- Arpit, D.; Jastrzębski, S.; Ballas, N.; Krueger, D.; Bengio, E.; Kanwal, M.S.; Maharaj, T.; Fischer, A.; Courville, A.; Bengio, Y.; et al. A Closer Look at Memorization in Deep Networks. Available online: https://arxiv.org/abs/1706.05394v2 (accessed on 15 June 2025).
- Hasson, U.; Honey, C.J. Future Trends in Neuroimaging: Neural Processes as Expressed within Real-Life Contexts. Neuroimage 2012, 62, 1272–1278. [Google Scholar] [CrossRef] [PubMed]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 Oct 2017; pp. 2980–2988. [Google Scholar]
- Kumar, M.; Packer, B.; Koller, D. Self-paced learning for latent variable models. In Advances in Neural Information Processing Systems (NIPS 2010), Vancouver, BC, Canada, 6–9 Dec 2010; Volume 23.
- Gold, J.I.; Shadlen, M.N. The neural basis of decision making. Annual Review of Neuroscience 2007, 30, 535–574. [Google Scholar] [CrossRef]
- Ratcliff, R.; Smith, P.L. A Comparison of Sequential Sampling Models for Two-Choice Reaction Time. Psychol Rev 2004, 111, 333–367. [Google Scholar] [CrossRef]
- Thayer, J.F.; Ahs, F.; Fredrikson, M.; Sollers, J.J.; Wager, T.D. A Meta-Analysis of Heart Rate Variability and Neuroimaging Studies: Implications for Heart Rate Variability as a Marker of Stress and Health. Neurosci Biobehav Rev 2012, 36, 747–756. [Google Scholar] [CrossRef]
- Kop, W.J.; Synowski, S.J.; Newell, M.E.; Schmidt, L.A.; Waldstein, S.R.; Fox, N.A. Autonomic Nervous System Reactivity to Positive and Negative Mood Induction: The Role of Acute Psychological Responses and Frontal Electrocortical Activity. Biol Psychol 2011, 86, 230–238. [Google Scholar] [CrossRef] [PubMed]
- Ioannou, S.; Ebisch, S.; Aureli, T.; Bafunno, D.; Ioannides, H.A.; Cardone, D.; Manini, B.; Romani, G.L.; Gallese, V.; Merla, A. The Autonomic Signature of Guilt in Children: A Thermal Infrared Imaging Study. PLOS ONE 2013, 8, e79440. [Google Scholar] [CrossRef]
- Brisinda, D.; Picerni, M.; Fenici, P.; Fenici, R. Automatic Calculation of the Parasympathetic, Sympathetic and Baevsky Stress Indexes Provides a More Comprehensive Assessment of Cardiac Autonomic Modulation. European Heart Journal 2024, 45, ehae666. [Google Scholar] [CrossRef]
- Amhia, H.; Wadhwani, A.K. Stability and Phase Response Analysis of Optimum Reduced-Order IIR Filter Designs for ECG R-Peak Detection. Journal of Healthcare Engineering 2022, 2022, 9899899. [Google Scholar] [CrossRef] [PubMed]





| Method Category | Model / Algorithm | Pre-trained on | ||
|---|---|---|---|---|
| Deep Learning | PhysFormer | UBFC-rPPG | 5.37 | 7.32 |
| PhysFormer | PURE | 7.01 | 9.72 | |
| PhysMamba* | UBFC-rPPG* | 5.38* | 7.43* | |
| PhysMamba | PURE | 6.16 | 8.75 | |
| Traditional | ICA | - | 13.02 | 16.34 |
| POS | - | 12.86 | 15.96 | |
| CHROM | - | 13.31 | 16.80 | |
| GREEN | - | 16.90 | 20.44 | |
| LGI | - | 15.76 | 19.28 | |
| PBV | - | 16.12 | 19.58 | |
| OMIT | - | 15.78 | 19.30 |
| Model | ) ↑ | SNR (db) ↑ | VRAM (GiB) | |||
|---|---|---|---|---|---|---|
| PhysFormer | 0.53 | 1.18 | 0.76 | 0.9994 | 0.25 | 10.1 / 16.3 |
| PhysMamba* | 0.35* | 0.79* | 0.51* | 0.9997* | 3.67* | 9.8 / 16.3* |
| Method | Accuracy (%) | F1 of Positive (%) | Weighted F1 (%) |
|---|---|---|---|
| CNN-LSTM [23] | 61.31 | 50.96 | 59.46 |
| Ours* | 66.04* | 74.29* | 61.97* |
| Method | Accuracy (%) | F1 of Positive (%) | Weighted F1 (%) |
|---|---|---|---|
| CNN-LSTM [23] | 73.50 | 76.23 | 73.14 |
| Ours* | 62.26* | 66.67* | 62.26* |
| Method | Accuracy (%) | Weighted F1 (%) |
|---|---|---|
| Ours (MTDE + Gated Pooling) * | 66.04* | 61.97* |
| MTDE + Attention Pooling | 50.94 | 47.56 |
| MTDE + Average Pooling | 50.94 | 39.07 |
| Feature Extractor Architecture | Accuracy (%) | Weighted F1 (%) |
|---|---|---|
| Ours (3-branch MTDE)* | 66.04* | 61.97* |
| 4-branch MTDE | 54.72 | 54.72 |
| 2-branch MTDE | 62.26 | 60.42 |
| TCN (Single-branch) | 58.49 | 52.05 |
| Training Strategy | Accuracy (%) | Weighted F1 (%) |
|---|---|---|
| Full Curriculum (Phase 0→2)* | 66.04* | 61.97* |
| Full Curriculum (w/o Top-K) | 64.15 | 63.89 |
| Phase 0 → 2 (w/o Auxiliary Classifier) | 58.49 | 58.48 |
| Phase 1 → 2 (w/o SupCon) | 54.72 | 55.19 |
| Phase 2 (Direct training) | 50.94 | 34.18 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).