Submitted:
19 February 2025
Posted:
20 February 2025
You are already at the latest version
Abstract
With advances in digital technology, including deep learning and big data analytics, new methods have been developed for autism diagnosis and intervention. Emotion recognition and the detection of autism in children are prominent subjects in autism research. Typically using single-modal data to analyze the emotional states of children with autism, previous research has found that the accuracy of recognition algorithms must be improved. Our study creates datasets on the facial and speech emotions of children with autism in their natural states. A convolutional vision transformer-based emotion recognition model is constructed for the two distinct datasets. The findings indicate that the model achieves accuracies of 79.12% and 83.47% for facial expression recognition and Mel spectrogram recognition, respectively. Consequently, we propose a multimodal data fusion strategy for emotion recognition and construct a feature fusion model based on an attention mechanism, which attains a recognition accuracy of 90.73%. Ultimately, by using gradient-weighted class activation mapping, a prediction heat map is produced to visualize facial expressions and speech features under four emotional states. This study offers technical direction for the use of intelligent perception technology in the realm of special education and enriches the theory of emotional intelligence perception of children with autism.
Keywords:
1. Introduction
2. Related Works
2.1. Characteristics of Emotional Expressions of Children with Autism
2.2. Emotional Perception of Children with Autism
2.3. Multimodal Fusion in Emotion Recognition
3. Dataset
3.1. Data Collection
3.1.1. Participants
3.1.2. Methods
3.2. Data Processing
3.2.1. Facial Expression Image Processing
3.2.2. Speech Feature Processing
3.3. Data Annotation
3.4. Data Division
4. Proposed Methodology
4.1. ViT-based facial expression recognition model
- 1.
- Divide the facial image into several patches, apply linear projection to each patch, and then incorporate positional encoding. Specifically, the entire token sequence is as follows:where Nis the number of patches, is the vector of each patch after linear projection, is the initial part of the input sequence used for classification, and is the vector carrying the positional information.
- 2.
- The token sequence z is processed by L Encoder blocks, where each block is composed of three components: Multiple Self-attention module (MSA), Layer Normalisation (LN), and Multi-Layer Perceptron (MLP). The computation of these components is as follows:
- 3.
- The final token sequence after several stacked Encoders is represented as , with the categorization information contained in the vector . The output y, obtained after processing by , is the ultimate classification result, i.e.,
4.2. CvT-based Models for Facial Expression, Mel Spectrogram Recognition
4.3. Bi-LSTM-based MFCCs Recognition Model
4.4. Expression and Speech Feature Fusion Model with Attention Mechanism
5. Model Experiment
5.1. Evaluation Metrics
5.2. Parameter Settingss
5.3. Results
5.3.1. Results of Different Emotion Recognition Models on the Single-modal Dataset
5.3.2. Results of Facial Expression and Speech Feature Fusion Model
5.4. Discussion and Analysis
5.4.1. CvT Model Incorporating Convolution Exhibits Excellent Performance in Facial Expression Recognition
5.4.2. Advantages of Mel Spectrogram for Analyzing Speech Characteristics of Autistic Children
5.4.3. Feature Fusion Model Leverage Complementary Benefits of Multimodal Data
5.4.4. Visualization of Emotional Features in Autistic Children
6. Conclusion and Limitations
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| ASD | autism spectrum disorder |
| MFCCs | Mel-frequency cepstral coefficients |
| CNN | convolutional neural network |
| MS-CAM | Multiscale Channel Attention Module |
References
- Garcia-Garcia, J.M.; Penichet, V.M.; Lozano, M.D.; Fernando, A. Using emotion recognition technologies to teach children with autism spectrum disorder how to identify and express emotions. Universal Access in the Information Society 2022, 21, 809–825. [Google Scholar] [CrossRef]
- Sarmukadam, K.; Sharpley, C.F.; Bitsika, V.; McMillan, M.M.; Agnew, L.L. A review of the use of EEG connectivity to measure the neurological characteristics of the sensory features in young people with autism. Reviews in the Neurosciences 2019, 30, 497–510. [Google Scholar] [CrossRef] [PubMed]
- Shi, J.; Liu, C.; Ishi, C.T.; Ishiguro, H. Skeleton-based emotion recognition based on two-stream self-attention enhanced spatial-temporal graph convolutional network. Sensors 2020, 21, 205. [Google Scholar] [CrossRef] [PubMed]
- Zhai, X.; Xu, J.; Wang, Y. Research on Learning Affective Computing in Online Education: From the Perspective of Multi-source Data Fusion. Journal of East China Normal University (Educational Sciences) 2022, 40, 32. [Google Scholar]
- Donahue, J.; Anne Hendricks, L.; Guadarrama, S.; Rohrbach, M.; Venugopalan, S.; Saenko, K.; Darrell, T. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 2625–2634.
- Zhang, K.; Huang, Y.; Du, Y.; Wang, L. Facial expression recognition based on deep evolutional spatial-temporal networks. IEEE Transactions on Image Processing 2017, 26, 4193–4203. [Google Scholar] [CrossRef]
- Samad, M.D.; Bobzien, J.L.; Harrington, J.W.; Iftekharuddin, K.M. Analysis of facial muscle activation in children with autism using 3D imaging. In Proceedings of the 2015 IEEE International Conference on Bioinformatics and Biomedicine (BIBM); IEEE, 2015; pp. 337–342. [Google Scholar]
- Metallinou, A.; Grossman, R.B.; Narayanan, S. Quantifying atypicality in affective facial expressions of children with autism spectrum disorders. In Proceedings of the 2013 IEEE international conference on multimedia and expo (ICME); IEEE, 2013; pp. 1–6. [Google Scholar]
- Guha, T.; Yang, Z.; Grossman, R.B.; Narayanan, S.S. A computational study of expressive facial dynamics in children with autism. IEEE transactions on affective computing 2016, 9, 14–20. [Google Scholar] [CrossRef]
- American Psychiatric Association, D.; American Psychiatric Association, D.; et al. Diagnostic and statistical manual of mental disorders: DSM-5; American psychiatric association: Washington, DC, 2013; Vol. 5. [Google Scholar]
- Jacques, C.; Courchesne, V.; Mineau, S.; Dawson, M.; Mottron, L. Positive, negative, neutral—or unknown? The perceived valence of emotions expressed by young autistic children in a novel context suited to autism. Autism 2022, 26, 1833–1848. [Google Scholar] [CrossRef]
- Bone, D.; Black, M.P.; Lee, C.C.; Williams, M.E.; Levitt, P.; Lee, S.; Narayanan, S.S. Spontaneous-Speech Acoustic-Prosodic Features of Children with Autism and the Interacting Psychologist. In Proceedings of the InterSpeech; 2012; pp. 1043–1046. [Google Scholar]
- Bone, D.; Black, M.P.; Ramakrishna, A.; Grossman, R.B.; Narayanan, S.S. Acoustic-prosodic correlates of’awkward’prosody in story retellings from adolescents with autism. In Proceedings of the Interspeech; 2015; pp. 1616–1620. [Google Scholar]
- Diehl, J.J.; Paul, R. Acoustic differences in the imitation of prosodic patterns in children with autism spectrum disorders. Research in autism spectrum disorders 2012, 6, 123–134. [Google Scholar] [CrossRef]
- Yankowitz, L.D.; Schultz, R.T.; Parish-Morris, J. Pre-and paralinguistic vocal production in ASD: Birth through school age. Current psychiatry reports 2019, 21, 1–22. [Google Scholar] [CrossRef]
- Winczura, B. Dziecko z autyzmem: terapia deficytów poznawczych a teoria umysłu. Psychologia Rozwojowa 2009, 14. [Google Scholar]
- Jaklewicz, H. Autyzm wczesnodzieciÄ™ cy: diagnoza, przebieg, leczenie; GdaĹ „skie Wyd-wo Psychologiczne, 1993. [Google Scholar]
- Mehrabian, A. Communication without words. In Communication theory; Routledge, 2017; pp. 193–200. [Google Scholar]
- Rozga, A.; King, T.Z.; Vuduc, R.W.; Robins, D.L. Undifferentiated facial electromyography responses to dynamic, audio-visual emotion displays in individuals with autism spectrum disorders. Developmental science 2013, 16, 499–514. [Google Scholar] [CrossRef]
- Jarraya, S.K.; Masmoudi, M.; Hammami, M. A comparative study of Autistic Children Emotion recognition based on Spatio-Temporal and Deep analysis of facial expressions features during a Meltdown Crisis. Multimedia Tools and Applications 2021, 80, 83–125. [Google Scholar] [CrossRef]
- Talaat, F.M. Real-time facial emotion recognition system among children with autism based on deep learning and IoT. Neural Computing and Applications 2023, 35, 12717–12728. [Google Scholar] [CrossRef]
- Landowska, A.; Karpus, A.; Zawadzka, T.; Robins, B.; Erol Barkana, D.; Kose, H.; Zorcec, T.; Cummins, N. Automatic emotion recognition in children with autism: a systematic literature review. Sensors 2022, 22, 1649. [Google Scholar] [CrossRef] [PubMed]
- Ram, C.S.; Ponnusamy, R. Assessment on speech emotion recognition for autism spectrum disorder children using support vector machine. World Applied Sciences J 2016, 34, 94–102. [Google Scholar]
- Sukumaran, P.; Govardhanan, K. Towards voice based prediction and analysis of emotions in ASD children. Journal of Intelligent & Fuzzy Systems 2021, 41, 5317–5326. [Google Scholar]
- Geetha, A.; Mala, T.; Priyanka, D.; Uma, E. Multimodal Emotion Recognition with deep learning: advancements, challenges, and future directions. Information Fusion 2024, 105, 102218. [Google Scholar]
- Minotto, V.P.; Jung, C.R.; Lee, B. Multimodal multi-channel on-line speaker diarization using sensor fusion through SVM. IEEE Transactions on Multimedia 2015, 17, 1694–1705. [Google Scholar] [CrossRef]
- Zhang, Y.; Sidibé, D.; Morel, O.; Mériaudeau, F. Deep multimodal fusion for semantic image segmentation: A survey. Image and Vision Computing 2021, 105, 104042. [Google Scholar] [CrossRef]
- Jun, H.; Caiqing, Z.; Xiaozhen, L.; Dehai, Z. Survey of research on multimodal fusion technology for deep learning. Computer Engineering 2020, 46, 1–11. [Google Scholar]
- Dai, Y.; Gieseke, F.; Oehmcke, S.; Wu, Y.; Barnard, K. Attentional feature fusion. In Proceedings of the Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 3560–3569.
- Gepner, B.; Godde, A.; Charrier, A.; Carvalho, N.; Tardif, C. Reducing facial dynamics’ speed during speech enhances attention to mouth in children with autism spectrum disorder: An eye-tracking study. Development and psychopathology 2021, 33, 1006–1015. [Google Scholar] [CrossRef] [PubMed]
- Wodajo, D.; Atnafu, S. Deepfake video detection using convolutional vision transformer. arXiv arXiv:2102.11126 2021.
- Meng, H.; Yan, T.; Yuan, F.; Wei, H. Speech emotion recognition from 3D log-mel spectrograms with deep learning network. IEEE access 2019, 7, 125868–125881. [Google Scholar] [CrossRef]
- Bulatović, N.; Djukanović, S. Mel-spectrogram features for acoustic vehicle detection and speed estimation. In Proceedings of the 2022 26th International Conference on Information Technology (IT). IEEE; 2022; pp. 1–4. [Google Scholar]
- Li, J.; Bhat, A.; Barmaki, R. A two-stage multi-modal affect analysis framework for children with autism spectrum disorder. arXiv arXiv:2106.09199 2021.














| Emotion type | Appearance characteristics | Example 1 | Example 2 |
|---|---|---|---|
| Calm | The eyebrows are in their natural state; the facial muscles are stretched; the eyes are naturally open, with the eyeballs gazing in various directions; and the mouth is relaxed. | ![]() |
![]() |
| Happy | The mouth curves up or opens, the brows tend to curve, the eye muscles contract, and the nasolabial folds emerge. | ![]() |
![]() |
| Sad | The brows are furrowed, the gaze is downcast, and the lips are clenched and projecting. | ![]() |
![]() |
| Angry | The eyebrows are elevated, the eyes are wide open, the mouth is open or closed in an arch, and the facial muscles show general drooping. | ![]() |
![]() |
| Emotion type | Training set | Validation set | Test set | Total |
|---|---|---|---|---|
| Calm | 11,328 | 2834 | 1574 | 15,736 |
| Happy | 7002 | 1752 | 974 | 9728 |
| Sad | 9106 | 2278 | 1266 | 12,650 |
| Angry | 8258 | 2066 | 1148 | 11,472 |
| Total | 35,694 | 8930 | 4962 | 49,586 |
| Emotion type | Training set | Validation set | Test set | Total |
|---|---|---|---|---|
| Calm | 160,474 | 96,672 | 39,444 | 296,590 |
| Happy | 138,168 | 29,982 | 19,494 | 187,644 |
| Sad | 203,224 | 22,192 | 14,592 | 240,008 |
| Angry | 176,662 | 20,482 | 20,748 | 217,892 |
| Total | 678,528 | 169,328 | 94,278 | 942,134 |
| Module Type | Layer Name | CvT-13 | Output Size | Params# | |
|---|---|---|---|---|---|
| Stage 1 | Convolutional Token Embedding | Conv.Embed | 7×7,64,stride 4 | [64,56,56] | 9600 |
| Convolutional Projection |
Conv.Proj | [3136,64] | 52096 | ||
| MHSA | |||||
| MLP | |||||
| Stage 2 | Rearrge | (b(h w)c)->(b c h w) | [64,56,56] | – | |
| Convolutional Token Embedding | Conv.Embed | 3×3,192,stride 2 | [192,28,28] | 111168 | |
| Convolutional Projection |
Conv.Proj | [784,192] | 902400 | ||
| MHSA | |||||
| MLP | |||||
| Stage 3 | Rearrge | (b(h w)c)->(b c h w) | [192,28,28] | – | |
| Convolutional Token Embedding | Conv.Embed | 3×3,384,stride 2 | [384,14,14] | 664704 | |
| Convolutional Projection |
Conv.Proj | [197,384] | 17871360 | ||
| MHSA | |||||
| MLP | |||||
| Head | Linear | [4] | 2308 |
| Module Type | Layer Name | Output Size | Params# |
|---|---|---|---|
| Initial Feature Fusion | Summation | [384,14,14] | – |
| Local Attention Branch | Conv2d | [96,14,14] | 75168 |
| ReLU | [96,14,14] | ||
| Conv2d | [384,14,14] | ||
| Global Attention Branch | AdaptiveAvgPool2d | [384,1,1] | 75168 |
| Conv2d | [96,1,1] | ||
| ReLU | [96,1,1] | ||
| Conv2d | [384,1,1] | ||
| Feature Fusion Layer | Sigmoid | [384,14,14] | – |
| Fully Connected Layer | Linear | [4] | 1540 |
| Module Type | Optimizer | Lr | loss function | Epochs | Batch size |
Random Shuffle |
|---|---|---|---|---|---|---|
| ViT-based Facial Expression Recognition Model | SGD | 1e-3 | Cross Entropy Loss | 500 | 32 | True |
| CvT-based Facial Expression Recognition Model | Adam | 1e-5 | Cross Entropy Loss | 500 | 32 | True |
| CvT-based Mel Spectrogram Recognition Model | Adam | 1e-4 | Cross Entropy Loss | 300 | 32 | True |
| Bi-LSTM-based MFCCs Recognition Model | Adam | 1e-5 | MSE | 500 | 64 | False |
| Expression and Speech Feature Fusion Model based on Attention Mechanism | Adam | 1e-6 | Cross Entropy Loss | 30 | 8 | False |
| Model Type | Overall Accuracy | Calm | Happy | Sad | Angry |
|---|---|---|---|---|---|
| ViT(Facial Expression) | 48% | 62.90% | 52.57% | 36.81% | 36.06% |
| CvT(Facial Expression) | 79.12% | 73.82% | 90.14% | 75.67% | 80.84% |
| CvT(Mel spectrogram) | 83.47% | 84.12% | 83.78% | 82.15% | 83.80% |
| Bi-LSTM(MFCCs) | 25.72% | 29.48% | 21.64% | 38.28% | 13.55% |
| Evaluation Metrics | Calm | Happy | Sad | Angry |
| Precision | 91.75% | 91.33% | 91.14% | 88.39% |
| Recall | 90.47% | 93.02% | 91.00% | 88.85% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).







