Submitted:
12 February 2026
Posted:
13 February 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction: Towards Brain-Inspired Biomimetic Systems for Audio-, Visual- and Audio-Visual Pattern Recognition
1.1. Problem Definition
1.2. Why Use Brain-Inspired SNN and the NeuCube Architecture for Audio-Visual Data?
- Temporal inputs (features) are converted into spike trains.
- An output classifier/regressor SNN is connected to neurons from the SNNcube, e.g., deSNN [31].
- The SNNcube structure is initialized as a small world connectivity 3D structure of spiking neurons.
- Unsupervised learning is performed in the SNNcube using STDP.
- Supervised learning is performed in the output SNN module, e.g. deSNN for classification.
- The model is further trained and adapted on new data, where new connections are evolved in the SNNcube and new output neurons are evolved in the deSNN classifier to capture new patterns and new classes different from those previously used.
1.3. Experimental Data
- Class 0 = Low emotional arousal: neutral, calm, sad;
- Class 1 = High emotional arousal: happy, angry, fearful, disgust, surprised.
2. Methods: A general eXCube2 Framework and Models for Emotion Recognition Based on Audio-, Visual- and Multimodal Audiovisual Data
2.1. The General eXCube2 Framework
2.2. Audio Feature Extraction and Feature Encoding
2.3. Tonotopic Mapping of Features into a 3D SNNcube
2.4. Feature Extraction and Topographic Mapping of Visual data
- Brow (browDownLeft/Right, browInnerUp, browOuterUpLeft/Right), 5 features.
- Eye (eyeBlinkLeft/Right, eyeSquintLeft/Right, eyeWideLeft/Right), 8 features.
- Cheek (cheekPuff, cheekSquintLeft/Right), 3 features.
- Nose (noseSneerLeft/Right), 2 features.
- Jaw (jawOpen, jawForward, jawLeft/Right), 4 features.
- Mouth (mouthSmileLeft/Right, mouthFrownLeft/Right, etc), 28 features.
- Tongue (tongueOut), 1 feature
- Neutral (neutral (always class 0), 1 feature.
- The Occipital Face Area (OFA) is responsible for early face detection (right hemisphere).
- The Fusiform Face Area (FFA) is responsible for face identity (right occipital, coordinates around (40, -55, -15).
- The Superior Temporal Sulcus (STS) is responsible for dynamic facial expressions and gaze (approximately (50, -45, 10)).
- Dorsal area (top, higher Z) encodes brow features (browDown, browInnerUp), eye features (eyeBlink, eyeSquint), and nose features (noseSneer);
- Ventral area (bottom, lower Z) encodes mouth-related features (mouthSmile, jawOpen).
2.5. Mapping Multimodal Audio-Visual Features into an eXCube2 Model
2.6. Training of an SNNcube for eXCube2 Models on Audio-, Visual- and Audio-Visual Data
2.7. State Vector Extraction from a Trained SNNcube and Their Classification
- (a)
- Spike Count: This method sums the total number of spikes per neuron across all timesteps for each sample:
- (b) DeSNN weight-based state vectors: Alternatively, state vectors can be extracted from the connectivity weights of a DeSNN classifier (see [31]). These weights are determined by the first spike time and the total number of spikes:
3. Experimental Results
3.1. Classification Results on the Experimental Data
3.2. Explainability of the eXCube2 Models
4. Conclusions, Discussions and Future Work
6. Patents
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Dedication
Conflicts of Interest
Appendix A
A.1. Encoding and Mapping of the 24 linear_fft Features


A.2 Experimental Details of the Tonotopic Mapping of Audio Features into the SNNcube
- The following setup was used for the tonotopic mapping of audio features:
- The HCP-MMP1 atlas (Glasser et al., 2016) on MNI152 template was used.
- A hybrid downsampling scheme was applied: 2.5 mm resolution for the auditory cortex and 8.1 mm for the rest of the brain, resulting in 4176 neurons in total (840 auditory + 3336 other neurons).
- Input neurons were placed only in A1 (primary auditory cortex) in both hemispheres.
- A PCA-based gradient was used to distribute neurons evenly along the tonotopic axis (high → low → high frequencies).
- A direct 1:1 mapping was applied: sample column i → neuron i
- A1 (Core): Primary auditory cortex with sharp frequency tuning and direct sensory input.
- Belt: Surrounding A1, integrates frequency channels and supports phonemes/timbre processing
- Parabelt: Higher-level auditory processing, including speech and music categories.
- Biological plausibility: Mimics how the real auditory cortex receives input.
- Spatial learning: Enables the SNN to learn spatial relationships between frequency bands.
- Interpretability: Neuron activations correspond to known brain regions.
- Emergent organization: As shown in TopoAudio (29), spatially constrained networks can develop brain-like organization without explicit supervision.
- Only A1 (core) receives direct input, not belt or parabelt regions.
- A1 neurons are frequency-selective, similar to mel spectrogram or fft bands.
- Belt regions receive processed output from A1 via learned SNN connections.
- Parabelt regions receive output from belt regions.
- This mapping matches the biological processing hierarchy.
- Auditory cortex: 2.5 mm voxel size (high resolution for input neurons).
- Other regions: 8.1 mm voxel size (standard NeuCube resolution).
- Result: 840 auditory neurons + 3336 other neurons = 4176 total neurons.
A.3. Experimental Details of the Training/Testing Parameters of the eXCube2 Models
- Training: 360 samples (6 actors × 60 recordings)
- Testing: 120 samples (2 actors × 60 recordings)
- Low arousal (neutral, calm, sad): 120 samples
- High arousal (happy, angry, fearful, disgust, surprised): 240 samples
- Ratio: 1:2 (imbalanced)
- Low arousal: 240 samples (duplicated from 120)
- High arousal: 240 samples (unchanged)
- Total training set: 480 samples (balanced 1:1)
- 40 low arousal + 80 high arousal = 120 samples
- 234–501 timesteps after silence trimming
- Shortest: 234 timesteps × 10 ms = 2.34 s
- Longest: 501 timesteps × 10 ms = 5.01 s
- 10 ms hop size, 25 ms window length (standard in speech processing)
- Training: 1520 samples (760 class 0, 760 class 1), 19 actors (3–12, 15–23)
- Testing: 1820 samples (860 class 0, 960 class 1), 5 actors (1, 2, 13, 14, 24)
References
- Chern, I.C.; Hung, K.H.; Chen, Y.T.; Hussain, T.; Gogate, M.; Hussain, A.; Tsao, Y.; Hou, J.C. Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2023; pp. 1–5. [Google Scholar]
- Saraceno, G. Deep Learning and Memorizing of Spectro-Temporal Data (Music) in the Spatio-Temporal Brain. Master’s Thesis, University of Trento, Trento, Italy, 2017. [Google Scholar]
- Zhang, H.; Zhang, B.; Huang, W.; Tian, Q. Gabor wavelet associative memory for face recognition. IEEE Trans. Neural Netw. 2005, 16, 275–278. [Google Scholar] [CrossRef]
- Liu, W.; Quan, Y.; Liu, Y.; Yan, D.-M. Bi-directional modality fusion network for audio-visual event localization. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022; pp. 4868–4872. [Google Scholar]
- Lacheze, L.; Guo, Y.; Benosman, R.; Gas, B.; Couverture, C. Audio/video fusion for objects recognition. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2009; pp. 652–657. [Google Scholar]
- Su, R.; Wang, L.; Liu, X. Multimodal learning using 3D audio-visual data for audio-visual speech recognition. In Proceedings of the International Conference on Asian Language Processing (IALP), 2017; pp. 40–43. [Google Scholar]
- Zheng, X.; Wei, Y. Audio-visual event and sound source localization based on spatial-channel feature fusion. In Proceedings of the International Conference on Signal and Image Processing (ICSIP), 2022; pp. 106–110. [Google Scholar]
- Kasabov, N.; Postma, E.; van den Herik, J. AVIS: A connectionist-based framework for integrated auditory and visual information processing. Inf. Sci. 2000, 123, 127–148. [Google Scholar] [CrossRef]
- Wysoski, S.G.; Benuskova, L.; Kasabov, N. Evolving spiking neural networks for audiovisual information processing. Neural Netw. 2010, 23, 819–835. [Google Scholar] [CrossRef]
- Beal, M.; Jojic, N.; Attias, H. A graphical model for audiovisual object tracking. IEEE Trans. Pattern Anal. Mach. Intell. 2003, 25, 828–836. [Google Scholar] [CrossRef]
- Yue, Q.; Wu, X.; Gao, J. Audio-visual event localization based on cross-modal interacting guidance. In Proceedings of the IEEE International Conference on Artificial Intelligence and Knowledge Engineering (AIKE), 2021; pp. 104–107. [Google Scholar]
- Chakraborty, S.; Aich, S.; Joo, M.I.; Sain, M.; Kim, H.C. A multichannel convolutional neural network architecture for the detection of the state of mind using physiological signals from wearable devices. J. Healthc. Eng. 2020, 2020, 5467936. [Google Scholar] [CrossRef] [PubMed]
- Chatterjee, D.; Hegde, S.; Thaut, M.H. Neural plasticity: The substratum of music-based interventions in neurorehabilitation. NeuroRehabilitation 2021, 48, 155–166. [Google Scholar] [CrossRef]
- Krautz, A.E.; Langner, J.; Helmhold, F.; Volkening, J.; Hoffmann, A.; Hasler, C. Bridging AI innovation and healthcare: Scalable clinical validation methods for voice biomarkers. Front. Digit. Health 2025, 7, 1575753. [Google Scholar] [CrossRef] [PubMed]
- Rao, A.; Salehi, M.-J.; Vajargah, S.H.; Bourque, J.L. Neural correlates of auditory predictive timing are linked to human vocal pitch stability. Sci. Rep. 2019, 9, 45105. [Google Scholar] [CrossRef]
- Reddy, V. PPINtonus: Early detection of Parkinson’s disease using deep-learning tonal analysis. arXiv 2022, arXiv:2406.02608. [Google Scholar] [CrossRef]
- Kasabov, N. Time-Space, Spiking Neural Networks and Brain-Inspired AI; Springer Nature: Cham, Switzerland, 2019. [Google Scholar]
- Kasabov, N.K. NeuCube: A spiking neural network architecture for mapping, learning and understanding of spatio-temporal brain data. Neural Netw. 2014, 52, 62–76. [Google Scholar] [CrossRef]
- Izhikevich, E.M. Polychronization: Computation with spikes. Neural Comput. 2006, 18, 245–282. [Google Scholar] [CrossRef]
- Abeles, M. Corticonics: Neural Circuits of the Cerebral Cortex; Cambridge University Press: Cambridge, UK, 1991. [Google Scholar]
- Kasabov, N.K. STAM-SNN: Spatio-temporal associative memory in brain-inspired spiking neural networks: Concepts and perspectives. In Recent Advances in Intelligent Engineering; Kovács, L., Haidegger, T., Szakál, A., Eds.; Springer: Cham, Switzerland, 2024; pp. 1–XX. [Google Scholar] [CrossRef]
- Kasabov, N.K. Life-long learning and evolving associative memories in brain-inspired spiking neural networks. MOJ Appl. Bio. Biomech. 2024, 8, 56–57. [Google Scholar] [CrossRef]
- Kasabov, N.K. Spatio-temporal associative memories in brain-inspired spiking neural networks: Concepts and perspectives. TechRxiv 2023. [Google Scholar] [CrossRef]
- Kasabov, N.; Bahrami, H.; Doborjeh, M.; Wang, A. Brain inspired spatio-temporal associative memories for neuroimaging data: EEG and fMRI. Bioengineering Preprints 2023. [Google Scholar] [CrossRef]
- Gao, C.; Green, J.J.; Yang, X.; Oh, S.; Kim, J.; Shinkareva, S.V. Audiovisual integration in the human brain: A coordinate-based meta-analysis. Cereb. Cortex 2023, 33, 5574–5584. [Google Scholar] [CrossRef]
- Kasabov, N. Neucube evospike architecture for spatio-temporal modelling and pattern recognition of brain signals. In Artificial Neural Networks in Pattern Recognition; Mana, N., Schwenker, F., Trentin, E., Eds.; Springer: Berlin, Germany, 2012; pp. 225–243. [Google Scholar] [CrossRef]
- Talairach, J.; Tournoux, P.; Rayport, M. Co-planar stereotaxic atlas of the human brain: 3-dimensional proportional system. J. Laryngol. Otol. 1988, 104, 72–73. [Google Scholar]
- Glasser, M.F.; et al. A multi-modal parcellation of human cerebral cortex. Nature 2016, 536, 171–178. [Google Scholar] [CrossRef] [PubMed]
- Al-Tahan, H.; et al. End-to-end topographic auditory models replicate signatures of human auditory cortex. arXiv 2025, arXiv:2509.24039. [Google Scholar]
- Moerel, M.; et al. An anatomical and functional topography of human auditory cortical areas. Front. Neurosci. 2014, 8, 225. [Google Scholar] [CrossRef]
- Kasabov, N.K.; Dhoble, K.; Nuntalid, N.; Indiveri, G. Dynamic evolving spiking neural networks for online spatio- and spectro-temporal pattern recognition. Neural Netw. 2013, 41, 188–201. [Google Scholar] [CrossRef]
- Kumarasinghe, K.; Kasabov, N.; Taylor, D. Deep learning and deep knowledge representation in spiking neural networks for brain–computer interfaces. Neural Netw. 2020, 121, 169–185. [Google Scholar] [CrossRef]
- Kasabov, N.K.; Tan, Y.; Doborjeh, M.; Tu, E.; Yang, J.; Goh, W.; Lee, J. Transfer learning of fuzzy spatio-temporal rules in the NeuCube brain-inspired spiking neural network. IEEE Trans. Fuzzy Syst. 2023, 31, 4542–4552. [Google Scholar] [CrossRef]
- Swanson, R.; Livingstone, S.R.; Russo, F.A. RAVDESS facial landmark tracking (Version 1.0.0) [Data set]. Zenodo 2019. [Google Scholar] [CrossRef]
- Livingstone, S.R.; Russo, F.A. The Ryerson audio-visual database of emotional speech and song (RAVDESS). PLoS ONE 2018, 13, e0196391. [Google Scholar] [CrossRef]
- Cao, F.; Vogel, A.P.; Gharahkhani, P.; Rentería, M.E. Speech and language biomarkers for Parkinson’s disease prediction. npj Parkinsons Dis. 2025, 11, 57. [Google Scholar] [CrossRef]
- Tan, C.; Šarlija, M.; Kasabov, N. Spiking neural networks: Background, recent development and the NeuCube architecture. Neural Process. Lett. 2020, 52, 1675–1701. [Google Scholar] [CrossRef]
- Chen, C.; Al-Halah, Z.; Grauman, K. Semantic audio-visual navigation. 2021. [Google Scholar]
- Guo, L.; et al. Transformer-based spiking neural networks for multimodal audiovisual classification. IEEE Trans. Cogn. Dev. Syst. 2024, 16, 1077–1086. [Google Scholar] [CrossRef]
- Furber, S.B.; Brown, G.; Bose, J.; Cumpstey, J.M.; Marshall, P.; Shapiro, J.L. Sparse distributed memory using rank-order neural codes. IEEE Trans. Neural Netw. 2007, 18, 648–659. [Google Scholar] [CrossRef]
- Behrenbeck, J.; Tayeb, Z.; Bhiri, C.; Richter, C.; Rhodes, O.; Kasabov, N.; Espinosa-Ramos, J.; Furber, S.; Cheng, G.; Conradt, J. Classification and regression of spatio-temporal signals using NeuCube. J. Neural Eng. 2019, 16, 026019. [Google Scholar] [CrossRef] [PubMed]
- James, R.; Garside, J.; Hopkins, M.; Plana, L.A.; Temple, S.; Davidson, S.; Furber, S. Parallel distribution of an inner hair cell and auditory nerve model. In Proceedings of the IEEE BioCAS, 2017; pp. 1–4. [Google Scholar]
- Furber, S.B.; Galluppi, F.; Temple, S.; Plana, L.A. The SpiNNaker project. Proc. IEEE 2014, 102, 652–665. [Google Scholar] [CrossRef]
- Paulun, L.; Wendt, A.; Kasabov, N.K. A retinotopic spiking neural network system for accurate recognition of moving objects. Front. Comput. Neurosci. 2018, 12, 1–15. [Google Scholar] [CrossRef]
- Song, Q.; Kasabov, N. NFI: A neuro-fuzzy inference method for transductive reasoning. IEEE Trans. Fuzzy Syst. 2005, 13, 799–808. [Google Scholar] [CrossRef]
- AbouHassan, I.; Kasabov, N. NeuDen: A framework for the integration of neuromorphic evolving spiking neural networks. Evolving Syst. 2025, 16, 3. [Google Scholar] [CrossRef]
- Kumarasinghe, K.; Kasabov, N.; Taylor, D. Brain-inspired spiking neural networks for decoding muscle activity. Sci. Rep. 2021, 11, 2486. [Google Scholar] [CrossRef]
- AbouHassan, I.; Kasabov, N.; Bankar, T.; Garg, R.; Bhattacharya, B. ePAMeT: Evolving predictive associative memory for time series. Evolving Syst. 2025, 16, 6. [Google Scholar] [CrossRef]
- Gong, Y.; Chung, Y.-A.; Glass, J. AST: Audio spectrogram transformer. arXiv 2021, arXiv:2104.01778. [Google Scholar] [CrossRef]
- NeuCubePy. Available online: https://github.com/KEDRI-AUT/NeuCube-Py.
- Kasabov, N. Global, local and personalised modelling and profile discovery in Bioinformatics: An integrated approach. Pattern Recognition Letters 2007, Vol. 28(Issue 6), 673–685. [Google Scholar] [CrossRef]











| Feature | Number | Frequencies | In Both Sides | Previous Usage |
| mel_spectrogram | 40 | 50–8000 Hz (mel) | 80 | SOTA emotion recognition |
| mel_fft | 12 | 50–8000 Hz (mel) | 24 | Biologically plausible |
| linear_fft | 12 | 50–8000 Hz (linear) | 24 | Technical analysis |
| mfcc | 12 | Cepstral coefficients | 24 | Speaker-independent ASR |
|
For each hemisphere: 1. Extract full-resolution of A1 coordinates of the template (≈858 left, ≈588 right voxels) 2. Identify downsampled neurons within A1 (≈123 left, ≈96 right) 3. Apply PCA to estimate the principal axis of A1 (tonotopic gradient direction) 4. Project neurons onto this axis to obtain a normalised tonotopic position in [0,1] 5. Select neurons evenly spaced along the gradient to match the number of frequency bands 6. Map each audio feature column to one A1 neuron: - Left channel (columns 1 to N) → Left A1 neurons - Right channel (columns N+1 to 2N) → Right A1 neurons 7. Neurons are ordered by tonotopic position, so: - Band 1 (lowest frequency, ≈50 Hz) → maps to the center of A1 - Band N (highest frequency, ≈8000 Hz) → maps to the ends of A1 8. Perform direct mapping where feature column i → neuron i (1-indexed bands), as summarized below: | ||
| Method | Left Channel Mapping | Right Channel Mapping |
| mel_spectrogram | bands 1-40 → neurons 1-40 | bands 1-40 → neurons 41-80 |
| mel_fft | bands 1-12 → neurons 1-12 | bands 1-12 → neurons 13-24 |
| linear_fft | bands 1-12 → neurons 1-12 | bands 1-12 → neurons 13-24 |
| mfcc | bands 1-12 → neurons 1-12 | bands 1-12 → neurons 13-24 |
| Feature | Audio | Visual |
|---|---|---|
| Features | 80 mel_spectrogram | 52 facial blendshapes |
| Brain region | Bilateral auditory cortex (A1) | Right FFA/STS |
| Input map | Tonotopic (low → high freq) | Topographic |
| Reservoir | 3108 neurons | 3108 neurons |
| Train samples | 160 | 160 |
| Test samples | 120 | 120 |
| Method | Mathematical Formulation | Description |
| (a) SVM (RBF Kernel) |
Gamma set automatically. Maximum-margin hyperplane in kernel space. | |
| (b) Weighted Weighted KNN (WWKNN) (Kasabov, 2010)[51] |
Feature-wise SNR weighting, where SNR_f = variance_between(f) / variance_within(f). Downweights noisy features, emphasises discriminative ones. |
|
| (c) Centroid Prototype |
Class represented by a normalised centroid; classification by maximum cosine similarity. | |
| (d) Learned Prototype |
Prototypes optimised via gradient descent on cross-entropy loss. Adam optimiser, lr=0.01, 300 epochs, tau=0.1. |
| Experiment | Method (see the legend below*) | Max Accuracy |
| Multimodal audio-visual | R: STS py | 82.0% |
| I: split | 82.7% | |
| I: hybrid | 82.3% | |
| Audio (mel_spectrograms) | I+R: SC | 80.3% |
| I+R: A1 SC | 80.3% | |
| I+R: A1 hybrid | 79.3% | |
| Video (blendshapes) | I+R: STS hybrid | 80.7% |
| I+R: split STS hybrid | 80.7% | |
| I+R: split STS py | 81.0% |
| Model | Confidence Threshold | Coverage | Accuracy | Accuracy (with conf. threshold) | Correct | Wrong | Rejected | Lost Correct |
Errors Avoided |
| Video | 65% | 71.3% | 81.0% | 87.9% | 188 | 26 | 86 | 55 | 31 |
| Audio | 55% | 72.3% | 80.3% | 85.3% | 185 | 32 | 83 | 56 | 27 |
| Multimodal | 60% | 81.0% | 82.7% | 87.9% | 213 | 30 | 57 | 35 | 22 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).