Submitted:
25 October 2023
Posted:
26 October 2023
You are already at the latest version
Abstract

Keywords:
1. Introduction
2. Related work
- "To what extent can we enhance information acquisition from transcriptions in elementary school classes by eliminating interferences and noises inherent to such settings, using an affordable UHF microphone transmitting to a smartphone located 8 meters away at the back of the classroom?"
- "To what extent can we ensure that, in our pursuit to enhance transcription quality by eliminating noises and interferences, valuable and accurate information is not inadvertently filtered out or omitted in the process?"
3. Materials and Methods
3.1. Data Collection
3.2. Preprocessing and Quality Control
3.3. Data labeling and Augmentation
3.4. Audio Features
3.5. Model architecture and training
3.6. Training results


Metrics of the training results
4. Results
4.1. Suppression of hallucinations in transcripts
5. Discussion
5.1. General Analysis of the Results
5.2. Damage Control
5.3. Benefits of Using the Filter
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Sample Availability
Appendix A
Appendix A.1. List of parameters and their correlations
| Feature | Brief Description | Detailed Description |
|---|---|---|
| mean | Mean of the absolute magnitude. | The mean of the absolute value of all samples in the audio segment was computed using np.abs(segment_samples).mean(). |
| std | Standard deviation of the absolute magnitude. | The standard deviation of the absolute value of the samples was computed using np.abs(segment_samples).std(). |
| pitch_mean | Mean of pitch frequencies. | Pitch frequencies were detected in the Mel spectrogram and their mean was obtained with np.mean(pitches). |
| pitch_std | Standard deviation of pitch frequencies. | The standard deviation of the detected pitch frequencies was computed with np.std(pitches). |
| pitch_confidence | Confidence in pitch detection. | The mean of the magnitudes associated with pitch frequencies was computed using np.mean(magnitudes). |
| spectral_centroid | "Gravity" center of the spectrum. | librosa.feature.spectral_centroid was used to compute the spectral centroid of the signal, giving a measure of the perceived "brightness". |
|---|---|---|
| spectral_bandwidth | Spectrum width. | librosa.feature.spectral_bandwidth was used to obtain the width of the frequency in which most of the energy is concentrated. |
| spectral_rolloff | Spectrum cutoff frequency. | librosa.feature.spectral_rolloff was used to determine the frequency below which a specified percentage of the total energy is located. |
| zero_crossing_rate | Zero crossing rate. | librosa.feature.zero_crossing_rate was used to compute how many times the signal changes sign within a period. |
| mfcc_x | Mel frequency cepstral coefficients for component x. | This was computed using librosa.feature.mfcc. The ’x’ index refers to the x-th component of the MFCC coefficients. It describes the shape of the power spectrum of an audio signal. |
|---|---|---|
| spectral_contrast_x | Amplitude difference between peaks and valleys of the spectrum for band x. | librosa.feature.spectral_contrast was used to measure the amplitude difference between peaks and valleys in the spectrum. The ’x’ index indicates that it is the average spectral contrast for band x. |
| quantile_x | Deciles of the absolute magnitude. | Deciles of the absolute magnitude of audio samples were computed using np.quantile(np.abs(segment_samples), q), for various q values. |
| Feature | Value | Status |
|---|---|---|
| spectral_contrast_6 | 0.397998 | Included |
| mfcc_1 | 0.381041 | Included |
| std | 0.380255 | Included |
| mfcc_3 | 0.373912 | Included |
| mfcc_10 | 0.350538 | Included |
| mfcc_8 | 0.336277 | Included |
| quantile90 | 0.324888 | Included |
| mean | 0.274777 | Included |
| quantile80 | 0.271496 | Included |
| mfcc_0 | 0.271207 | Included |
| spectral_contrast_3 | 0.267938 | Included |
| mfcc_9 | 0.237584 | Included |
| spectral_bandwidth | 0.233203 | Included |
| quantile70 | 0.224501 | Included |
| mfcc_7 | 0.203561 | Included |
| spectral_contrast_4 | 0.197853 | Included |
| quantile60 | 0.183169 | Included |
| mfcc_4 | 0.169382 | Included |
| mfcc_5 | 0.164113 | Included |
| pitch_mean | 0.159703 | Included |
| quantile50 | 0.146121 | Included |
| pitch_std | 0.128296 | Included |
| spectral_contrast_5 | 0.122807 | Included |
| mfcc_6 | 0.117541 | Included |
| mfcc_12 | 0.113921 | Included |
| quantile40 | 0.107676 | Included |
| Feature | Value | Status |
|---|---|---|
| quantile30 | 0.067156 | Not Included |
| mfcc_11 | 0.061327 | Not Included |
| mfcc_2 | 0.044224 | Not Included |
| zero_crossing_rate | 0.040466 | Not Included |
| pitch_confidence | 0.035399 | Not Included |
| quantile20 | 0.028609 | Not Included |
| Segmento | 0.015662 | Not Included |
| time | 0.015662 | Not Included |
| spectral_contrast_2 | 0.009942 | Not Included |
| spectral_centroid | 0.006812 | Not Included |
| quantile10 | 0.002807 | Not Included |
| spectral_rolloff | 0.001632 | Not Included |
| spectral_contrast_0 | 0 | Not Included |
| spectral_contrast_1 | 0 | Not Included |
References
- Li, H., Wang, Z., Tang, J., Ding, W., & Liu, Z. Siamese Neural Networks for Class Activity Detection; In: Bittencourt, I., Cukurova, M., Muldner, K., Luckin, R., Millán, E. (eds) Artificial Intelligence in Education. AIED 2020. Lecture Notes in Computer Science, vol 12164. Springer, Cham, 2020. [CrossRef]
- Li, H., Kang, Y., Ding, W., Yang, S., Yang, S., Huang, G. Y., & Liu, Z. Multimodal Learning for Classroom Activity Detection; In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 9234-9238, 2020. [CrossRef]
- Cosbey, R., Wusterbarth, A., Hutchinson, B. Deep Learning for Classroom Activity Detection from Audio. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2019; pp. 3727–3731. DOI: 10.1109/ICASSP40776.2020.9054407.
- Zhang, X.-L., Wu, J. Denoising Deep Neural Networks Based Voice Activity Detection. In arXiv preprint arXiv:1303.0663; 2013.
- Thomas, S., Ganapathy, S., Saon, G., Soltau, H. Analyzing Convolutional Neural Networks for Speech Activity Detection in Mismatched Acoustic Conditions. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2014; pp. 2519–2523.
- Ma, Y.; Wiggins, J.B.; Celepkolu, M.; Boyer, K.E.; Lynch, C.; Wiebe, E. The Challenge of Noisy Classrooms: Speaker Detection During Elementary Students’ Collaborative Dialogue. Lecture Notes in Artificial Intelligence 2021, 12748, 268–281. [Google Scholar]
- Kingma, D. P., & Ba, J. Adam: A Method for Stochastic Optimization. 2017; arXiv preprint arXiv:1412.6980. arXiv:1412.6980. https://arxiv.org/abs/1412.6980.
- Abreu Araujo, F., Riou, M., Torrejon, J., et al. Role of non-linear data processing on speech recognition task in the framework of reservoir computing. Sci Rep 10, 328 (2020). DOI: 10.1038/s41598-019-56991-x.
- Dong, H.-Y. Speech intelligibility improvement in noisy reverberant environments based on speech enhancement and inverse filtering; EURASIP Journal on Audio, Speech, and Music Processing, 2018.
- Radford, A., Kim, J., Xu, T., Brockman, G., McLeavey, C., Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision; 2022.
- Sabri, M., Alirezaie, J., Krishnan, S. Audio noise detection using hidden Markov model. In Statistical Signal Processing, 2003 IEEE Workshop on; IEEE: 2003; pp. 637–64.
- Rangachari, S.; Loizou, P.C. A noise estimation algorithm with rapid adaptation for highly nonstationary environments. In Speech Communication; Elsevier: Amsterdam, Netherlands, 2006; pp. 220–231. [Google Scholar]
- Çakır, E.; Heittola, T.; Huttunen, H.; Virtanen, T. Polyphonic sound event detection using multi label deep neural networks. In Proceedings of the International Joint Conference on Neural Networks (IJCNN); IEEE: Killarney, Ireland, 2015; pp. 1–7. 10.1109/IJCNN.2015.7280624.
- Dinkel, H., Wang, S., Xu, X., Wu, M., Yu, K. Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training. In IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING; IEEE: Piscataway, NJ, USA, 2021; Vol. 29. https://myw19.github.io/AboutPageAssets/papers/hedi7-dinkel-taslp2021-2.pdf. IEEE.
- Rashmi, S.; Hanumanthappa, M.; Gopala, B. Training Based Noise Removal Technique for a Speech-to-Text Representation Model. In Journal of Physics: Conference Series, Volume 1142, Second National Conference on Computational Intelligence (NCCI 2018); IOP Publishing: Bangalore, India, 2018; pp. 1–2, doi:10.1088/1742-6596/1142/1/012019. IOP Publishing. [CrossRef]
- Kartik, P., & Jeyakumar, G. A Deep Learning Based System to Predict the Noise (Disturbance) in Audio Files; In Advances in Parallel Computing, Volume 37; ISBN 9781643681023, 2020; DOI: 10.3233/APC20013. DOI: 10.3233/APC200135.
- NVIDIA Corporation. Matrix Multiplication Background User’s Guide; NVIDIA Docs; © 2020-2023 NVIDIA Corporation & affiliates. All rights reserved. Last updated on Feb 1, 2023. Available at: https://docs.nvidia.com/deeplearning/performance/pdf/Matrix-Multiplication-Background-User-Guide.pdf.
- Heaton, J. Artificial Intelligence for Humans: Deep Learning and Neural Networks, Volume 3 of Artificial Intelligence for Humans Series; Heaton Research, Incorporated., 2015; ISBN 1505714346.
- nal Adil, M., Ullah, R., Noor, S. et al. Effect of number of neurons and layers in an artificial neural network for generalized concrete mix design. Neural Comput & Applic, Volume 34, 8355–8363, 2022. DOI: 10.1007/s00521-020-05305-8.
- Schlotterbeck, D., Uribe, P., Araya, R., Jimenez, A., Caballero, D. What Classroom Audio Tells About Teaching: A Cost-effective Approach for Detection of Teaching Practices Using Spectral Audio Features. Pages 132-140, 2021. DOI: 10.1145/3448139.3448152.
- Schlotterbeck, D., Jiménez, A., Araya, R., Caballero, D., Uribe, P., & Van der Molen Moris, J. “Teacher, Can You Say It Again?" Improving Automatic Speech Recognition Performance over Classroom Environments with Limited Data. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Volume 13355 LNCS, Pages 269-280, 23rd International Conference on Artificial Intelligence in Education, AIED 2022; Durham; United Kingdom; 27 July 2022 through 31 July 202.
- Smith, M. K., Jones, F. H. M., Gilbert, S. L., & Wieman, C. E. The Classroom Observation Protocol for Undergraduate STEM (COPUS): A New Instrument to Characterize University STEM Classroom Practices. Published Online: 13 Oct 2017. 1: Online, DOI: 10.1187/cbe.13-08-0154.








| Equipment | Cost (Chilean Pesos) | Cost (USD) |
|---|---|---|
| Redmi 9A Mobile Phone | 70,000 | 90 |
| Lavalier UHF Lapel Microphone | 17,000 | 20 |
| Mobile Phone Tripod | 5,000 | 6 |
| Total per teacher | 92,000 | 116 |
| Datos | Precision | Accuracy | Recall | F1 | AUC |
|---|---|---|---|---|---|
| train | 0.703 | 0.823 | 0.751 | 0.726 | 0.835 |
| test | 0.631 | 0.782 | 0.646 | 0.639 | 0.810 |
| Teaching session 1 | |||
|---|---|---|---|
| Label | Time(HH:MM:SS) | Percentage of Transcription | Total Percentage |
| the dummy | |||
| Good | 00:04:08.23 | 51.61% | 47.14% |
| Regular | 00:01:12.71 | 15.12% | 13.81% |
| Bad | 00:02:40.04 | 33.27% | 30.39% |
| No transcription | 00:00:45.63 | 8.66% | |
| Total | 00:08:46.62 | 100.0% | |
| our model | |||
| Good | 00:04:15.62 | 93.12% | 48.54% |
| Regular | 00:00:14.07 | 5.13% | 2.67% |
| Bad | 00:00:04.82 | 1.76% | 0.92% |
| No transcription | 00:04:12.12 | 47.87% | |
| Total | 00:08:46.63 | 100.0% | |
| Teaching session 2 | |||
|---|---|---|---|
| Label | Time(HH:MM:SS) | Percentage of Transcription | Total Percentage |
| the dummy | |||
| Good | 00:07:07.85 | 54.34% | 53.92% |
| Regular | 00:05:31.97 | 42.16% | 41.84% |
| Bad | 00:00:27.55 | 3.5% | 3.47% |
| No transcription | 00:00:06.04 | 0.76% | |
| Total | 00:13:13.44 | 100.0% | |
| our model | |||
| Good | 00:05:00.35 | 91.16% | 37.85% |
| Regular | 00:00:29.12 | 8.84% | 3.67% |
| Bad | 00:00:00.00 | 0.0% | 0.0% |
| No transcription | 00:07:43.95 | 58.47% | |
| Total | 00:13:13.43 | 100.0% | |
| Teaching session 3 | |||
|---|---|---|---|
| Label | Time(HH:MM:SS) | Percentage of Transcription | Total Percentage |
| the dummy | |||
| Good | 00:13:16.33 | 91.16% | 86.83% |
| Regular | 00:00:25.01 | 2.86% | 2.73% |
| Bad | 00:00:52.17 | 5.97% | 5.69% |
| No transcription | 00:00:43.55 | 4.75% | |
| Total | 00:15:17.06 | 100.0% | |
| our model | |||
| Good | 00:12:21.01 | 99.73% | 80.8% |
| Regular | 00:00:02.00 | 0.27% | 0.22% |
| Bad | 00:00:00.00 | 0.0% | 0.0% |
| No transcription | 00:02:54.06 | 18.98% | |
| Total | 00:15:17.08 | 100.0% | |
| (Z-tests) | ||
|---|---|---|
| Lecture | Z-value | p |
| 1 | -0.15 | 0.5309 |
| 2 | 7.23 | 5.01e-13 |
| 3 | 3.64 | 2.71e-04 |
| (Z-tests Good) | (Z-tests Regular) | (Z-tests Bad) | ( Overall) | ||||
|---|---|---|---|---|---|---|---|
| Lecture | Z-value | p | Z-value | p | Z-value | p | p |
| 1 | -6.5624 | 5.29e-11 | 2.3417 | 0.0192 | 5.8619 | 4.58e-09 | 1.55e-10 |
| 2 | -5.8475 | 4.99e-09 | 5.4056 | 6.46e-08 | 1.8874 | 0.0591 | 3.09e-08 |
| 3 | -2.9063 | 3.66e-03 | 1.4755 | 0.1401 | 2.4807 | 0.0131 | 0.0143 |
| Lecture | Duration | Word Count (Segment) | Word Count (Hallucination) | Percentage |
|---|---|---|---|---|
| 1 | 01:18:40 | 3668 | 119 | 3.24% |
| 2 | 01:19:26 | 4309 | 3 | 0.07% |
| 3 | 01:26:10 | 8989 | 38 | 0.42% |
| Lecture | Duration | Word Count (Segment) | Word Count (Hallucination) | Percentage |
|---|---|---|---|---|
| 1 | 01:18:40 | 2175 | 22 | 1.01% |
| 2 | 01:19:26 | 4281 | 0 | 0.00% |
| 3 | 01:26:10 | 7191 | 0 | 0.00% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).