Submitted:
09 December 2024
Posted:
10 December 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Works
3. Methodology
- Coarse onset detection: Localize the coarse attack-transient pairs with the arched waveform determined by analyzing the ZC intervals or envelope zero-crossings;
- Onset fine-tuning: Remove the attack-transient pair if the transient doesn’t meet the conditions of sudden change to improve the precision. Fine-tune the remaining transients and attacks by searching extrema and inflections;
- Vibration signal extraction: Determine the offset between transient and next attack (the next attack may coincide with offset), extract the the voiced area between transient and offset for pitch analysis;
- Pitch marking: Select an initial pitch marker pair, track forward and backward the pitch markers via the continuous time-period mapping (CTPM);
- Manual note segmentation: Manually select an attack-offset pair as the boundaries of a note;
- Pitch interpolation: Smooth the pitch curve within a note, interpolate the pitch curve clips to fill the gaps within tremolo notes.
3.1. Pitch Estimation
3.2. Annotation Process and Processing of the Pitch Curves in Tremolo Notes
4. Experiments
4.1. Datasets and Evaluation Metrics
4.2. Experimental Results
4.2.1. Boundary Results
| Methods | Jasmine Flower | Nanni Bay | The Love of the Wei River | Ambush from Ten Sides | ||||
|---|---|---|---|---|---|---|---|---|
| F/P/R (%) | MAE (ms) | F/P/R (%) | MAE (ms) | F/P/R (%) | MAE (ms) | F/P/R (%) | MAE (ms) | |
| SpecFlux | 80/91/72 | 21.9 | 84/83/86 | 21.3 | 30/22/46 | 19.6 | 57/60/55 | 18.3 |
| SuperFlux | 70/96/55 | 21.7 | 80/88/73 | 21.3 | 27/27/26 | 21.4 | 60/72/52 | 18.0 |
| ComplexFlux | 72/96/58 | 21.4 | 78/88/70 | 21.3 | 25/23/28 | 21.2 | 59/76/49 | 17.7 |
| Log-Energy | 80/71/92 | 2.3 | 58/42/93 | 2.9 | 22/13/92 | 3.0 | 47/43/54 | 5.6 |
| ZC | 88/85/92 | 1.8 | 89/87/92 | 2.4 | 55/40/90 | 4.0 | 52/81/38 | 6.5 |
| Extrema w/o padding | 90/85/96 | 2.4 | 89/84/94 | 2.5 | 54/39/88 | 2.9 | 51/69/40 | 6.8 |
| Extrema w/ padding | 83/73/95 | 2.5 | 81/70/95 | 2.7 | 42/27/92 | 3.7 | 54/43/73 | 7.0 |
| Methods | Nanni Bay | The Love of the Wei River | Ambush from Ten Sides | ||||||
|---|---|---|---|---|---|---|---|---|---|
| TNR (%) | PD (%) | MAE (ms) | TNR (%) | PD (%) | MAE (ms) | TNR (%) | PD(%) | MAE (ms) | |
| SpecFlux | 77/62/100 | 83/97/73 | 20.3 | 57/40/100 | 67/99/51 | 17.2 | 64/88/50 | 60/98/43 | 17.9 |
| SuperFlux | 87/77/100 | 69/97/53 | 20.1 | 67/50/100 | 45/100/29 | 17.8 | 55/95/39 | 49/98/33 | 17.8 |
| ComplexFlux | 91/83/100 | 62/97/46 | 19.5 | 67/50/100 | 46/99/30 | 17.4 | 55/90/39 | 43/98/28 | 17.7 |
| Log-Energy | 45/29/100 | 86/96/79 | 3.2 | 53/36/100 | 68/98/52 | 5.5 | 63/68/59 | 67/96/51 | 6.9 |
| ZC | 100/100/100 | 98/100/96 | 3.0 | 89/80/100 | 93/99/87 | 6.5 | 69/100/52 | 65/99/48 | 8.0 |
| Extrema w/o padding | 95/91/100 | 99/100/99 | 2.7 | 89/80/100 | 98/99/96 | 4.7 | 78/97/65 | 74/98/60 | 7.5 |
| Extrema w/ padding | 56/38/100 | 99/98/99 | 2.8 | 62/44/100 | 98/97/100 | 5.0 | 82/73/93 | 89/95/85 | 7.1 |
4.2.2. Pitch Results
4.2.3. Running Speed
5. Discussion
6. Conclusion and Future work
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Conflicts of Interest
References
- Palmer, C.; Hutchins, S. What Is Musical Prosody? Psychology of Learning and Motivation 2016, 46, 245–278. [Google Scholar]
- Computer-aided melody note transcription using the Tony software: Accuracy and efficiency. Proc. TENOR, 2015, p. 23–31.
- Ting-Wei, S.; Yuan-Ping, C.; Li, S.; Yi-Hsuan, Y. Tent: Technique-embedded note tracking for real-world guitar solo recordings. Transactions of the International Society for Music Information Retrieval 2019, 2, 15–28. [Google Scholar]
- Yu, Z.; Ziya, Z.; Xiaobing, L.; Feng, Y.; Sun, M. CCOM-HuQin: An Annotated Multimodal Chinese Fiddle Performance Dataset. Transactions of the International Society for Music Information Retrieval 2023. [Google Scholar]
- Bochen, L.; Xinzhao, L.; Karthik, D.; Zhiyao, D.; Gaurav, S. Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications. IEEE Trans. Multimedia 2019, 21, 522–535. [Google Scholar]
- Helena, C.; Emilia, G. Voice assignment in vocal quartets using deep learning models based on pitch salience. Transactions of the International Society for Music Information Retrieval 2022, 5, 99–112. [Google Scholar]
- Pipa Professional Committee of Shanghai Musicians Association. The Collection for Pipa Grade Examination of China, Shanghai Music Publishing House, 2012.
- Xi, Q.; Bittner, R.; Pauwels, J.; Ye, X.; Bello, J. GuitarSet: A dataset for guitar transcription. 19th International Society for Music Information Retrieval Conference (ISMIR), 2018, p. 453–460.
- Wang, Y.; Jing, Y.; Wei, W.; Cazau, D.; Adam, O.; Wang, Q. PipaSet and TEAS: A Multimodal Dataset and Annotation Platform for Automatic Music Transcription and Expressive Analysis dedicated to Chinese Traditional Plucked String Instrument Pipa. IEEE ACCESS 2022, 10, 113850–113864. [Google Scholar] [CrossRef]
- Drugman, T.; Stylianou, Y.; Kida, Y.; Akamine, M. Voice Activity Detection: Merging Source and Filter-based Information. IEEE Signal Processing Letters 2016, 23, 252–256. [Google Scholar] [CrossRef]
- Quinn, B.G.; Thomson, P.J. Estimating the frequency of a periodic function. Biometrika 1991, 78, 65–74. [Google Scholar] [CrossRef]
- Shi, L.; Nielsen, J.K.; Jensen, J.R.; others. Robust Bayesian Pitch Tracking Based on the Harmonic Model. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2019, 27, 1737–1751, Code available: https://github.com/LimingShi/Bayesian-Pitch-Tracking-Using-Harmonic-model.. [Google Scholar] [CrossRef]
- Shimamura, T.; Kobayashi, H. Weighted autocorrelation for pitch extraction of noisy speech. IEEE transactions on speech and audio processing 2001, 9, 727–730. [Google Scholar] [CrossRef]
- Mauch, M.; Dixon, S. pYIN: A Fundamental Frequency Estimator using Probabilistic Threshold Distributions. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014. Code v1.1.1 available: https://code.soundsoftware.ac.uk/projects/pyin/files.
- Freire, S.; Nézio, L. Study of the tremolo technique on the acoustic guitar: Experimental setup and preliminary results on regularity. Int. Conf. Sound and Music Computing, 2013, pp. 329–334.
- Bello, J.P.; Daudet, L.; Abdallah, S.; Duxbury, C.; Davies, M.; Sandler, M. A tutorial on onset detection in music signals. IEEE Transactions on speech and audio processing 2005, 13, 1035–1047. [Google Scholar] [CrossRef]
- Bock, S.; Widmer, G. Maximum Filter Vibrato Suppression for Onset Detection. International Conference on Digital Audio Effects (DAFx), 2013.
- Bock, S.; Widmer, G. Local Group Delay based Vibrato and Tremolo Suppression for Onset Detection. 14th International Society for Music Information Retrieval Conference (ISMIR), 2013.
- Nord, H.; Zheng, S.; Steven, L.; Manli, W.; Hsing, S.; Quanan, Z.; Nai-Chyuan, Y.; Chi Chao, T.; Henry, L. The Empirical Mode Decomposition and the Hilbert Spectrum for Nonlinear and Non-Stationary Time Series Analysis. Proceedings of the Royal Society of London. Series A: Mathematical, Physical and Engineering Sciences, 1998, Vol. 454, pp. 903–995.
- Manman, L.; Hongwu, Y.; Weitong, G.; Dong, P.; Hongying, S. Endpoint Detection of Noisy Speech using Empirical Mode Decomposition. International Journal of Digital Content Technology and its Application(JDCTA) 2012, 6, 196–203. [Google Scholar]
- Huang, H.; Pan, J. Speech pitch determination based on Hilbert-Huang transform. Signal Processing 2006, 86, 792–803. [Google Scholar] [CrossRef]
- Hess, W. Pitch and Voicing Determination of Speech with an Extension Toward Music Signals. In Springer Handbook of Speech Processing; 2008. Available online: https://api.semanticscholar.org/CorpusID:16346607.
- Morise, M.; Kawahara, H.; Katayose, H. Fast and Reliable F0 Estimation Method Based on the Period Extraction of Vocal Fold Vibration of Singing Voice and Speech. Physics 2009. [Google Scholar]
- Wishart, T. Audible Design: A Plain and Easy Introduction to Practical Sound Composition; Orpheus books, 1994.
- Moulines, E.; Charpentier, F. Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones. Speech Commun. 1990, 9, 453–467. [Google Scholar] [CrossRef]
- Pratzlich, T.; Bittner, R.M.; Liutkus, A.; Müller, M. Kernel additive modeling for interference reduction in multi-channel music recordings. International Conference on Acoustic, Speech and Signal Processing (ICASSP), 2015, p. 584–588. Code available: https://members.loria.fr/ALiutkus/kamir/.
- Cheng-Yuan, L.; Jyh-Shing, R.J. A two-phase pitch marking method for TD-PSOLA synthesis. INTERSPEECH, 2004.
- Sen, B. Introduction to the Nonparametric statistics, n.d. Columbia University. Available online: https://www.stat.columbia.edu/~bodhi/Talks/Intro&NP-Stat.pdf.
- Cleveland, W.S. Robust Locally Weighted Regression and Smoothing Scatterplots. Journal of the American Statistical Association 1978, 74, 829–833. [Google Scholar] [CrossRef]
- Raffel, C.; McFee, B.; Humphrey, E.J. ; others. mir_eval: A transparent implementation of common mir metrics. Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR, 2014, pp. 367–372.
- Chen, Y.H.; Huang, C.F. Sound Synthesis of the Pipa Based on Computed Timbre Analysis and Physical Modeling. IEEE Journal of Selected Topics in Signal Processing, 1070. [Google Scholar]






| Metrics | Methods | |||
|---|---|---|---|---|
| pYIN (%) | BNLS (%) | ZC (%) | Extrema (%) | |
| VAD | 42.6 | 34.7 | 60.1 | 51.3/50.5 |
| PES | 90.9 | 89 | N/A | N/A |
| SPE | 29.2 | 29.6 | 87.7 | 83.9 |
| 1 | pYIN code available: https://code.soundsoftware.ac.uk/projects/pyin
|
| 2 | BNLS code available: https://github.com/LimingShi/Bayesian-Pitch-Tracking-Using-Harmonic-model
|
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).