Submitted:
12 July 2026
Posted:
15 July 2026
You are already at the latest version
Abstract
This paper proposes a switching mode beamforming (SMVB) technique for a two-microphone array setup in a multi-source acoustic environment with one dominant target and two interferers. The proposed method adaptively switches between a robust minimum variance beamformer (RMVB) and a linearly constrained minimum variance (LCMV) beamformer based on the eigenvalue characteristics of the spatial covariance matrix. The eigenvalue distribution is used to infer the signal environment and select the appropriate beamforming strategy. Simulation results show that the proposed SMVB method outperforms conventional MVDR, RMVB, and LCMV beamformers in terms of interference suppression and target signal preservation, achieving improved output signal-to-interference-plus-noise ratio (SINR) and overall signal quality.
Keywords:
speech enhancement
; robust beamforming
; switching mode beamforming
1. Introduction
Beamforming is a spatial signal processing technique widely used in microphone array systems to enhance a desired signal arriving from a specific direction while suppressing interference and noise from other directions [1,2]. By exploiting spatial diversity, beamforming enables improved signal quality in applications such as speech enhancement, hearing aids, teleconferencing, and human–machine interaction [3,4]. Beamforming techniques can be broadly categorized into conventional and adaptive methods. Conventional beamformers, such as delay-and-sum [5], employ fixed weights based on array geometry, resulting in low computational complexity but limited interference suppression.
Adaptive beamformers exploit signal statistics to dynamically adjust their weights, enabling improved performance in complex acoustic environments [6]. Prominent adaptive approaches include the Minimum Variance Distortionless Response (MVDR) beamformer [7,8], the Linearly Constrained Minimum Variance (LCMV) beamformer, and the Robust Minimum Variance Beamformer (RMVB) [9]. In recent years, neural network (NN)–based beamforming approaches have gained attention, where deep learning models such as convolutional neural networks (CNNs), are used to estimate beamforming weights or time–frequency masks for speech enhancement [10,11,12,13,14,15]. Despite the success of neural-network based approaches, adaptive beamforming methods remain attractive due to their lower computational complexity, no preprrocessing, interpretability, and robustness.
Although adaptive beamformers are effective for speech enhancement in microphone arrays, their performance is inherently constrained in small-aperture configurations. The primary constraint arises from reduced spatial diversity and limited inter-microphone spacing, which restrict the achievable spatial resolution and the ability to discriminate between closely spaced sources. This limitation becomes particularly evident in a two-microphone array configuration with a dominant target and multiple interferers. In such scenarios, the performance of conventional adaptive beamformers is fundamentally bounded by the restricted spatial degrees of freedom [16]. For instance, the widely used MVDR beamformer has only degrees of freedom (where N is the number of microphones) [17]. In a two-microphone configuration, this effectively suppresses only one dominant interference source. However, when multiple interferers have comparable power, the beamformer cannot selectively cancel any single source. Instead, it provides only partial suppression of both, leading to noticeable residual noise in the output. The LCMV beamformer extends MVDR by introducing multiple linear constraints to control the beamformer response in different directions. While LCMV can perform effectively in a two-microphone configuration when one interferer is dominant—by allocating its limited degrees of freedom to suppress it—its performance degrades significantly in the presence of two equal-power interferers. This is due to insufficient degrees of freedom and ambiguity in enforcing multiple constraints simultaneously [18]. The RMVB, A robust variant of MVDR beamformer, attempts to mitigate steering vector mismatch and covariance estimation errors. It generally demonstrates improved performance over MVDR and slightly lower performance compared to LCMV in scenarios with unequal interferer strengths [19]. However, in the presence of equal-power interferers, it often yields more stable performance than LCMV, as the latter becomes overconstrained under limited degrees of freedom, whereas the RMVB maintains resilience despite its limited interference suppression capability [20].
From the above discussion, it can be noted that a beamformer that performs well when one interferer is dominant may not be optimal when multiple interferers of comparable strength are present. This variability highlights the limitation of relying on a single beamforming strategy. Therefore, it is more effective to adopt adaptive approaches that combine or switch between beamformers based on the underlying signal conditions, enabling improved performance across a wider range of acoustic environments. Motivated by this, we propose a switching mode beamforming (SMVB) approach that dynamically selects between RMVB and LCMV based on the eigenvalue characteristics of the spatial covariance matrix.
The remainder of this paper is organized as follows: Section II describes the signal model and the experimental setup. Section III presents the proposed eigenvalue-based Switching Mode Beamforming (SMVB) technique. Section IV discusses the experimental results, and Section V concludes the study.
2. System Model and Experimental Setup
2.1. Signal Model
Consider a compact microphone array of M omnidirectional sensors. In the Short-Time Fourier Transform (STFT) domain, with time-frame index t and frequency-bin f, the received signal vector is modeled as:
Here, denotes the clean target speech signal and is the Relative Transfer Function (RTF), or steering vector [3,13], corresponding to the target source. The terms and represent the k-th interference signal and its associated steering vector, respectively. K denotes the number of interfering sources. The vector models spatially uncorrelated sensor noise. From the received signal, , the spatial covariance matrix (SCM) is constructed as:
The SCM provides the spatial statistical information required to design, analyze, and optimize beamformers for interference suppression and signal enhancement. Note that including sensor noise ensures the spatial covariance matrix remains strictly positive definite, thereby guaranteeing its invertibility and improving numerical stability in subsequent beamforming computations.
2.2. Simulation Setup and Experimental Conditions
In general, M can take different values depending on the array geometry and application requirements. In this study, however, we use a two-microphone () setup to reflect the practical constraints of mobile edge devices. The target source is fixed at broadside (), ensuring a known and consistent steering vector.Additionally, two interfering sources (K = 2) are assumed to arrive from varying azimuths with equal average power over the entire utterance, although their power may vary at the frame level. This setup forms a challenging, low DoF multi-interferer scenario that highlights the fundamental vulnerabilities of static spatial filters.
To systematically evaluate the performance of the beamformers under the described multi-interferer scenario, the acoustic environment is simulated at a sampling frequency of Hz. The setup employs a two-microphone array with an inter-element spacing of cm, representative of a standard compact geometry. To strictly isolate the performance of the beamforming logic against spatial geometric vulnerabilities without the confounding effects of multipath propagation, the environment is modeled under purely anechoic conditions. The Acoustic Transfer Functions (ATFs) are generated utilizing the free-space propagation model of the Image Source Method (ISM) [21,22] with zero reflections. The target source remains fixed at , while interfering sources are distributed across the azimuth at a fixed radial distance of m.
The final observed mixture, , is constructed by summing the spatialized target signal, the aggregate spatialized interference, and spatially white sensor noise. The sensor noise is explicitly added to ensure a strictly positive-definite spatial covariance matrix. The target is chosen from selected from the LJSpeech [23] and distinct interferences are drawn from LibriSpeech [24], and MUSAN [25] datasets respectively.
3. Proposed Approach
In practical speech enhancement scenarios, a single utterance is often characterized by highly dynamic interference conditions. As illustrated in Figure 2, the instantaneous power of multiple interferers varies significantly across time frames. In some frames, the interferers exhibit comparable (equal) power levels, while in others, their powers differ substantially. Additionally, there are instances in which one of the interferers is inactive or absent. Such non-stationary and frame-dependent interference patterns pose a significant challenge to conventional fixed beamforming approaches.
Figure 1.
Block diagram of the proposed Switching Minimum Variance Beamformer (SMVB).

Figure 2.
Instantaneous normalized envelope power of two interferers across time, illustrating dynamic scenarios including equal-power regions, power imbalance, and intermittent absence within a single utterance.
Figure 2.
Instantaneous normalized envelope power of two interferers across time, illustrating dynamic scenarios including equal-power regions, power imbalance, and intermittent absence within a single utterance.

Motivated by these observations, we propose a switching beamforming strategy that adaptively selects the beamformer based on instantaneous interference characteristics, enabling robust target enhancement under varying conditions. The proposed beamforming strategy is shown in Figure 1.
3.1. Spatial Anisotropy
The first failure mode to address is spatial rank deficiency. In an anechoic setup, a single dominant interferer yields a rank-1 interference covariance matrix. However, when multiple spatially distinct interferers possess comparable power, the interference subspace expands, exhausting the available spatial DoF. We perform an eigendecomposition on the estimated interference-plus-noise sample covariance matrix :
where are the eigenvalues sorted in descending order. To quantify the spatial coherence of the noise field, we define the eigenvalue anisotropy ratio :
Since the observation model includes spatially white intrinsic sensor noise with variance , the total spatial covariance matrix is strictly positive definite, naturally bounding . A large indicates a highly directional, single-dominant interference subspace where the array possesses the necessary DoF to construct a stable spatial null using LCMV. Conversely, as multiple interferers achieve comparable power, the spatial eigenvalues equalize, driving . Attempting an unconstrained LCMV inversion in this exhausted state severely degrades the WNG [16,20,26].
3.2. Spatial Coherence
The second fundamental failure mode is purely geometric and occurs when a directional interferer physically encroaches upon the target’s look direction. In a dynamic multi-interferer scenario, the principal eigenvector captures the instantaneous spatial signature of the most dominant interfering component within that specific time-frequency bin. To detect spatial encroachment without requiring explicit Direction of Arrival (DOA) estimation, we evaluate the spatial correlation coefficient between the target steering vector and this dominant interference subspace:
We utilize this subspace projection to establish a threshold-based target-exclusion zone. As an interferer moves closer to the broadside target, the spatial vectors become increasingly collinear ( ). Enforcing a strict spatial null on such an encroaching interference vector rapidly inflates the weight norm [20]. In a compact array constrained to a single spatial DoF, this forces the null into direct mathematical conflict with the distortionless constraint, drastically amplifying sensitivity to slight steering vector mismatches and causing severe target signal self-cancellation [27].
3.3. Switching Minimum Variance Beamformer
Because continuous blending and dimensional expansions face severe practical limitations on compact arrays, we formulate the SMVB as a strictly state-dependent diagnostic router. To account for the non-stationarity of speech, the routing logic is evaluated per time-frequency bin . Let denote the instantaneous binary spatial state indicator, governed by the predefined safety thresholds and :
In practice, these spatial thresholds are empirically tuned using a simulated validation set, yielding robust default values of and . The SMVB functions as a strict logical gate that assigns the final weight vector based on an evaluated routing state :
When , the architecture exploits its single spatial DoF to carve a precise null using the standard LCMV formulation. However, when (indicating DoF starvation via diffusion or impending signal distortion due to null-constraint conflict), standard inversion is bypassed, and the system safely routes to a regularized state.
3.4. Temporal Smoothing and Regularization
To prevent rapid chattering of the routing state across adjacent frames due to transient speech overlaps, a first-order recursive moving average is applied to the instantaneous state :
where is the temporal smoothing factor. The final routing decision utilized in the SMVB gate is determined by thresholding the smoothed state at .
Upon routing to the RMVB state (), the beamformer utilizes a dynamically loaded covariance matrix to stabilize the inversion:
To avoid the blindness of static regularization, the dynamic loading parameter is bounded by the minimum eigenvalue of the total sample covariance matrix [28]:
where is a scaling constant and provides an absolute upper bound. We explicitly anchor the loading factor to the total observation covariance rather than the noise covariance because the total covariance is directly observable and strictly positive definite. This cross-matrix anchoring provides a highly stable, signal-aware upper bound that prevents excessive loading when the target speech dominates the frame. Because the SMVB relies only on the eigendecomposition and inversion of a matrix per frequency bin, its computational complexity is strictly bounded by . For minimal array geometries (), this introduces negligible overhead compared to a standard LCMV, mathematically guaranteeing its viability for strict edge-compute latency budgets.
4. Experimental Setup and Results
To empirically validate the routing architecture, we evaluate the proposed SMVB against conventional spatial filters under the specific theoretical failure modes defined in Section 3. All evaluations are conducted using the dual-microphone ( ) anechoic simulated acoustic datasets detailed in Section 2.1.
We first aim to test the beamforming logic itself, rather than the accuracy of the systems that estimate the acoustic environment. Therefore, we bypass standard estimation steps and compute the target steering vector and interference covariance matrices as perfect Oracle quantities, derived directly from the clean ground-truth signals. We note that Oracle quantities represent an unachievable upper bound in practice, which is why all of the masked results should be interpreted as the best SCM estimations made. We evaluate the proposed SMVB against three baselines, (a) an Unmasked standard LCMV, (b) an Oracle Masked LCMV, and (c) a Multi-Frame SMVB (MF-SMVB) utilizing a temporal tap length of . The proposed SMVB operates with empirically derived routing thresholds and . The dynamic loading parameter is governed by the scaling constant and bounded by . Performance is quantified using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) to track spatial nulling energy dynamics, and Short-Time Objective Intelligibility (STOI) [29] to assess perceptual quality and speech intelligibility preservation.
4.1. Mitigation of Interferer Collinearity
Oracle quantities represent a theoretical bound. Thus, "Masked" results denote best-case performance. We compare our proposed SMVB ( , , , ) against three baselines: Unmasked LCMV, Oracle Masked LCMV, and Multi-Frame SMVB ( ). System performance is evaluated using SI-SDR for spatial nulling and STOI [29] for speech intelligibility.
However, as the interferer enters the physical exclusion zone ( to ), the spatial covariance matrix becomes ill-conditioned due to spatial collinearity. Consequently, the standard Masked LCMV places a spatial null in the look direction, causing target self-cancellation and a drop to nearly 0 dB SI-SDR. In contrast, the SMVB detects the threshold crossing and transitions to the regularized RMVB state. As shown in Figure 3, this switching logic prevents severe signal degradation within the exclusion zone, improving SI-SDR by up to 15 dB over the standard LCMV. Qualitatively, this preserves target speech intelligibility and prevents the muffling typically caused by nearby interferers.
Notably, the MF-SMVB ( ) yields consistently lower performance across all angles. In an anechoic STFT framework, stacking consecutive frames causes temporal smearing of phase information. This expands the covariance matrix dimensions but distorts the precise inter-microphone phase relationships required for spatial nulling, demonstrating that spatio-temporal expansion degrades overall interference rejection in compact arrays.
4.2. Overloading of Interferers
The second experiment is motivated by the physical limits of the array’s degrees of freedom. We evaluate the spatial filters in an overloaded acoustic state by incrementally increasing the number of active, equal-power interferers from 1 to 4. Because an array possesses exactly one spatial DoF, any scenario with interferers induces total DoF starvation, causing the exact spatial nulling constraint of a standard LCMV to physically fail and leading to sensor noise amplification.
As shown in Figure 4, when the environment is fully determined ( ), the SMVB actually outperforms the Masked LCMV baseline by approximately 5 dB in SI-SDR. This performance gain is directly attributed to the temporal smoothing mechanism ( ) of the routing state and the dynamic loading bounds, which stabilize frame-by-frame weight fluctuations and mitigate the covariance estimation variance inherently introduced by harsh ideal ratio masks. As the number of interferers increases and the spatial eigenvalues equalize, the SMVB Anisotropy metric ( ) accurately flags the exhausted spatial state and bounds the inversion process via dynamic loading. The data confirms that the SMVB safely navigates this DoF starvation, achieving graceful degradation and maintaining parity or a slight superiority over the baselines in both SI-SDR and STOI. Crucially, this validates that the diagnostic switching logic does not introduce secondary mathematical penalties or catastrophic loss in worst-case overloaded scenarios. The SMVB effectively preserves speech naturalness and provides a highly robust, computationally efficient spatial filtering solution for unpredictable multi-source edge environments.
5. Conclusion
In this work, we proposed the Switching Minimum-Variance Beamformer (SMVB) to overcome the fundamental spatial limitations of compact microphone arrays. By evaluating spatial eigenvalue anisotropy and target-interferer coherence, the SMVB dynamically routes between an unconstrained LCMV and a regularized RMVB. Anechoic experiments demonstrate that this architecture prevents target self-cancellation during spatial encroachment—yielding up to a 15 dB relative SI-SDR gain—and gracefully manages total degrees-of-freedom (DoF) starvation. By avoiding computationally expensive multi-frame expansions, the SMVB provides a highly robust, practical solution for low-latency edge deployments.
While this study establishes the theoretical performance bounds of the SMVB using Oracle covariance matrices, practical deployment requires robustness against estimation errors. Future work will integrate this diagnostic router with lightweight, causal neural network-based Voice Activity Detectors (VAD) and covariance estimators to evaluate real-time tracking capabilities in highly non-stationary, uncooperative acoustic environments.
Author Contributions
All authors have contributed equally to this work. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
No new data were created or analyzed in this study. Data sharing is not applicable to this article.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Veen, B.D.V.; Buckley, K.M. Beamforming: A versatile approach to spatial filtering. IEEE ASSP Mag. 1988, 5. [Google Scholar] [CrossRef]
- Benesty, J.; Chen, J.; Huang, Y. Microphone Array Signal Processing; Springer Science & Business Media, 2008. [Google Scholar]
- Gannot, S.; Vincent, E.; Markovich-Golan, S.; Ozerov, A. A consolidated perspective on multimicrophone speech enhancement and source separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2017; 25. [Google Scholar]
- Doclo, S.; Kellermann, W.; Makino, S.; Nordholm, S.E. Multichannel signal enhancement algorithms for assisted listening devices. IEEE Signal Process. Mag. 2015, 32. [Google Scholar] [CrossRef]
- Johnson, D.H.; Dudgeon, D.E. Array Signal Processing: Concepts and Techniques; PTR Prentice Hall, 1993. [Google Scholar]
- Widrow, B.; Mantey, P.E.; Griffiths, L.J.; Goode, B.B. Adaptive antenna systems. In Proceedings of the IEEE, 1967; 55. [Google Scholar]
- Capon, J. High-resolution frequency-wave number spectrum analysis. In Proceedings of the IEEE, 1969; 57. [Google Scholar]
- Souden, M.; Benesty, J.; Affes, S. On optimal frequency-domain multichannel linear filtering for noise reduction. IEEE Transactions on Audio, Speech, and Language Processing, 2010; 18. [Google Scholar]
- Li, J.; Stoica, P.; Wang, Z. On robust Capon beamforming and diagonal loading. IEEE Trans. Signal Process. 2003, 51. [Google Scholar] [CrossRef]
- Heymann, J.; Drude, L.; Haeb-Umbach, R. Neural network based spectral mask estimation for acoustic beamforming. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016. [Google Scholar]
- Erdogan, H.; Hershey, J.R.; Watanabe, S.; Mandel, M.I.; Roux, J.L. Improved MVDR beamforming using single-channel mask prediction networks. In Proceedings of the Interspeech, 2016. [Google Scholar]
- Meng, Z.; Watanabe, S.; Hershey, J.R.; Erdogan, H. Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017. [Google Scholar]
- Li, Y.; et al. Neural RTF estimation for multi-microphone speech enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022; 30. [Google Scholar]
- Zhang, X.; et al. Design of a robust MVDR beamforming method with low-latency by reconstructing covariance matrix for speech enhancement. Appl. Acoust. 2023, 211. [Google Scholar] [CrossRef]
- Olivieri, M.; Comanducci, L.; Pezzoli, M.; Balsarri, D.; Menescardi, L.; Buccoli, M.; Pecorino, S.; Grosso, A.; Antonacci, F.; Sarti, A. Real-time multichannel speech separation and enhancement using a beamspace-domain-based lightweight CNN. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023. [Google Scholar]
- Elko, G.W. Microphone array systems for electronic hearing aids. In Advances in Acoustics, Noise, and Vibration-2000; CRC Press, 2000. [Google Scholar]
- Trees, H.L.V. Optimum Array Processing: Part IV of Detection, Estimation, and Modulation Theory; John Wiley & Sons, 2002. [Google Scholar]
- Frost, O.L. An algorithm for linearly constrained adaptive array processing. In Proceedings of the IEEE, 1972; 60. [Google Scholar]
- Vorobyov, S.A.; Gershman, A.B.; Luo, Z.Q. Robust adaptive beamforming using worst-case performance optimization: A solution to the signal mismatch problem. IEEE Trans. Signal Process. 2003, 51. [Google Scholar] [CrossRef]
- Cox, H.; Zeskind, R.; Owen, M. Robust adaptive beamforming. IEEE Trans. Acoust. Speech Signal Process. 1987, 35. [Google Scholar] [CrossRef]
- Allen, J.B.; Berkley, D.A. Image method for efficiently simulating small-room acoustics. J. Acoust. Soc. Am. 1979, 65. [Google Scholar] [CrossRef]
- Scheibler, R.; Bezzam, E.; Dokmanić, I. Pyroomacoustics: A python package for audio room simulation and array processing algorithms. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018. [Google Scholar]
- Ito, K.; et al. The LJ Speech Dataset. 2017. Available online: https://keithito.com/LJ-Speech-Dataset/.
- Panayotov, V.; Chen, G.; Povey, D.; Khudanpur, S. Librispeech: an ASR corpus based on public domain audio books. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. [Google Scholar]
- Snyder, D.; Chen, G.; Povey, D. MUSAN: A Music, Speech, and Noise Corpus. arXiv 2015, arXiv:1510.08484. [Google Scholar]
- Liu, S.; et al. Microphone Array Signal Processing and Deep Learning for Speech Enhancement. arXiv 2025, arXiv:2501.07215. [Google Scholar]
- Wang, Y.; et al. Robust Adaptive Beamforming Based on Interference Covariance Matrix Reconstruction and Steering Vector Estimation. IEEE Transactions on Signal Processing, 2026. [Google Scholar]
- Mestre, X.; Lagunas, M.A. Diagonal loading for finite sample size beamforming: An asymptotic approach. IEEE Trans. Signal Process. 2005, 54. [Google Scholar]
- Taal, C.H.; Hendriks, R.C.; Heusdens, R.; Jensen, J. An algorithm for intelligibility prediction of time–frequency weighted noisy speech. IEEE Transactions on Audio, Speech, and Language Processing, 2011; 19. [Google Scholar]
Figure 3.
Spatial collinearity stress test. The proposed SMVB prevents the self-cancellation suffered by the Masked LCMV within the exclusion zone (shaded).
Figure 3.
Spatial collinearity stress test. The proposed SMVB prevents the self-cancellation suffered by the Masked LCMV within the exclusion zone (shaded).

Figure 4.
System robustness under increasing number of interfering sources.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.