Robust Estimation of Direction of Arrival and Time Frequency Masks for Speech Enhancement

Abstract

Nowadays, we use speech-enabled smart devices to improve human-machine interaction. One of the newlineprimary tasks in these speech-enabled devices is speech enhancement. Speech enhancement is the extraction newlineof the desired speech signal from the noisy and reverberant signals recorded by the microphones. The noise newlinecould be background noise or interfering speakers. The performance of the speech-enabled devices relies newlinesignificantly on the performance of the speech enhancement methods. Speech enhancement is also a primary newlinetask in devices that improve the comfort of human-human interaction, such as hearing aids and hands-free newlinecommunication devices. newlineAmong several speech enhancement methods, beamformers and Time-Frequency (TF) mask-based newlinemethods are widely used. Beamformers are linear spatial filters that aim to boost the signal coming from a newlinespecific direction by appropriate configuration of the microphone array, and in doing so attenuates interfering newlinesignals from other directions. A source TF mask identifies the time-frequency regions where the source is newlinedominant and can be applied on the mixture TF representation to extract the desired source. Recently, newlineenhancement with mask-aided beamformers has become popular, where TF masks are used to estimate newlineSecond order Statistics (SoS) required for computing the weights of the beamformer. newlineSeveral beamformers and a few TF mask estimators need the Direction of Arrival (DoA) of the sources, newlinewhich must be estimated from the microphone data, if not known a priori. DoA refers to the azimuth and newlineelevation angles of arrival of the sound sources with respect to the microphone array axis. For example, to newlinedesign a set of beamformers called the data-independent beamformers, DoA of the source to be enhanced newlineis necessary. Also, the DoA of the desired and undesired sources are required to set the constraints newlinewhile designing data-dependent beamformers. In literature, a few TF masks are estimated based on prior newlineknowledge of the DoA of the sources. In all the enhancement methods that utilize the DoA

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced