Continuous Speech Separation Using Speaker Inventory for Long Recording

Han, Cong; Luo, Yi; Li, Chenda; Zhou, Tianyan; Kinoshita, Keisuke; Watanabe, Shinji; Delcroix, Marc; Erdogan, Hakan; Hershey, John R.; Mesgarani, Nima; Chen, Zhuo

doi:10.21437/Interspeech.2021-338

Continuous Speech Separation Using Speaker Inventory for Long Recording

Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by using the target speaker’s voice snippet and jointly separating all participating speakers by using a pool of additional speaker signals, which is known as speech separation using speaker inventory (SSUSI). However, all these systems ideally assume that the pre-enrolled speaker signals are available and are only evaluated on simple data configurations. In realistic multi-talker conversations, the speech signal contains a large proportion of non-overlapped regions, where we can derive robust speaker embedding of individual talkers. In this work, we adopt the SSUSI model in long recordings and propose a self-informed, clustering-based inventory forming scheme for long recording, where the speaker inventory is fully built from the input signal without the need for external speaker signals. Experiment results on simulated noisy reverberant long recording datasets show that the proposed method can significantly improve the separation performance across various conditions.

doi: 10.21437/Interspeech.2021-338

Cite as: Han, C., Luo, Y., Li, C., Zhou, T., Kinoshita, K., Watanabe, S., Delcroix, M., Erdogan, H., Hershey, J.R., Mesgarani, N., Chen, Z. (2021) Continuous Speech Separation Using Speaker Inventory for Long Recording. Proc. Interspeech 2021, 3036-3040, doi: 10.21437/Interspeech.2021-338

@inproceedings{han21d_interspeech,
  author={Cong Han and Yi Luo and Chenda Li and Tianyan Zhou and Keisuke Kinoshita and Shinji Watanabe and Marc Delcroix and Hakan Erdogan and John R. Hershey and Nima Mesgarani and Zhuo Chen},
  title={{Continuous Speech Separation Using Speaker Inventory for Long Recording}},
  year=2021,
  booktitle={Proc. Interspeech 2021},
  pages={3036--3040},
  doi={10.21437/Interspeech.2021-338}
}