基于核函数领域自适应的语音声源DOA分类
PDF下载 (140)刘明民,章联军,叶庆卫,陆志华*.基于核函数领域自适应的语音声源DOA分类[J].宁波大学学报(理工版),2025,38(3):43-51.DOI:10.20098/j.cnki.1001-5132.2023.1127
LIU Mingmin,ZHANG Lianjun,YE Qingwei,LU Zhihua.DOA classification of single speech sound sources based on kernel function domain adaptation[J].Journal of Ningbo University(Natural Science & Engineering Edition),2025,38(3):43-51.DOI:10.20098/j.cnki.1001-5132.2023.1127
| Title: | DOA classification of single speech sound sources based on kernel function domain adaptation |
| 作者: | 刘明民, 章联军, 叶庆卫, 陆志华* |
| Author(s): | LIU Mingmin, ZHANG Lianjun, YE Qingwei, LU Zhihua |
| 关键词: | 语音信号处理; 分类; 核函数; 领域自适应; 机器学习 |
| Keywords: | speech signal processing; classification; kernel function; domain adaptation; machine learning |
| 分类号: | TN912.3 |
| DOI: | 10.20098/j.cnki.1001-5132.2023.1127 |
| 文献标识码: | A |
| 摘要: | 鉴于训练和测试阶段存在不同的噪声或混响环境, 并且由于真实数据的稀缺会降低语音声源到达方向(Direction of Arrival, DOA)的分类准确性, 因此提出一种基于核函数领域自适应的机器学习DOA分类算法。通过优化结构风险函数和减小域之间的条件分布差异, 实现对训练数据的适应性学习, 从而提升测试数据的分类准确率。实验结果证明在中小型数据集中, 新算法在各种声学条件下均明显优于对比的深度学习算法。 |
| Abstract: | Given the existence of different noise or reverberant environments in the training and testing phases, and due to the scarcity of real data, the classification accuracy of the Direction of Arrival (DOA) of speech sound sources shall be reduced. In this study, we propose a machine learning DOA classification algorithm based on kernel function domain adaptation, which achieves adaptive learning of training data by optimizing the structural risk function and reducing the conditional distribution difference between domains, thus improving the classification accuracy of test data. The experimental results demonstrate that the algorithm significantly outperforms the contrasting deep learning algorithms under various acoustic conditions in small and medium-sized datasets. |
| 参考文献 /References: | [1].GANNOT S, VINCENT E, MARKOVICH-GOLAN S, et al. A consolidated perspective on multimicrophone speech enhancement and source separation[J]. IEEE/ ACM transactions on audio, speech, and language processing, 2017, 25(4):692-730. [2].WANG D L, CHEN J T. Supervised speech separation based on deep learning: an overview[J]. IEEE/ACM transactions on audio, speech, and language processing, 2018, 26(10):1702-1726. [3].SCHMIDT R. Multiple emitter location and signal parameter estimation[J]. IEEE transactions on antennas and propagation, 1986, 34(3):276-280. [4].DIBIASE J H, SILVERMAN H F, BRANDSTEIN M S. Robust localization in reverberant rooms[M]// BRANDSTEIN M, WARD D. Microphone arrays. Berlin: Springer, 2001:157-180. [5].HE W P, MOTLICEK P, ODOBEZ J M. Adaptation of multiple sound source localization neural networks with weak supervision and domain-adversarial training[C]// IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Brighton, UK. IEEE, 2019: 770-774. [6].MUANDET K, FUKUMIZU K, SRIPERUMBUDUR B, et al. Kernel mean embedding of distributions: a review and beyond[J]. Foundations and trends in machine learning, 2017, 10(1/2):1-141. [7].LONG M S, WANG J M, DING G G, et al. Adaptation regularization: a general framework for transfer learning [J]. IEEE transactions on knowledge and data engineering, 2014, 26(5):1076-1089. [8].GANNOT S, BURSHTEIN D, WEINSTEIN E. Signal enhancement using beamforming and nonstationarity with applications to speech[J]. IEEE transactions on signal processing, 2001, 49(8):1614-1626. [9].COHEN I. Relative transfer function identification using speech signals[J]. IEEE transactions on speech and audio processing, 2004, 12(5):451-459. [10].MÜLLER K, MIKA S, RÄTSCH G, et al. An introduction to kernel-based learning algorithms[J]. IEEE transactions on neural networks, 2001, 12(2):181-201. [11].VAPNIK V N. An overview of statistical learning theory [J]. IEEE transactions on neural networks, 1999, 10(5): 988-999. [12].GHIFARY M, BALDUZZI D, KLEIJN W B, et al. Scatter component analysis: a unified framework for domain adaptation and domain generalization[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 39(7):1414-1430. [13].SCHÖLKOPF B, SMOLA A, MÜLLER K R. Nonlinear component analysis as a kernel eigenvalue problem[J]. Neural computation, 1998, 10(5):1299-1319. [14].BELKIN M, NIYOGI P, SINDHWANI V. Manifold regularization: a geometric framework for learning from labeled and unlabeled examples[J]. Journal of machine learning research, 2006, 7(1):2399-2434. [15].LONG M S, WANG J M, DING G G, et al. Transfer feature learning with joint distribution adaptation[C]// IEEE International Conference on Computer Vision. Sydney, Australia. IEEE, 2013:2200-2207. [16].COVER T, HART P. Nearest neighbor pattern classifi- cation[J]. IEEE transactions on information theory, 1967, 13(1):21-27. [17].SCHÖLKOPF B, PLATT J, HOFMANN T. Analysis of representations for domain adaptation[C]//Advances in Neural Information Processing Systems: Proceedings of the 2006 Conference, Neural Information Processing Systems. 2007:137-144. [18].QUANZ B, HUAN J. Large margin transductive transfer learning[C]//Proceedings of the 18th ACM Conference on Information and Knowledge Management. Hong Kong, China. ACM, 2009:1327-1336. [19].PANAYOTOV V, CHEN G G, POVEY D, et al. Librispeech: an ASR corpus based on public domain audio books[C]//IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). South Brisbane, Australia. IEEE, 2015:5206-5210. [20].HADAD E, HEESE F, VARY P, et al. Multichannel audio database in various acoustic environments[C]//IEEE 14th International Workshop on Acoustic Signal Enhancement (IWAENC). Juan-les-Pins, France. IEEE, 2014:313-317. [21].VARGA A, STEENEKEN H J M. Assessment for automatic speech recognition: Ⅱ. NOISEX-92: a database and an experiment to study the effect of additive noise on speech recognition systems[J]. Speech communication, 1993, 12(3):247-251. [22].THIEMANN J, ITO N, VINCENT E. The diverse environments multi-channel acoustic noise database (demand): a database of multichannel environmental noise recordings[C]//Proceedings of Meetings on Acoustics. Montreal, Canada. ASA, 2013:19-33. [23].BIANCO M J, GANNOT S, FERNANDEZ-GRANDE E, et al. Semi-supervised source localization in reverberant environments with deep generative modeling[J]. IEEE access, 2021, 9:84956-84970. [24].VERBURG-RIEZU S A, FERNANDEZ-GRANDE E. Room impulse response dataset - ACT, DTU Elektro (011, IEC; plane, sphere)[DB/OL]. [2022-07-12]. https://data.dtu. dk/articles/dataset/Room_Impulse_Response_Dataset_-_ ACT_DTU_Elektro_011_IEC_plane_sphere_/14320166/1. |
| 备注/Memo: | 收稿日期: 2023−11−23 宁波大学学报(理工版)网址: http://journallg.nbu.edu.cn/ 基金项目: 国家自然科学基金(61801255) 第一作者: 刘明民, 硕士研究生, 主要研究方向: 语音声源定位。E-mail: 2723144048@qq.com *通信作者: 陆志华, 副教授, 主要研究方向: 语音信号处理。E-mail: luzhihua@nbu.edu.cn 宁波大学学报(理工版)网址:http://journallg.nbu.edu.cn/ |