| 0 | 0 | 63 |
| 下载次数 | 被引频次 | 阅读次数 |
提出了一种基于 ECAPA⁃TDNN 模型的轨道交通车站声纹识别方法,将该方法应用于复杂背景噪声下的安防预警监测。针对传统视频监控在声音事件检测中的不足,本研究构建了包含尖叫声、爆炸声及环境噪声的 4384 个音频样本的领域专用数据集,并提出融合通道注意力机制(SE 模块)与动态统计池化(ASP)的优化策略,显著提升了模型在噪声环境下的鲁棒性。试验表明,模型在测试集上的准确率达 0.8742,其中爆炸声识别召回率超过 96.17%,验证了其在轨道交通场景中的实用价值。
Abstract:This paper proposes a voiceprint recognition method for rail transit stations based on the ECAPA⁃TDNN model, and applied this method to security early warning monitoring under complex background noise. To address the limitations of traditional video surveillance in sound event detection,this paper constructs a domain⁃specific dataset comprising 4384 audio samples, including screams, explosion sounds, and environmental noises. An optimization strategy that integrates the Squeeze⁃and⁃Excitation (SE) module for channel attention mechanism with Attentive Statistical Pooling (ASP) is proposed, significantly enhancing the model's robustness in noisy environments. Experiments demonstrate that the module achieves an accuracy of 0.8742 on the test set, with an explosion sound recognition recall rate exceeding 96.17%, validating its practical value in rail transit scenarios.
[ 1] 裴莹玲,罗晖,张诗慧,等 . 基于改进 Faster R⁃CNN的高铁扣件检测算法[J]. 华东交通大学学报,2023,40(1):75⁃81.
[ 2] 赵力 . 语音信号处理[M]. 2 版 . 北京:机械工业出版社,2009:72.
[ 3] LUCK J E. Automatic speaker verification using cepstral measurements[J]. The Journal of the Acoustical Society of America,1969,46(4B):1026⁃1032.
[ 4] ATAL B S. Effectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification[J]. The Journal of the Acoustical Society of America,1974,55(6):1304⁃1312.
[ 5] DAVIS S, MERMELSTEIN P. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences[J]. IEEE Trans⁃
actions on Acoustics, Speech, and Signal Processing,1980,28(4):357⁃366.
[ 6] HOCHREITER S, SCHMIDHUBER J. Long shortterm memory[J]. Neural Computation,1997,9(8):1735⁃1780.
[ 7] VARIANI E, LEI X, MCDERMOTT E, et al. Deep neural networks for small footprint text⁃dependent speaker verification[C]//Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE,2014:4052⁃4056.
[ 8] CHEN N X, QIAN Y M, YU K. Multi⁃task learning for text⁃dependent speaker verification[C]//Proceedings of Interspeech 2015. ISCA,2015:185⁃189.
[ 9] SNYDER D, GARCIA⁃ROMERO D, SELL G, et al.X⁃vectors: robust DNN embeddings for speaker recognition[C]//Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE,2018:5329⁃5333.
[10] WAIBEL A, HANAZAWA T, HINTON G, et al.Phoneme recognition using time⁃delay neural networks
[J]. IEEE Transactions on Acoustics, Speech, and Signal Processing,1989,37(3):328⁃339.
[11] RAVANELLI M, BENGIO Y. Speaker recognitionfrom raw waveform with SincNet[C]//Proceedings of IEEE Spoken Language Technology Workshop
(SLT). IEEE,2019:1021⁃1028.
[12] ZEINALI H, WANG S, SILNOVA A, et al. BUT system description to VoxCeleb speaker recognition challenge 2019 [EB/OL]. (2019⁃10⁃16). https://arXiv.org/abs/1910.12592.
[13] HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]//Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE,2016:770⁃778.
[14] DESPLANQUES B, THIENPONDT J, DEMUYNCK K. ECAPA⁃TDNN: emphasized channel attention,propagation and aggregation in TDNN based speaker verification [C]//Proceedings of Interspeech 2020.ISCA,2020:3830⁃3834.
[15] OKABE K, KOSHINAKA T, SHINODA K. Attentive statistics pooling for deep speaker embedding[EB/OL].(2018⁃03⁃29).https://arXiv.org/abs/1803.10963.
[16] 梁少铭,李建鑫 . 基于 ECAPA⁃TDNN 模型的声纹识别系统设计[J]. 物联网技术,2025,15(7):16⁃19.
[17] 张家良,张强 . 基于 ECAPA⁃TDNN 网络改进的说话人 确 认 方 法[J]. 电 脑 知 识 与 技 术 ,2024,20(1):25⁃28.
[18] 张敏敏,马骏,龚晨晓,等 . 基于 MATLAB 的声纹识别系统软件的设计[J]. 科技视界,2013,3(22):7.
[19] 杨宇奇 . 基于多分支聚合网络的短语音说话人确认方法研究[D]. 哈尔滨:哈尔滨工业大学,2021.
[20] 林玲惠,张永富,张馨月,等 . 基于声纹识别的安全保障系统设计[J]. 信息通信,2020,33(6):112⁃114.
[21] 叶 田 田 . 声 纹 识 别 系 统 设 计[J]. 工 业 控 制 计 算 机 ,2012,25(6):88⁃89.
基本信息:
引用信息:
[1]谢烨,陈捷,陈昌进,等.基于ECAPA‑TDNN模型的轨道交通车站安防预警监测声纹识别方法[J],2026(04):77-84.
基金信息:
浙江省交通投资集团有限公司科技计划项目(202301)
2026-07-15