跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.66) 您好!臺灣時間:2026/08/15 21:53
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:賴辰瑋
研究生(外文):Chen-Wei Lai
論文名稱:強健性語音辨認之研究:語音前端端點偵測與語音強化法
論文名稱(外文):The Research on the Voice Activity Detection and Speech Enhancement for Noisy Speech Recognition
指導教授:洪志偉洪志偉引用關係
指導教授(外文):Jeih-Weih Hung
學位類別:碩士
校院名稱:國立暨南國際大學
系所名稱:電機工程學系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2005
畢業學年度:93
語文別:中文
論文頁數:84
中文關鍵詞:語音辨識語音端點偵測語音強化處理
外文關鍵詞:Speech RecognitionVoice Activity DetectionSpeech Enhancement Method
相關次數:
  • 被引用被引用:2
  • 點閱點閱:449
  • 評分評分:
  • 下載下載:108
  • 收藏至我的研究室書目清單書目收藏:1
自動語音辨識,在實際系統應用中,語音信號經常受到環境雜訊的影響而降低其辨識率。為了提升系統的效能,許多研究語音辨識的人員歷年來不斷的研究語音的強健技術,期望達到系統的最佳化。而本論文將針對環境中的加成性雜訊做估測,並在線性頻譜領域提出一些強健的補償方式。
在語音強健處理中,雜訊之估測為相當重要的步驟之一,因此本論文討論了幾種語音端點偵測(Voice Activity Detection)之方法,包括階層統計濾波器(Order Statistic Filter,OSF)、分頻階層統計濾波器(Subband Order Statistics Filter,SOSF)、長期頻譜分散法(Long-Term Spectrum Divergence,LTSD)、Kullback-Leibler量測法(KL)、能量偵測法(Energy)及亂度偵測法(Entropy)。對於一段輸入之語音訊號,透過端點偵測,將包含語音之起始端點偵測出來。其中,以Kullback-Leibler量測法具有最佳的端點偵測率,因此,在結合強健性技術後,也具有較佳的辨識結果。
本論文所討論的語音強化技術中,在線性頻譜下所使用的方法包含非線性頻譜消去法(Nonlinear Spectral Subtraction,NSS),以及韋納濾波器(Winner Filter,WF),而梅爾頻譜消去法(Mel Spectral Subtraction,MSS)則是處理於梅爾刻度的頻譜上,另外我們也提出了倒頻譜補償法(Cepstral Statistics Compensation,CSC),藉此有效地拉近訓練語料與測試語料的統計特性。
When a speech recognizer is applied in a real environment, its performance is often degraded seriously due to the existence of additive noise. In order to improve the robustness of the recognition system under noisy conditions, various approaches have been proposed, one direction of these approaches is attempt to detect the presence the presence of noise, to estimate the characteristics of the noise and then to remove or alleviate the noise in speech signals.
In the thesis, we first study several voice activity detection (endpoint detection) approaches, which may detect the noise-only portions in a speech sequence. Then the noise statistics can be estimated via these noise portions. These approaches include order statistic filter (OSF), subband order statistic filter(SOSF), long-term spectrum divergence(LTSD), Kullback-Leibler distance(KL),energy and entropy, experimental results show that K-L distance method performs the best. That is, it gives the endpoints of noise-only portions closest to those obtained manually.
Secondly, the speech enhancement approaches are studied, which try to reduce the noise component within the speech signal in different domains. For example, Nonlinear Spectral Subtraction(NSS) and Wiener Filter(WF) perform in linear spectral domain, Mel Spectral Subtraction(MSS) performs in mel spectral domain. Furthermore, we propose the Cepstral Statistics Compensation(CSC) method, which performs in cepstral domain, it is found that the effect of these back-end speech enhancement approaches in general depends on the accuracy of the front-end VAD, and CSC gives the optimal recognition rates among all approaches. CSCeven performs better than two popular temporal filtering approaches, Cepstaral mean subtraction(CMS) and Cepsral normalization(CN).
In conclusion, robust VAD and speech enhancement approaches can effectively improve the noisy speech recognition, and have one special advantage. That is since they just perform on the speech to be recognized, it is no need to adjust the recognition models.
目 錄
誌謝 i
摘要 ii
Abstract iii
目錄 v
圖目錄 vii
表目錄 ix
第一章 緒論………………………………………………………………………….. 1
1-1 研究動機………………………………………………………………… 1
1-2 研究方法簡介……………………………………………………………. 2
第二章 實驗背景與基礎系統的建立…………………………………………………. 6
2-1 語音資料庫...................................................................................... 6
2-2 語音資料庫與加成性雜訊 7
2-3 語音特徵參數抽取 7
2-4 語音聲學模型的建立 14
2-5 辨識效能評估 15
2-6 基礎系統之實驗結果 15
2-7 本章結論 16
第三章 語音強健技術 17
3-1 非線性頻譜消去法(Nonlinear Spectral Subtraction,NSS) 17
3-2 韋納濾波器(Winner Filter,WF) 20
3-3 梅爾頻譜消去法(Mel Spectral Subtraction,MSS) 24
3-4 倒頻譜統計補償法(Cepstral Statistics Compensation,CSC) 24
3-5 本章結論 27
第四章 語音端點偵測與雜訊估測法 28
4-1 階層統計濾波器(Order Statistic Filter,OSF) 28
4-2 分頻階層統計濾波器(Subband Order Statistics Filter,SOSF) 32
4-3 長期頻譜分散法(Long-Term Spectrum Divergence,LTSD) 35
4-4 Kullback-Leibler量測法 37
4-5 能量偵測法(Energy) 39
4-6 亂度偵測法(Entropy) 40
4-7 本章結論 42
第五章 語音前端偵測與改良式頻譜消去法實驗結果 43
5-1 前端端點偵測精確度實驗結果 43
5-2 頻譜消去法基本實驗結果 48
5-3 階層統計濾波器之語音強健實驗結果 49
5-4 分頻階層統計濾波器之語音強健實驗結果 51
5-5 長期頻譜分散法之語音強健實驗結果 54
5-6 Kullback-Leibler距離量測法之語音強健實驗結果 55
5-7 能量偵測法之語音強健實驗結果 57
5-8 亂度偵測法之語音強健實驗結果 58
5-9 本章結論 60
第六章 語音前端偵測與韋納濾波器實驗結果 62
6-1 韋納濾波器基本實驗結果 62
6-2 端點偵測法結合韋納濾波器之實驗結果 63
6-3 本章結論 66
第七章 語音前端偵測與梅爾頻譜消去法實驗結果 67
7-1 梅爾頻譜消去法基本實驗結果 67
7-2 端點偵測法結合梅爾頻譜消去法之實驗結果 68
7-3 本章結論 71
第八章 倒頻譜統計補儅法之實驗結果 73
8-1 倒頻譜統計補儅法基本實驗結果
(Cepstral Statistics Compensation,CSC) 73
8-2 Kullback-Leibler量測法結合倒頻譜統計補償法之實驗結果 74
8-3 各種語音強健技術之實驗結果 75
8-4 本章結論 77
第九章 結論與展望 79
9-1 結論 79
9-2 展望 81
參考文獻 82








圖目錄
圖1-1 語音聲學辨識基本架構圖………………………………………………....... 2
圖1-2 語音訊號與摺積性雜訊和加成性雜訊之關係……………………………... 3
圖2-1 語音特徴參數的抽取流程………………………………………………….. 9
圖2-2 漢明窗……………………………………………………………………….. 10
圖3-1 乾淨之語音訊號前端為語音不存在的部份……………………………….. 19
圖3-2 頻譜消去法在抽取特徴參數流程中插入之位置…………………………... 20
圖3-3 韋納濾波器…………………………………………………………………... 21
圖3-4 韋納濾波器在抽取特徴參數流程中插入之位置…………………………... 23
圖3-5 倒頻譜統計補償法(CSC)之流程圖…………………………………………. 26
圖4-1 階層統計濾波器之架構圖………………………………………………….. 29
圖4-2 OSF端點偵測流程圖……………………………………………………….. 30
圖4-3 階層統計濾波器處理20dB飛機場雜訊語音之前端端點偵測…………… 31
圖4-4 SOSF端點偵測流程圖……………………………………………………… 32
圖4-5 分頻階層統計濾波器處理20dB人聲雜訊語音之前端端點偵測………… 34
圖4-6 長期頻譜分散法流程圖……………………………………………………... 35
圖4-7 長期頻譜分散法處理20dB飛機場雜訊語音之前端端點偵測…………… 36
圖4-8 KL端點偵測流程圖…………………………………………………………. 38
圖4-9 KL距離量測法處理20dB飛機場雜訊語音之前端端點偵測…………….. 39
圖4-10 能量偵測法處理20dB飛機場雜訊語音之前端端點偵測………………… 40
圖4-11 亂度偵測法分辨語音與非語音…………………………………………….. 41
圖5-1 測試語料中0dB與20dB下目測之起始音框與LTSD偵測法所測得之起
始音框之差距………………………………………………………………. 47
圖5-2 SOSF與NSS取回授流程圖……………………………………………….. 52
圖5-3 各種端點偵測法結合非線性頻譜消去法之平均辨識率(%)……………. 61
圖6-1 各種端點偵測法結合韋納濾波器之平均辨識率(%)……………………. 66
圖7-1 各種端點偵測法結合梅爾頻譜消去法之平均辨識率(%)……………… 72
圖8-1 各種語音強健法辨識率之比較(%)……………………………………… 78





















表目錄

表2-1 NUM-100A錄製的環境...................................................................................... 6
表2-2 NUM-100A訓練及測試語料的統計表.............................................................. 7
表2-3 梅爾濾波器之中心頻率與頻寬對照表............................................................. 11
表2-4 語音特徵參數之抽取設定................................................................................ 14
表2-5 雜訊語音未經處理之實驗................................................................................ 16
表5-1 OSF前端語音端點偵測率................................................................................. 44
表5-2 SOSF前端語音端點偵測率............................................................................... 44 表5-3 LTSD前端語音端點偵測法門檻值設定 45
表5-4 LTSD前端語音端點偵測率.............................................................................. 45
表5-5 KL距離量測法前端語音端點偵測率............................................................... 46
表5-6 能量偵測法前端語音端點偵測率..................................................................... 46
表5-7 雜訊語音未經處理之辨識精確率(%)........................................................ 48
表5-8 取前七個音框平均做頻譜消去法之辨識精確率(%)................................. 48
表5-9 目測法之端點做頻譜消去法之辨識精確率(%)......................................... 49
表5-10 OSF估測雜訊加NSS之辨識精確率(%)................................................. 50
表5-11 SOSF估測雜訊加NSS之辨識精確率(%)............................................... 52
表5-12 取回授之SOSF與回授前SOSF辨識精確率(%)之比較........................ 53
表5-13 LTSD估測雜訊加NSS之辨識精確率(%)............................................... 54
表5-14 KL估測雜訊加NSS之辨識精確率(%)................................................... 56
表5-15 Energy估測雜訊加NSS之辨識精確率(%)............................................. 57
表5-16 Entropy估測雜訊加NSS之辨識精確率(%)............................................ 59
表6-1 取前七個音框之平均結合韋納濾波器之基本實驗(%)............................. 62
表6-2 目測法之端點結合韋納濾波器之實驗結果(%)......................................... 63
表6-3 各種端點偵測結合韋納濾波器辨識率之比較(%)..................................... 64
表7-1 取前七個音框之平均結合梅爾頻譜消去法之基本實驗(%).................... 67
表7-2 目測法之端點結合梅爾頻譜消去法之實驗結果(%)................................. 68
表7-3 各種端點偵測結合梅爾頻譜消去法辨識率之比較(%)............................. 69
表8-1 倒頻譜統計補償法之基本實驗(%)............................................................. 73
表8-2 Kullback-Leibler距離量測法結合倒頻譜統計補償法之辨識率(%)........ 74
表8-3 各種語音強健技術之實驗結果(%)............................................................. 75
[1]王小川,”語音訊號處理”,全華科技圖書,2004.
[2] Y. Gong, “Speech Recognition in Noisy Environments: A Survey”, Speech Communication 16, 1995.
[3] M.J.F. Gales, “Model-based Techniques for Noise Robust Speech Recognition ”, University of Cambridge, Sep. 1995.
[4] Boll, S. F, “Suppression of Acoustic Noise in Speech Using Spectral Subtraction”,IEEE Trans. on ASSP, Vol. 27, No. 2, pp.113-120.1979.
[5] P. Lockwood and J. Boudy, “Experiments with a Nonlinear Spectral Subtractor (NSS) , Hidden Markov Models and the Projection, for Robust Speech Recognition in Cars”, Eurospeech 1991.
[6] ITU-T Recommendation G.729 – Annex B: A silence compression sceme for G. 729 optimized for terminals conforming to Recommendation V.70.
[7] B.A. Mellor and A.P. Varga, “Noise Masking in the MFCC Domain for the Recognition of Speech in Background Noise”, ICASSP 1992.
[8] Y. Ephraim and H.L. Van Trees, “A Signal Subspace Approach for Speech Enhancement”, IEEE Trans. on Speech and Audio Processing, 1995.
[9] S. Furui, "Cepstral Analysis Technique for Automatic Speaker Verification", IEEE Trans. Acoust. Speech Signal Process. 1981.
[10] O. Viikki and K. Laurila, “Noise Robust HMM-based Speech Recognition Using Segmental Cepstral Feature Vector Normalization”, in ESCA NATO Workshop Robust Speech Recognition Unknown Communication Channels, Pont-a-Mousson, France, 1997, pp. 107–110.
[11]H. Hermansky and N. Morgan, “RASTA Processing of Speech”. IEEE Trans. on Speech and Audio Processing. 2, pp. 578-589, 1994 .
[12]Kuo-Hwei Yuo and Hsiao-Chuan Wang, “Robust Features for Noisy Speech Recognition Based on Temporal Trajectory Filtering of Short-Time Autocorrelation Sequences”, Speech Communication 28, 1999.
[13]J.W. Hung, J.L. Shen, L.S. Lee, “New Approaches for Domain Transformation and Parameter Combination for Improved Accuracy in Parallel Model Combination ( PMC) Techniques”, IEEE Trans. on Speech and Audio Processing, Nov. 2001.
[14]J.L. Gauiain and C.H.Lee, “Maximum a Posteriori Estimation for Multivariate Gaussian Mixture Observations of Markov Chains”, IEEE Trans. on Speech and Audio Processing, 1994.
[15]C.J. Leggetter and P.C. Woodland, “Maximum Likelihood Linear Regression for Speaker Adaptation of Continuous Density Hidden Markov Models”, Computer Speech and Language, 1995.
[16]呂麗如, “Improved Techniques for Continuous Mandarin Speech Recognition Under Telephone Environment”,國立台灣大學碩士論文,June 1999.
[17]ITU-T Recommendation G.729 (Annex B): A Silence Compression Scheme for G.729, Optimized for Terminals Conforming to Recommendation V.70, ITU,1996.
[18]Hemant Misra_, Shajith Ikbal_, Herv´e Bourlard_, Hynek Hermansky, “Spectral Entropy Based feature for Robust ASR”,ICASSP 2004.
[19]郭正雄, “Robust Speech Recognition: Improved Spectral Subtraction”,國立暨南國際大學碩士論文,June 2004.
[20]Harold Gene Longbotham,Alan Conrad Bovik, “Theory of Order Statistic Filter and Their Relationship to Linear FIR Filters”,IEEE TRANSACTIONS ON ACOUSTICS. SPEECH. AND SIGNAL PROCESSING. VOL. 37. NO. 2. FEBRUARY 1989.
[21]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “Feature Extraction Combining Spectral Noise Reduction and Cepstral Histogram Equalization for Robust ASR”,ICSLP 2002.
[22]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “A New Voice Activity Detector Using Subband Order-Statistics Filters for Robust Speech Recognition”,ICASSP 2004.
[23]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “Improved Feature Extraction Based on Spectral Noise Reduction and Nonlinear Feature Normalization”,EUROSPEECH 2003.
[24]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “Voice Activity Detection With Noise Reduction and Long-Term Spectral Divergence Estimation”,ICASSP 2004.
[25]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “A New Adaptive Long-Term Spectral Estimation Voice Activity Detector”, EUROSPEECH 2003-GENEVA.
[26]Jos´e C. Segura, Javier Ram´ırez, Carmen Ben´ıtez, Angel de la Torre, Antonio Rubio, “Improved Voice Activity Detection Combining Noise Reduction and Subband Divergence Measures”, INTERSPEECH 2004 – ICSLP.
[27]Beena Ahmed and W. Harvey Holmes,“A Voice Activity Detector Using The Chi-Square Test”,ICASLP 2004.
[28] R. Stern, A. Acero, F.-H. Liu, and Y. Ohshima, "Signal processing for robust speech recognition," Automatic Speech and Speaker Recognition. Advanced Topics. Kluwer Academic Pub., pp. 357--384, 1997.
[29] Ananthakrishnan, K. S. "A comparison of modified k-means(MKM) and NN based real time adaptive clustering algorithms for articulatory space codebook formation", In ICSLP-1996, 1253-1256.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top