跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.66) 您好!臺灣時間:2026/08/15 21:53
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:杜文祥
研究生(外文):Wen Hsiang Tu
論文名稱:端點偵測技術在強健語音參數擷取之研究
論文名稱(外文):Study on the Voice Activity Detection Techniques for Robust Speech Feature Extraction
指導教授:洪志偉洪志偉引用關係
指導教授(外文):Jeih-weih Hung
學位類別:碩士
校院名稱:國立暨南國際大學
系所名稱:電機工程學系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2007
畢業學年度:95
語文別:中文
論文頁數:92
中文關鍵詞:端點偵測法能量特徵頻譜消去法自動語音辨認
外文關鍵詞:voice activity detectionspectral magnitudespectral subtractionspeech recognition
相關次數:
  • 被引用被引用:2
  • 點閱點閱:251
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
由於發展環境和應用環境兩者之間的不匹配,導致於語音辨識系統效能經常會下降,而引起這不匹配的主要原因之一是加成性雜訊,處理加成性雜訊的方法我們可以分成三類,語音強化法、強健性語音特徵參數、以及語音模型調適法,而本論文所討論的方法主要是屬於強健性語音特徵參數之技術。
在本論文中,我們主要的重點在於探討不同的語音特徵對於語音端點偵測的影響,所利用的特徵分別為低頻帶頻譜強度、全頻帶頻譜強度、累積量化頻譜、以及高通對數能量等。利用以上這些不同的特徵進行語音之端點偵測,所得之純雜訊的位置資訊可以提供頻譜消去法與靜音對數能量正規化法中所需的雜訊頻譜或能量的估測。
在實驗環境上我們採用Aurora2語料庫,在八種背景雜訊以及訊雜比0~20dB下做實驗。在第五章中所呈現的實驗數據與分析可證明以上所述的各種特徵顯然可用以有效的鑑別出一段語音中純雜訊部分與語音部分,使之後所使用的頻譜消去法與靜音對數能量正規化法等強健性語音特徵技術,得以明顯提升在雜訊環境下語音辨識的精確度,增加語音辨識系統的強健性。
The performance of a speech recognition system is often degraded due to the mismatch between the environments of development and application. One of the major sources that give rises to this mismatch is additive noise. The approaches for handling the problem of additive noise can be divided into three classes: speech enhancement, robust speech feature extraction, and compensation of speech models. In this thesis, we are focused on the second class, robust speech feature extraction.
The approaches of speech robust feature extraction are often together with the voice activity detection in order to estimate the noise characteristics. A voice activity detector (VAD) is used to discriminate the speech and noise-only portions within an utterance. This thesis primarily investigates the effectiveness of various features for the VAD. These features include low-frequency spectral magnitude (LFSM), full-band spectral magnitude (FBSM), cumulative quantized spectrum (CQS) and high-pass log-energy. The resulting VAD offers the noise information to two noise-robustness techniques, spectral subtraction (SS) and silence log-energy normalization (SLEN), in order to reduce the influence of additive noise in speech recognition.
The recognition experiments are conducted on Aurora-2 database. Experimental results show that the proposed VAD is capable of providing accurate noise information, with which the following processes, SS and SLEN, significantly improve the speech recognition performance in various noise-corrupted environments. As a result, we confirm that an appropriate selection of features for VAD implicitly improves the noise robustness of a speech recognition system.
目 錄

誌謝...................................................... i
摘要.....................................................iii
Abstract..................................................iv
目錄.......................................................v
圖目錄..................................................viii
表目錄.....................................................x

第一章 緒論
1.1研究動機及主............................................1
1.2 強健式語音辨識方法的分類...............................3
1.3 研究方法簡介...........................................5
1.4 論文架構...............................................6

第二章 語音訊號特徵參數及模型之建立
2.1 語音特徵參數之抽取.....................................7
2.2 語音聲學模型的建立....................................18

第三章 基本系統之建立
3.1 語音資料庫之簡介......................................20
3.2 辨識效能評估..........................................27
3.3 基本系統的訓練和結果..................................27
3.4 本章結論..............................................30

第四章 強健性語音特徵參數技術與端點偵測法
4.1 端點偵測法............................................31
4.1.1 頻譜特徵判斷法......................................32
4.1.2 累積量化頻譜法......................................38
4.1.3 高通對數能量法......................................41
4.2 強健之語音技術........................................45
4.2.1 非線性頻譜消去法....................................45
4.2.2 靜音對數能量正規化法................................49

第五章 端點偵測法配合強健性語音技術之實驗結果
5.1 梅爾倒頻譜特徵參數之實驗結果..........................52
5.1.1 本頻譜消去法之實驗結果..............................52
5.1.2 低頻域頻譜強度之端點偵測法配合頻譜消去法之實驗結果..54
5.1.3 全頻帶頻譜強度之端點偵測法配合頻譜消去法之實驗結果..55
5.1.4 累積量化頻譜法配合頻譜消去法之實驗結果..............57
5.1.5 高通對數能量法配合頻譜消去法之實驗結果..............58
5.1.6 實驗結果綜合討論....................................60
5.2 梅爾倒頻譜系數與對數能量之特徵參數的實驗結果..........61
5.2.1 基本頻譜消去法之實驗結果............................62
5.2.2 靜音對數能量正規化法................................65
5.2.3 低頻域頻譜強度之端點偵測法配合靜音音框對數能量正規化法
之頻譜消去法實驗結果................................66
5.2.4 全頻域頻譜強度之端點偵測法配合靜音音框對數能量正規化法
之頻譜消去法實驗結果................................67
5.2.5 累積量化頻譜法配合頻譜消去法之實驗結果..............69
5.2.6 高通對數能量法......................................71
5.2.7 實驗結果綜合討論....................................72

第六章 結論與未來展望
6.1結論與未來展望.........................................75

參考文獻..................................................77
[1] 王小川, "語音訊號處理" , 全華科技圖書, 2004.
[2] 黃志楠, "Improved Techniques for Speech Recognition Under Additive and Covolutional Noisy Environment" , 國立台灣大學碩士論文, June 2001.
[3] 呂麗如, "Improved Techniques for Continuous Mandarin Speech Recognition Under Telephone Environment" , 國立台灣大學碩士論文, June 1999.
[4] Yifan Gong, "Speech Recognition in Noisy Environments; A Survey", Speech Communication 16, 1995.
[5] S.F. Boll, “Suppression of Acoustic Noise in Speech Using Spectral Subtraction” IEEE Trans. on Acoustics, speech, and Processing, VOL. ASSP-27, NO. 2, April 1979.
[6] P. Lockwood and J. Boudy, "Experiments with a Nonlinear Spectral Subtractor ( NSS ), Hidden Markov Models and the Projection, for Robust Speech Recognition in Cars", Eurospeech 1991
[7] ITU-T Recommendation G.729-Annex B: A silence compression sceme for G.729 optimized for terminals conforming to Recommendation V.70
[8] B.A. Mellor and A.P. Varga, "Noise Masking in the MFCC Domain for the Recognition of Speech in Background Noise", ICASSP 1992
[9] S. Furui, "Cepstral Analysis Technique for Automatic Speaker Verification", IEEE Trans. Acoust. Speech Signal PROCES, 1981
[10] O. Viikki and K. Laurla, "Noise Robust HMM-based Speech Recognitio Using Segmental Cepstral Feature Vector Normalization", in ESCA NATO Workshop Robust Speech Recognition Unknown Communication Channels, Pont-a-Mousson, France, 19997, pp. 107-110
[11] H. Hermansky and N. Morgan, "RASTA Processing of Speech", IEEE Trans. on Speech and Audio Processing. 2, pp. 578-589, 1994
[12] J.L. Gauiain and C.H. Lee, "Maximum a Posteriori Estimation for Multivariate Gaussian Mixture Observations of Markov Chains", IEEE Trans. on Speech and Audio Processing, 1994
[13] J.W. Hung, I.L. Shen, L.S. Lee, "New Approaches for Domain Transformation and Parameter Combination for Improved Accuracy in Parallel Model Combination ( PMC ) Techniques", IEEE Trans. o Speech and Audio Processing, Nov. 2001
[14] C.J. Leggetter and P.C. Woodland, "Maximum Likelihood Linear Regression for Speaker Adaptation of Continuous Density Hidden Markov Models", Computer Speech and Language, 1995
[15] K. Yamashita,; T. Shimamura; " Nonstationary noise estimation using low-frequency regions for spectral subtraction", Signal Processing Letters, IEEE Volume 12, Issue 6, June 2005
[16] E.L. Bocchieri , and J.G. Wilpon, " Discriminative analysis for feature reduction in automatic speech recognition", in Proc. IEEE ICASSP, vol. 1, March 1992, pp.501-504
[17] E. Julien and H.C. Choi, " An Energy Search Approach to Variable Frame Rate Front-End Processing for Robust ASR", INTERSPEECH 2005-EUROSPEECH, 2613-2616
[18] C.F. Tai; J.W. Hung, " Silence Energy Normalization for Robust Speech Recognition in Additive Noise Environments", INTERSPEECH 2006 – ICSLP
[19] S. Ayat; M.T. Manzuri; R. Dianat; J. Kabudian; "An improved spectral subtraction speech enhancement system by using an adaptive spectral estimator", Electrical and Computer Engineering, 2005. Canadian Conference on 1-4 May 2005
[20] Y. Fan; Yi Li; C. Wu; "Speech Endpoint Detection Based on Speech Time-Frequency Enhancement and Spectral Entropy", Engineering in Medicine and Biology Society, 2005, IEEE-EMBS 2005.
[21] R. Gemello; F. Mana,; R.D. Mori, " A modified Ephraim-Malah noise suppression rule for automatic speech recognition", Acoustics, Speech, and Signal Processing, 2004. Proceedings. (ICASSP 2004).
[22] 戴仲甫, "The Improved Techniques of Energy Feature Enhancement and Frame Selection for Robust Speech Recognition", 暨南國際大學碩士論文, June 2006.
[23] Y. Chen, L.S. Lee, " Energy-Based Frame Selection for Reliable Feature Normalization and Transformation in Robust Speech Recognition", INTERSPEECH 2005-EUROSPEECH, 385-388
[24] S.M. Ahadi; H. Sheikhzadeh; R.L. Brennan; G.H. Freeman, "An Energy Normalization Scheme for Improved Robustness in Speech Recognition", INTERSPEECH 2005, EUROSPEECH, 2613-2616
[25] P. Krishnamoorthy, S.R.M. Prasanna; "Modified Spectral Subtraction Method for Enhancement of Noisy Speech", Intelligent Sensing and Information Processing, 2005. ICISIP 2005.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top