跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.202) 您好!臺灣時間:2026/08/29 01:32
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:黃義政
研究生(外文):Huang, Yi-Cheng
論文名稱:適用於華語數位助聽器之低延遲且類ANSI S1.11 1/3-octave規範濾波器組的音高式噪音消除與語音偵測輔助之廣泛動態範圍壓縮技術設計
論文名稱(外文):Design of Pitch Based Noise Reduction Adopting Low Latency Quasi ANSI S1.11 1/3 Octave Filter Bank and VAD-based Wide Dynamic Range Compression for Mandarin Digital Hearing Aid System
指導教授:周世傑
指導教授(外文):Jou, Shyh-Jye
學位類別:碩士
校院名稱:國立交通大學
系所名稱:電子工程學系 電子研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2013
畢業學年度:101
語文別:英文
論文頁數:102
中文關鍵詞:助聽器雜訊消除動態範圍壓縮語音區間偵測華語音高動態背景環境
外文關鍵詞:hearing aidsnoise reductiondynamic range compressionvoice activity detectionmandarinpitchnon-stationary background environment
相關次數:
  • 被引用被引用:0
  • 點閱點閱:375
  • 評分評分:
  • 下載下載:27
  • 收藏至我的研究室書目清單書目收藏:1
在本論文中,我們提出一套採用低延遲的類ANSI 1/3 octave濾波器組且適合實現於助聽器系統的音高式雜訊消除系統與語音偵測基準之廣泛動態範圍壓縮技術。所提出的音高式雜訊消除系統包含一個音高式語音偵測器與仰賴子音起始的雜訊抑制器,而且使用語音的特性如音高與相對應之和諧音、子音起始和單音節字長度的時間。由於quasi ANSI濾波器組有低解析度的缺點,提出的音高式語音偵測器將音高與子音起始特性跟彈性和諧音偵測器整合在一起來提升語音偵測器的準度,而提出的仰賴子音起始的雜訊抑制器是設計來克服濾波器組的低解析度。除此之外,一個長期平均能量更新機制被使用來增進子音起始特性的偵測率,模擬的結果顯示,提出的音高式雜訊消除系統能同時在靜態背景雜訊環境與高動態背景雜訊環境有好的表現,提出的音高式語音偵測器的準度結果是可以與採用高解析度ANSI濾波器組的音高式語音偵測器相比的,平均準度可以分別在靜態與動態背景雜訊環境裡達到83.70%與85.70%。而提出的仰賴子音起始的雜訊抑制器的語音區段訊雜比和語音訊雜比在靜態背景雜訊環境中,平均改進5.95dB和9.12dB,在動態背景雜訊環境裡平均改進6.49dB和9.47dB。另外,語音品質(PESQ)在靜態與動態背景雜訊環境裡平均改進0.19和0.22。
再來是提出的語音偵測基準之廣泛動態範圍壓縮技術,它可以提升語音與雜訊之間的能量差。由於廣泛動態範圍壓縮技術演算法通常是在沒有考慮背景雜訊的乾淨語音環境中設計的,被廣泛動態範圍壓縮技術衰減的高能量語音,程度可能會比低能量的背景雜訊多,當雜訊消除系統與廣泛動態範圍壓縮技術一起使用時,這會造成交互干擾的影響,雜訊消除系統的效能可能會因為廣泛動態範圍壓縮技術而衰減,因而縮減的語音與雜訊間的能量差,導致語音辨識度降低。有了來自雜訊消除系統的語音偵測結果幫忙,廣泛動態範圍壓縮技術可以針對語音區段和雜訊區段做不同的處理來增加語音辨識度,由模擬的結果可以看出語音偵測基準之廣泛動態範圍壓縮技術,對於減少雜訊消除系統與廣泛動態範圍壓縮技術之間的交互干擾影響是有益處的。
所提出的音高式雜訊消除系統與語音偵測基準之廣泛動態範圍壓縮技術的運算複雜度是低的,而且一點改良的代價可以換來很好的效能。最後,提出的演算法包含類ANSI濾波器組的總延遲只有11.3ms,這是符合助聽器系統的要求而且適合應用在助聽器系統上。

In this thesis, we propose a pitch based noise reduction (NR) system and a VAD-based wide dynamic range compression (WDRC) which adopts a quasi-ANSI 1/3 octave filter bank with low group delay for realistic implementation in hearing aids (HA) systems. The proposed pitch based NR includes a pitch based voice activity detection (VAD) and onset-depended noise attenuation (ONA). The characteristics of speech such as pitch and corresponding harmonics, onset, and time of monosyllable word length are utilized by the proposed pitch based NR. Due to the drawback of low resolution resulted from quasi ASNI filter bank, the proposed pitch based VAD integrates the pitch and onset features with the flexible harmonics detection to improve the accuracy of VAD. The proposed ONA is designed to conquer the poor resolution of the filter bank. In addition, an update mechanism of long-term average magnitude is employed to enhance the detection of onset feature. The simulation results show that the proposed pitch based NR can perform well in both stationary (the situation that user is still) background noise environment and highly dynamic (the situation that user is moving) background noise environment. The accuracy results of proposed pitch based VAD are comparable with the pitch based VAD adopting ANSI filter bank which has high resolution. The average accuracy of proposed pitch based VAD is about 83.70% and 85.70% in stationary and dynamic noise situations respectively. And the average improvement of segmental signal-noise-ratio (SNRseg) and signal-noise-ratio (SNR) of the proposed ONA is 5.95dB and 9.12dB in stationary noise environment and 6.49dB and 9.47dB in dynamic noise environment. Moreover, the average improvement of sound quality (PESQ) is 0.19 and 0.22 in stationary and dynamic noise environments respectively.
The proposed VAD-based WDRC enhances the energy difference between speech and noise. Because the WDRC algorithms are usually developed on clean speech scenarios without considering the presence of background noise, the high energy of speech may be suppressed more than low energy of background noise due to the characteristic of WDRC. This incurs the undesired interaction effect when NR and WDRC are connected. The performance of NR might be degraded by WDRC block. Thus, the energy difference between speech and noise is decreased and degrades the speech intelligibility. With the help of VAD information from NR block, WDRC can perform different operations to speech regions and noise regions and increases the speech intelligibility. The simulation results show that the proposed VAD-based WDRC has benefit to reduce the undesired interaction effect between NR and WDRC.
For the proposed pitch based NR and VAD-based WDRC, the computational complexity of the proposed algorithms is low and the slight cost of modifications could exchange the outstanding performance. Finally, the total latency of the proposed algorithm including the quasi ANSI filter bank is only 11.3ms which matches the requirement of HA system and is suitable for the HA applications.

Content VI
Chapter 1 Introduction 1
1.1 Digital Hearing Aid System 1
1.2 Design Motivation and Goal 2
1.3 Thesis Organization 4
Chapter 2 Overview of Noise Reduction and Wide Dynamic Range Compression 6
2.1 Voice Activity Detection and Noise Reduction Algorithms 6
2.1.1 Voice Activity Detection 6
2.1.2 Noise Attenuation 8
2.2 Wide Dynamic Range Compression Algorithms 10
2.3 Design Concepts 11
2.4 Summary 13
Chapter 3 Pitch Based Voice Activity Detection and Onset-depended Noise Attenuator 14
3.1 Introduction 14
3.2 Characteristics of Speech and Human Hearing System 16
3.2.1 Characteristics of Speech 16
3.2.2 Characteristics of Human Hearing System 19
3.3 Pitch Based Noise Reduction 21
3.3.1 Characteristics of Quasi ANSI S1.11 1/3-Octave Filter Bank for Digital Hearing Aids System 21
3.3.2 Short/Long-Term Average Magnitude Calculation and Nonlinear Energy Operation 25
3.3.3 Pitch Detection and Onset Detection 28
3.3.4 VAD Decision 31
3.3.5 Onset-depended Noise Attenuator 33
3.3.6 Threshold Update 42
3.4 Summary 48
Chapter 4 Wide Dynamic Range Compression with Clues of VAD 50
4.1 Low Complexity Wide Dynamic Range Compression Algorithm 51
4.2 VAD-based WDRC 55
4.2.1 VAD-based Gain Estimation 55
4.2.2 IO-curve for Noise 61
4.3 Summary 63
Chapter 5 Simulation Results and Analysis 65
5.1 Performance Index 65
5.1.1 The Performance Indices for NR 65
5.1.2 The Performance Indices for WDRC 67
5.2 Simulation Environments: Noise Environment and Speech Database 67
5.3 Simulation Results of Pitch Based Noise Reduction 70
5.3.1 The Performance Comparison between the Proposed and Original Pitch Based NR 70
5.3.2 Mandarin Database with Stationary Noise Environment 73
5.3.3 Mandarin Database with Dynamic Noise Environment 79
5.4 Simulation Results of VAD-based Wide Dynamic Range Compression 82
5.5 Comparison with Noise Reduction Algorithms 87
5.6 Summary 94
Chapter 6 Conclusion and Future Work 96
6.1 Conclusion 96
6.1.1 Noise Reduction 96
6.1.2 Wide Dynamic Range Compression 97
6.2 Future Work 98
6.2.1 Noise Reduction 98
6.2.2 Wide Dynamic Range Compression 98
Reference 100

[1] J. Kates, Digital Hearing Aids, Plural, San Diego, Calif, USA, 2008.
[2] Y. J. Chen and S. J. Jou, "Design and Implementation of Neuromorphic Pitch Based Noise Reduction for Mandarin Digital Hearing Aid System," Master Thesis, Department of Electronics Engineering &; Institute of Electronics, National Chiao Tung University, 2011.
[3] M. A. Stone and B. C. J. Moore, "Tolerable hearing aid delays. II. Estimation of limits imposed during speech production," J. Ear and Hearing, vol. 23, no. 4, pp. 325-338, 2002.
[4] J. Agnew and J. M. Thornton, “Just noticeable and objectionable group delays in digital hearing aids,” J. American Academy of Audiology, vol. 11, no. 6, pp. 330–336, 2000.
[5] W. H. Chang, " Complexity-Effective Multi-Channel Dynamic Range Compression (DRC) for Digital Hearing Aids," Master Thesis, Department of Electronics Engineering &; Institute of Electronics, National Chiao Tung University, 2008.
[6] M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Processing, vol. 4, pp. 208-211, 1979.
[7] J. C. Junqua, B. Reaves, and B. Mak, “A study of endpoint detection algorithms in adverse conditions: incidence on a DTW and HMM recognize,” in Proc. Conf. Eurospeech, pp. 1371-1374, 1991.
[8] R. Tucker, "Voice activity detection using a periodicity measure," Proc. Inst. Electr. Eng. I, vol. 139, no. 4, pp.377 -380, 1992
[9] P. K. Ghosh, A. Tsiartas and S. Narayanan, "Robust Voice Activity Detection Using Long-Term Signal Variability," IEEE Transactions on Audio, Speech, and Language Processing, Volume: 19, Issue: 3, pp. 600 – 613, 2011.
[10] P. C. Loizou, Speech Enhancement: Theory and Practice, Boca Raton, Florida: CRC Press, 2007.
[11] K. Ngo., "Digital signal processing algorithms for noise reduction, dynamic range compression and feedback cancellation in hearing aids," PhD thesis, ESAT, Katholieke Universiteit Leuven, Belgium, Jul. 2011.
[12] H. Yi and P. C. Loizou, "A generalized subspace approach for enhancing speech corrupted by colored noise," IEEE Trans. Speech and Audio Processing, vol. 11, no. 4, pp. 334-341, 2003.
[13] F. Jabloun and B. Champagne, “Incorporating the human hearing properties in the signal subspace approach for speech enhancement,” IEEE Trans. Speech and Audio Processing, vol. 11, no. 6, pp. 700–708, 2003.
[14] S. F. Boll, "Suppression of acoustic noise in speech using spectral subtraction," IEEE Trans. Acoustics, Speech, Signal Processing, vol. 27, no.2, pp. 113-120, 1979.
[15] S. Kamath and P. C. Loizou, “A multi-band spectral subtraction method for enhancing speech corrupted by colored noise,” in Proc. IEEE Int. Conf. Acoustic, Speech, Signal Processing, vol. 4, pp. 4164-4164, 2002.
[16] C. C. Tsai and T. S. Chang, “Low Power Noise Reduction Design for Hearing Aids Application,” Master Thesis, Department of Electronics Engineering &; Institute of Electronics, National Chiao Tung University, 2009.
[17] C. W. Wei, C. C. Tsai, T. S. Chang and S. J. Jou, “ Perceptual multiband spectral subtraction for noise reduction in hearing aids,” in Proc. IEEE Int. Conf. APCCAS, pp. 692-695, May 2010.
[18] Y. T. Kuo, T. J. Lin, Y. T. Li, and C. W. Liu, "Design&;implementation of low-power ANSI S1.11 filter bank for digital hearing aids," IEEE Tran. Circuits Syst. I, Reg. Papers, vol. 57, no. 7, pp. 1684–1696, Jul. 2010.
[19] C. W. Liu, K. C. Chang, M. H. Chuang, C. H. Lin, "10-ms 18-Band Quasi-ANSI S1.11 1/3-Octave Filter Bank for Digital Hearing Aid," IEEE Tran. Circuits Syst. I, Reg. Papers, vol. 60, no. 3, pp. 638-649, Mar. 2013.
[20] D. F. Rosenthal and H. G. Okuno, Computational Auditory Scene Analysis, Mahwah, NJ: Lawrence Erlbaum, 1998.
[21] D. Wang and G. J. Brown, Eds., Computational Auditory Scene Analysis: Principles, Algorithms and Applications. New York: Wiley-IEEE Press, 2006.
[22] T. Y. Chang and B. F. Wu, "Research and Implementation of MP3 Encoding Algorithm," Master Thesis, Department of Electrical and Control Engineering, National Chiao Tung University, National Chiao Tung University, July, 2002.
[23] J. Chalupper, “Aural Exciter and Loudness Maximizer: What's Psychoacoustic about" Psychoacoustic Processors?",“ In Proc. 109th AES Convention, Los Angeles, Sep. 2000.
[24] N. Roman, D. L. Wang, and G. J. Brown, “Speech segregation based on sound localization,” J. Acoustical Society of America, vol. 114, pp. 2236–2252, 2003.
[25] E. James, A. K. Barros, T. Yoshinori, D. Mandic and N. Ohnishi, ”Speech enhancement by lateral inhibition and binaural masking,” in Proc. IEEE Machine Learning for Signal Processing, pp. 365-370, 2004.
[26] Specification for Octave-Band and Fractional-Octave-Band Analog and Digital Filters, ANSI Standard S1.11-2004.
[27] T. Fawcett. An introduction to ROC analysis. Pattern Recognition Letters, 27(8):861 – 874, 2006.
[28] Y. T. Kuo, " Low-Power Auditory Compensation for Digital Hearing Aids," PhD thesis, Department of Electronics Engineering &; Institute of Electronics, National Chiao Tung University, 2011.
[29] J. N. Mitchell, “Computer multiplication and division using binary logarithms,” IRE Trans. Electron. Computers, vol. 11, pp. 512–517, Aug. 1962.
[30] P. Y. Lin, "Feasibility Study of the Implementation of Hearing Aid Signal Processing Algorithms on the TI TMS320C6713 DSK", Master thesis, Institute of Biomedical Engineering, National Yang Ming University, 2004.
[31] A. Rix, J. Beerends, M. Hollier, and A. Hekstra, "Perceptual evaluation of speech quality (PESQ) ─ a new method for speech quality assessment of telephone networks and codecs," IEEE Int. Conf. Acoustic, Speech, Signal Processing, pp. 749-752, 2001.
[32] Perceptual evaluation of speech quality (PESQ), an objective method forend-to-end speech quality assessment of narrowband telephone networks and speech codecs. ITU-T Draft Recommendation, pp.862, May 2000.
[33] J. H. Chang and S. T. Young, “Effect comparison of hearing aids prescriptions on Mandarin speech perception,” Master Thesis, Institute of Biomedical Engineering, National Yang Ming University, 2005.
[34] A. Varga, H. J. M. Steenneken, M. Tomlinson, D. Jones. (1992). NOISEX-92. Available: http://spib.rice.edu/spib/select_noise.html
[35] Y. Hu and P. C. Loizou, ”Speech enhancement based on wavelet thresholding the multitaper spectrum,” IEEE Trans. Speech and Audio Processing, vol. 12, no. 1, pp. 59-67, 2004.
[36] Y. Ephraim and D.Malah, “Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,” IEEE Trans. Acoustics, Speech, Signal Processing, vol. 33, no.2, pp. 443-445, 1985.
[37] M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Processing, vol. 4, pp. 208-211, 1979.

連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊