跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.60) 您好!臺灣時間:2026/08/01 02:40
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:黃盈彰
研究生(外文):Huang, Ying-Zhang
論文名稱:基於時頻調變之適用於聽損病患的中文語音理解度客觀量測指標
論文名稱(外文):Spectro-Temporal Modulations Based Objective Mandarin Speech Intelligibility Measure for Hearing-Impaired Patients
指導教授:冀泰石
指導教授(外文):Chi, Tai-Shih
口試委員:李沛群劉奕汶冀泰石
口試委員(外文):Li,Pei-CyunLiou, Yi-WenChi, Tai-Shih
口試日期:2016-01-14
學位類別:碩士
校院名稱:國立交通大學
系所名稱:電信工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2016
畢業學年度:104
語文別:中文
論文頁數:61
中文關鍵詞:客觀理解度指標時頻域調變聽力受損
外文關鍵詞:objective intelligibility measurementspectro-temporal modulationhearing-impaired
相關次數:
  • 被引用被引用:0
  • 點閱點閱:215
  • 評分評分:
  • 下載下載:15
  • 收藏至我的研究室書目清單書目收藏:0
為了驗證助聽器演算法的效果,我們發展一套中文語音理解度量測指標來預估聽損者的語音理解度。為了符合不同聽損者的聽損程度,我們使用了一套能夠模擬聽損者耳蝸的聽損模型,模擬最小可聽水平提升、響度聚集、以及分頻解析度降低等聽損現象。許多文獻經由時域調變分析萃取出語音特徵來計算並預估語音理解度。在本論文中,我們採用時頻調變分析方法,擷取出影響語音理解度的語音特徵,來預估聽損者的語音理解度。時頻調變分析分為兩個聽覺感知階段,第一階段為人耳至中腦的頻譜預估,第二階段為中腦至大腦皮質聽覺區對時頻域調變的分析,同時考慮時間上與頻率上的變化。為了能考慮分頻解析度降低對語音理解度的影響,我們將參考Tiago H. Falk提出的基於時域調變封包為架構的理解度指標演算法,開發基於時頻調變的非侵入式語音理解度指標。設計實驗為量測聽損者在兩種雜訊環境下的中文單字語音理解度,並比較訊噪比對語音理解度的影響,最後比較Tiago H. Falk所提出的演算法及開發的理解度指標兩者與中文語音理解度的相關性,以評估兩種演算法的效果。未來將透過此開發的語音理解度客觀量測指標來開發語音增強演算法,並將其實際應用在助聽器。
In order to verify the performance of algorithms developed for hearing aids, we developed a measure that predicts the Mandarin speech intelligibility of the hearing impaired people. We constructed a model that simulates the cochlear of the hearing impaired to fit different patients. This model solves threshold elevation, loudness recruitment, and reduced frequency selectivity of hearing impaired. In order to predict speech intelligibility, many studies utilized the modulation spectral signal representation, which is obtained by an auditory-inspired filterbank analysis of the speech signal. In this study, a joint spectro-temporal auditory model was utilized to assess speech quality objectively. In this auditory model, the first stage is to simulate cochlear function of the spectrum estimation. The second stage is to simulate cortical function of the multi-dimensional spectrum analysis. In order to consider effecting speech intelligibility due to reduced frequency selectivity of hearing impaired, we developed a spectro-temporal modulations based non-intrusive intelligibility measure through referring to the construction of SRMR proposed by Tiago H. Falk. SRMR is a temporal modulation envelope based intelligibility measure. To validate our proposed measure, the performance of the proposed measure is compared to the SRMR intelligibility measurement algorithms under several noisy conditions. We can utilize the proposed measure to assess the performance of speech enhancement algorithms, and develop a speech enhancement algorithm in the future.
中文摘要 I
英文摘要 II
誌 謝 IV
目 錄 V
表 目 錄 VII
圖 目 錄 VIII
第一章 緒論 1
1.1 研究背景 1
1.2 聽損現象簡介 2
1.2.1 響度聚集與最小可聽水平提升 2
1.2.2 分頻解析度降低 3
1.3 研究動機 4
1.4 章節大綱 5
第二章 個人化聽損模型 6
2.1 響度模型 6
2.2 頻譜模糊化模型 7
2.2.1 濾波器變寬程度計算 8
2.2.2 模糊化增益計算 10
2.3 混合模型 11
第三章 聽覺感知模型介紹 13
3.1 生理聽覺感知運作方式 13
3.1.1 耳朵基本構造簡介 13
3.1.2 耳蝸生理聲學現象 15
3.1.3 聲音強度及響度 17
3.2 常見的語音訊號處理方式 19
3.2.1 短時間傅立葉轉換(STFT) 19
3.2.2 聽覺濾波器組(auditory filter bank) 22
3.3 聽覺感知模型 23
3.3.1 初期耳蝸階段 24
3.3.2 大腦皮質階段 26
第四章 理解度的主觀量測與客觀指標 29
4.1 中文語音資料庫 29
4.2 主觀中文語音理解度測量 30
4.2.1 測試環境與使用者介面 30
4.2.2 測試條件與評分方式 30
4.2.3 聽力補償 31
4.3 客觀理解度量測指標 32
4.3.1 侵入式客觀量測指標 32
4.3.2 非侵入式客觀量測指標 34
第五章 基於時頻調變之理解度指標 38
5.1 背景知識 38
5.2 研究方法 41
5.2.1 適用於客觀量測指標之聽損模型 41
5.2.2 基於時頻調變之理解度指標演算法 46
5.3 研究結果 48
5.3.1 相關性 48
5.3.2 訓練參數時的誤差 50
5.3.3 頻域調變的範圍選取 54
第六章 結論與未來展望 56
參考文獻 58

[1]Edgar Villchur. "Simulation of the effect of recruitment on loudness relationships in speech." The Journal of the Acoustical Society of America 56.5 (1974): 1601-1611.
[2]Brian R. Glasberg, Brian C.J. Moore, and Sid P. Bacon. "Gap detection and masking in hearing‐impaired and normal‐hearing subjects." The Journal of the Acoustical Society of America 81.5 (1987): 1546-1556.
[3]Peter J. Fitzgibbons and Frederic L. Wightman. "Gap detection in normal and hearing‐impaired listeners." The Journal of the Acoustical Society of America72.3 (1982): 761-765.
[4]Brian R. Glasberg, and Brian C.J. Moore. "Auditory filter shapes in subjects with unilateral and bilateral cochlear impairments." The Journal of the Acoustical Society of America 79.4 (1986): 1020-1033.
[5]Richard S. Tyler, et al. "Auditory filter asymmetry in the hearing impaired." The Journal of the Acoustical Society of America 76.5 (1984): 1363-1368.
[6]Humes, Larry E., and Lisa Roberts. "Speech-Recognition Difficulties of the Hearing-Impaired ElderlyThe Contributions of Audibility." Journal of Speech, Language, and Hearing Research 33.4 (1990): 726-735.
[7]Kathryn Hopkins, Brian C.J. Moore, and Michael A. Stone. "Effects of moderate cochlear hearing loss on the ability to benefit from temporal fine structure information in speech." The Journal of the Acoustical Society of America 123.2 (2008): 1140-1153.
[8]Brian C.J. Moore. "Perceptual consequences of cochlear hearing loss and their implications for the design of hearing aids." Ear and hearing 17.2 (1996): 133-161.
[9]Thomas Baer, and Brian C.J. Moore. "Effects of spectral smearing on the intelligibility of sentences in the presence of interfering speech." The Journal of the Acoustical Society of America 95.4 (1994): 2277-2280.
[10]Brian C.J. Moore, and Brian R. Glasberg. "Simulation of the effects of loudness recruitment and threshold elevation on the intelligibility of speech in quiet and in a background of speech." The Journal of the Acoustical Society of America94.4 (1993): 2050-2062.
[11]Yoshito Nejime, and Brian C.J. Moore. "Simulation of the effect of threshold elevation and loudness recruitment combined with reduced frequency selectivity on the intelligibility of speech in noise." The Journal of the Acoustical Society of America 102.1 (1997): 603-615.
[12]Hongmei Hu, et al. "Simulation of hearing loss using compressive gammachirp auditory filters." Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on. IEEE, 2011.
[13]Tiago H. Falk, et al. "Objective Quality and Intelligibility Prediction for Users of Assistive Listening Devices: Advantages and limitations of existing tools." Signal Processing Magazine, IEEE 32.2 (2015): 114-124.
[14]Tiago H. Falk, Chenxi Zheng, and Wai-Yip Chan. "A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech." Audio, Speech, and Language Processing, IEEE Transactions on 18.7 (2010): 1766-1774.
[15]Tiago H. Falk and Wai-Yip Chan. "A non-intrusive quality measure of dereverberated speech." Proc. Int. Workshop on Acoustic Echo and Noise Control (IWAENC). 2008.
[16]David Suelzle, Vijay Parsa, and Tiago H. Falk. "On a reference-free speech quality estimator for hearing aids." The Journal of the Acoustical Society of America 133.5 (2013): EL412-EL418.
[17]T. S. Chi, class notes of Auditory and Acoustical Information Processing, Department of Communication Engineering, National Chiao-Tung University, Taiwan, 2013.
[18]Stanley Finger (1994). Origins of neuroscience : a history of explorations into brain function (N.e. ed.). Oxford: Oxford University Press. ISBN 0-19-5146948.
[19]H. Fletcher, Speech and Hearing in Communication, 2^nded., Bell Telephone Laboratories Series, Van Nosetrand, Princeton, NJ, 1953.
[20]Stanley S Stevens. "The measurement of loudness." The Journal of the Acoustical Society of America 27.5 (1955): 815-829.
[21]G. Heinzel, A. Rudiger, and R. Schilling (2002). Spectrum and spectral density estimation by the Discrete Fourier transform (DFT), including a comprehensive list of window functions and some new flat-top windows (Technical report). Max Planck Institute (MPI) fur Gravitationsphysik / Laser Interferometry & Gravitational Wave Astronomy. 395068.0. Retrieved 2013-02-10.
[22]Roy D. Patterson and Anne Cutler. "Auditory preprocessing and recognition of speech." Research directions in cognitive science: A european perspective: Vol. 1. Cognitive psychology. Erlbaum, 1989. 23-60.
[23]E. De Boer and Chr Kruidenier. "On ringing limits of the auditory periphery." Biological cybernetics 63.6 (1990): 433-442.
[24]Toshio Irino and Roy D. Patterson. "A time-domain, level-dependent auditory filter: The gammachirp." The Journal of the Acoustical Society of America101.1 (1997): 412-419.
[25]T. S. Chi, Powen Ru, and Shihab A. Shamma. "Multiresolution spectrotemporal analysis of complex sounds." The Journal of the Acoustical Society of America 118.2 (2005): 887-906.
[26]Mounya, Elhilali, T. S. Chi, and Shihab A. Shamma. "A spectro-temporal modulation index (STMI) for assessment of speech intelligibility." Speech communication 41.2 (2003): 331-348.
[27]Pei-Chun Tsai, et al. "A hearing model to estimate mandarin speech intelligibility for the hearing impaired patients." Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, 2015.
[28]http://www.hearingreview.com/2000/09/concave-curvilinear-wdrc-optimizing-the-shape-of-compression/
[29]P. M. Sellick, R. Patuzzi, and B. M. Johnstone. "Measurement of basilar membrane motion in the guinea pig using the Mössbauer technique." The journal of the acoustical society of America 72.1 (1982): 131-141.
[30]Mario A. Ruggero and Nola C. Rich. "Furosemide alters organ of Corti mechanics: evidence for feedback of outer hair cells upon the basilar membrane." The Journal of neuroscience 11.4 (1991): 1057-1067.
[31]Brian C.J. Moore, et al. "Effects of flanking noise bands on the rate of growth of loudness of tones in normal and recruiting ears." The Journal of the Acoustical Society of America 77.4 (1985): 1505-1513.
[32]Roy D. Patterson and Ian Nimmo‐Smith. "Off‐frequency listening and auditory‐filter asymmetry." The Journal of the Acoustical Society of America67.1 (1980): 229-245.
[33]Roy D. Patterson, et al. "The deterioration of hearing with age: Frequency selectivity, the critical ratio, the audiogram, and speech threshold." The Journal of the Acoustical Society of America 72.6 (1982): 1788-1803.
[34]Brian R. Glasberg and Brian CJ Moore. "Derivation of auditory filter shapes from notched-noise data." Hearing research 47.1 (1990): 103-138.
[35]Auditory Perception Group University of Cambridge provides Auditory demonstrations and useful software.
[36]David R. Soderquist and John W. Lindsey. "Physiological noise as a masker of low frequencies: the cardiac cycle." The Journal of the Acoustical Society of America 52.4B (1972): 1216-1220.
[37]Victor Nedzelnitsky. "Sound pressures in the basal turn of the cat cochlea." The Journal of the Acoustical Society of America 68.6 (1980): 1676-1689.
[38]Thomas J. Lynch III, Victor Nedzelnitsky, and William T. Peake. "Input impedance of the cochlea in cat." The Journal of the Acoustical Society of America 72.1 (1982): 108-130.
[39]J. J. Zwislocki. "The role of the external and middle ear in sound transmission." The nervous system 3 (1975): 45-55.
[40]James M. Kates and Kathryn H. Arehart. "The hearing-aid speech perception index (HASPI)." Speech Communication 65 (2014): 75-93.
[41]James Kates. "An auditory model for intelligibility and quality predictions." Proceedings of Meetings on Acoustics. Vol. 19. No. 1. Acoustical Society of America, 2013.
[42]Brian CJ Moore, et al. "Inter-relationship between different psychoacoustic measures assumed to be related to the cochlear active mechanism." The Journal of the Acoustical Society of America 106.5 (1999): 2761-2778.
[43]Shih-Ting Lin and T. S. Chi. "Combining Spectral and temporal Speech Enhancement to Improve Speech Intelligibility." A Thesis for the Degree of Master of Science in Communication Engineering, National Chiao-Tung University, 2014.
[44]Pei-Chun Tsai and T. S. Chi "A Study of Perceptual Effects of Spectral Sharpening on the Hearing-impaired." A Thesis Submitted to Master Program of Sound and Music Innovation Technologies College of Engineering, National Chiao-Tung University, 2014.
[45]Cees H. Taal, et al. "An algorithm for intelligibility prediction of time–frequency weighted noisy speech." Audio, Speech, and Language Processing, IEEE Transactions on 19.7 (2011): 2125-2136.
[46]Fei Chen, Oldooz Hazrati, and Philipos C. Loizou. "Predicting the intelligibility of reverberant speech for cochlear implant listeners with a non-intrusive intelligibility measure." Biomedical signal processing and control 8.3 (2013): 311-314.
[47]Donald D Greenwood. "A cochlear frequency‐position function for several species—29 years later." The Journal of the Acoustical Society of America87.6 (1990): 2592-2605.
[48]Malcolm Slaney. "An efficient implementation of the Patterson-Holdsworth auditory filter bank." Apple Computer, Perception Group, Tech. Rep 35 (1993): 8.
[49]Kuen-Shian Tsai, et al. "Development of a mandarin monosyllable recognition test." Ear and hearing 30.1 (2009): 90-99.
[50]T. S. Chi, et al. "Spectro-temporal modulation transfer functions and speech intelligibility." The Journal of the Acoustical Society of America 106.5 (1999): 2719-2732.
[51]Taffeta M. Elliott and Frédéric E. Theunissen. "The modulation transfer function for speech intelligibility." PLoS comput biol 5.3 (2009): e1000302-e1000302.
[52]Brian C.J. Moore and Brian R. Glasberg. "Use of a loudness model for hearing-aid fitting. I. Linear hearing aids." British journal of audiology 32.5 (1998): 317-335.
[53]Jing Chen, Thomas Baer, and Brian CJ Moore. "Effect of spectral change enhancement for the hearing impaired using parameter values selected with a genetic algorithm." The Journal of the Acoustical Society of America 133.5 (2013): 2910-2920.

連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top