|
[Aubert 2002] X. L. Aubert, “An Overview of Decoding Techniques for Large Vocabulary Continuous Speech Recognition,” Computer Speech and Language, January 2002. [Boll 1979] S.F. Boll, “Supperssion of Acoutstic Noise in Speech Using Spectral Subtraction,” IEEE Trans. on ASSP, Vol. 27, No. 2, pp. 133-120, 1979. [Bocchieri and Wilpon 1992] EL Bocchieri, JG Wilpon "Discriminative analysis for feature reduction in automatic speechrecognition," Acoustics, Speech, and Signal Processing, ICASSP 1992. [Chen et al. 2004] Berlin Chen, Jen-Wei Kuo, Wen-Hung Tsai, “Lightly Supervised and Data-Driven Approaches to Mandarin Broadcast News Transcription,” in Proc. ICASSP 2004. [Chen et al. 2005] Berlin Chen, Jen-Wei Kuo, Wen-Huang Tsai, “Lightly Supervised and Data-Driven Approaches to Mandarin Broadcast News Transcription,” International Journal of Computational Linguistics and Chinese Language Processing, Vol. 10, No. 1, pp. 1-18, March 2005. [Davis et al. 1980] Davis, S. B. and P Mermelstein, "Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences." IEEE Transactions on Acoustics, Speech, and Signal Processing 28(4); pp. 357-366, 1980. [ETSI 2000] H. G. Hirsch, D. Pearce, “The AURORA Experimental Framework for the Performance Evaluations of Speech Recognition Systems under Noisy Conditions,” in Proc. ISCA ITRW ASR 2000. [Furui 1981] S. Furui, “Cepstral Analysis Techniques for Automatic Speaker Verification,” IEEE Trans. on ASSP, 1981. [Gauian and Lee 1994] J.L. Gauian and C.H. Lee, “Maximum a Posteriori Estimation for Multivariate Gaussian Mixture Observations of Markov Chains,” IEEE Trans. on Speech and Audio Processing, 1994. [Gomez et al. 2004] R. Gomez, A. Lee, K. Shikano, “Robust Speech Recognition with Spectral Subtraction in low SNR,” in Proc. ICSLP 2004. [Gillick and Cox 1989] L. Gillick and S. Cox, "Some Statistical Issues in the Comparison of Speech Recognition Algorithms", in Proc. ICASSP 89, pp. 532-535. [Gong 1995] Gong, Y., "Speech Recognition in Noisy Environments:A Survey," Speech Communication 16(3); pp. 261-291. [Górriz et al. 2006] J.M. G´orriz, J. Ram´ırez, C.G. Puntonet, J.C. Segura, “An Efficient Bispectrum Phase Entropy-based Algorithm for VAD,” in Proc. ICSLP 2006. [Gillick and Cox 1989] L. Gillick and S. Cox, "Some Statistical Issues in the Comparison of Speech Recognition Algorithms", in Proc. ICASSP 89, pp. 532-535. Matched Pairs Sentence-Segment Word Error (MAPSSWE) Test http://www.nist.gov/speech/tests/sigtests/mapsswe.htm. [Hermansky 1998] Hynek Hermansky, “Should Recognizers Have Ears?”, Speech Communication, 1998. [Huang and Hon 2001] X. Huang, A. Acero and H. Hon, “Spoken Language Processing: A Guide to Theory, Algorithm and System Development,” Prentice Hall PTR Upper Saddle River, NJ, USA, 2001. [HTK 2006] S. Young et al., “The HTK Book Version 3.4,” 2006. [Katz 1987] S. M. Katz, “Estimation of Probabilities from Sparse Data for Other Language Component of a Speech Recognizer,” IEEE Trans. Acoustics, Speech and Signal Processing, Vol. 35, No. 3, pp. 400-401, 1987. [Leggetter and Woodland 1995] C.J. Leggetter and P.C. Woodland, “Maximum Likelihood Linear Regression for Speaker Adaptation of Continuous Density Hidden Markov Models,” Computer Speech and Language, 1995. [LDC] Linguistic Data Consortium: http://www.ldc.upenn.edu. [Lin et al. 2006] Shih-Hsiang Lin, Yao-Ming Yeh, Berlin Chen, "Exploiting Polynomial-Fit Histogram Equalization and Temporal Average for Robust Speech Recognition," the 9th International Conference on Spoken Language Processing (Interspeech - ICSLP 2006), Pittsburgh PA, USA, September 17-21, 2006. [Misra et al. 2004] H Misra, S Ikbal, H Bourlard, H Hermansky, “Spectral Entropy Based Feature For Robust ASR,” Acoustics, Speech, and Signal Processing, 2004. [NIST] National Institute of Standards and Technology. http://www.nist.gov/. [Ramírez 2004] Juan Manuel Górriz, Javier Ramírez, Carlos G. Puntonet, and José Carlos Segura, ”Generalized LRT-Based Voice Activity Detector,” IEEE Signal Processing Letters, Vol. 13, No. 10, October 2006. [SRILM] A. Stolcke, “SRI language Modeling Toolkit, ” version 1.3.3, http://www.speech.sri.com/projects/srilm/. [Tai and Hung 2006] Chung-fu Tai and Jeih-weih Hung, “Silence Energy Normalization for Robust Speech Recognition in Additive Noise Environments,” in Proc. ICSLP 2006. [Viikki and Laurila 1998] O. Viikki, K. Laurila, “Cepstral Domain Segmental Feature Vector Normalization for Noise Robust Speech Recognition,” Speech Communication, Vol. 25, pp. 133-147, August 1998. [Weizhong and Douglas 2005] Weizhong Zhu and Douglas O’Shaughnessy, ” Log-Energy Dynamic Range Normalizaton for Robust Speech Recognition,” in Proc. ICASSP 2005 pp. 245- 248. [Wang et al. 2005] Hsin-min Wang, Berlin Chen, Jen-Wei Kuo, and Shih-Sian Cheng, “MATBN: A Mandarin Chinese Broadcast News Corpus,” International Journal of Computational Linguistics & Chinese Language Processing, Vol. 10, No. 2, June 2005, pp. 219-236. [戴仲甫 2006] 戴仲甫, “強健性語音辨認中能量特徵強化及音框選擇之改進技術的研究,” 2006.
|