|
[1] Lawrence R Rabiner, “A tutorial on hidden markov models and selected applications in speech recognition,” Proceedings of the IEEE, vol. 77, no. 2, pp. 257–286, 1989. [2] Navdeep Jaitly, Patrick Nguyen, AndrewWSenior, and Vincent Vanhoucke, “Application of pretrained deep neural networks to large vocabulary speech recognition.,” in INTERSPEECH, 2012. [3] George E Dahl, Dong Yu, Li Deng, and Alex Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” Audio, Speech, and Language Processing, IEEE Transactions on, vol. 20, no. 1, pp. 30–42, 2012. [4] Thomas Kemp and Alex Waibel, “Unsupervised training of a speech recognizer: recent experiments.,” in Eurospeech, 1999. [5] Lori Lamel, Jean-Luc Gauvain, and Gilles Adda, “Lightly supervised and unsupervised acoustic model training,” Computer Speech & Language, vol. 16, no. 1, pp. 115–129, 2002. [6] Alex S Park and James R Glass, “Unsupervised pattern discovery in speech,” Audio, Speech, and Language Processing, IEEE Transactions on, vol. 16, no. 1, pp. 186– 197, 2008. [7] Gautam K Vallabha, James L McClelland, Ferran Pons, Janet F Werker, and Shigeaki Amano, “Unsupervised learning of vowel categories from infant-directed speech,” Proceedings of the National Academy of Sciences, vol. 104, no. 33, pp.13273–13278, 2007. [8] Fang Zheng, Guoliang Zhang, and Zhanjiang Song, “Comparison of different implementations of mfcc,” Journal of Computer Science and Technology, vol. 16, no.6, pp. 582–589, 2001. [9] Tara N Sainath, Brian Kingsbury, and Bhuvana Ramabhadran, “Auto-encoder bottleneck features using deep belief networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2012 IEEE International Conference on. IEEE, 2012, pp. 4153–4156. [10] Ruslan Salakhutdinov and Geoffrey E Hinton, “Deep boltzmann machines,” in International Conference on Artificial Intelligence and Statistics, 2009, pp. 448–455. [11] Aren Jansen and Benjamin Van Durme, “Efficient spoken term discovery using randomized algorithms,” in Automatic Speech Recognition and Understanding (ASRU), 2011 IEEE Workshop on. IEEE, 2011, pp. 401–406. [12] Armando Muscariello, Guillaume Gravier, and Fr´ed´eric Bimbot, “Unsupervised motif acquisition in speech via seeded discovery and template matching combination,” Audio, Speech, and Language Processing, IEEE Transactions on, vol. 20, no.7, pp. 2031–2044, 2012. [13] Cheng-Tao Chung, Chun-an Chan, and Lin-shan Lee, “Unsupervised discovery of linguistic structure including two-level acoustic patterns using three cascaded stages of iterative optimization,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013, pp. 8081–8085. [14] Xuedong Huang, Alex Acero, Hsiao-Wuen Hon, and Raj Foreword By-Reddy, Spoken language processing: A guide to theory, algorithm, and system development, Prentice Hall PTR, 2001. [15] Kristen Precoda, “Non-mainstream languages and speech recognition: Some challenges,” CALICO journal, vol. 21, no. 2, pp. 229–243, 2013. [16] Peter W Jusczyk, The discovery of spoken language, MIT press, 2000. [17] Steven Pinker, The language instinct: The new science of language and mind, vol.7529, Penguin UK, 1994. [18] Jenny R Saffran, “Constraints on statistical language learning,” Journal of Memory and Language, vol. 47, no. 1, pp. 172–196, 2002. [19] James Glass, “Towards unsupervised speech processing,” in Information Science, Signal Processing and their Applications (ISSPA), 2012 11th International Conference on. IEEE, 2012, pp. 1–4. [20] Alex Park and James R Glass, “Towards unsupervised pattern discovery in speech,” in Automatic Speech Recognition and Understanding, 2005 IEEE Workshop on. IEEE, 2005, pp. 53–58. [21] Armando Muscariello, Guillaume Gravier, and Fr´ed´eric Bimbot, “Audio keyword extraction by unsupervised word discovery,” in INTERSPEECH 2009: 10th Annual Conference of the International Speech Communication Association, 2009. [22] Yaodong Zhang and James R Glass, “Towards multi-speaker unsupervised speech pattern discovery,” in Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on. IEEE, 2010, pp. 4366–4369. [23] Aren Jansen, Kenneth Church, and Hynek Hermansky, “Towards spoken term discovery at scale with zero resources.,” in INTERSPEECH, 2010, pp. 1676–1679. [24] Man-hung Siu, Herbert Gish, Steve Lowe, and Arthur Chan, “Unsupervised audio patterns discovery using hmm-based self-organized units,” in Twelfth Annual Conference of the International Speech Communication Association, 2011. [25] Armando Muscariello, Guillaume Gravier, and Fr´ed´eric Bimbot, “Zero-resource audio-only spoken term detection based on a combination of template matching techniques,” in INTERSPEECH 2011: 12th Annual Conference of the International Speech Communication Association, 2011. [26] Man-Hung Siu, Herbert Gish, Arthur Chan, and William Belfield, “Improved topic classification and keyword discovery using an hmm-based speech recognizer trained without supervision,” in Eleventh Annual Conference of the International Speech Communication Association, 2010. [27] Aren Jansen, Samuel Thomas, and Hynek Hermansky, “Weak top-down constraints for unsupervised acoustic model training.,” in ICASSP, 2013, pp. 8091–8095. [28] Leonardo Badino, Claudia Canevari, Luciano Fadiga, and Giorgio Metta, “An autoencoder based approach to unsupervised learning of subword units,” in Acoustics,Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 7634–7638. [29] Cheng-Tao Chung, Chun-an Chan, and Lin-shan Lee, “Unsupervised spoken term detection with spoken queries by multi-level acoustic patterns with varying model granularity,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 7814–7818. [30] Geoffrey E Hinton, “Learning distributed representations of concepts,” in Proceedings of the eighth annual conference of the cognitive science society. Amherst, MA, 1986, vol. 1, p. 12. [31] Joseph Turian, Lev Ratinov, Yoshua Bengio, and Dan Roth, “A preliminary evaluation of word representations for named-entity recognition,” in NIPS Workshop on Grammar Induction, Representation of Language and Language Learning, 2009, pp. 1–8. [32] Richard Socher, Christopher D Manning, and Andrew Y Ng, “Learning continuous phrase representations and syntactic parsing with recursive neural networks,” in Proceedings of the NIPS-2010 Deep Learning and Unsupervised Feature Learning Workshop, 2010, pp. 1–9. [33] William Blacoe and Mirella Lapata, “A comparison of vector-based representations for semantic composition,” in Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. Association for Computational Linguistics, 2012, pp. 546–556. [34] Duyu Tang, Furu Wei, Nan Yang, Ming Zhou, Ting Liu, and Bing Qin, “Learning sentiment-specific word embedding for twitter sentiment classification,” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, 2014, pp. 1555–1565. [35] Wei Xu and Alexander I Rudnicky, “Can artificial neural networks learn language models?,” 2000. [36] Andriy Mnih and Geoffrey Hinton, “Three new graphical models for statistical language modelling,” in Proceedings of the 24th international conference on Machine learning. ACM, 2007, pp. 641–648. [37] Tomas Mikolov, Martin Karafi´at, Lukas Burget, Jan Cernock`y, and Sanjeev Khudanpur, “Recurrent neural network based language model.,” in INTERSPEECH 2010, 11th Annual Conference of the International Speech Communication Association, Makuhari, Chiba, Japan, September 26-30, 2010, 2010, pp. 1045–1048. [38] Ronan Collobert, JasonWeston, L´eon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa, “Natural language processing (almost) from scratch,” The Journal of Machine Learning Research, vol. 12, pp. 2493–2537, 2011. [39] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems, 2013, pp. 3111–3119. [40] Jeffrey Pennington, Richard Socher, and Christopher D Manning, “Glove: Global vectors for word representation,” Proceedings of the Empiricial Methods in Natural Language Processing (EMNLP 2014), vol. 12, 2014. [41] Peter F Brown, Peter V Desouza, Robert L Mercer, Vincent J Della Pietra, and Jenifer C Lai, “Class-based n-gram models of natural language,” Computational linguistics, vol. 18, no. 4, pp. 467–479, 1992. [42] Yoshua Bengio, R´ejean Ducharme, Pascal Vincent, and Christian Janvin, “A neural probabilistic language model,” The Journal of Machine Learning Research, vol. 3, pp. 1137–1155, 2003. [43] Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, “On the difficulty of training recurrent neural networks,” arXiv preprint arXiv:1211.5063, 2012. [44] Kaisheng Yao, Baolin Peng, Geoffrey Zweig, Dong Yu, Xiaolong Li, and Feng Gao, “Recurrent conditional random field for language understanding,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 4077–4081. [45] Yun-Chiao Li, “Enhenced semantic retrieval of spoken content with query expansion and automatically discovered acoustic patterns,” 2014. [46] Tomas Mikolov, Stefan Kombrink, Anoop Deoras, Lukar Burget, and Jan Cernocky, “Rnnlm-recurrent neural network language modeling toolkit,” in Proc. of the 2011 ASRU Workshop, 2011, pp. 196–201. [47] Sainbayar Sukhbaatar and Rob Fergus, “Learning from noisy labels with deep neural networks,” arXiv preprint arXiv:1406.2080, 2014. [48] M.A. Pitt, L. Dilley, K. Johnson, S. Kiesling,W. Raymond, E. Hume, and E. Fosler-Lussier, “Buckeye corpus of conversational speech (2nd release),” 2007, Columbus, OH: Department of Psychology, Ohio State University (Distributor). [49] Nic J De Vries, Marelie H Davel, Jaco Badenhorst,Willem D Basson, Febe DeWet, Etienne Barnard, and Alta De Waal, “A smartphone-based asr data collection tool for under-resourced languages,” Speech communication, vol. 56, pp. 119–131, 2014. [50] Thomas Schatz, Vijayaditya Peddinti, Francis Bach, Aren Jansen, Hynek Hermansky, and Emmanuel Dupoux, “Evaluating speech features with the minimal-pair abx task: Analysis of the classical mfc/plp pipeline,” in INTERSPEECH 2013: 14th Annual Conference of the International Speech Communication Association, 2013, pp. 1–5. [51] Thomas Schatz, Vijayaditya Peddinti, Xuan-Nga Cao, Francis Bach, Hynek Hermansky, and Emmanuel Dupoux, “Evaluating speech features with the minimal-pair abx task (ii): Resistance to noise,” in Fifteenth Annual Conference of the International Speech Communication Association, 2014. [52] Bogdan Ludusan, Maarten Versteegh, Aren Jansen, Guillaume Gravier, Xuan-Nga Cao, Mark Johnson, and Emmanuel Dupoux, “Bridging the gap between speech technology and natural language processing: an evaluation toolbox for term discovery systems,” in Language Resources and Evaluation Conference, 2014. [53] Yun-Chiao Li, Hung-yi Lee, Cheng-Tao Chung, Chun-an Chan, and Lin-shan Lee, “Towards unsupervised semantic retrieval of spoken content with query expansion based on automatically discovered acoustic patterns,” in Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on. IEEE, 2013, pp. 198–203. [54] Thomas L Griffiths and Mark Steyvers, “Finding scientific topics,” Proceedings of the National Academy of Sciences, vol. 101, no. suppl 1, pp. 5228–5235, 2004. [55] David M Blei, Andrew Y Ng, and Michael I Jordan, “Latent dirichlet allocation,” the Journal of machine Learning research, vol. 3, pp. 993–1022, 2003. [56] Gregor Heinrich, “Parameter estimation for text analysis,” Tech. Rep., Technical report, 2005. [57] Christian Plahl, Ralf Schluter, and Hermann Ney, “Hierarchical bottle neck features for lvcsr.,” in Interspeech, 2010, pp. 1197–1200. [58] Frantisek Grezl and Petr Fousek, “Optimizing bottle-neck features for lvcsr,” in Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on. IEEE, 2008, pp. 4729–4732. [59] Dong Yu and Michael L Seltzer, “Improved bottleneck features using pretrained deep neural networks.,” in INTERSPEECH, 2011, vol. 237, p. 240. [60] Ahilan Kanagasundaram, Robbie Vogt, David B Dean, Sridha Sridharan, and Michael W Mason, “I-vector based speaker recognition on short utterances,” in Proceedings of the 12th Annual Conference of the International Speech Communication Association. International Speech Communication Association (ISCA), 2011, pp. 2341–2344. [61] Vishwa Gupta, Patrick Kenny, Pierre Ouellet, and Themos Stafylakis, “I-vectorbased speaker adaptation of deep neural networks for french broadcast audio transcription,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 6334–6338. [62] Lawrence Davis et al., Handbook of genetic algorithms, vol. 115, Van Nostrand Reinhold New York, 1991.
|