|
[1] G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition, in IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, 2012. [2] A. Graves, A.-R. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks, in ICASSP, 2013. [3] Speech Recognition on LibriSpeech test-clean. Available: https://paperswithcode.com/sota/speech-recognition-on-librispeech-test-clean [4] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books, in ICASSP, 2015. [5] J. Rajnoha, and P. Pollák, “ASR systems in noisy environment: analysis and solutions for increasing noise robustness, in Radioengineering, vol. 20, no. 1, pp. 74–84, 2011. [6] V. Peddinti, V. Manohar, Y. Wang, D. Povey, and S. Khudanpur, “Far-field ASR without parallel data, in Interspeech, 2016. [7] S.-L. Chen, “Code-switched word recognition by Taiwanese-Mandarin bilinguals, 2000. [8] 葉高華, “99年人口及住宅普查, 2010. [9] J. Yi, J. Tao, Z. Wen,and Y. Bai, “Adversarial multilingual training for low-resource speech recognition, in ICASSP, 2018. [10] Z. Zeng, Y. Khassanov, V. T. Pham, H. Xu, E. S. Chng, and H. Li, “On the end-to-end solution to Mandarin-English code-switching speech recognition, arXiv:1811.00241, 2018. [11] E. Yılmaz, H. Van den Heuvel, and D. A. Van Leeuwen, “Acoustic and textual data augmentation for improved asr of code-switching speech, in Interspeech, 2018. [12] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks, in ICML, 2006. [13] W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell, arXiv:1508.01211, 2015. [14] C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, “State-of-the-art speech recognition with sequence-to-sequence models, in ICASSP, 2018. [15] D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI, in Interspeech, 2016. [16] V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts, in Interspeech, 2015. [17] J. Kunze, L. Kirsch, I. Kurenkov, A. Krug, J. Johannsmeier, and S. Stober, “Transfer learning for speech recognition on a budget, in ACL, 2017. [18] P. Ghahremani, V. Manohar, H. Hadian, D. Povey, and S. Khudanpur, “Investigation of transfer learning for ASR using LF-MMI trained neural networks, in ASRU, 2017. [19] Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky. , “Domain-adversarial training of neural networks, in JMLR, 2016. [20] S. Sun, C.-F. Yeh, M.-Y. Hwang, M. Ostendorf, and L. Xie, “Domain adversarial training for accented speech recognition, in ICASSP, 2018. [21] S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid ctc/attention architecture for end-to-end speech recognition, in IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 8, pp. 1240–1253, 2017. [22] iCorpus 臺華平行新聞語料庫. Available: http://icorpus.iis.sinica.edu.tw [23] 台語文語料庫蒐集及語料庫為本台語書面語音節詞頻統計. Available: http://ip194097.ntcu.edu.tw/giankiu/keoe/KKH/guliau-supin/guliau-supin.asp [24] 台語文數位典藏資料庫. Available: http://ip194097.ntcu.edu.tw/nmtl/dadwt/pbk.asp [25] 新約聖經語料. Available: https://bible.fhl.net [26] 臺語國校仔課本. Available: https://github.com/Taiwanese-Corpus/kok4hau7-kho3pun2 [27] King-ASR-044. Available: https://kingline.speechocean.com/exchange.php?id=766&act=view [28] King-ASR-360. Available: https://kingline.speechocean.com/category.php?id=120 [29] H.-M. Wang, B. Chen, J.-W. Kuo, and S.-S. Cheng, “MATBN: A Mandarin Chinese broadcast news corpus, International Journal of Computational Linguistics and Chinese Language Processing, vol. 10, no. 2, pp. 219–236, 2005. [30] TCC-300. Available: http://www.aclclp.org.tw/use_mat_c.php [31] Tagged Chinese Gigaword Version 2.0. Available: https://catalog.ldc.upenn.edu/LDC2009T14 [32] PTT 八卦版問答中文語料. Available: https://github.com/zake7749/Gossiping-Chinese-Corpus [33] M. Mohri, F. Pereira, and M. Riley, “Weighted finite-state transducers in speech recognition, in Computer, Speech and Language, vol. 16, no. 1, pp. 69–88, 2002. [34] ChhoeTaigi 找台語:台語字詞資料庫. Available: https://github.com/ChhoeTaigi/ChhoeTaigiDatabase [35] D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit, in ASRU, 2011. [36] G. Lample, M. Ott, A. Conneau, L. Denoyer, and M. A. Ranzato, “Phrase-based & neural unsupervised machine translation, arXiv:1804.07755, 2018. [37] A. Stolcke, “SRILM - an extensible language modeling toolkit, in Proc. ICSLP, pp. 901–904, 2002
|