跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.127) 您好!臺灣時間:2026/07/30 09:25
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:康乃人
研究生(外文):KANG, NAN-RAN
論文名稱:使用卷積神經網路進行語者識別
論文名稱(外文):Speaker Verification using Convolution Neural Network
指導教授:王益文
指導教授(外文):WANG,YI-WEN
口試委員:王益文陳德生劉怡芬
口試委員(外文):WANG,YI-WENCHEN, DE-SHENGLIU, YI-FEN
口試日期:2018-07-18
學位類別:碩士
校院名稱:逢甲大學
系所名稱:資訊工程學系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2018
畢業學年度:106
語文別:中文
論文頁數:33
中文關鍵詞:卷積神經網路語者辨識
外文關鍵詞:Convolution Neural NetworkSpeaker Verification
相關次數:
  • 被引用被引用:0
  • 點閱點閱:265
  • 評分評分:
  • 下載下載:27
  • 收藏至我的研究室書目清單書目收藏:0
生物辨識技術在日常生活中早已不是新鮮事,在近年來越來越流行,從指紋辨識、虹膜辨識、聲紋辨識、I-phone的Face ID等等都是,而語者辨識也是其中的一種。
在處理語者識別的問題中,語者辨識系統可分為兩個部分:特徵擷取、比較分類。過去的文獻中都是將兩個部分使用不同的方法解決,由於近年來深度學習(Deep Learning)發展快速,使用類神經網路進行系統辨識獲得廣度的發展,而在語者辨識的部分較常使用遞歸神經網路(Recurrent Neural Network)來解決問題,較少看到使用卷積神經網路(Convolution Neural Network)解決語者辨識的問題,由於卷積神經網路(Convolution Neural Network)對於擷取特徵細節有著不錯的效果,本文決定使用卷積神經網路技術作為基礎,嘗試同時解決語者辨識系統的兩個部分並且提出一個能夠成功辨識的方法。

Biometric system is no longer a new thing in daily life, and it has become more and more popular in recent years, fingerprint recognition, iris recognition, voiceprint recognition, and I-phone's Face ID are all biometric system, and speaker verification is one of them. Speaker recognition can be divided into two parts: feature extraction and classification. In the past, the two parts were solved by different methods, due to the rapid development of deep learning, the neural network for speaker recognition has gained breadth of development. In the part of the speaker verification, the Recurrent Neural Network is often used to solve the problem. It is less common to use the Convolution Neural Network to solve the problem of speaker verification. Since the Convolution Neural Network has a good effect on extracting feature details, this paper decided to use convolutional neural network as the basis to try to solve the two parts of the speaker recognition and propose a successful Methods to solve speaker verification.
誌謝 i
摘要 ii
Abstract iii
目錄 iv
圖目錄 vi
表目錄 vii
第一章 緒論 1
1.1 研究動機 1
1.2 章節介紹 2
第二章 相關研究 3
2.1前饋式神經網路(FEED FORWARD NEURAL NETWORK) 4
2.1.1激活函數(Activation Function) 4
2.1.2反向傳導演算法(BACKPROPAGATION ALGORITHM) 5
2.2卷積神經網路(CONVOLUTIONAL NEURAL NETWORKS) 6
2.2.1卷積層(Convolutional Layer) 6
2.2.2池化層(Pooling Layer) 7
2.3 梅爾頻率倒譜係數 (MEL-FREQUENCY CEPSTRAL COEFFICIENTS) 8
2.3.1梅爾頻率能量係數 (Mel-Frequency energy coefficients) 8
第三章 研究方法 9
3.1 系統架構 9
3.2 聲音訊號處理 10
3.3 卷積神經網路架構 14
3.4 計算準確率 15
第四章 實驗結果 17
4.1 系統環境及實作細節 17
4.2 模型評估 17
4.3 實驗結果 19
4.3.1 模型辨識能力檢驗 19
4.3.2 模型拒絕能力檢驗 19
4.3.3 聲音相似度對於模型驗證身分能力之影響 21
第五章 結論 24
參考文獻 25


[1]Xinhui Hu, Xugang Lu, and Chiori Hori."Mandarin “Speech Recognition Using Convolution Neural Network with Augmented Tone Features” International Symposium on Chinese Spoken Language Processing (ISCSLP), 2014.
[2]Takuya Yoshioka, Shigeki Karita, and Tomohiro Nakatani. “Far-Field Speech Recognition using CNN-DNN-HMM with Convolution in time” IEEE ICASSP, 2015.
[3]Mirco Ravanellix, Philemon Brakely, Maurizio Omologox, and Yoshua Bengioy. “Batch-Normalized Joint Training for DNN-Based Distant Speech Recognition” IEEE GlobalSIP, 2016.
[4]Yuan Liu, Yanmin Qian, Nanxin Chen, Tianfan Fu, Ya Zhang, and Kai Yu. “Deep feature for text-dependent speaker verification” ScienceDirect.Speech.Communication.73, July 2015.
[5]Omid Ghahabi and Javier “Hernando Deep Learning Backend for Single and Multi-Session i-Vector Speaker Recognition” IEEE/ACM Transactions on Audio, Speech, and Language Processing. , 2017.
[6]Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez “Deep Neural Networks for Small Footprint text-dependent Speaker Vverification,” IEEE.International Conference on Acoustic, Speech and Signal Processing., 2014.
[7]Panu Somervuo, Aki Härmä, and Seppo and Fagerlund “Parametric Representations of Bird Sounds for Automatic Species Recognition” IEEE Transactions on Audio, Speech, and Language Processing, 2006
[8]Amirsina Torfi, Jeremy Dawson, and Nasser M. Nasrabadi “Text-Independent Speaker Verification Using 3D Convolutional Neural Networks” cornell university Computer Science > Computer Vision and Pattern Recognition, 2017
[9]IEEE Subcommittee, IEEE Recommended Practice for Speech Quality Measurements. IEEE Trans. Audio and Electroacoustics, AU-17(3), 225-246, 1969.

QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top