跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.173) 您好!臺灣時間:2026/10/06 15:35
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:周瑋傑
研究生(外文):Chou, Wei-Chieh
論文名稱:以深層類神經網路標記中文階層式多標籤語意概念
論文名稱(外文):Hierarchical Multi-Label Chinese Word Semantic Labeling using Deep Neural Network
指導教授:陳信宏陳信宏引用關係
指導教授(外文):Chen, Sin-Horng
口試委員:王逸如、江振宇
口試委員(外文):Wang, Yih-Ru、Chiang, Chen-Yu
學位類別:碩士
校院名稱:國立交通大學
系所名稱:電機工程學系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2018
畢業學年度:106
語文別:中文
論文頁數:46
中文關鍵詞:詞向量、類神經網路、最小分類誤差、廣義知網、階層式分類、多標籤分類
外文關鍵詞:Word2Vec、neural network、minimum classification error、E-HowNet、hierarchical classification、multi-label classification
相關次數:
  • 被引用被引用:1
  • 點閱點閱:371
  • 評分評分:
  • 下載下載:33
  • 收藏至我的研究室書目清單書目收藏:1
傳統上對超過100個階層式標籤分類可以使用其扁平 (flatten) 標籤做分類,但如此會喪失架構樹 (taxonomy) 的階層資訊。本研究旨在對廣義知網中文詞彙做概念分類與標記並提出考慮廣義知網架構樹階層關係之深層類神經網路訓練方法,此方法將上一階層神經網路的輸出結果作為下一階層神經網路訓練的輔助,該架構之輸入為詞彙樣本點的詞向量,而本研究亦希望在繁體中文詞彙上汲取更深層的詞彙意義,故詞向量方面本研究亦提出考慮上下文前後關係之2-Bag Word2Vec,而各階層的訓練結果有不同的重要性,所以在模型的最後使用最小分類誤差法以賦予各階層在測試階段時不同的權重。實驗結果顯示階層式 (hierarchical) 分類預測正確率會比扁平分類還高。
Traditionally, classifying over 100 hierarchical multi-labels could use flatten classification, but it will lose the taxonomy structure information. This paper aimed to classify the concept of word in E-HowNet and proposed a deep neural network training method with hierarchical relationship in E-HowNet taxonomy. This method takes neural network output of the upper level as the auxiliary neural network input in the next level. The input of neural network is word embedding. About word embedding, this paper proposed order-aware 2-Bag Word2Vec. Experiment results shown hierarchical classification will achieved higher accuracy than flatten classification.
中文摘要 I
ABSTRACT II
誌謝 III
目錄 IV
圖目錄 VI
表目錄 VII
第一章 緒論 1
1.1研究動機 1
1.2文獻回顧 2
1.3研究方法 2
1.4 章節概要說明 3
第二章 建立詞向量 4
2.1文字語料庫簡介 4
2.2文本前處理: 5
2.2.1 CRF斷詞 5
2.2.2 斷詞後處理 6
2.3詞向量 8
2.4 WORD2VEC架構 9
2.4.1連續詞袋模型(Continuous Bag-of-Word Model) 9
2.4.2跳躍式模型(Skip-gram Model) 11
2.5 2-BAG WORD2VEC 12
2.6 餘弦相似度 13
第三章 扁平分類 (FLATTEN) 15
3.1《廣義知網》(E-HOWNET) 15
3.2 最近鄰居演算法(K-NEAREST NEIGHBORS ALGORITHM, KNN) 17
3.3 類神經網路(NEURAL NETWORK, NN) 18
3.4 實驗流程 22
3.5.1 文字語料庫與概念抽取 23
3.5.2 訓練詞向量 25
3.5.3 不平衡類別資料處理 25
3.5.4 評測效能與KNN的平票問題 26
3.5.5 實驗結果與討論 27
第四章 階層式多標籤分類(HIERARCHICAL) 32
4.1 廣義知網架構樹 32
4.2 階層式多標籤分類模型 33
4.3 更正矛盾情況 34
4.4 效能評估方式 35
4.6 實驗流程 36
4.6.1 加入架構樹資訊 36
4.6.2 賦予各階層不同權重 38
4.6.3 實驗結果與討論 40
第五章 結論與未來展望 42
5.1 結論 42
5.2 未來展望 42
參考文獻 43
[1] Huang, Shu-Ling, You-Shan Chung, and Keh-Jiann Chen. "E-HowNet: the expansion of HowNet." Proceedings of the First National HowNet Workshop. 2008.
[2] 蘇偉峰、李紹滋, “一个基于概念的中文文本分类模型,” 廈門大學計算機科學系, 2002.
[3] 劉群、李素建,“基於《知網》的辭彙語義相似度計算”, Computational Linguistics and Chinese Language Processing, Vol. 7, No. 2, August 2002, pp. 59-76。
[4] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. “ Efficient Estimation of Word Representations in Vector Space, “, In Proceedings of Workshop at ICLR, 2013.
[5] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. “ Distributed Representations of Words and Phrases and their Compositionality”, In Proceedings of NIPS, 2013.
[6] Ling, Wang and Dyer, Chris and Black, Alan and Trancoso, Isabel, "Two/Too Simple Adaptations of word2vec for Syntax Problems", Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2015.
[7] Wang, Yih-Ru, et al. "Conditional random field-based parser and language model for traditional Chinese spelling checker." Proceedings of the 7th SIGHAN Workshop on Chinese Language Processing (SIGHAN’13). 2013.
[8] Huang, Shu-Ling, You-Shan Chung, and Keh-Jiann Chen. "E-HowNet: the expansion of HowNet." Proceedings of the First National HowNet Workshop. 2008.
[9] 陳克健, et al. "多層次概念定義與複雜關係表達-繁體字知網的新增架構." (2004).
[10] Dong, Zhendong, and Qiang Dong. "HowNet." (2000).
[11] Altman,N.S. “An introduction to kernel and nearest –neighbor nonparametricregression . The American Statistician”. 1992, 46(3):175-185
[12] code.google.com(2015,September, 1).Google Code Archive – word2vec(2016).
Available: https://code.google.com/archive/p/word2vec/
[13] Everitt, B. S., Landau, S., Leese, M. and Stahl, D. (2011) MiscellaneousClustering Methods, in Cluster Analysis, 5th Edition, John Wiley & Sons, Ltd, Chichester, UK.
[14] Mikolov, Tomas, Quoc V. Le, and Ilya Sutskever. "Exploiting similarities among languages for machine translation." arXiv preprint arXiv:1309.4168 (2013).
[15] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman. “Indexing by latent semantic analysis”, Journal of the American Society for Information Science, 41:391–407, September, 1990.
[16] Levy, Omer, and Yoav Goldberg. "Neural word embedding as implicit matrix factorization." Advances in neural information processing systems. 2014.
[17] Miller, George A. "WordNet: a lexical database for English." Communications of the ACM 38.11 (1995): 39-41.
連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊