跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.73) 您好!臺灣時間:2026/07/22 17:28
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:蔡承翰
研究生(外文):Cheng-han Tsai
論文名稱:非監督式文件分類法基於維基百科內容及連結資訊
論文名稱(外文):Unsupervised Text Categorization Method Using Wikipedia Content and Linking Information
指導教授:許中川許中川引用關係
指導教授(外文):Chung-chian Hsu
學位類別:碩士
校院名稱:國立雲林科技大學
系所名稱:資訊管理系碩士班
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2013
畢業學年度:101
語文別:英文
論文頁數:39
中文關鍵詞:監督式學習演算法自動擷取關鍵字非監督式學習演算法文件分類
外文關鍵詞:Automatic keyword extractionSupervised learning algorithmUn-supervised learning algorithmText categorization
相關次數:
  • 被引用被引用:0
  • 點閱點閱:279
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
隨著網際網路的發展,有越來越多的文字資訊被產生出來,但如何將如此大量的文字資料進行正確的分類卻是一項困難且複雜的工程。現今已有相當多的監督式學習演算法被應用在文件分類領域上,但是監督式學習法有幾項缺點,其一是需要大量的已標記文件來當作訓練資料,用於計算字詞之間的相似度。另一項缺點是收集已標記文件是相當耗費時間以及人力,因此本論文提出一項非監督式學習法應用於自動從未標記文件擷取關鍵字,我們利用維基百科當作知識庫,藉由維基百科龐大的文字資料來輔助抓取關鍵字,提升關鍵字擷取的正確度、深度及廣度。從實驗結果可得知,我們的方法相較於傳統基於頻率的方法有著更佳的績效。
With the rapid growth of the Internet, huge amount of text documents has been generated. How to classify the huge quantity of text documents into correct categories is a complex task. Scores of supervised learning algorithms have been proposed for text categorization. However, the supervised learning algorithms have some weaknesses. One of them is that they need a large number of labeled training documents for computing term similarity in order to obtain high accuracy performance. Generally, collecting labeled documents is difficult and costly. In this paper, we proposed an unsupervised learning method, which automatically extracts keywords from unlabeled documents for text categorization. We regard Wikipedia as knowledge resource, and extract words which appeared with keywords for enhancing the keyword list. The experimental results show that our method is better than other term weighting methods based on frequency in text classification.
中文摘要 i
Abstract ii
致謝 iii
1. Introduction 1
2. Literature Review 3
2.1 Supervised learning algorithm 3
2.2 Unsupervised learning algorithm 4
2.3 Term weighting 5
3. Method 6
3.1 Framework 6
3.2 Extracting first-order words 7
3.2.1 Preprocessing and creating first-order words list 7
3.3 Extracting document keywords from Wikipedia 8
3.3.1 Filtering unimportant words 9
3.3.2 Deleting ambiguous words 10
3.4 Extracting the keywords list of each category 11
3.5 The Relationship between document keywords and category titles 13
3.5.1 Wikipedia Relation 13
3.5.2 A Hybrid method 15
4. Experiment 18
4.1 Data Corpora 18
4.1.1 Reuters-10 18
4.2 Comparison methods 18
4.2.1 Comprehensively Measure Feature Selection (CMFS) 18
4.2.2 Replacing the acronyms 20
4.3 Performance measures 20
4.4 Result 21
4.5 The influence of parameter K 27
5. Conclusion 29
References 30
[1]Alencar, R. O. d., Clodoveu Augusto Davis, J., Andr, M., Gon, and alves 2010. "Geographical classification of documents using evidence from Wikipedia," in Proceedings of the 6th Workshop on Geographic Information Retrieval, ACM: Zurich, Switzerland, pp. 1-8.
[2]Chen, P.-I., and Lin, S.-J. 2010. "Automatic keyword prediction using Google similarity distance," Expert Systems with Applications (37:3) 3/15/, pp 1928-1938.
[3]Chen, P.-I., and Lin, S.-J. 2011. "Word AdHoc Network: Using Google Core Distance to extract the most relevant information," Knowledge-Based Systems (24:3) 4//, pp 393-405.
[4]Garcia Esparza, S., O’Mahony, M. P., and Smyth, B. 2012. "Mining the real-time web: A novel approach to product recommendation," Knowledge-Based Systems (29:0) 5//, pp 3-11.
[5]Jingbo, Z., Wenliang, C., and Tianshun, Y. 2004. "Using seed words to learn to categorize Chinese text," Advances in Natural Language Processing), pp 464-473.
[6]Ko, Y., Park, J., and Seo, J. 2004. "Improving text categorization using the importance of sentences," Information Processing &; Management (40:1) 1//, pp 65-79.
[7]Ko, Y., and Seo, J. 2000. "Automatic text categorization by unsupervised learning," in Proceedings of the 18th conference on Computational linguistics - Volume 1, Association for Computational Linguistics: Saarbr;cken, Germany, pp. 453-459.
[8]Ko, Y., and Seo, J. 2009. "Text classification from unlabeled documents with bootstrapping and feature projection techniques," Information Processing &; Management (45:1) 1//, pp 70-83.
[9]Man, L., Chew Lim, T., Jian, S., and Yue, L. 2009. "Supervised and Traditional Term Weighting Methods for Automatic Text Categorization," Pattern Analysis and Machine Intelligence, IEEE Transactions on (31:4), pp 721-735.
[10]Milne, D., and Witten, I. H. 2013. "An open-source toolkit for mining Wikipedia," Artificial Intelligence (194:0) 1//, pp 222-239.
[11]Ming, L., Calvo, R. A., Aditomo, A., and Pizzato, L. A. 2012. "Using Wikipedia and Conceptual Graph Structures to Generate Questions for Academic Writing Support," Learning Technologies, IEEE Transactions on (5:3), pp 251-263.
[12]Romero, M., Moreo, A., Castro, J. L., and Zurita, J. M. 2012. "Using Wikipedia concepts and frequency in language to extract key terms from support documents," Expert Systems with Applications (39:18) 12/15/, pp 13480-13491.
[13]Salton, G., and Buckley, C. 1988. "Term-weighting approaches in automatic text retrieval," Information Processing &; Management (24:5) //, pp 513-523.
[14]Wan, C. H., Lee, L. H., Rajkumar, R., and Isa, D. 2012. "A hybrid text classification approach with low dependency on parameter by integrating K-nearest neighbor and support vector machine," Expert Systems with Applications (39:15) Nov 1, pp 11880-11888.
[15]Wang, P., Hu, J., Zeng, H. J., and Chen, Z. 2009. "Using Wikipedia knowledge to improve text classification," Knowledge and Information Systems (19:3), pp 265-281.
[16]Xiaojun, Q., Liu, W., and Qiu, B. 2011. "Term Weighting Schemes for Question Categorization," Pattern Analysis and Machine Intelligence, IEEE Transactions on (33:5), pp 1009-1021.
[17]Yang, J., Liu, Y., Zhu, X., Liu, Z., and Zhang, X. 2012. "A new feature selection based on comprehensive measurement both in inter-category and intra-category for text categorization," Information Processing &; Management (48:4) 7//, pp 741-754.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top