跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.2) 您好!臺灣時間:2026/09/21 18:18
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:陳朝龍
研究生(外文):Chao-long Chen
論文名稱:以字詞分群法及摘要為基礎之文件分類法
論文名稱(外文):A Text Categorization Method Based on Term Distributional Clustering and Automatic Summarization
指導教授:蕭文峰蕭文峰引用關係
指導教授(外文):Wen-Feng Hsiao
學位類別:碩士
校院名稱:國立屏東商業技術學院
系所名稱:資訊管理系
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2008
畢業學年度:96
語文別:中文
論文頁數:81
中文關鍵詞:字詞擴展法字詞分群法自動摘要文件分類
外文關鍵詞:Automatic SummarizationText CategorizationTerm Distributional ClusteringTerm Expansion
相關次數:
  • 被引用被引用:2
  • 點閱點閱:371
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:1
近年來由於電子文件量的指數成長,投入自動摘要的研究也相當多。此可由Computational Linguistics的第28卷第4期(2002年),以及Information Processing & Processing的第43卷第6期(2007年)皆是以摘要為專刊主題得到印證。但這些自動摘要除了提供文件搜尋者快速知曉文件內容外少有其它用途,因此本研究提議以文件摘要來進行特徵字選取(term selection),以替代現行之維度縮減法。
另外,為解決類似的概念在不同文件間的用詞不同造成分類上的偏誤,本研究亦探討字詞擴展(term expansion)對分類的影響。相較於過去研究使用字詞分群法(term distributional clustering)來進行特徵字擷取(term extraction),本研究提議以字詞分群法進行字詞擴展,並探討其對分類正確性之影響。
最後,本研究由實驗不同屬性的4個資料集(包括中英文新聞資料集、長度較短的醫學摘要),及不同的分類演算法(KNN、Naive Bayes、SVM)來瞭解所提方法的可行性。
由實驗結果顯示,文件摘要是一個有效的維度縮減法,由摘要所形成的特徵字向量可有效的進行文件分類,其分類正確性優於傳統的TFIDF及Information Gain等維度縮減法。同時,字詞分群法亦能用於維度擴充,進而提高分類的正確性,尤其在當所使用的屬性之數目較少時。
本研究所提方法除可減少特徵向量的維度及選取出較具代表性的特徵字外,更重要的是可節省運算資源,亦即不需為了分類而再次從原文件進行特徵向量之萃取;同時,所產生的指示性(indicative)摘要,可供使用者在瀏覽分類完的資訊時快速知曉其文件之概念。
Owing to the exponential growth of electronic documents, research on automatic summarization is flourishing in the last decade. This is evident from the fact that Computational Linguistics (Vo. 28, No. 4, 2002) and Information Processing & Processing (Vol. 43, No. 6, 2007) both have published a special issue on automatic summarization. However, the summaries generated by these sophisticated methods have little uses except for affording document searchers a glimpse of what the document is about. Therefore, this study proposes to use the automatic summarization for term selection as a way of dimension reduction.
Furthermore, to solve the problem that similar concepts with different term representations might cause the deficiency of classification, we also investigate the effect of term expansion on the classification accuracy. Contrary to using term distributional clustering for feature extraction, we propose use it for expanding the feature terms.
Finally, we compare four data sets with different attributes (including Chinese and English news stories, longer articles like academic research papers, and short articles like medical abstracts), and different classification algorithms (KNN, Naive Bayes, and SVM) to understand the feasibility of the proposed method.
The results show that text summarization is an effective way for dimension reduction. The classification accuracy from the summarization performs better than the traditional TFIDF and Information Gain term weighting schemes. Also term distributional clustering can also be applied to term expansion, and further improve the classification accuracy, especially when the size of feature terms is small.
The proposed method will not only reduce the dimensionality of the term vector and select more representative terms; it can also save the computation resources. That is, one need not redo the feature selection process to cope with the task of text categorization.
Finally, a by-product of our proposed method is that it can generate indicative summaries of those documents. Thus, readers can easily grasp the concepts of those documents by our method when browsing the classification results.
目錄
1 緒論 1
1.1 研究目的 2
1.2 研究流程 3
1.3 研究貢獻 3
2 文獻探討 4
2.1 文件分類 4
2.1.1 預處理 4
2.1.2 產生向量空間模型(Vector Space Model) 5
2.1.3 維度縮減(Dimensionality Reduction, DR) 7
2.1.4 分類演算法 9
2.2 文件摘要 11
2.3 文件摘要方法 12
2.3.1 統計方法 13
2.3.2 語言學方法 17
2.4 字詞擴展法 20
2.4.1 相關詞 20
3 研究方法 22
3.1 以摘要與自動分群法為基礎之分類法 22
3.1.1 預處理 23
3.1.2 建立指示性摘要並且形成特徵向量 23
3.1.3 建構分類器 26
3.1.4 字詞分群 26
3.1.5 字詞擴展 29
3.1.6 文件分類 29
3.2 系統雛型介紹 30
3.2.1 Yahoo!中文新聞擷取 30
3.2.2 中英文剖析 31
3.2.3 建立摘要 35
3.2.4 文件分類 35
3.2.5 分類資訊檢索 36
3.3 文件分類的效能指標 37
4 實驗結果與討論 39
4.1 指示性摘要之結果 39
4.2 實驗 40
4.2.1 實驗一 41
4.2.2 實驗二 50
4.2.3 實驗三 56
4.2.4 實驗四 61
5 結論及未來研究方向 64
6 參考文獻 65
7 附錄 69
7.1 各資料集 69
7.1.1 Retuers21578-90Cat之平均正確性結果 69
7.1.2 Retuers21578-115Cat之平均正確性結果 71
7.1.3 20_newsgroup-20Cat之平均正確性結果 72
7.1.4 Yahoo!中文新聞-13Cat之平均正確性結果 73
7.1.5 ohsumed-23Cat之平均正確性結果 74
7.2 使用套件 74

表目錄
表2 1特徵字選取方法[39] 8
表3 1 混淆矩陣(Confusion Matrix) 38
表4 1中文新聞全文一例 39
表4 2中文新聞10%摘要一例 39
表4 3 英文文件全文一例 40
表4 4 英文文件10%摘要一例 40
表4 5實驗資料集 41
表4 6各實驗之主要目的 41
表4 7Chen and Chen’s IDF與一般的IDF比較結果(10%摘要比例) 42
表4 8 不同K值之平均正確性 43
表4 9 實驗一,不同摘要比例KNN(以文件為單位)之分類績效 43
表4 10實驗一,不同摘要比例KNN(以類別為單位)之分類績效 45
表4 11實驗一,不同摘要比例Naïve Bayes分類器之分類績效 46
表4 12 實驗一,10%摘要比例與維度設定為1000、3000、5000與不設定維度 48
表4 13實驗一,30%摘要比例與維度設定為1000、3000、5000與不設定維度 48
表4 14實驗一,50%摘要比例與維度設定為1000、3000、5000與不設定維度 48
表4 15 不同群數下之時間結果 50
表4 16 實驗二,10%摘要, KNN以文件為單位字詞擴展之比較 51
表4 17實驗二,10%摘要, KNN以類別為單位字詞擴展之比較 51
表4 18 實驗二,10%摘要, Naïve Bayes分類器進行字詞擴展之比較 51
表4 19實驗二,30%摘要,KNN以文件為單位字詞擴展之比較 52
表4 20實驗二,30%摘要,KNN以類別為單位字詞擴展之比較 52
表4 21實驗二,30%摘要,Naïve Bayes分類器進行字詞擴展之比較 52
表4 22實驗二,50%摘要,KNN以文件為單位字詞擴展之比較 53
表4 23實驗二,50%摘要,KNN以類別為單位字詞擴展之比較 53
表4 24實驗二,50%摘要,Naïve Bayes分類器進行字詞擴展之比較 54

圖目錄
圖 1 1 研究流程 3
圖 2 1 文件分類流程圖 4
圖 2 2 向量空間模型示意圖 5
圖 2 3 KNN (K個最鄰近法)示意圖[25] 9
圖 2 4支援向量機分類模型 11
圖 2 5 詞頻與鑑別力的相關圖 [29] 15
圖 2 6 計算句子的重要性 [29] 15
圖 2 7 高層次資料濾淨之架構[37] 16
圖 2 8 詞彙鏈結產生摘要的步驟[11] 17
圖 2 9字詞分佈情形[10] 20
圖 3 1 文件向量產生詳細步驟 23
圖 3 2 文件相似度計算(未考慮相關詞) 27
圖 3 3分割式之資訊理論分群法之演算法 28
圖 3 4文件相似度計算(利用相關詞) 30
圖 3 5新聞擷取 31
圖 3 6中英文剖析頁次 32
圖3 7未斷句剖析前之中文資料 33
圖3 8斷句剖析後之中文資料 33
圖3 9未斷句剖析前之英文資料 34
圖3 10斷句剖析後之英文資料 34
圖 3 11建立摘要 35
圖3 12 系統操作流程 36
圖 3 13文件分類 36
圖 3 14分類資訊檢索 37
圖 4 1實驗一,Yahoo!10類中文新聞,不同摘要比例各方法之正確率比較 48
圖 4 2實驗一,10%摘要比例與維度設定為1000、3000、5000與不設定維度之結果 49
圖 4 3實驗一,30%摘要比例與維度設定為1000、3000、5000與不設定維度之結果 49
圖 4 4實驗一,50%摘要比例與維度設定為1000、3000、5000與不設定維度之結果 49
圖 4 5實驗二,KNN以文件為單位-摘要比例10%之字詞擴展比較 51
圖 4 6實驗二,KNN以類別為單位-摘要比例10%之字詞擴展比較 51
圖 4 7實驗二,Naïve Bayes分類器-摘要比例10%之字詞擴展比較 51
圖 4 8實驗二,KNN以文件為單位-摘要比例30%之字詞擴展比較 52
圖 4 9實驗二,KNN以類別為單位-摘要比例30%之字詞擴展比較 52
圖 4 10實驗二,Naïve Bayes分類器-摘要比例30%之字詞擴展比較 53
圖 4 11實驗二,KNN以文件為單位-摘要比例50%之字詞擴展比較 54
圖 4 12實驗二,KNN以類別為單位-摘要比例50%之字詞擴展比較 54
圖 4 13實驗二,Naïve Bayes分類器-摘要比例50%之字詞擴展比較 54
圖 4 14 實驗二,摘要比例10%應用字詞擴展法,不同維度之比較 55
圖 4 15實驗二,摘要比例30%應用字詞擴展法,不同維度之比較 55
圖 4 16實驗二,摘要比例50%應用字詞擴展法,不同維度之比較 55
圖 4 17 實驗三,Reuters21578-90Cat,10%摘要比例之結果 56
圖 4 18實驗三,Reuters21578-90Cat,30%摘要比例之結果 56
圖 4 19實驗三,Reuters21578-90Cat,50%摘要比例之結果 56
圖 4 20實驗三,Reuters21578-115Cat,10%摘要比例結果 57
圖 4 21實驗三,Reuters21578-115Cat,30%摘要比例結果 57
圖 4 22實驗三,Reuters21578-115Cat,50%摘要比例結果 58
圖 4 23實驗三,20_newsgroup,10%摘要比例結果 58
圖 4 24實驗三,20_newsgroup,30%摘要比例結果 59
圖 4 25實驗三,20_newsgroup,50%摘要比例結果 59
圖 4 26實驗三,Yahoo!13類新聞文件,10%摘要比例 59
圖 4 27實驗三,Yahoo!13類新聞文件,30%摘要比例 60
圖 4 28實驗三,Yahoo!13類新聞文件,50%摘要比例 60
圖 4 29實驗三,ohsumed,摘要全文之實驗結果 61
圖 4 30 實驗四,Yahoo!中文資料集13類(各分類器)之比較結果 62
圖 4 31實驗四,Reuters21578之英文資料集10類(各分類器)之比較結果 63
1.《同義詞詞林(擴展版)》(2005),哈爾濱工業大學資訊實驗室提供。
2.《同義詞詞林(擴展版)》說明文件(2005),哈爾濱工業大學資訊實驗室提供。
3.中文斷詞系統網址(中研院),http:// ckipsvr.iis.sinica.edu.tw/.
4.陳信希(2000),「自動摘要方法之研究:單一中文文本之摘要」,行政院國家科學委員會研究計畫,計劃編號:NSC89-2213-E002-064。
5.劉群、李素建(2002),「基於《知網》的詞彙語義相似度計算」,第三屆漢語詞彙語義學研討會論文集,臺北:,pp. 59-76。
6.蕭文峰、張德民、胡國信,「以遞增式分群為基之分類方法過濾具偏斜類別及概念漂移之垃圾郵件」,資訊管理學報(已接受,2007/12)。
7.蕭文峰、劉凱帆,2006,「以自動摘要為基礎之中文文件分類器」,第十七屆國際資訊管理學術研討會論文集,義守大學,高雄。
8.Aas, K. and Eikvil, L. (1999), “Text Categorization: A Survey,” Technical report, Norwegian Computing Center, Junho
9.Angheluta, R., De Busser, R., & Moens, M.F. (2002). “The use of topic segmentation for automatic summarization,” In U. Hahn & D. Harman (Eds.), Proceedings of the workshop on automatic summarization, Philadelphia, Pennsylvania, USA, pp. 66-70.
10.Baker, L.D. and Mccallum, A.K. (1998), "Distributional clustering of words for text classification,” In Proceedings of the 21th Ann Int ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’98), pp. 96-103.
11.Brunn, M., Chali, Y., Pinchak, C. J. (2001). “Text Summarization Using Lexical Chains,” Workshop on Text Summarization, ACM SIGIR Conference. New Orleans, Louisiana USA.
12.Chen, H.H., Kuo, J.J., Huang, S.J., Lin, C.J., and Wung, H.C. (2003), "A Summarization System for Chinese News from Multiple Sources," Journal of American Society for Information Science and Technology, 54(13), November 2003, pp. 1224-1236.
13.Chen, K.H. (1995), “Topic Identification in Discourse,” In Proceedings of the 7th Conference of the European Chapter of Association for Computational Linguistics, pp. 267-271, Dublin, Ireland
14.Chen, K.H., Chen, H.H. (1995), “A Corpus-Based Approach to Text Partition,” In Proceedings of the Workshop of Recent Advances in Natural Language Processing , pp. 152-161, Sofia, Bulgaria
15.Chen, K.H., Huang, S.J., Lin, W.C., and Chen, H.H. (1998), “An NTU-Approach to Automatic Sentence Extraction for Summary Generation,” In Proceedings of the First Automatic Text Summarization Conference (SUMMAC-1), pp. 163-170, Virginia, May
16.Ferrier, L. (2001), A Maximum Entropy Approach to Text Summarization, School of Artificial Intelligence, Division of Informatics, University of Edinburgh.
17.Goldstein, J., Kantrowitz, M., Mittal, V., and Carbonell, J. (1999), “Summarizing Text Documents: Sentence Selection and Evaluation Metrics,” In Proceedings of ACM-SIGIR'99, Berkeley, CA
18.Hearst, M.A. (1999), “Untangling text data mining,” Proceedings of ACL’99: the 37th annual meeting of the association for computational linguistics, University of Maryland.
19.Hsiao, W.F. and Chang, T.M. (2007), “An Incremental Cluster-based Approach to SPAM Filtering,” Expert Systems with Applications (in press, Available online 28 January 2007)
20.Hand, T.F., Sundheim, B. (1998). "TIPSTER-SUMMAC Summarization Evaluation," Proceedings of the TIPSTER Text Phase III Workshop, Washington DC, USA, pp. 353-340.
21.Inderjit S., D., Subramanyam, M., and Rahul, K. (2003), “A Divisive Information-Theoretic Feature Clustering Algorithm for Text Classification”
22.Joachims, T. (1998), “Text categorization with support vector machines: learning with many relevant features,” In Proceedings of ECML-98, 10th European Conference on Machine Learning (Chemnitz, Germany, 1998), pp. 137–142.
23.Kan, M.Y., Klavans, J. (2002), "Using librarian techniques in automatic text summarization for information retrieval," Proceedings of the 2nd ACM/IEEE-CS joint conference on Digital libraries, pp.36-45.
24.Katz S., M. (1995), “Distribution of content words and phrases in text and language modeling,” Natural Language Engineering, Vol. 2(1), pp. 15–59.
25.K Nearest Neighbor Classifier (2008), http://en.wikipedia.org/wiki/Nearest_neighbor_ (pattern_ recognition)
26.Lang, K. (1995), NEWSWEEDER: learning to filter netnews. In Proceedings of ICML-95, 12th International Conference on Machine Learning (Lake Tahoe, CA, 1995), 331-339.
27.Li, S.J., Zhang, J., Huang, X., Bai, S. and Liu, Q. (2002), “Semantic Computation in a Chinese Question-Answering System,” Journal of Computer Science & Technology(JCST), vol.17, No.6, pp. 933 – 939.
28.Liang, C.Y., Guo, L., Xia, Z.J., Nie, F.G., Li, X.X., Su, L., and Yang, Z.Y. (2006), “Dictionary-based text categorization of chemical web pages,” Information Processing and Management Volume: 42, Issue: 4.
29.Luhn, H.P. (1958), “The automatic creation of literature abstracts,” I.B.M. Journal of Research and Development, 2 (2), pp. 159-165.
30.Lingpipe, http://alias-i.com/lingpipe/index.html
31.Ma, L.P., Shepherd, J. and Zhang, Y.C.h. (2003), “Enhancing text classification using synopses extraction,” In Proceeding of the fourth international conference on web information systems engineering, pp. 115–124.
32.Nakao, Y. (2000), An Alogrithm for One-page Summarization of a Long Text Based on Thematic Hierarchy Detection, Fujitsu Laboratories Ltd.
33.Nigam, K., Mccallum, A.K., Thrun, S., and Mitchell, T. (2000), “Text Classification from Labeled and Unlabeled Documents using EM,” Machine Learning, Vol. 39, pp. 103-134.
34.Noam, S. and Naftali, T. (2001), “Agglomerative Information Bottleneck”
35.Naïve Bayes Classifier, http://nlp.stanford.edu/IR-book/html/htmledition/naive-bayes- text -classification-1.html
36.Porter, M. (1980), “An algorithm for suffix stripping,” Automated Library and Information Systems, Vol. 14, No. 3, pp. 130-137
37.Saravanan, M. and Raman, S. (2002), “The term distribution model for summarization of multiple documents,” In Proc. of the Indo European Conference on Multilingual Communication Technologies (IEMCT 2002), pp. 182–192.
38.Saravanan, M., Reghuraj P., C., and Raman, S. (2003), “Summarization and categorization of text data in high-level data cleaning for information retrieval,” Applied Artificial Intelligence, Vol. 17, pp. 461–474.
39.Sebastiani, F. (2002), “Machine learning in automated text categorization,” ACM Computing Surveys, Vol. 34, No.1, pp.1-47.
40.Seidl, T. and Kriegel, H. (1998), “Optimal Multi-Step k-Nearest neighbor search,” Proceedings of ACM SIGMOD Internet Conference On Management of Data, pp. 154-165.
41.Silla Jr., C. N., Kaestner, C. A. A., and Freitas, A. A. (2003), “A non-linear topic detection method for text summarization using WordNet,” http://citeseer.ist.psu.edu/703670.html.
42.Spärck Jones, K. (1999), “Automatic Summarizing: Factors and Directions,” in I. Mani and M. Maybury (eds.), Advances in Automatic Text Summarization, MIT Press, Cambridge, MA.
43.Spärck Jones, K. (2007), “Automatic summarising: The state of the art,” Information Processing and Management, Vol. 43(6), pp. 1449-1481.
44.Tsay, J.J. and Wang, J.D. (2000), “Design and Evaluation of Approaches to Automatic Chinese Text Categorization,” International Journal of Computational Linguistics & Chinese Language Processing, Vol. 5, No. 2, pp. 43-58.
45.Toutanova, K., Klein, D., Manning, C., and Singer, Y. (2003), “Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network,” Proceedings of Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL 2003), pp. 252-259
46.Wang, Z.W., Wong, S.K.M. and Yao, Y.Y. (1992),”An Analysis of Vector Space Models Based on Computational Geometry,” Proceedings of the 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 152-160.
47.Wei, C., Hu, P., Huang, C.N., Yang, C.S., and Tai, C.H. (2004), “Managing Word Mismatch Challenge in Information Retrieval: A Clustering-Based Query Expansion Method,” Proceedings of the Third Workshop on e-Business (WEB 2004), Washington D.C., pp.82-92.
48.Witten, I.H. and Frank, E. (2005), Data Mining: Practical Machine Learning Tools and Techniques, 2nd Edition, Morgan Kaufmann Series in Data Management Systems.
49.Yang, Y.M. and Liu, X. (1999), "A re-examination of text categorization methods," Proceedings of ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'99), pp. 42-49.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top