跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.79) 您好!臺灣時間:2026/09/02 15:56
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:謝英琳
研究生(外文):Ying-Ling Hsieh
論文名稱:以約略集合理論為基礎的減量式替代規則擷取演算法
論文名稱(外文):A Decremental Alternative Rule Extract Algorithm Based on the Rough Set Theory
指導教授:黃俊哲黃俊哲引用關係
指導教授(外文):Chun-Che Huang
口試委員:柳永青梁文耀
口試日期:2011-07-15
學位類別:碩士
校院名稱:國立暨南國際大學
系所名稱:資訊管理學系
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2011
畢業學年度:99
語文別:英文
論文頁數:112
中文關鍵詞:資料探勘動態資料庫約略集合理論減量式演算法規則歸納法
外文關鍵詞:Data miningDynamic DatabaseRough Set TheoryDecremental AlgorithmRule Induction
相關次數:
  • 被引用被引用:0
  • 點閱點閱:417
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
近年來資料探勘已發展成為在知識管理領域中的一項重要研究,主要用以挖掘隱含在龐大的資料庫中不易觀察且有意義的知識。而動態資料庫是目前常見的一種資料庫型態,在資料庫的管理上,經常需要使用到資料刪除的操作。然而,在現有的資料探勘演算法中,探勘的目標資料庫大多假設為靜態資料庫,因此在每次更新完資料庫後,若要獲得更新後的資料庫規則,則需要重新完整的掃描新資料庫後,才可進行資料探勘和規則擷取,而減量式演算法是被用來解決此問題的技術。
在資料探勘中,約略集合理論 (RST, Rough Set Theory) 被視為適合用來協助處理質性資料的方法,能從資料中發現潛在的重要事實,然而過去的文獻中指出傳統的約略集合理論卻無法產生包含優先順序的規則,且無法確保決策表的分類是可信的,易造成規則不集中且可信度較低的結果。因此,學者Bill Tseng於2008年提出了替代性規則歸納演算法 (AREA, Alternative Rule Extraction Algorithm),以偏好權重及強度指數SI (Strength Index) 為基礎來解決上述所提到的相關問題,並同時可處理當產生之規則有同等價值時,以替代性做法保留規則。因此,本研究以AREA為基礎發展了減量式替代性規則歸納演算法 (DAREA, Decremental Alternative Rule Extract Algorithm) 來處理當資料庫在進行刪除資料後,規則重新擷取的情況。不需重新對更新後的資料庫做完整的掃描,便可有效率的進行規則歸納。最後,本研究以自行開發的應用程式來驗證其減量式演算法效率將優於傳統演算法。
Data mining that explores useful information and helpful knowledge from large databases has evolved into an important research area in recent years. A dynamic database is a common type in many business databases, and the data deletion operations that are also frequently used in database management activities. Unfortunately, most existing data-mining algorithms assume that the database is static, and updating a database requires to re-compute all the patterns by scanning the updated database after deleting data for data mining and rule extraction. The decremental technique is a way to solve the issue of removed-out object without re-implementing the DM algorithm in a dynamic database. One of the promising approaches of DM to knowledge discovery, pattern recognition, decision analysis, and etc. is the Rough Set Theory (RST). The RST is a knowledge discovery tool that can be used to help induce logical patterns hidden in massive data.
However, previous RS approaches cannot produce rules containing preference order, namely, cannot achieve to generate more meaningful and general rules. Also, induction based on RS often generates too many rules without focuses and cannot guarantee that the classification of a decision table is credible. Tseng (2008) proposed the AREA (Alternative Rule-Extraction Algorithm) to solve above mentioned problems with discovering preference-based rules according to the reducts with the maximum of strength index (SI), specifially the case that the desired reducts are not necessarily unique since several reducts could include the same value of SI. Thus, in this study based on AREA, DAREA (Decremental Alternative Rule Extract Algorithm) is proposed to solve issue of removed-out objects from database. The algorithm is unnecessary to re-compute rule sets from the very beginning that can quickly generate and complete rules. The experiments are made to validate the proposed approach to be superior to the traditional RS approach.
LIST OF CONTENTS
誌 謝 i
摘 要 ii
Abstract iii
1. Introduction 1
1.1 Motivation and Background 1
1.2 Objectives 4
1.3 Research Rationale 4
1.4 Assumption and Limitation 5
1.5 Research Organization 6
2. Literature Review 6
2.1 Approximation concepts in the rough set theory 7
2.2 Rough set based rule generation 10
2.2.1 Attribute reduction process 10
2.2.2 Rule induction process 12
2.3 The decremental techniques 17
2.4 Observation from literature review 19
3. Solution Approach 20
3.1 Structure of the proposed decremental rough set based rule induction 20
3.2 Lemma of the decremental algorithm 22
3.3 Procedure of decremental algorithm 25
3.4 Illustrative examples for the cases 31
3.4.1 Case 1: No rule Generated; No rule Replaced; No rule Deleted 31
3.4.2 Case 2: No rule Generated; No rule Replaced; Rule Deleted 35
3.4.3 Case 3: No rule Generated; Rule Replaced; No rule Deleted 39
3.4.4 Case 4: Rule Generated; No rule Replaced; No rule Deleted 43
3.4.5 Case 5: Hybrid case 47
3.4.5.1 Case 5-1: Rule Generated; Rule Replaced; No rule Deleted 47
3.4.5.2 Case 5-2: Rule Generated; No rule Replaced; Rule Deleted 51
3.4.5.3 Case 5-3: No rule Generated; Rule Replaced; Rule Deleted 55
3.4.5.4 Case 5-4: Rule Generated ; Rule Replaced ; Rule Deleted 59
3.5 The time complexity of the proposed decremental algorithm 63
4. Experiments 64
4.1. Experiments for decrement for algorithm AREA 65
4.2. Experiments for decrement for algorithm DAREA 67
4.3. Comparison with AREA and DAREA 69
5. Conclusions 71
References 74
Appendix A 82
Appendix B 92
Appendix C 95
Appendix D 98
Appendix E 101
Appendix F 104
Appendix G 107
Appendix H 110


LIST OF FIGURE
Figure 3-1. The procedure of the proposed decremental algorithm. 26
Figure 3-2. The procedure of Case-1 32
Figure 3-3. The procedure of Case-2 36
Figure 3-4. The procedure of Case-3 40
Figure 3-5. The procedure of Case-4 44
Figure 3-6. The procedure of Case 5-1 48
Figure 3-7. The procedure of Case 5-2 52
Figure 3-8. The procedure of Case-5-3 56
Figure 3-9. The procedure of Case 5-4 60
Figure 4-1. Show the system re-extracts the AREA rules 66
Figure 4-2. Show the system re-extracts the DAREA rules 68
Figure 4-3. Execution time of CPU against AREA and DAREA 69
Figure 4-4. Execution time of CPU ratio against AREA and DAREA 70
Figure 4-5. Execution time of CPU ratio against AREA and DAREA 71


LIST OF TABLES
Table 3-1. The structure of case of the decremental algorithm 21
Table 3-2. The original information 31
Table 3-3. The difference set of table 33
Table 3-4. The reduct set of table 33
Table 3-5. The final decision rules and alternative rules 34
Table 3-6. The resulting concise rules 34
Table 3-7. The original information 35
Table 3-8. The difference set of table 37
Table 3-9. The reduct set of table 38
Table 3-10. The final decision rules and alternative rules 38
Table 3-11. The resulting concise rules 38
Table 3-12. The original information 39
Table 3-13. The difference set of table 41
Table 3-14. The reduct set of table 41
Table 3-15. The final decision rules and alternative rules 42
Table 3-16. The resulting concise rules 42
Table 3-17. The original information 43
Table 3-18. The difference set of table 45
Table 3-19. The reduct set of table 45
Table 3-20. The final decision rules and alternative rules 46
Table 3-21. The resulting concise rules 46
Table 3-22. The original information 47
Table 3-23. The difference set of table 49
Table 3-24. The reduct set of table 49
Table 3-25. The final decision rules and alternative rules 50
Table 3-26. The resulting concise rules 50
Table 3-27. The original information 51
Table 3-28. The difference set of table 53
Table 3-29. The reduct set of table 53
Table 3-30. The final decision rules and alternative rules 54
Table 3-31. The resulting concise rules 54
Table 3-32. The original information 55
Table 3-33. The difference set of table 57
Table 3-34. The reduct set of table 57
Table 3-35. The final decision rules and alternative rules 58
Table 3-36. The resulting concise rules 58
Table 3-37. The original information 59
Table 3-38. The difference set of table 61
Table 3-39. The reduct set of table 61
Table 3-40. The final decision rules and alternative rules 62
Table 3-41. The resulting concise rules 62
Table 3-42. The complexity of this proposed algorithm is composed of five cases 63
Table 3-43. The complexity of the proposed decremental algorithm 64
Table 4-1. Using AREA to re-extract rules (with attributes=5) 67
Table 4-2. Using DAREA to re-extract rules (with attributes=5) 68
Ahn, B.S., Cho, S.S. and Kim, C.Y., 2000, "The integrated methodology of rough set theory and artificial neural network for business failure prediction," Expert Systems with Applications, Vol. 18, No. 2, pp. 65–74.
Asharaf, S., Murty, M. Narasimha and Shevade, S.K., 2006, “Rough set based incremental clustering of interval data,” Pattern Recognition Letters, Vol. 27, No. 6, pp. 515-519.
Aumann, Yonatan, Feldman, Ronen, Lipshtat, Orly and Manilla, Heikki, 1999, “Borders: An Efficient Algorithm for Association Generation in Dynamic Databases,” Journal of Intelligent Information Systems, Vol. 12, No. 1, pp.61-73.
Blaszczynski, Jerzy and Słowiński, Roman, 2003, “Incremental Induction of Decision Rules from Dominance-based Rough Approximations,” Electronic Notes in Theoretical Computer Science, Vol. 82, No. 4, pp.40-51.
Breault, Joseph L., 2001, “Data mining diabetic databases: are rough sets a useful addition?,” Proceedings of the Computing Science and Statistics, Vol. 33.
Cheng, Ching-Hsue, Chen, Tai-Liang and Wei, Liang-Ying, 2010, “A hybrid model based on rough sets theory and genetic algorithms for stock price forecasting,” Information Sciences, Vol. 180, No. 9, pp.1610-1629.
Choi, Hyun-Seon, Kim, Ji-Su and Lee, Dong-Ho, 2011, “Real-time scheduling for reentrant hybrid flow shops: A decision tree based mechanism and its application to a TFT-LCD line,” Expert Systems with Applications, Vol. 38, No. 4, pp. 3514-3521.
Chou, Hsin-Chuan, Cheng, Ching-Hsue and Chang, Jing-Rong, 2007, “Extracting drug utilization knowledge using self-organizing map and rough set theory,” Expert Systems with Applications, Vol. 33, No. 2, pp.499-508.
Chu, Xue-Zheng, Gao, Liang, Qiu, Hao-Bo, Li, Wei-Dong and Shao, Xin-Yu, 2009, “An expert system using rough sets theory for aided conceptual design of ship’s engine room automation,” Expert Systems with Applications, Vol. 36, No. 2, pp.3223-3233.
Crespo, Fernando and Weber, Richard, 2005, “A methodology for dynamic data mining based on fuzzy clustering,” Fuzzy Sets and Systems, Vol. 150, pp.267-284.
Dong, Haiying, Zhang, Yubo and Xue, Junyi, 2002, “Hiearchical fault diagnosis for substation based on rough set,” Proceedings of the Power System Technology, Vol. 4, pp.2318–2321.
Doumpos, M., Marinakis, Y., Marinaki, M. and Zopounidis, C., 2009, “An evolutionary approach to construction of outranking models for multicriteria classification: The case of the ELECTRE TRI method,” European Journal of Operational Research, Vol. 199, No. 2, pp. 496-505.
Fan, Yu-Neng, Tseng, Tzu-Liang (Bill), Chern, Ching-Chin and Huang, Chun-Che, 2009, "Rule induction based on an incremental rough set," Expert Systems with Applications , Vol. 36, No. 9, pp.11439-11450.
Gaudreault, Jonathan, Frayret, Jean-Marc and Pesant, Gilles, 2009, “Distributed search for supply chain coordination,” Computers in Industry, Vol. 60, No. 6, pp.441-451.
Goh, Carey and Law, Rob, 2003, "Incorporating the rough sets theory into travel demand analysis," Tourism Management, Vol. 24, No. 5, pp.511-517.
Gorsevski, Pece V. and Jankowski, Piotr, 2008, “Discerning landslide susceptibility using rough sets,” Computers, Environment and Urban Systems, Vol. 32, No. 1, pp.53-65.
Greco, Salvatore, Matarazzo, Benedetto and Slowinski, Roman, 2001, "Rough sets theory for multicriteria decision analysis," European Journal of Operational Research, vol. 129, No. 1, pp.1-47.
Guan, Lihe, 2009, "An incremental updating algorithm of attribute reduction set in decision tables," Proceedings of 6th International Conference on Fuzzy Systems and Knowledge Discovery, Tianjin, China, pp.421-425.
Hassanien, Aboul-Ella, 2004, “Rough set approach for attribute reduction and rule generation: a case of patients with suspected breast cancer,” Journal of the American Society for Information Science and Technology, Vol. 55, No. 11, pp. 954–962.
Hong, Tzung-Pei, Wang, Ching-Yao and Tseng, Shian-Shyong, 2011, “An incremental mining algorithm for maintaining sequential patterns using pre-large sequences ,” Expert Systems with Applications, Vol. 38, No. 6, pp.7051-7058.
Huang, Chun-Che, Tseng, Tzu-Liang (Bill), Chuang, Horng-Fu and Liang, Hui-Fen, 2006/06, “Data mining special No.: A rough set based approach to manufacturing process document retrieval,” International Journal of Production Research, Vol. 44, No. 14, pp.2889-2911.
Jensen, Richard and Shen, Qiang, 2004, “Fuzzy–rough attribute reduction with application to web categorization,” Fuzzy Sets and Systems, Vol. 141, No. 3, pp.469-485.
Jiang, Jian-jun, Zhang, Li, Wang, Yi-qun, Zhang, Kun, Yang, Da-Xin and He, Wen, 2011, “Association rules analysis of human factor events based on statistics method in digital nuclear power plant,” Safety Science, Vol. 49, No. 6, pp.946-950.
Kashani, Moein Navvab and Shahhosseini, Shahrokh, 2010, “A methodology for modeling batch reactors using generalized dynamic neural networks,” Chemical Engineering Journal, Vol. 159, No. 1-3, pp.195-202.
Kusiak, Andrew, 2001, “Feature transformation methods in data mining”, IEEE Transaction on Electronics Packaging Manufacturing, Vol. 24, No.3, pp.214-221.
Li, Tianrui, Ruan, Da, Geert, Wets, Song, Jing and Xu, Yang, 2007, “A rough sets based characteristic relation approach for dynamic attribute generalization in data mining,” Knowledge-Based Systems, Vol. 20, No. 5, pp.485-494.
Li, Peng, Wang, Xiao-long and Guan, Yi, 2008, "Question classification with incremental rule learning algorithm based on rough set," Journal of Electronics & Information Technology, Vol. 30, No. 5, pp.1127-1130.
Liang, Wen-Yau and Huang, Chun-Che, 2006, “Agent-based demand forecast in multi-echelon supply chain,” Decision Support Systems, Vol. 42, No. 1, pp. 390-407.
Lin, Chiun-Sin, Tzeng, Gwo-Hshiung and Chin, Yang-Chieh, 2011, “Combined rough set theory and flow network graph to predict customer churn in credit card accounts,” Expert Systems with Applications, Vol. 38, No. 1, pp.8-15.
Lingras, Pawan, Hogo, Mofreh, Snorek, Miroslav and West, Chad, 2005, “Temporal analysis of clusters of supermarket customers: conventional versus interval set approach,” Information Sciences, Vol. 172, No. 1-2, pp.215-240.
Liu, Yong, Xu, Congfu and Pan, Yunhe, 2004, “An Incremental Rule Extracting Algorithm Based on Pawlak Reduction,” Proceedings of IEEE International Conference on Systems, Man and Cybernetics, Vol. 6, pp. 5964-5968.
Liu, Min, Shao, Mingwen, Zhang, Wenxiu and Wu, Cheng, 2007, “Reduction method for concept lattices based on rough set theory and its application,” Computers & Mathematics with Applications, Vol. 53, No. 9, pp.1390-1410.
Otey, Matthew Eric, Wang, Chao, Parthasarathy, Srinivasan, Veloso, Adriano and Meira, Wagner, 2003, “Mining Frequent Itemsets in Distributed and Dynamic Databases,” IEEE International Conference on Data Mining, Melbourne, Florida, pp.617-620.
Pattaraintakorn, Puntip and Cercone, Nick, 2008, "Integrating rough set theory and medical applications," Applied Mathematics Letters, Vol. 21, No. 4, pp.400-403.
Pawlak, Zdzisław, 1982, “Rough Sets,” International Journal of Computer and Information Sciences, Vol. 11, No 5, pp. 341-356.
Pawlak, Zdzisław, 1991, Rough Sets: Theoretical Aspects of Reasoning about Data, Kluwer Academic Publishers, Boston.
Petitjean, Francois, Ketterlin, Alain and Gancarski, Pierre, 2011, “A global averaging method for dynamic time warping, with applications to clustering,” Pattern Recognition, Vol. 44, No. 3, pp.678-693.
Phuong, Nguyen Hoang, Phong, Le Linh, Santiprabhob, P. and Baets, B. De, 2001, “Approach to generating rules for expert systems using rough set theory,” Proceedings of IEEE International Conference on IFSA World Congress and 20th NAFIPS, Vol. 2, Vancouver, BC , Canada, pp.877–882.
Questier, F., Rollier, I. A., Walczak, B. and Massart, D.L., 2002, “Application of rough set theory to feature selection for unsupervised clustering,” Chemometrics and Intelligent Laboratory Systems, Vol. 63, No. 2, pp.155–167.
Shen, Qiang and Jensen, Richard, 2004, “Selecting informative features with fuzzy-rough sets and its application for complex systems monitoring,” Pattern Recognition, Vol. 37, No.7, pp.1351–1363.
Shan, Ning and Ziarko, Wojciech, 2007, "Data-based acquisition and incremental modification of classification rules," Computational Intelligence, Vol. 11, No.2, pp.357-370.
Shyng, Jhieh-Yu, Wang, Fang-Kuo, Tzeng, Gwo-Hshiung and Wu, Kun-Shan, 2007, "Rough set theory in analyzing the attributes of combination values for the insurance market," Expert Systems with Applications, Vol. 32, No. 1, pp.56-64.
Shyng, Jhieh-Yu, Shieh, How-Ming, Tzeng, Gwo-Hshiun and Hsieh, Shu-Huei, 2010, “Using FSBT technique with Rough Set Theory for personal investment portfolio analysis,” European Journal of Operational Research, Vol. 201, No. 2, pp. 601-607.
Shyng, Jhieh-Yu, Shieh, How-Ming and Tzeng, Gwo-Hshiung, 2011, “Compactness rate as a rule selection index based on Rough Set Theory to improve data analysis for personal investment portfolios,” Applied Soft Computing, Vol. 11, No. 4, pp.3671-3679.
Sohel, Ferdous Ahmed and Rahman, Chowdhury Mofizur, 2003, “Association Rule Mining in Dynamic Database Using the Concept of Border sets,” Asian Journal of Information Technology, Vol. 3, No. 7, pp.508-515.
Sohel, Ferdous Ahmed and Rahman, Chowdhury Mofizur, 2004, “Association Rule Mining in Dynamic Database Using the Concept of Border sets,” Asian Journal of Information Technology, Vol. 3, No. 7 pp.508-515.
Sun, Cheng-Min, Liu, Da-You, Sun, Shu-Yang, Li, Jia-Fei and Zhang, Zhao-Hui, 2005, “Containing order rough set methodology,” Proceedings of 2005 International Conference on Machine Learning and Cybernetics, Vol. 3, Guangzhou, China, pp.1722–1727.
Swiniarski, Roman W. and Skowron, Andrzej, 2003, “Rough set methods in feature selection and recognition,” Pattern Recognition Letters, Vol. 24, No. 6, pp.833–849.
Tan, Raymond R., 2005, "Rule-based life cycle impact assessment using modified rough set induction methodology," Environmental Modelling & Software, Vol. 20, No. 5, pp.509-513.
Tasoulis, Dimitris K. and Vrahatis, M.N., 2005, “Unsupervised clustering on dynamic databases,” Pattern Recognition Letters, Vol. 26, No. 13, pp.2116-2127.
Thangavel, K. and Pethalakshmi, A., 2009, “Dimensionality reduction based on rough set theory: A review Review Article,” Applied Soft Computing, Vol. 9, No. 1, pp.1-12.
Tseng, Tzu-Liang (Bill), 1999, “Quantitative Approaches for Information Modeling,” Ph.D. Dissertation, University of Iowa.
Tseng, Tzu-Liang (Bill), Huang, Chun-Che, Jiang, Fuhua and Ho, Johnny C., 2006, “Applying a hybrid data mining approach to prediction problems: A case of preferred suppliers prediction,” International Journal of Production Research, Vol. 44, No. 14, pp.2935-2954.
Tseng, Tzu-Liang (Bill) and Huang, Chun-Che, 2007/08, "Rough set based approach to feature selection in customer relationship management," The International Journal of Management Science, OMEGA journal, Vol. 35, No. 4, pp.365-383.
Tseng, Tzu-Liang (Bill), Huang, Chun-Che and Ho, Johnny C., 2008, "Autonomous Decision Making in Customer Relationship Management: A Data Mining Approach," Proceeding of the Industrial Engineering Research 2008 Conference, Vancouver, British Columbia, Canada.
Wang, Qing Hui and Li, Jing Rong, 2004, “A rough set-based fault ranking prototype system for fault diagnosis,” Engineering Applications of Artificial Intelligence, Vol. 17, No. 8, pp.909-917.
Wu, Wei-Zhi, Mi, Ju-Sheng and Zhang, Wen-Xiu, 2003, "Generalized fuzzy rough sets," Information Sciences, Vol. 151, pp.263-282.
Xiao, Zhi, Chen, Ling and Zhong, Bo, 2010, “A model based on rough set theory combined with algebraic structure and its application: Bridges maintenance management evaluation,” Expert Systems with Applications, Vol. 37, No. 7, pp.5295-5299.
Yin, Xuri, Zhou, Zhihua, Li, Ning and Chen, Shifu, 2001, “An Approach for Data Filtering Based on Rough Set Theory,” Lecture Notes in Computer Science, Vol. 2118/2001, pp. 367-374.
Zabłocka-Malicka, Monika, Ciechanowski, Bartłomiej, Szczepaniak, Włodzimierz and Gaweł, Wiesław, 2008, “Internal cation mobility in molten LiCl–NdCl3 system,” Electrochimica Acta, Vol. 53, No. 5, pp.2081-2086.
Zhong, Ning, Dong, Ju-Zhen, Ohsuga, Setsuo and Lin, Tsau Young, 1998, "An incremental, probabilistic rough set approach to rule discovery," IEEE International Conference on Fuzzy Systems, Vol. 2, Anchorage, AK , USA, pp.933-938.
Zhang, Shichao and Liu, Li, 2003, “Mining dynamic databases by weighting,” Acta Cybernetica, Vol. 16, No. 1.
Zhang, Shichao, Zhang, Jilian and Zhang, Chengqi, 2007, “EDUA: An efficient algorithm for dynamic database mining,” Information Sciences, Vol. 177, No. 13, pp.2756-2767.
Zhang, Shichao, Zhang, Jilian and Jin, Zhi, 2009, “A decremental algorithm of frequent itemset maintenance for mining updated databases ,” Expert Systems with Applications, Vol. 36, No. 8, pp.10890-10895.
Ziarko, Wojciech P. and Van Rijsbergen, C. J., 1994, Rough Sets, Fuzzy Sets and Knowledge Discovery, Springer-Verlag, New York.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top