跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.174) 您好!臺灣時間:2026/08/09 07:11
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:謝百恩
研究生(外文):Bai-En Shie
論文名稱:有型樣限制的序列型樣探勘演算法
論文名稱(外文):Mining Sequential Patterns with Pattern Constraints
指導教授:顏秀珍顏秀珍引用關係
指導教授(外文):Show-Jane Yen
學位類別:碩士
校院名稱:銘傳大學
系所名稱:資訊工程學系碩士班
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2006
畢業學年度:94
語文別:中文
論文頁數:76
中文關鍵詞:序列型樣常規表示式型樣限制序列型樣探勘資料探勘
外文關鍵詞:Data miningSequential pattern miningPattern constraintRegular expressionSequential pattern
相關次數:
  • 被引用被引用:3
  • 點閱點閱:377
  • 評分評分:
  • 下載下載:28
  • 收藏至我的研究室書目清單書目收藏:1
序列型樣探勘(Mining Sequential Patterns)是要從交易資料庫中找出大部份客戶常常依序購買商品的循序行為,而此行為又稱為序列型樣(Sequential Pattern)。以往已有很多論文提出找尋序列型樣的方法,然而使用者可能只需要一些包含特定項目或行為的序列型樣。若用以往的方法找出所有的序列型樣後,再額外篩選出使用者感興趣的部分,將會浪費很多時間。若在探勘前就能先讓使用者自行設定其所感興趣的型樣,並在探勘過程中應用使用者自定的內容直接找出相關的結果,則找出來的資訊不但可以完全符合使用者的需求,也可省去很多找尋使用者不感興趣的資訊的時間。若使用者限制序列型樣探勘結果所找出的序列型樣必須包含其所指定的型樣,則我們將此限制稱為型樣限制(Pattern Constraint)。本論文提出有效率的演算法,從交易資料庫中直接找出符合型樣限制的所有序列型樣。在實驗的章節中,我們分別使用真實資料集與虛擬資料集將本論文提出的方法與SPIRIT(R)演算法及Bit-String演算法比較演算法的執行時間與使用的記憶體空間。實驗結果證明本論文提出的演算法雖然使用的記憶體空間較SPIRIT(R)演算法為多,但執行速度較快;也證明無論在執行速度或使用記憶體空間上,本論文所提出的演算法皆較Bit-String演算法為佳。
Sequential pattern mining is to find sequential behaviors which most customers frequently do in a transaction database. These behaviors are called sequential patterns. There were many papers proposed algorithms for finding all sequential patterns. However, there is a new problem: users may only need some special sequential patterns, for example, the sequential patterns which include certain items or behaviors. If we let users set the items or patterns which they are interested in before mining process, we will save much execution time and the sequential patterns we found can fit the users'' need. The items or patterns which are preset by users are "pattern constraints." We propose an effective algorithm to find all sequential patterns which fit the constraint from the transaction database. In the experimental results cheaper, we use real dataset and synthetic dataset to compare our algorithm with SPIRIT(R) algorithm and Bit-String algorithm. The results show that although our algorithm used more memory than SPIRIT(R) algorithm during the mining process, our algorithm was faster than SPIRIT(R) algorithm. The results also show that our method not only used less memory space than Bit-String algorithm but also was faster than Bit-String algorithm.
摘 要  I
ABSTRACT II
致 謝  III
目 錄  IV
表目錄  V
圖目錄  VI
第壹章 導論        1
第貳章 相關工作      7
第參章 有型樣限制的序列型樣探勘演算法  15
 第一節 型樣限制的格式  15
 第二節 我們的演算法   16
 第三節 範例       29
第肆章 實驗結果      34
第伍章 結論與未來工作   66
第陸章 參考文獻      67
1. R. Agrawal and R. Srikant. (1994), Fast Algorithms for Mining Association Rules, in Proceedings of the 20th International Conference on Very Large Data Bases, September 1994, pp. 487-499.
2. R. Agrawal and R. Srikant. (1995), Mining Sequential Patterns, in Proceedings of the 11th International Conference on Data Engineering, Match 1995, pp. 3-14.
3. H. Albert-Lorincz, J-F. Boulicaut, (2003), A Framework for Frequent Sequence Mining under Generalized Regular Expression Constraints. KDID 2003, pp: 2-16.
4. H. Albert-Lorincz, J-F. Boulicaut, (2003) Mining Frequent Sequential Patterns under Regular Expressions: A Highly Adaptive Strategy for Pushing Constraints. SDM 2003.
5. M. N. Garofalakis, R. Rastogi, and K. Shim. (1999), SPIRIT: Sequential Pattern Mining with Regular Expression Constraints, in Proceedings of the 25th International Conference on Very Large Data Bases, September 1999, pp. 223-234.
6. M. N. Garofalakis, R. Rastogi, and K. Shim. (2002), Mining Sequential Patterns with Regular Expression Constraints, IEEE Transactions on Knowledge and Data Engineering, Vol. 14, No.3, May/June 2002. (TKDE 2002).
7. J. Han, J. Pei, B. Mortazavi-Asl, Q. Chen, U. Dayal, and M. C. Hsu. (2000), FreeSpan: Frequent Pattern-Projected Sequential Pattern Mining, in Proceedings of ACM SIGKDD International Conference on Knowledge Discovery in Databases, August 2000, pp. 355-359. (KDD’00)
8. J. Han, J. Pei, and Y. Yin. (2000), Mining Frequent Patterns without Candidate Generation, in Proceedings of the ACM-SIGMOD International Conference on Management of Data, May 2000, pp. 1-12. (SIGMOD’00).
9. R. Kohavi, C. Brodley, B. Frasca, L. Mason, and Z. Zheng (2000), KDD-Cup 2000 Organizers Report: Peeling the Onion, Proc. SIGKDD Explorations, vol. 2, pp. 86-98, 2000.
10. J. S. Park, Ming-Syan Chen, Philip S. Yu (1995), An Effective Hash Based Algorithm for Mining Association Rules, SIGMOD Conference 1995, pp. 175-186.
11. J. Pei, J. Han (2002), Constrained Frequent Pattern Mining: A Pattern-Growth View, SIGKDD Explorations, 4(1): 31-39, 2002.
12. J. Pei, J. Han, H. Lu, S. Nishio, S. Tang, D. Yang (2001), H-Mine: Hyper-Structure Mining of Frequent Patterns in Large Databases, ICDM 2001, pp. 441-448
13. J. Pei, J. Han, R. Mao. (2000), Closet: An efficient algorithm for mining frequent closed itemsets, in SIGMOD Int’l Workshop on Data Mining and Knowledge Discovery, May 2000.
14. J. Pei, J. Han, B. Mortazavi-Asl, H. Pinto, Q. Chen, U. Dayal, M. C. Hsu. (2001), PrefixSpan: Mining Sequential Patterns Efficiently by Prefix-Projected Pattern Growth, in Proceedings of the 17th International Conference on Data Engineering, April, 2001, pp.215-224.
15. J. Pei, J. Han, B. Mortazavi-Asl, H. Pinto, Q. Chen, U. Dayal, M. C. Hsu. (2004) Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach, IEEE Transactions on Knowledge and Data Engineering, Vol.16, No.10, October 2004.
16. J. Pei, J. Han, W. Wang (2002), Mining Sequential Patterns with Constraints in Large Databases, in CIKM’02, November 4-9, 2002.
17. A. Savasere, E. Omiecinski, S. B. Navathe, (1995), An Efficient Algorithm for Mining Association Rules in Large Databases. VLDB 1995, pp: 432-444.
18. Show-Jane Yen, Yue-Shi Lee and Bai-En Shie, "Mining Sequential Patterns with Pattern Constraints from Transaction Databases," Proceedings of 11th Conference on Information Management and Practice (IMP''2005), pp. 584-598, December 10, 2005.
19. R. Srikant, R. Agrawal. (1996), Mining Sequential Patterns: Generalizations and Performance Improvements, in Proceedings of the 5th International Conference on Extending Database Technology, March 1996, pp. 3-17, (EDBT’96).
20. R. Srikant, Q. Vu, and R. Agrawal. (1997), Mining Association Rules with Item Constraints, in Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining, August 1997, pages 67-73, (KDD’97).
21. Show-Jane Yen (2005), Mining Interesting Sequential Patterns for Intelligent Systems, International Journal of Intelligent Systems, Vol. 20, Issue 1, Jan. 2005, pp. 73-87
22. M. J. Zaki. (2001), SPADE: An Efficient Algorithm for Mining Frequent Sequences, in Machine Learning Journal, special issue on Unsupervised Learning, Vol. 42 Nos. 1/2, Jan/Feb 2001, pp. 31-60.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top