跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.73) 您好!臺灣時間:2026/07/23 00:33
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:李立雅
研究生(外文):Li-Ya Li
論文名稱:跨序列關聯性規則資料探勘
論文名稱(外文):Inter-sequence Association Rules Mining
指導教授:李瑞庭李瑞庭引用關係
指導教授(外文):Anthony J. T. Lee
學位類別:碩士
校院名稱:國立臺灣大學
系所名稱:資訊管理研究所
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2003
畢業學年度:91
語文別:英文
論文頁數:51
中文關鍵詞:序列資料探勘跨交易關聯性
外文關鍵詞:inter-transactionassociation rule miningsequence
相關次數:
  • 被引用被引用:0
  • 點閱點閱:237
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:1
序列樣式分析已經應用在許多領域上,而且也已經有許多尋找序列樣式的研究。但是目前的研究都把各序列的出現視為獨立事件,而沒有探討到序列間的關係,因此這篇論文將會討論尋找跨序列關聯性規則。
尋找跨序列關聯性規則中最花時間的地方在於尋找跨序列樣式,因此我們設計了一個演算法來尋找跨序列樣式。第一步,我們會利用既有的演算法找出序列樣式,然後第二步才去尋找跨序列樣式。在第一步中,我們會針對每個序列樣式去紀錄它出現的時點,因此每個序列樣式都有一個相對應的時間列表,顯示出該序列樣式出現的時點。我們會對這些時間列表做好分類,以方便之後的查詢。在第二步中我們會由短到長,一層一層產生可能的跨序列樣式。透過第一步中所紀錄的時點,我們可以有效率的找到需要的時間列表,並且透過時間列表計算出某個序列樣式所出現的次數是否夠多。一但產生一個可能的跨序列樣式,我們就可以立刻算它出現的次數。
透過時間列表以及對於時間列表的分類,這個演算法會比Apriori-like的方式尋找跨序列樣式來的快很多,因為它在計算跨序列樣式出現的次數上較有效率。實驗結果也顯示出這個演算法很有效率。

There are many algorithms proposed to find sequential patterns in sequence databases where each transaction contains one sequence. Previously proposed algorithms treat each sequence as an independent one. This kind of mining belongs to intra-transaction sequential patterns mining.
In this paper, we propose an algorithm, ProbSif, to mine inter-sequence association rules. Our proposed algorithm consists of three phases. First, we find all large intra-sequence patterns. For each large pattern found, all the time points at which the pattern occurs are recorded in a time point list. Second, those time point lists are hashed into L-buckets. Third, we use a level-wise candidate generation-and-test method to generate candidate patterns across different sequences and check if a candidate is large. Once we generate a candidate, we count its support by reading relevant time point lists from L-buckets. By using the L-buckets, our proposed algorithm requires fewer database scans than the Apriori-like approach. Therefore, our proposed algorithm is more efficient. The experimental results show that our proposed algorithm outperforms the Apriori-like approach by several orders of magnitude.

Table of Contents i
List of Figures iii
List of Tables iv
Chapter 1 Introduction 1
Chapter 2 Literature Survey 4
2.1 Association Rules Mining 4
2.1.1 The Overview of the Apriori Algorithm 4
2.1.2 Candidate Generation in the Apriori Algorithm 5
2.1.3 Counting Supports of Candidates in the Apriori Algorithm 5
2.1.4 Buffer Management in Apriori 6
2.2 Inter-transaction Association Rules Mining 6
2.2.1 Inter-transaction Association Rules 7
2.2.2 Large Extended Itemsets Discovery Phases 8
2.2.2.1 Candidate Generation 8
2.2.2.2 Counting Supports of Candidates 9
2.3 Sequential Patterns Mining 10
2.3.1 PrefixSpan Steps 11
2.3.2 Optimization in Projected Databases 12
2.3.3 Variations of the PrefixSpan Algorithm 13
2.3.3.1 PrefixSpan with Bi-level Projection 13
2.3.3.2 PrefixSpan with Pseudo Projection 14
2.4 Mining Confident Rules without Support Requirement 14
2.4.1 Mining Confident Rules 15
2.4.2 The Disk-based Implementation 16
2.4.3 Counting Confidences of Candidates 17
2.5 Discussion 17
Chapter 3 Mining Inter-sequence Association Rules 19
3.1 Finding Large Extended Sequence Sets 21
3.1.1 Function: PrefixSpan 22
3.1.2 Candidate Generation 23
3.1.3 Counting Supports of Candidates 24
3.2 An Example of Mining Large Extended Sequence Sets 25
3.3 Buffer Management 29
Chapter 4 Performance Evaluation 31
4.1 Generation of Synthetic Data 31
4.2 The Comparing Algorithm 34
4.3 Experiments on Synthetic Data 34
4.3.1 Basic Experiment 35
4.3.2 Scale-up Experiment 35
4.3.3 Effect of the Maxspan 38
Chapter 5 Conclusions and Future Work 40
References 41

[1] Agrawal, R. and Srikant, R., Fast algorithms for mining association rules, In Proc. 1994 Int. Conf. Very Large Data Bases (VLDB’94), Santiago, Chile, pp. 487-499, September 1994.
[2] Agrawal, R. and Srikant, R., Mining sequential patterns, In Proc. 1995 Int. Conf. Data Engineering (ICDE’95), Taipei, Taiwan, pp. 3-14, March 1995.
[3] Agrawal, R. and Srikant, R., Mining sequential patterns: Generalizations and performance improvements, In Proc. 5th Int. Conf. Extending Database Technology (EDBT’96), Avignon, France, pp. 3-17, March 1996.
[4] Chen, Q., Dayal, U., Han, J., Hsu, M.-C., Mortazavi-Asl, B., Pei, J., and Pinto, H., PrefixSpan: Mining sequential patterns efficiently by prefix-projected pattern growth, In Proc. 2001 Int. Conf. Data Engineering (ICDE ’01), Heidelberg, Germany, pp. 215-224, Aprial 2001.
[5] Chen, Q., Dayal, U., Han, J., Pei, J., Pinto, H., and Wang, K., Multi-dimensional sequential pattern mining, In Proc. Tenth Int'l Conf. on Information and Knowledge Management (CIKM 2001), Atlanta, Georgia, USA, pp. 81-88, November 2001.
[6] Cheung, D. W., Chin, F. Y. L., He, Y., and Wang, K, Mining confident rules without support requirement, In Proc. Tenth Int'l Conf. on Information and Knowledge Management (CIKM 2001), Atlanta, Georgia, USA, pp. 89-96, November 2001.
[7] Feng, L., Han, J., and Lu, H, Beyond intratransaction association analysis: Mining multidimensional inter-transaction association rules, In ACM Transactions on Information Systems, Vol. 18, No. 4, pp. 423-454, October 2000.
[8] Han, J. and Kamber, M, Data Mining: Concepts and Techniques, Morgan Kaufmann, San Francisco, 2000.
[9] He, Y., Wang, K., and Zhou, S, Growing decision trees on support-less association rules, In Conference on Knowledge Discovery in Data (KDD), Boston, pp. 265-269, August 2000.
[10] Roddick, J. F. and Spiliopoulou, M, A survey of temporal knowledge discovery paradigms and methods, In IEEE Transactions on Knowledge and Data Engineering, Vol. 14, No. 4, pp. 750-767, July 2002

QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top