跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.73) 您好!臺灣時間:2026/07/23 00:33
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:黃明華
研究生(外文):Minghua Huang
論文名稱:大型資料庫中關聯式規則平行發掘演算法
論文名稱(外文):Algorithms for Parallel Association Rules Mining
指導教授:呂永和
指導教授(外文):Yungho Leu
學位類別:碩士
校院名稱:國立臺灣科技大學
系所名稱:管理研究所資訊管理學程
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:1999
畢業學年度:87
語文別:中文
論文頁數:48
中文關鍵詞:資料擷取知識粹取關聯式規則平行演算法菜籃分析
外文關鍵詞:data miningKnowledge Discovery in Databaseassociation rulesparallel algorithmmarket basket analysis
相關次數:
  • 被引用被引用:1
  • 點閱點閱:588
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:1
資料擷取(Data Mining)中發掘關聯式規則(Association rules)是一項重要的工作。相對於以Apriori演算法為基礎的各種平行演算法,我們提出了一個架構在磁碟共享環境下的平行演算法PBSM,並且將之實作於nCUBE2平行電腦上。此方法包含了有效地利用多處理器來減少產生高頻項目組(frequent itemsets)所需的時間,以及利用所產生的高頻項目組生成關連式規則兩步驟。
我們主要的研究重點在於第一個步驟的效率改善。我們的方法是以布林運算為基礎作表格處理,在產生高頻項目組的時候可以於各個處理器獨立運作,不需透過處理器間的網路來傳遞項目組或支持度等訊息。較之前人所提的平行演算法,我們的方法減少了許多計算與訊息傳遞的動作,因此在效能上得到了許多的改進。
Mining association rules is an important task. Many parallel algorithms have been proposed to expedite the execution of the mining process. In this thesis, we propose a parallel algorithm called ''PBSM'' for shared-disk environments, and implement the PBSM algorithm on an nCUBE parallel computer.
In the PBSM algorithm, mining process is divided into two steps. In the first step, multiple processors are used to generate frequent itemsets. Then, in the second phase, a chosen processor is used to generate the related association rules.
Through boolean-based table operations, the PBSM algorithm needs not generate candidate itemsets─which constitute the major part of execution time in the previous Apriori-based mining algorithms. Further-more, in the PBSM algorithm, each processor works independently in generating frequent itemsets. There is no need to send messages for itemsets, supports or counts between processors. As a result, our PBSM algorithm shows a superb performance compared to the existing parallel mining algorithms.
第一章 簡介1
1.1 何謂Data Mining1
1.2 關聯式規則的定義2
1.3 論文結構3
第二章 相關研究4
2.1 發掘關聯式規則的序列化演算法4
2.1.1 Apriori4
2.1.2 DHP (Direct Hash-based Pruning)7
2.1.3 DLG (Direct Large itemsets Generation)9
2.1.4 BSM (Boolean-based algorithm with Spare Matrix implementation)10
2.2 發掘關聯式規則的平行演算法10
2.2.1 CD (Count Distribution)10
2.2.2 DD (Data Distribution)12
2.2.3 FDM (Fast Distributed Mining)13
第三章 PBSM 演算法14
3.1 產生高頻項目組的效率瓶頸14
3.2 PBSM平行演算法16
3.3 使用nCUBE平行電腦環境實作23
3.4 平行演算法的實作與測試28
第四章 實驗結果與分析34
4.1 人造資料庫的建立34
4.2 各演算法執行效率的比較35
4.2.1 序列演算法的比較35
4.2.2 平行演算法的比較38
第五章 結論44
參考文獻46
作者簡介49
[1] R. Agrawal, T. Imielinski, and A. Swami, "Mining Association Rules between Sets of Items in Large Databases," Proceedings of the ACM SIGMOD International Conference on Management of Data, pp. 207-216, May 1993.
[2] M. Houtsma and A. Swami, "Set-Oriented Mining for Association Rules in Relational Databases," IEEE 11th International Conference on Data Engineering, pp. 25-33, 1995.
[3] R. Agrawal and R. Srikant, "Fast Algorithms for Mining Association Rules in Large Databases," Proceedings of the 20th International Conference on Very Large Data Bases, September, 1994.
[4] Joseph P. Bigus, "Data Mining with Neural Networks” , pp. 9-16 , 1996 ,McGraw-Hill.
[5] Alex Berson Stephen J. Smith, "Data Warehousing ,Data Mining & OLAP” , pp. 333-349 ,1997 ,McGraw-Hill.
[6] Jong Soo Park, Ming-Syan Chen, and Philip S. Yu, "Using a Hash-Based Method with Transaction Trimming for Mining Association Rules", IEEE Transactions on Knowledge and Data Enginerring, pp. 813-825, Vol.9 ,No.5 September/October 1995.
[7] S. Brin, R. Motwani, J. D. Ullman, and S. Tsur, "Dynamic Itemset Counting and
Implication Rules for Market Basket Data," Proceedings of the ACM SIGMOD International Conference on Management of Data, pp. 255-264, 1997.
[8] R. Agrawal, John C. Shafer, "Parallel Mining of Association Rules," IEEE Transactions on Knowledge and Data Engineering, vol. 8 , no.6 , December 1996 .
[9] David W. Cheung , Jiawei Han ,etc., "A Fast Distributed Algorithm for Mining Association Rules," Proceedings of 4th international Conference on Parallel and Distributed Information System , December 1996.
[10] Show-Jane Yen , Arbee L.P. Chen, "An Efficient Approach to Discovering Knowledge from Large Databases," Proceedings of the international Conference on Parallel and Distributed Information System, pp. 8-18, 1996.
[11] Ming-Syan Chen, Jiawei Han, and Philip S. Yu, "Data Mining: An Overview from a Database Perspective," IEEE Transactions on Knowledge and Data Engineering, Vol. 8, No. 6, pp. 866-882, December 1996.
[12] Vipin Kumar, Ananth grama , "Parallel Computing design and analysis of algorithms" , pp. 15-48 , pp. 65-106 , Benjamin/Cummings Pub. ,1994.
[13] Eui-Hong Han, George Karypis, Vipin Kumar, "Scalable Parallel Data Mining for Association Rules, " ACM SIGMOD Conference, pp. 277-288 , 1997
[14] nCUBE , " nCUBE2 Programmer’s Guide" , Chapter4. Interprocess communication, 1992
[15] Suh-Ying Wur and Yungho Leu, "An Effective Boolean Algorithm for Mining Association Rules in Large Databases," 6th International Conference on Database Systems for Advanced Applications (DASFAA), 1999
[16] Yungho Leu and Minghua Huang, "Algorithms for Parallel Association Rules Mining in Large Databases" 1999 Workshop on Distributed System Technologies & Applications, R.O.C., 1999.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊