跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.66) 您好!臺灣時間:2026/08/15 02:24
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:邱耀慶
研究生(外文):Chiu, Yao-Ching
論文名稱:一個針對動態資料驅動應用系統概念飄移的平行偵測與預測方法
論文名稱(外文):A Parallel Detection and Prediction Method for Concept Drift in Dynamic Data Driven Application System
指導教授:羅濟群羅濟群引用關係黃興進黃興進引用關係
指導教授(外文):Lo, Chi-ChunHwang, Hsin-Ginn
口試委員:林熙偵王平黃興進羅濟群
口試委員(外文):Shi-Jen LinWang, PingHwang, Hsin-GinnLo, Chi-Chun
口試日期:2015-05-56
學位類別:碩士
校院名稱:國立交通大學
系所名稱:資訊管理研究所
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2015
畢業學年度:103
語文別:英文
論文頁數:54
中文關鍵詞:大資料動態資料驅動系統概念飄移機器學習
外文關鍵詞:Big DataDynamic-Data-Driven-Application SystemConcept Driftmachine LearningMap-Reduce
相關次數:
  • 被引用被引用:0
  • 點閱點閱:300
  • 評分評分:
  • 下載下載:5
  • 收藏至我的研究室書目清單書目收藏:1
傳統的資料分析與預測方法,其預測模型都假設資料是穩定分佈的,所以藉由參照歷史資料、學習資料之間的關係,能夠很準確地預測(分類)尚未標記的資料的標記。然而,在今天多變性的大資料環境下,預測模型因為太過於依賴歷史的資料,而無法正確地推測出隨著情境而改變的資料關聯性的現象(概念飄移)。本研究提出一個針對動態資料驅動應用系統概念飄移的平行偵測與預測方法。所提出的方法快速偵測資料概念的改變,並即時的將概念飄移回饋給系統,進而調整預測模型來提高即時預測的準確率。同時,我們利用平行運算,透過區域性預測來計算出全域性預測,有效的提高預測準確率,、並減少了整體運算的時間。我們利用Map-Reduce的分散式平臺和分類演算法來實作。結果顯示,在兩個實驗案例中,平均預測的準確率較以往的預測方法分別提升了 14% 和 35%;在運算效能部分,較傳統計算方式分別節省了近 45% 和 29% 的時間。
The traditional data analysis and prediction method assumes that data distribution is stable. Therefore, it can predict unlabeled data precisely by analyzing the historical data. However, in today’s big-data environment, which is changing frequently, the traditional approach can no longer be effective; it cannot handle concept drift in a Dynamic Data Driven Application System (DDDAS). This thesis proposes a parallel detection and prediction method for concept drift in DDDAS. The proposed method can detect changing data and then feedback to the prediction model for better subsequent predictions. Furthermore, this method computes a global prediction by aggregating local predictions. Therefore, prediction accuracy is increased and computation time is decreased. In simulation, Map-Reduce is used for parallel processing. Two cases are tested. Results show that prediction accuracy is raised by 14% and 35% for these two cases, respectively. The execution time is improved by almost 45% and 29%, respectively.
Contents
Chapter 1 Introduction 1
1.1 Research Background and Motivation 1
1.2 Research Objective 3
1.3 Organization 4
Chapter 2 Literature Review 5
2.1 Dynamic Data Driven Application System (DDDAS) 5
2.1.1 The Architecture of DDDAS 6
2.1.2 Dynamic Data-Driven Application System Related Research 7
2.2 Concept Drift 8
2.3 Map-Reduce 10
2.4 Na#westeur048#ve Bayes classifiers 16
2.4.1 Gaussian Na#westeur048#ve Bayes Classifier 17
2.5 Dynamic Weighted Majority 17
2.5.1 Weighted Majority Algorithm (WMA) 18
2.5.2 Dynamic Weighted Majority Algorithm (DWM) 19
Chapter 3 A Parallel Detection and Prediction Method for Concept Drift in DDDAS 21
3.1 Design Issues 21
3.2 Definition and Notations 22
3.2.1 Definition 22
3.2.2 Notations 24
3.3 The proposed method 24
3.3.1 Overview 24
3.3.2 Detailed Discussion 27
3.4 Discussions 33
Chapter 4 Simulation and Analyses of Results 35
4.1 Simulation 35
4.1.1 Simulation Assumption 35
4.1.2 Simulation Environment 35
4.1.3 Spark Setup and Parameters Settings 36
4.2 Results and Analyses 40
4.2.1 Case 1 41
4.2.2 Case 2 45
4.3 Discussions 46
Chapter 5 Conclusions and Future work 49
5.1 Conclusions 49
5.2 Future work 50
References …………………………………………………………………………..51

1. Series, I.-T.T.W.B.R., Ubiquitous Sensor Networks (USN). 2008.
2. Ashton, K., That ‘internet of things’ thing. RFiD Journal, 2009. 22(7): p. 97-114.
3. Mell, P. and T. Grance, The NIST definition of cloud computing. National Institute of Standards and Technology, 2009. 53(6): p. 50.
4. McAfee, A. and E. Brynjolfsson. Big Data: The Management Revolution. 2012.
5. Tsymbal, A., The problem of concept drift: definitions and related work. Computer Science Department, Trinity College Dublin, 2004. 106.
6. Chen, Y., et al. Emerging topic detection for organizations from microblogs. in Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval. 2013. ACM.
7. Darema, F., Dynamic data driven applications systems: A new paradigm for application simulations and measurements, in Computational Science-ICCS 2004. 2004, Springer. p. 662-669.
8. Kolter, J.Z. and M.A. Maloof, Dynamic weighted majority: An ensemble method for drifting concepts. The Journal of Machine Learning Research, 2007. 8: p. 2755-2790.
9. Darema, F., Introduction to the ICCS 2007 workshop on dynamic data driven applications systems, in Computational Science–ICCS 2007. 2007, Springer. p. 955-962.
10. Douglas, C.C., et al. DDDAS approaches to wildland fire modeling and contaminant tracking. in Simulation Conference, 2006. WSC 06. Proceedings of the Winter. 2006. IEEE.
11. Rodr#westeur046#guez, R., A. Cort#westeur042#s, and T. Margalef. Data Injection at Execution Time in Grid Environments Using Dynamic Data Driven Application System for Wildland Fire Spread Prediction. in Proceedings of the 2010 10th IEEE/ACM International Conference on Cluster, Cloud and Grid Computing. 2010. IEEE Computer Society.
12. Douglas, C.C. and Y. Efendiev, A dynamic data-driven application simulation framework for contaminant transport problems. Computers &; Mathematics with Applications, 2006. 51(11): p. 1633-1646.
13. Douglas, C.C., et al., Dynamic Data-Driven Application Systems for empty houses, contaminat tracking, and wildland fireline prediction, in Grid-Based Problem Solving Environments. 2007, Springer. p. 255-272.
14. Allen, G., Building a dynamic data driven application system for hurricane forecasting, in Computational Science–ICCS 2007. 2007, Springer. p. 1034-1041.
15. Hirschfeld, R. and K. Kawamura. Dynamic service adaptation. in Distributed Computing Systems Workshops, 2004. Proceedings. 24th International Conference on. 2004. IEEE.
16. Wang, S., S. Schlobach, and M. Klein, What is concept drift and how to measure it?, in Knowledge Engineering and Management by the Masses. 2010, Springer. p. 241-256.
17. Harries, M.B., C. Sammut, and K. Horn, Extracting hidden context. Machine learning, 1998. 32(2): p. 101-126.
18. Widmer, G. and M. Kubat, Learning in the presence of concept drift and hidden contexts. Machine learning, 1996. 23(1): p. 69-101 %@ 0885-6125.
19. Street, W.N. and Y. Kim. A streaming ensemble algorithm (SEA) for large-scale classification. in Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining. 2001. ACM.
20. Zliobaite, I., Learning under concept drift: an overview. 2009, Overview”, Technical report, Vilnius University, 2009 techniques, related areas, applications Subjects: Artificial Intelligence.
21. Rangari, S.R., S. Dongre, and L. Malik, A new classifier for handling concept drifting data stream. International Jour-nal of Science and Research, 2013. 2(5): p. 441-444.
22. Dean, J. and S. Ghemawat, MapReduce: simplified data processing on large clusters. Communications of the ACM, 2008. 51(1): p. 107-113.
23. Chu, C., et al., Map-reduce for machine learning on multicore. Advances in neural information processing systems, 2007. 19: p. 281.
24. Yang, H.-c., et al. Map-reduce-merge: simplified relational data processing on large clusters. in Proceedings of the 2007 ACM SIGMOD international conference on Management of data. 2007. ACM.
25. Zaharia, M., et al. Spark: cluster computing with working sets. 2010.
26. Murphy, K.P., Naive bayes classifiers. University of British Columbia, 2006.
27. Littlestone, N. and M.K. Warmuth, The weighted majority algorithm. Information and computation, 1994. 108(2): p. 212-261 %@ 0890-5401.
28. Gama, J., et al., Learning with drift detection, in Advances in Artificial Intelligence–SBIA 2004. 2004, Springer. p. 286-295 %@ 3540232370.
29. Andrzejak, A. and J.B. Gomes. Parallel Concept Drift Detection with Online Map-Reduce. in Data Mining Workshops (ICDMW), 2012 IEEE 12th International Conference on. 2012. IEEE.
30. Murthy, A. Apache Hadoop YARN – Background and an Overview. 2012; Available from: http://hortonworks.com/blog/apache-hadoop-yarn-background-and-an-overview/.
31. foundation, A. and G. contributors. Apache Spark. 2010; Available from: https://github.com/apache/spark.
32. Forest Covertype Data Set. Available from: http://moa.cs.waikato.ac.nz/datasets/.

連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊