跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.79) 您好!臺灣時間:2026/09/02 15:57
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:黃國閔
研究生(外文):Kuo-Min Huang
論文名稱:約略集合論之Apache Spark實現與應用
論文名稱(外文):Realizing Rough Set Theory with Spark for Large Scale Information Systems
指導教授:熊甘霖
指導教授(外文):Kan-Lin Hsing
口試委員:李俊賢賴建良
口試委員(外文):Jin-Shyan LeeJian-Liang Lai
口試日期:2016-07-14
學位類別:碩士
校院名稱:元智大學
系所名稱:電機工程學系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2016
畢業學年度:104
語文別:英文
論文頁數:36
中文關鍵詞:約略集合論巨量資料資訊系統Spark
外文關鍵詞:Rough setBig dataInformation systemsSpark
相關次數:
  • 被引用被引用:0
  • 點閱點閱:398
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
作為一個替代Hadoop MapReduce的框架,Apache Spark是目前巨量資料領域中最活躍的開源計畫之一。
它與Hadoop最大的不同點,是在於它在進行疊代運算時可以將資料暫存在記憶體中。
由於它能提供現成的疊代演算法且兼具容錯的特性,在某些場合當中Spark已經被逐漸用來取代MapReduce。

約略集合論是一種資料探勘工具,適合用在處理資訊不一致的資訊系統上。
約略集合論最主要的優點之一是不需要許多額外的資訊(例如在模糊集合模型中所需的成員值,或是統計模式中所需的機率分布)。
約略集合演算法已經被運用在許多的領域之中,包含聲音辨識、影音和影像處理、程序控制,醫療與製藥、文字探勘和網路搜索,以及電力系統安全分析等等。
在這篇論文中我們會探討約略集合論於Apache Spark中之實現與應用。
Apache Spark, an alternative to Hadoop MapReduce,
is currently one of the most active open source projects in the big data world.
The major characteristic to differ Spark from Hadoop is that it is a cluster computing framework
that lets users perform in-memory computations, which catches data in memory during iterations, in a fault tolerant manner.
Supporting iterative algorithms out of the box, Spark has been adopted by many organizations to replace MapReduce.
One of the main advantages of rough set models is that they require no preliminary or additional information concerning data,
such as membership values in fuzzy set models or probability distribution in statistics.
Due to their versatility, rough set methods and algorithms have been widely used in various fields, including
voice recognition, audio and image processing, finance, process control, pharmacology and medicine,
text mining and exploration of the web, and power system security analysis.
In this thesis, a parallel and distributed implementation over Apache Spark to compute rough approximations in huge information systems is reported.
Table of Contents
書名頁. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . i
論文口試委員審定書. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ii
摘要. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . iii
Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . iv
Acknowledgements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . v
Table of Contents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . vi
List of Tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix
List of Figures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . x
1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1
2 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.1 Apache Spark . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.2 MapReduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2.3 Rough Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2.3.1 Information systems . . . . . . . . . . . . . . . . . . . . . . . 6
2.3.2 Equivalence relation . . . . . . . . . . . . . . . . . . . . . . . 7
2.3.3 Lower and upper approximations . . . . . . . . . . . . . . . . 7
2.3.4 Accuracy of approximation . . . . . . . . . . . . . . . . . . . . 8
2.3.5 Core and reduct of attributes . . . . . . . . . . . . . . . . . . 9
2.3.6 An example . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.4 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.5 Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
2.5.1 Rough Set Data Explorer (ROSE) . . . . . . . . . . . . . . . . 14
2.5.2 Rough Set Exploration System (RSES) . . . . . . . . . . . . . 15
2.5.3 Rough Set Toolkit for Analysis of Data (ROSETTA) . . . . . 15
2.5.4 Waikato Environment for Knowledge Analysis (Weka) . . . . . 16
2.5.5 RoughSets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
3 Spark Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
3.1 Data preparation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.2 Data processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.2.1 Map . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
3.2.2 Reduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
4 Simulation analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
4.1 Data sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
4.2 Numerical results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26
5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
5.1 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
5.2 Future work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
1. M. Zaharia, M. Chowdhury, M. J. Franklin, S. Shenker, and I. Stoica, “Spark: Cluster computing with working sets," in Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (HotCloud'10), 2010.
2. V. Agneeswaran, Big Data Analytics Beyond Hadoop: Real-Time Applications with Storm, Spark, and More Hadoop Alternatives. Pearson FT Press, 2014.
3. M. Zaharia, M. Chowdhury, T. Das, J. M. A. Dave, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica, “Spark: Cluster computing with working sets,” in Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation (NSDI'12), 2012.
4. Z. Pawlak and A. Skowron, “Rudiments of rough sets,” An International Journal of Information Sciences, vol. 177, no. 1, pp. 3--27, 2007.
5. ——,“Rough sets: Some extensions,” An International Journal of Information Sciences, vol. 177, no. 1, pp. 28--40, 2007.
6. ——,“Rough sets and Boolean reasoning,” An International Journal of Information Sciences, vol. 177, no. 1, pp. 41--73, 2007.
7. J. Stepaniuk, Rough -- Granular Computing in Knowledge Discovery and Data Mining, ser. Studies in Computational Intelligence. Springer, 2008, vol. 152.
8. Z. Pawlak, “Rough sets,” International Journal of Computer and Information Sciences, pp. 341--356, 1982.
9. ——, Rough Sets: Theoretical Aspects of Reasoning about Data. Norwell, MA, USA: Kluwer Academic Publishers, 1991.
10. J. Bazan, H. S. Nguyen, and M. Szczuka, “A view on rough set concept approximations," Fundamenta Informaticae, vol. 59, no. 2-3, pp. 107--118, Apr.2004.
11. A. Mitra, S. R. Satapathy, and S. Paul, “The static security analysis in power system based on Spark cloud computing platform,” in Proceedings of the 2013 IEEE International Advance Computing Conference, 2013, pp. 476--481.
12. H. Zhu, Y. Guo, M. Niu, G. Yang, and L. Jiao, “Distributed SAR image change detection based on spark,” in Proceedings of the 2015 IEEE International Geoscience and Remote Sensing Symposium, 2015, pp. 4149--4152.
13. H. Chen and F. Z. Wang, “Spark on entropy: A reliable and efficient scheduler for low-latency parallel jobs in heterogeneous cloud,” in Proceedings of the 2015 IEEE Local Computer Networks Conference Workshops, 2015, pp. 708--713.
14. G. Zhou, D. Zhao, K. Zou, W. Xu, X. Lv, Q. Wang, and W. Yin, “The static security analysis in power system based on Spark cloud computing platform,” in Proceedings of the 2015 IEEE Innovative Smart Grid Technologies, 2015, pp.1--6.
15. O. Spjuth, M. Capuccini, L. Carlsson, and U. Norinder, “Conformal prediction in Spark: Large-scale machine learning with con_dence,” in Proceedings of the 2015 IEEE/ACM International Symposium on Big Data Computing, 2015, pp. 61--67.
16. D. Harnie, A. E. Vapirev, J. K. Wegner, A. G., M. Steijaert, R. Wuyts, and W. D. Meuter, “Scaling machine learning for target prediction in drug discovery using Apache Spark,” in Proceedings of the 2015 IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, 2015, pp. 871--879.
17. R. Verma and C. Mattmann, “Extending Spark analytics through tika-based information extraction and retrieval,” in Proceedings of the 2015 IEEE International Conference on Information Reuse and Integration, 2015, pp. 215--218.
18. M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica, “Resilient distributed datasets: A fault- tolerant abstraction for in-memory cluster computing,” in Proceedings of the 7th International Conference, 2012, pp. 155--160.
19. R. Palamuttam, R. M. Mogrovejo, C. Mattmann, B. Wilson, K. Whitehall, R. Verma, L. McGibbney, and P. Ramirez, “SciSpark: Applying in-memory distributed computing to weather event detection and tracking,” in Proceedings of the 2015 IEEE International Conference on Big Data, 2015, pp. 2020--2026.
20. B. Amos and D. Tompkins, “Performance study of Spindle, a web analytics query engine implemented in Spark,” in Proceedings of the 2014 IEEE International Conference on Cloud Computing Technology and Science, 2014, pp. 505--510.
21. V. C. Dhande and B. V. Pawar, “A survey on parallel method for rough set using MapReduce technique for data mining,” International Journal of Science and Research, vol. 4, no. 1, pp. 423--426, 2013.
22. J. Zhang, J. S. Wong, T. Li, and Y. Pan, “A comparison of parallel large-scale knowledge acquisition using rough set theory on different MapReduce runtime systems,” International Journal of Approximate Reasoning, pp. 896--907, 2014.
23. W. Gromniak, “Scalability of attribute selection methods: Application of rough sets and MapReduce,” Master's thesis, University of Warsaw, 2015.
24. A. Dubewar, “Parallel rough set approximation using Map-Reduce technique in Hadoop,” International Journal of Scientific Research and Management, pp. 2149--2152, 2015.
25. T. Li, H. S. Nguyen, G. Wang, J. G. Busse, R. Janicki, A. E. Hassanien, and H. Yu, “Parallelized computing of attribute core based on rough set theory and MapReduce,” in Proceedings of the 7th International Conference, 2012, pp. 155--160.
26. Q. He, X. Cheng, F. Zhuang, and Z. Shi, “Parallel feature selection using positive approximation based on MapReduce,” Proceedings of the 2014 International Conference on Fuzzy Systems and Knowledge Discovery, pp. 397--402, 2014.
27. D. Y. Li and B. Q. Hu, “Query by example for large-scale video data by parallelizing rough set theory based on MapReduce,” International Conference on Fuzzy Systems and Knowledge Discovery, 2007.
28. K. Shirahama, Y. Lin, Y. Matsuoka, and K. Uehara, “Query by example for large-scale video data by parallelizing rough set theory based on MapReduce,” International Conference on Science and Social Research, pp. 390--395, 2010.
29. Y. Yang, Z. Chen, Z. Liang, and G. Wang, “Attribute reduction for massive data based on rough set theory and MapReduce,” in Rough Set and Knowledge Technology, J. Y. Greco, P. Lingras, G. Wang, and A. Skowron, Eds. Berlin: Springer, 2010, pp. 627--678.
30. J. Zhang, T. Li, and Y. Pan, “Parallel rough set based knowledge acquisition using MapReduce from big data,” in Proceedings of the 1st International Workshop on Big Data, Streams and Heterogeneous Source Mining: Algorithms, Systems, Programming Models and Applications. New York, NY, USA: Association for Computing Machinery, August 2012, pp. 20--27.
31. Q. He, X. Cheng, F. Zhuang, and Z. Shi, “Parallel feature selection using positive approximation based on MapReduce,” in Proceedings of the the 11th International Conference on Fuzzy Systems and Knowledge Discovery (FSKD), 2014, pp. 397--402.
32. M. Bal, “Rough sets theory as symbolic data mining method: An application on complete decision table,” An International Journal of Information Science Letters, pp. 35--47, 2012.
33. S. V. Nandgaonkar and A.B. Raut, “Parallel rough set approximation using MapReduce technique in Hadoop,” International Journal of Advanced Technology in Engineering and Science, vol. 3, no. 1, pp. 152--159, 2015.
34. A. Dubewar, “Parallel rough set approximation using Map-Reduce technique in Hadoop,” International Journal of Scientific Research and Management, vol. 3, no. 2, pp. 2149--2152, 2015.
35. H. Asfoor, “Fuzzy rough set approximations in large scale information systems,” Master's thesis, University of Washington, 2015.
36. J. Zhang, T. Li, D. Ruan, Z. Gao, and C. Zhao, “A parallel method for computing rough set approximations,” Intelligent Knowledge-Based Models and Methodologies for Complex Information Systems, pp. 209--223, 2012.
37. S. Y. Jing, J. Yang, and K. She, “A parallel method for rough entropy computation using MapReduce,” International Conference on Computational Intelligence and Security, pp. 707--770, 2014.
38. K. S. Tiwari and A. G. Kothari, “Design and implementation of rough set algorithms on FPGA: A survey,” International Journal of Advanced Research in Artificial Intelligence, vol. 3, no. 9, pp. 13--23, 2014.
39. F. Hu, G. Wang, and Y. Xia, “Attribute core computation based on divide and conquer method,” in Proceedings of the International Conference on Rough Sets and Intelligent Systems Paradigms, 2007, pp. 310--319.
40. G. Y.Wang, J. F. Peters, A. Skowron, and Y. Yao, “A new discernibility matrix and function,” in Proceedings of the First International Conference, 2006, pp. 114--121.
41. M. Yao, Y. Li, and D. Yang, “Application of rough set in diagnostics of supply chain integration,” in Proceedings of the 2009 International Conference on Management and Service Science, 2009, pp. 1--4.
42. A. Butalia, M. Dhore, and G. Tewani, “Applications of rough sets in the field of data mining,” in The First International Conference on Emerging Trends in Engineering and Technology, 2008, pp. 498--503.
43. Y. Zhi and Y. Zhao, “Application rough sets to safety risk index screening in ATC,” in Proceedings of the 2nd International Conference on Artificial Intelligence, Management Science and Electronic Commerce (AIMSEC), 2011, pp. 1955--1958.
44. T. Mittal, P. Gupta, and S. Chakraverty, “Application of rough sets in diagnosis of the depressive state of mind,” in Proceedings of the 2014 Recent Advances in Engineering and Computational Sciences (RAECS), 2014, pp. 1--6.
45. C. Chien and L. F. Chen, “Using rough set theory to recruit and retain high-potential talents for semiconductor manufacturing,” IEEE Transactions on Semiconductor Manufacturing, pp. 528--541, 2007.
46. A. Kusiak, “Rough set theory: A data mining tool for semiconductor manufacturing,” IEEE Transactions on Electronics Packaging Manufacturing, pp. 44--50, 2001.
47. J. G. Bazan, M. S. Szczuka, and J. Wroblewski, “A new version of rough set exploration system,” in Rough Sets and Current Trends in Computing, ser. Lecture Notes in Computer Science, J. J. Alpigini, J. F. Peters, A. Skowron, and N. Zhong, Eds., vol. 245. Springer, 2002, pp. 397--404.
48. C. F. Chien, K. H. Chang, and W. C. Wang, “An empirical study of design-of-experiment data mining for yield-loss diagnosis for semiconductor manufacturing,” Journal of Intelligent Manufacturing, vol. 25, pp. 961--972, 2014.


連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊