跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.197) 您好!臺灣時間:2026/09/15 19:21
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:孫毓夆
研究生(外文):Yu-Fong Sun
論文名稱:巨量資料分析平台之建置與評估:於vSphere上部署Hadoop生態圈
論文名稱(外文):Building and Evaluation of a Big Data Platform : Deploying a Hadoop Ecosystem on the vSphere
指導教授:江季翰江季翰引用關係
指導教授(外文):JI-HAN JIANG
學位類別:碩士
校院名稱:國立虎尾科技大學
系所名稱:資訊工程系碩士班
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2017
畢業學年度:105
語文別:中文
論文頁數:53
中文關鍵詞:虛擬化巨量資料雲端計算私有雲
外文關鍵詞:VirtualizationBig DataCloud computingPrivate Cloud
相關次數:
  • 被引用被引用:0
  • 點閱點閱:333
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
近年來網路平台發展快速,如社群網路、網路流量、搜尋引擎、線上影音內容、線上交易等資料的產生,快速地累積成巨量資料,若要對這些巨量資料進行運算分析,須借助於分散式運算的技術來完成。目前最熱門的雲端運算平台為Apache軟體基金會所推出的Hadoop開源軟體框架,以及Hadoop生態圈完整的基礎架構並提供分散式運算的環境,適合用於處理和儲存巨量資料。
而虛擬化技術作為雲端運算架構的關鍵技術,藉由虛擬化技術的資源排程,動態分配資源給虛擬機器,提高硬體資源的利用率,使資源的調派上更為靈活,而容錯機制可避免硬體在計畫外的停機而造成服務中斷。
本論文將使用VMware vSphere虛擬化軟體,在虛擬化環境下部署Hadoop生態圈,並規劃伺服器與儲存設備的網路架構,使用虛擬機器配置Hadoop運算節點,並且部署多組不同節點數的Hadoop叢集,使用資料集進行運算分析,評估各叢集運算的執行效率。
In recent years, internet and network platform fast-growing, such as social networks, network traffic, search engine, online audio and video content, online transactions and more data to generate the rapid growth of a large amount of data, called Big Data. If we want top process and analytics, it need to use cloud computing, and the Apache Software Foundation of Hadoop project is by far the most commonly used for macros of data analysis of open source cluster computing framework, Hadoop its full infrastructure for the cloud computing offers distributed computing technologies.
The virtualizations a cloud computing architecture of the key technology, dynamic resource adapter to the virtual machine by using virtualization, the resources of dispatch greater flexibility, and fault tolerance can avoid hardware in the planning of downtime and result in a disruption of service.
In this paper, we will use VMware vSphere virtualization software, deploy Hadoop ecosystem on the virtualization platform and design the network architecture for server cluster and storage cluster. Install the Hadoop nodes on the virtual machine, and deploy multi Hadoop cluster of different nodes, executive program for computing and analytics by using dataset and evaluate the cluster performance.
摘要...............i
Abstract...............ii
誌謝...............iii
目錄...............iv
表目錄...............v
圖目錄...............vi
第一章 簡介...............1
1.1 研究背景與動機...............1
1.2 研究目的...............1
1.3 論文架構...............2
第二章 文獻探討...............3
2.1 雲端運算...............3
2.2 巨量資料...............5
2.3 虛擬化平台VMware vSphere...............7
2.4 MapReduce 架構及工作機制...............10
2.5 Spark RDD架構與工作機制...............11
第三章 研究方法...............13
3.1 系統規劃...............13
3.2 系統架構...............15
第四章 巨量資料分析平台建置實作...............25
4.1 硬體環境...............25
4.2 虛擬化平台配置...............27
4.3 網路與儲存設備配置...............30
4.4 運算節點配置...............31
4.5 實驗結果...............35
第五章 結論與未來展望...............43
參考文獻...............44
附件一 中英文名稱與英文縮寫及軟體版本...............47
Extended Abstract...............48
簡歷(CV)...............53
[1]Z. Zheng, J. Zhu, and M. R. Lyu, 2013, "Service-generated big data and big data-as-a-service: An overview", IEEE International Congress on Big Data, IEEE, pp. 403–410, 27 June-2 July.
[2]P. Mell and T. Grance, 2009, "The NIST definition of cloud computing", in National Institute of Standards and Technology, Volume 53, Issus 6.
[3]J. Dean and S. Ghemawat, 2008,"MapReduce: Simplified Data Processing on Large Clusters", Communications of the ACM, vol. 51, no. 1, pp. 107–113, Jan.
[4]S. Ghemawat, H. Gobioff, and S.-T. Leung, 2003,"The Google file system," Proceedings of the nineteenth ACM symposium on Operating systems principles - SOSP ’03.
[5]K. Shvachko, H. Kuang, S. Radia, and R. Chansler, 2010,"The Hadoop distributed file system," 2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST).
[6]H.-c. Yang, A. Dasdan, R.-L. Hsiao, and D. S. Parker, 2007,” Map-reduce-merge: simplified relational data processing on large clusters,” In SIGMOD ’07, pages 1029–1040. ACM.
[7]M. Zaharia, M. Chowdhury, M. J. Franklin, and S. Shenker, 2010, "Spark: Cluster computing with working sets", Proceedings of the 2nd USENIX conference on Hot topics in cloud computing, p.10-10, June 22-25.
[8]C. Lucchese, S. Orlando, R. Perego, and F. Silvestri, 2004, "Webdocs: a real-life huge transactional dataset", In Proceedings of the ICDM Workshop in Frequent Itemset Mining Implementations.
[9]CDH Overview,
https://www.cloudera.com/documentation/enterprise/5-6-x/topics/cdh_intro.html.
[10]VMware, vNetwork Standard Switch Environment, 2016,
https://pubs.vmware.com/vsphere-50/index.jsp?topic=%2Fcom.vmware.wssdk.pg.doc_50%2FPG_Ch9_Networking.11.4.html.
[11]Overview of vNetwork Distributed Switch concepts
https://kb.vmware.com/selfservice/microsites/search.do?language=en_US&cmd=displayKC&externalId=1010555.
[12]vSphere Resource Management,
http://pubs.vmware.com/vsphere-60/index.jsp#com.vmware.vsphere.resmgmt.doc/GUID-98BD5A8A-260A-494F-BAAE-74781F5C4B87.html.
[13]Spark Cluster mode overview
http://spark.apache.org/docs/latest/cluster-overview.html.
[14]Hue - Hadoop user experienceHue,
http://gethue.com/.
[15]Apache Oozie Workflow Scheduler for Hadoop,
http://oozie.apache.org/.
[16]Apache hive,
https://hive.apache.org/.
[17]Apache pig,
https://pig.apache.org/.
[18]Apache Impala,
https://www.cloudera.com/documentation/enterprise/5-8-x/topics/impala.html.
[19]Apache Solr,
http://lucene.apache.org/solr/.
[20]Apache Hadoop YARN,
https://hadoop.apache.org/docs/r2.7.2/hadoop-yarn/hadoop-yarn-site/YARN.html.
[21]Apache ZooKeeper,
https://zookeeper.apache.org/.
[22]Apache HDFS,
https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-hdfs/HdfsUserGuide.html.
[23]Apache HBase,
https://hbase.apache.org/.
[24]Hypervisor,
https://en.wikipedia.org/wiki/Hypervisor.
[25]Backup and disaster recovery,
https://www.cloudera.com/documentation/enterprise/5-8-x/topics/cm_bdr_about.html.
[26]Cloudera Data replication,
https://www.cloudera.com/documentation/enterprise/5-8-x/topics/cm_bdr_replication_intro.html.
[27]Cloudera Snapshots,
https://www.cloudera.com/documentation/enterprise/5-8-x/topics/cm_bdr_snapshot_intro.html.
[28]Initial Placement and Ongoing Balancing, https://pubs.vmware.com/vsphere-51/index.jsp?topic=%2Fcom.vmware.vsphere.resmgmt.doc%2FGUID-E12EA49F-972A-4417-819F-D3B965EF0A48.html
[29]Network Architecture,
https://pubs.vmware.com/vsphere-4-esx-vcenter/index.jsp?topic=/com.vmware.vsphere.intro.doc_41/c_network_architecture.html
[30]VCenter server for vSphere management,
https://www.vmware.com/products/vcenter-server.html.
[31]ISCSI,
https://en.wikipedia.org/wiki/ISCSI.
[32]Wikipedia Entire,
http://prof.ict.ac.cn/BigDataBench/dowloads/.
[33]張淵仁, 段裘慶, 陳建中, 陳錦杏, 黃文增, 2010, "以虛擬化技術開發應用於遠距照護服務系統之雲端 運算環境研究與評估", International Journal of Advanced Information Technologies (IJAIT), vol. 4, no. 1, pp. 167–180, 1月.
[34]簡玠忠, 2013 ,"基於Hadoop 框架建立巨量資料分析處理模型研究",國立中興大學,碩士論文.
[35]Holden Karau, Andy Konwinski, Patrick Wendell, Matei Zaharia , 2016, Spark學習手冊, 許致軒, 碁峰資訊, 台北市.
[36]簡禎富, 許嘉裕, 2014, 資料挖礦與大數據分析, 前程文化事業有限公司,新北市.
[37]熊信彰, 2015,實戰雲端作業系統建置與維護|VMware vSphere 5.5虛擬化全面啟動, 碁峰資訊, 台北市.
[38]胡世忠, 2013, 雲端時代的殺手級應用:Big Data 海量資料分析, 天下雜誌,台北市.
[39]陸嘉桓, 挑戰大數據,Facebook、Google、Amazon怎麼處理Big Data?:用NoSQL搞定每年100億顆硬碟資料, 台北市: 佳魁資訊, 2014.
[40]陸嘉桓, 2012, Hadoop實戰技術手冊, 第二版, 佳魁資訊,台北市.
[41]王偉任, VMware vMotion運作架構及效能最佳建議, 2016, http://www.netadmin.com.tw/article_content.aspx?sn=1309180014&jump=1.
[42]王宏仁, 2016, Hadoop技術協助企業解決巨量資料難題 ,
http://www.ithome.com.tw/node/73977.
王宏仁, 2011, 巨量資料來襲 ,
http://online.ithome.com.tw/itadm/article.php?c=68277&s=5.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊