跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.66) 您好!臺灣時間:2026/08/16 22:18
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:王保翔
研究生(外文):PAO - HSIANG WANG
論文名稱:品質預測之位置相關特徵擷取與關聯分類
論文名稱(外文):Location-based Feature Extraction and Associative Classification for Quality Prediction
指導教授:吳宜鴻
指導教授(外文):Yi-Hung Wu
學位類別:碩士
校院名稱:中原大學
系所名稱:資訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2018
畢業學年度:106
語文別:中文
論文頁數:39
中文關鍵詞:特徵選擇分群關聯分類
外文關鍵詞:feature selectionclusteringassociative classification
相關次數:
  • 被引用被引用:0
  • 點閱點閱:156
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
本研究旨於尋找目標物件與其他不同類型物件在空間上的關聯,我們以飯店為目標物件,飯店周遭的其他物件類型及數量可能影響其評等,我們提出基於不同距離擷取環境特徵的方法,藉以統計各飯店周遭一定距離內不同類型的環境物件數量,形成該飯店的環境特徵,再從不同類型環境物件中各自挑選部分特徵以探勘針對飯店評等的關聯法則,排序及修剪這些法則之後即可建立分類器。實驗資料採用交通部觀光局公布的合法飯店評等以及Google地圖上50種不同類型的環境物件及其距離資訊,我們根據不同的距離或者不同的飯店評等相關程度分別擷取特徵,然後產生對應的關聯法則及分類器,藉此觀察距離或相關程度對分類準確度的影響,最佳特徵集合可以達到93%的平均準確率。為了驗證所發現的關聯法則是否符合人們對飯店品質影響因素的一般認知,我們透過人工檢視的方式標註分類器中的關聯法則,最佳特徵集合產生的結果顯示分類器中85%的關聯法則被認為是合理的。
This thesis aims at finding spatial relationships between target objects and surrounding objects of various types. We consider hotels as target objects. The types of other objects surrounding the hotels and their quantities may have impacts on the ratings of hotels. We propose an approach to extract environmental features based on different distances. In our approach, within a certain distance the number of objects surrounding each hotel is computed to form its environmental features. After that, for each type of surrounding objects, only a portion of features are chosen to discover association rules with respect to hotel rating. A classifier can be built after these rules are sorted and pruned. Experimental data are the ratings of legal hotels announced by the Tourism Bureau in Taiwan and the objects of 50 types on Google map together with their distance information. According to different distances, or various degrees of correlation with hotel ratings, we respectively extract features and then generate association rules and the corresponding classifier. In this way, we can observe the influence of the distance or correlation degree on the classification accuracy. The best set of features can achieve an average precision of 93%. In order to verify whether the discovered association rules are in line with people’s general cognition of the factors affecting hotel quality, we label the association rules in the classifier by manual inspection. The results produced by the best set of features show that 85% of the association rules in the classifier are considered reasonable.
摘要 ..................................................................................................................................... I
Abstract .............................................................................................................................. II
致謝 ................................................................................................................................... III
目錄 .................................................................................................................................... IV
附圖目錄 ............................................................................................................................. V
附表目錄 ............................................................................................................................ VI
第一章 緒論 ...................................................................................................................... 1
第二章 相關研究 .............................................................................................................. 6
2.1 資料分群 ................................................................................................................... 6
2.2 特徵選擇 ................................................................................................................... 7
2.3 關聯分類 .................................................................................................................. 10
第三章 主要方法 ............................................................................................................. 13
3.1 特徵擷取 .................................................................................................................. 13
3.2 關聯分類 .................................................................................................................. 18
第四章 實驗 ..................................................................................................................... 22
4.1 實驗設定 ….............................................................................................................. 22
4.2 實驗結果 ….............................................................................................................. 23
第五章 結論及未來展望 ................................................................................................. 31
參考文獻 ............................................................................................................................. 32

附圖目錄
圖 一、取得旅遊資訊來源管道..............................................................................1
圖 二、Google map 飯店評論.................................................................................2
圖 三、執行流程......................................................................................................3
圖 四、糕餅店候選半徑........................................................................................14
圖 五、候選半徑示意圖........................................................................................15
圖 六、從分群結果取得距離值的示意圖............................................................17
圖 七、以距離值取得各飯店數量特徵示意圖....................................................17
圖 八、圖七經過離散化後的結果........................................................................18
圖 九、較佳規則....................................................................................................21
圖 十、兩種離散化方法........................................................................................24
圖 十一、距離上限縮減........................................................................................25
圖 十二、距離分類檢驗(二分法) .........................................................................26
圖 十三、距離分類檢驗(資訊獲利) .....................................................................27
圖 十四、高低相關特徵與準確率(二分法) .........................................................28
圖 十五、高低相關特徵與準確率(資訊獲利) .....................................................28
圖 十六、正反規則正確性....................................................................................30

附表目錄
表格 一、策略代號對照表....................................................................................23
[1].R. Agrawal, and R. Srikant, “Fast Algorithms for Mining Association Rules in Large Databases,” VLDB pp. 487-499, 1994.
[2].E. Baralis, and P. Garza, “A Lazy Approach to Pruning Classification Rules,” ICDM pp. 35-42, 2002.
[3].J. A. Hartigan, “Clustering algorithms,” 1975.
[4].G. Kundu, S. Munir, Md. Faizul Bari, Md. Monirul Islam, and Kazuyuki Murase , “A Novel Algorithm for Associative Classification,” ICONIP (2) pp. 453 - 459, 2007.
[5].W. Li, J. Han, and J. Pei, “CMAR: Accurate and Efficient Classification Based on Multiple Class-Association Rules,” ICDM pp. 369-376, 2001.
[6].X. Li, D. Qin, and C. Yu, “ACCF: Associative Classification Based on Closed Frequent Itemsets,” FSKD (2) pp. 380-384, 2008.
[7].B. Liu, W. Hsu, and Y. Ma, “Integrating Classification and Association Rule Mining,” KDD pp. 80-86, 1998.
[8].B. Liu, Y. Ma, and C-K. Wong, “Classification Using Association Rules: Weakness and Enhancements,” In Vipin Kumar, et al, (eds) Data mining for scientific applications, 2001.
[9].D. n. d. Madigan, “Descriptive modeling,” New York, NY: Columbia University.
[10].N. Slonim, E. Aharoni, K. Crammer, “Hartigan''s K-Means Versus Lloyd''s K-Means - Is It Time for a Change? ,” IJCAI pp. 1677-1684, 2013.
[11].J. Tang, S. Alelyani, and H. Liu, “Feature Selection for Classification: A Review,” Data Classification: Algorithms and Applications pp. 37-64, 2014.
[12].F. Thabtah, P. Cowling, and Y. Peng, “MCAR: Multi-class Classification based on Association Rule,” AICCSA pp. 33, 2005.
[13].F. Thabtah, W. Hadi, N. Abdelhamid, A. Issa,”Prediction Phase in Associative Classification Mining,” International Journal of Software Engineering and Knowledge Engineering pp. 855-876, 2011.
[14].F. Thabtah, P. Cowling, and Y Peng, “MMAC: A New Multi-Class, Multi-Label Associative Classification Approach,” ICDM pp. 217-224, 2004.
[15].S. Wedyan, and F. Wedyan, “An Associative Classification Data Mining Approach for Detecting Phishing Websites”, Journal of Emerging Trends in Computing and Information Sciences 4(12) pp. 888-899, 2013.
[16].S. Wedyan, “Review and comparison of associative classification data mining approaches,” International Scholarly and Scientific Research & Innovation 8(1) pp. 34-45, 2014 .
[17].Z. Xiang, and I. Md Zahidul, “Hartigan''s Method for K-modes Clustering and Its Advantages,” AusDM, 2014.
[18].S. Lloyd, “Least squares quantization in PCM,” IEEE Trans. Information Theory 28(2) pp. 129-136, 1982.
[19].I. Guyon, A. Elisseeff,”An Introduction to Variable and Feature Selection,” Journal of Machine Learning Research 3 pp. 1157-1182, 2003.
[20].“2017 Visa旅遊意向調查,” https://www.visa.com.tw/about-visa/newsroom/press-releases/nr-tw-170724/
[21].“台灣旅宿網,” https://www.taiwanstay.net.tw/Directory/Prepare
[22].“Google map地方資訊程式庫,” https://developers.google.com/maps/?hl=zh-tw
[23]. “資訊獲利,” http://ccckmit.wikidot.com/st:mutualinformation
電子全文 電子全文(本篇電子全文限研究生所屬學校校內系統及IP範圍內開放)
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊