跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.43) 您好!臺灣時間:2026/09/01 08:17
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:李宜圜
研究生(外文):Yi-Huan. Lee
論文名稱:一個基於特徵相似度的特徵選取法
論文名稱(外文):A Feature Selection method Based on Feature Similarity
指導教授:余松年余松年引用關係
指導教授(外文):Sung-Nien Yu
學位類別:碩士
校院名稱:國立中正大學
系所名稱:電機工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2003
畢業學年度:91
語文別:中文
論文頁數:49
中文關鍵詞:特徵選擇圖形識別非監督式
外文關鍵詞:feature selectionpattern recognitionunsupervised
相關次數:
  • 被引用被引用:1
  • 點閱點閱:485
  • 評分評分:
  • 下載下載:66
  • 收藏至我的研究室書目清單書目收藏:1
本論文提出一個非監督式的特徵選擇方法,尤其是適用於中、高維度的資料集合。
在各種非監督式的特徵選擇法中,我們大概可以分成兩類方法。一是追求分群效能,因為這類方法通常伴隨著搜尋的過程,所以執行的速度比較慢。另一類是降低資料的冗餘,這類方法不需要搜尋的過程,所以相當快速,但是會降低分群的效能。在本論文,提出了一個混合式的方法。本方法主要分為兩個步驟:1. 利用特徵之間的相似性,以去除多餘的特徵;2. 加入基因演算法,以調整特徵的權重值。既可以快速找到適合的特徵,又能增強分類、分群的效能。例如,Iris資料集合中原有4個特徵,我們把它縮減成2個特徵後,用KNN分類器來做分類時,可達到 的精確度,再加入基因演算法調整特徵權重後,可達到 。
在實驗中,本論文利用一些指標和資料集合來評估特徵選擇的結果,資料集合包含了低、中和高維度三種類型。指標包括了對分類效能、分群效能以及資料的冗餘量進行評估。另外,也比較了本方法與利用主要成份分析的方法。
In this paper, an unsupervised feature selection algorithm is proposed. This algorithm is suitable especially for medium- and high-dimensional data sets.
The unsupervised feature selection algorithms can be classified into two categories. One is aimed at maximizing clustering performance. Since these methods usually need a searching process, the execute speed is usually slow. Another is aimed at reducing the redundancy in the data sets. These methods do not need a searching process, so the execute speed is faster than previous method. However, these methods usually have inferior clustering performance. This paper proposed a hybrid approach. The proposed algorithm has two steps, namely, elimination of redundant features by feature similarity and adjustment of remaining features by Genetic Algorithm (GA). The proposed algorithm can not only find suitable features but enhance clustering performance. For example, The data set Iris has 4 features originally. When we reduced it to 2-features dataset and used K-NN classifier to classify it, the accuracy is 87%. If we further adjust the features’ weighting coefficient by using GA, the accuracy can be as high as 93%.
In the experiment, we test the capability of the proposed method by different data sets and six indices were used to evaluate the results. Three categories of real-life public domain data sets were used, including low-dimensional , medium-dimensional , and high-dimensional . Six indices were used to measure the classification effectiveness, clustering performance, and the amount of redundancy of the reduced feature subset. In addition, we also compare the proposed method with PCA.
摘要 i
Abstract ii
目錄 iii
表目錄 v
圖目錄 vi
第一章 緒論 1
1.1 簡介 1
1.2 研究動機 2
1.3 研究目的 3
1.4 論文章節安排 3
第二章 研究背景 5
2.1 特徵選擇原理與文獻回顧 5
2.2 基因演算法原理與流程 10
第三章 本研究的方法 13
3.1 相似性量測工具 13
3.2 基於特徵相似度的特徵選取法 18
3.3 本論文所使用的方法 22
3.4應用基因演算法來調整特徵權重 26
3.4.1染色體表示法 26
3.4.2 適應函數 27
3.4.3 演算方法 28
第四章 實驗成果與討論 30
4.1 特徵評價指標 30
4.2 資料集合簡介 35
4.3 實驗結果與討論 36
4.3.1 與Mitra提出之方法的比較 36
4.3.2 與特徵提取方法之比較 43
第五章 結論 46
參考文獻 47
[1] M. Dash and H. Liu, “Unsupervised Feature Selection,” Proc. Pacific Asia Conf. Knowledge Discovery and Data Mining, pp. 110-121, 2000.
[2] J. Dy and C. Brodley, “Feature Subset Selection and Order Identification for Unsupervised Learning,” Proc. 17th Int’l. Conf. Machine Learning, 2000.
[3] S. Basu, C.A. Micchelli, and P. Olsen, “Maximum Entropy and Maximum Likelihood Criteria for Feature Selection from Multivariate Data,” Proc. IEEE Int’l. Symp. Circuits and Systems, vol. 3, pp. 267-270, 2000.
[4] S.K. Pal, R.K. De, and J. Basak, “Unsupervised Feature Evaluation: A Neuro-Fuzzy Approach,” IEEE Trans. Neural Network, vol. 11, pp. 366-376, 2000.
[5] F.C.-H. Rhee and Y.J. Lee, “Unsupervised Feature Selection using a Fuzzy-Genetic Algorithm,” IEEE International Conference on Fuzzy Systems, vol. 3, pp. 1266-1269, 1999
[6] P.A. Devijver and J. Kittler, Pattern Recognition: A Statistical Approach. Englewood Cliffs: Prentice Hall, 1982.
[7] P. Pudil, J. Novovicova, and J. Kittler, “Floating Search Methods in Feature Selection,” Pattern Recognition Letters, vol. 15, pp. 1119-1125, 1994.
[8] B. King, “Step-Wise Clustering Procedures,” J. Am. Statistical Assoc., pp. 86-101, 1967.
[9] P. Mitra, C.A. Murthy, and S.K. Pal, “Unsupervised Feature Selection using Feature Similarity,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 24, pp. 301-312, 2002
[10] C.L. Blake and C.J. Merz, UCI Repository of Machine Learning Databases, Univ. of California, Irvine, Dept. of Information and Computer Sciences, http://www.ics.uci.edu/~mlearn/MLRepository.html, 1998.
[11] 蘇木春, 張孝德 編著, “機器學習:類神經網路、模糊系統以及基因演算法,” 全華科技圖書股份有限公司, 1997
[12] M.L. Raymer, W.F. Punch, E.D. Goodman, L.A. Kuhn, A.K. Jain, “Dimensionality Reduction using Genetic Algorithms,” IEEE Trans. Evolutionary Computation, vol 4, pp. 164-171, 2000
[13] G. Biswas, J.B. Weinberg, D.H. Fisher, “ITERATE: A Conceptual Clustering Algorithm for Data Mining,” IEEE Trans. Systems, Man and Cybernetics, Part C, vol 28, pp. 219-230, 1998
[14] 張智星, “Matlab 程式設計與應用”, 清蔚科技, 2000
[15] A. Jain, D. Zongker, “Feature selection: evaluation, application, and small sample performance,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol 19, pp. 153-158, 1997
[16] L. J. Wang, X. Z. Wang, M. H. Ha and Y. S. Gu, “Mining the weights of similarity measure through Learning,” Proc. 2002 Int’l Conf. Machine Learning and Cybernetics, vol 4, pp. 1837-1841 2002
[17] K. Fukunaga, Introduction to statistical pattern recognition. Boston: Academic Press 1990
[18] M. Kirby, Geometric data analysis: an empirical approach to dimensionality reduction and the study of patterns. N.Y.: Wiley, 2001
[19] A. Gonzalez, P. Perez, “Selection of Relevant Features in a Fuzzy Genetic Learning Algorithm,” IEEE Trans. Systems, Man and Cybernetics, Part B, vol 31, pp. 417-425, 2001
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top