跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.177) 您好!臺灣時間:2026/08/26 04:57
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:何彥毅
研究生(外文):Yann-Yih Ho
論文名稱:群集技術與工程運用之探討
論文名稱(外文):Studies on Clustering Techniques on Construction Engineering
指導教授:呂守陞
指導教授(外文):Sou-Sen Leu
學位類別:碩士
校院名稱:國立臺灣科技大學
系所名稱:營建工程系
學門:工程學門
學類:土木工程學類
論文種類:學術論文
論文出版年:2003
畢業學年度:91
語文別:中文
論文頁數:103
中文關鍵詞:群集分析資料探勘小波轉換格柵式演算多變量統計分析知識萃取
外文關鍵詞:ClusteringData MiningWavelet transformGrid-based clusteringMultivariate Statistical AnalysisKnowledge Discovery
相關次數:
  • 被引用被引用:5
  • 點閱點閱:464
  • 評分評分:
  • 下載下載:33
  • 收藏至我的研究室書目清單書目收藏:5
群集技術與工程運用之探討
研 究 生:何彥毅
指導教授:呂守陞
時 間:92年7月
論 文 摘 要
營建產業於邁入新世代之際,將面臨迥然不同的經營環境,因此企業的組織、經營策略亦將因應經營環境的需求而有所改變。於營建管理的觀念導入後,營建領域為提升企業競爭力,而必需增進經營效率與決策效能。因此必需能夠掌握知識技術的能力、管理的能力及溝通的能力。近年來於各領域中蓬勃發展的資料探勘(Data mining),成為利用資訊科技以掌握知識技術能力的解決之道。
資料探勘之過程中,往往由於大量資料的累積,造成資料特性混雜、雜訊資料散佈與資料特徵不易突顯的現象。故而群集分析的有效性,具體影響資料探勘結果之成效。本研究所建構之群集分析模式(Wavelet and Grid-based Clustering, WGCLUS),藉由導入小波轉換技術與結合格柵式演算法(Grid-based Clustering),以精進格柵之切割效果與改善現有的試誤程序,並降低其不確定性與參數間之相依性。且透過多變量統計分析(Multivariate Statistical Analysis)概念之導入,以處理群集邊緣失真之問題。進而透過改良群集驗證指標並整合於模式中,使群集分析模式更為完備,更適切的扮演資料探勘承先啟後之角色。進而使得知識萃取流程更具有效性,以期使營建產業於導入知識管理系統時,得以有效掌握知識技術之能力,而改善經營效率與決策效能,獲致產業競爭力之提升。
Studies on Clustering Techniques and its application to Construction Engineering
Thesis Advisor : Sou-Sen Leu
Graduate Student : Yann-Yih Ho
Date : July, 2003
ABSTRACT
To adapt to the rapidly-changing environment and rapidly-growing information technologies, construction industries should keep up with the advance of information technologies. Data mining, a branch of information technologies, which developed recently has been utilized in many fields. However, due to tremendous amount of cumulative data during the process of mining implicit information, the data pattern could be hidden by mixed pattern data and scattered noisy data. Cluster analysis prior to detailed data mining will significantly affect the mining result. Therefore, the effectiveness of clustering will affect the result of data mining significantly. In this study, a new clustering model, Wavelet and Grid-based Clustering (WGCLUS), was built, which combines wavelet transform and grid-based clustering to increase the precision of subspace cutting procedure and lessen the trial-error procedure. Furthermore, this model can alleviate the uncertainties with parameter-setting and the dependencies between parameters. In addition, the multivariate statistical analysis is applied in this model to solve the information missing within grid boundary and to improve the cluster validity index. According to the improvements mentioned above, we can conclude that the functions of this proposed model operate more efficiently, and the procedure of knowledge discovery is more effective via this model.
目錄
摘要 Ⅰ
ABSTRACT Ⅱ
誌謝 Ⅲ
目錄 Ⅳ
圖目錄 Ⅶ
表目錄 Ⅹ
第一章 緒論 1-1
1.1 研究動機與目的 1-1
1.2 研究範疇與整體架構 1-3
1.3 研究方法與進行步驟 1-8
第二章 文獻回顧 2-1
2.1 資料探勘 2-1
2.2 群集分析 2-4
2.2.1 群集分析之目的 2-4
2.2.2 群集分析發展現況 2-5
2.2.3 群集驗證方法 2-10
2.3 小波轉換理論 2-14
2.3.1 小波轉換之簡介 2-15
2.3.2 小波轉換之特性 2-17
2.4 小結 2-19
第三章 模式之建構 3-1
3.1 模式架構 3-1
3.2 模式之基本假設 3-1
3.2.1 資料集之基本假設 3-2
3.2.2 群集之基本假設 3-3
3.2.3 異端值之基本假設 3-4
3.3 小波理論基礎之最適切割 3-4
3.3.1 子資料集群集分析之概念 3-4
3.3.2 小波轉換高低頻展現之應用 3-8
3.3.3 小波轉換多解析特性之應用 3-12
3.4 密集度準則 3-14
3.4.1 密集度準則之應用 3-14
3.4.2 密集單元之演算 3-16
3.5 關聯子資料集之鍵結 3-19
3.5.1 子資料集之關聯關係 3-20
3.5.2 關聯子資料集之鍵結演算 3-22
3.6 群集邊緣補遺與異端值偵測 3-23
3.6.1 群集補遺之目的 3-24
3.6.2 群集邊緣補遺之方法 3-25
3.6.3 群集邊緣補遺與異端值偵測之演算 3-28
3.7 群集分析模式演算流程 3-30
第四章 模式驗證與案例模擬 4-1
4.1 群集驗證指標之評估 4-1
4.2 群集驗證指標之決策準則 4-7
4.3 三維空間兩群集之驗證 4-10
4.3.1 兩群集大小關係之驗證 4-11
4.3.2 兩群集疏密關係之驗證 4-13
4.3.3 兩群集距離關係之驗證 4-16
4.3.4 兩群集資料分佈投影重疊之驗證 4-18
4.4 三維空間多群集之驗證 4-20
4.5 多維空間群集分析驗證 4-23
4.6 小結 4-26
第五章 模式效率與效能評估 5-1
5.1 模式效能評估 5-2
5.1.1 效能評估之方法 5-2
5.1.2 模式效能評估與分析 5-3
5.2 模式效率評估 5-6
5.2.1 效率評估之方法 5-6
5.2.2 模式效率評估與分析 5-7
5.3 評估結果綜合分析 5-10
第六章 結論與未來展望 6-1
6.1 結論 6-1
6.2 未來展望 6-3
參考文獻 A-1
附錄一 三維空間兩群集大小關係之測試 B-1
附錄二 三維空間兩群集疏密程度關係之測試 C-1
附錄三 三維空間兩群集距離關係測試 D-1
A.兩群集大小比例B/S=1(兩群集同為小) D-1
B.兩群集大小比例B/S=10 D-2
C.兩群集大小比例B/S=1(兩群集同為大) D-3
附錄四 三維空間兩群集重疊關係測試 E-1
附錄五 三維空間多群集測試 F-1
附錄六 多維空間群集分析驗證 G-1
附錄七 模式效能評估分析結果 H-1
A.測試資料1 N=2200 H-1
B.測試資料集2 N=550 H-2
中英文對照表 I-1
圖目錄
圖1.1 研究動機與目的示意圖 1-2
圖1.2 資料探勘主要功能 1-4
圖1.3 群集分析於資料探勘架構中之角色 1-6
圖1.4 研究範疇與整體架構 1-7
圖1.5 本研究之核心技術 1-8
圖1.6 研究方法步驟 1-10
圖2.1 資料探勘於知識萃取流程中之色 2-2
圖2.2 資料探勘技術為多種領域結合之應用 2-3
圖2.3 資料探勘之流程 2-3
圖2.4 群集分析發展現況 2-6
圖2.5 各種群集驗證評估理論相關性架構 2-12
圖2.6 可分程度指標Dens_bw(c)之概念 2-13
圖2.7 群集指標之應用 2-14
圖2.8 STFT與小波轉換之解析度示意圖 2-17
圖2.9 尺度函數與小波函數之向量空間 2-18
圖3.1 群集模式架構 3-2
圖3.2 資料集示意圖 3-3
圖3.3 格柵式演算利用子資料集演算之流程 3-5
圖3.4 格柵式演算法之概念 3-6
圖3.5 CLIQUE 與 MAFIA之概念比較 3-7
圖3.6 MAFIA之適切性切割模式 3-8
圖3.7 Haar 尺度函數與小波函數 3-9
圖3.8 導入小波轉換技術於適切性切割之構想 3-10
圖3.9 小波轉換之高、低頻展現 3-11
圖3.10 單解析轉換對原始訊號之切割示意圖 3-12
圖3.11 多解析提供之切割資訊與Windows size間之關係 3-14
圖3.12 密集度準則相關參數之關係 3-15
圖3.13 格柵式演算衍生密集單元 3-17
圖3.14 關聯法則之聚合演算方式 3-19
圖3.15 子資料集之關聯關係 3-21
圖3.16 關聯子資料集鍵結演算流程 3-22
圖3.17 密集單元儲存之向量格式 3-23
圖3.18 格柵式演算群集邊緣資訊遺失之示意 3-24
圖3.19 二次群集分配機制之概念 3-25
圖3.20 運用異端值偵測於分類分析之概念 3-27
圖3.21 群集邊緣補遺與異端值偵測之演算流程 3-29
圖3.22 群集分析模式演算流程 3-31
圖3.23 群集分析模式演算法 3-32
圖3.23 群集分析模式演算法(續) 3-33
圖4.1 兩群集資料群集驗證指標測試評估 4-2
圖4.2 三群集資料群集驗證指標測試評估 4-2
圖4.3 四群集資料群集驗證指標測試評估 4-3
圖4.4 五群集資料群集驗證指標測試評估 4-3
圖4.5 二維常態分佈輪廓概念 4-5
圖4.6 二維常態分佈之輪廓示意 4-6
圖4.7 格柵式子資料集切割資訊不完全示意 4-8
圖4.8 最適群集分析結果決策流程 4-10
圖4.9 資料集中之群集資料特性 4-11
圖4.10 三維空間兩群集大小關係示意 4-11
圖4.11 群集資料量大小關係之測試流程 4-12
圖4.12 異端值資料與測試資料群集大小比例之關係 4-13
圖4.13 三維空間兩群集疏密程度關係示意 4-14
圖4.14 群集間資料量疏密程度之驗證測試流程 4-15
圖4.15 異端值資料與測試資料群集疏密程度之關係 4-16
圖4.16 三維空間兩群集距離關係示意 4-17
圖4.17 群集間距離關係之驗證測試流程 4-17
圖4.18 資料分佈投影重疊測試資料示意 4-19
圖4.19 資料分佈重疊投影之測試結果 4-19
圖4.20 三維空間兩群集資料 4-21
圖4.21 三維空間三群集資料 4-21
圖4.22 三維空間四群集資料 4-21
圖4.23 三維空間五群集資料 4-22
圖4.24 多維空間群集分析驗證架構 4-24
圖4.25 四維空間二群集測試資料集資料分佈 4-25
圖4.26 五維空間二群集測試資料集資料分佈 4-25
圖4.27 六維空間二群集測試資料集資料分佈 4-25
圖5.1 模式評估架構圖 5-1
圖5.2 模式效能評估流程 5-2
圖5.3 群集模式效率評估流程 5-7
圖5.4 CLIQUE 與 WGCLUS群集分析模式效率比較 5-9
表目錄
表2.1 現有群集分析方法比較 2-10
表4.1 群集可分程度指標Dens_bw(c)與Sep(c)之比較 4-7
表4.2 群集分析結果實例 4-9
表4.3 三維空間多群集資料測試結果匯整 4-22
表5.1 模式效能評估測試資料1分析結果 5-4
表5.2 模式效能評估測試資料2分析結果 5-6
表5.3 WGCLUS測試資料分析結果效率評估 5-8
表5.4 CLIQUE測試資料分析結果效率評估 5-9
附表1.1 兩群集大小比例不同之測試結果 B-1
附表2.1 兩群集疏密比例不同之測試結果 C-1
附表3.1 兩群集距離不同之測試結果 D-1
附表3.2 兩群集距離不同之測試結果 D-2
附表3.3 兩群集距離不同之測試結果 D-3
附表4.1 兩群集大小B/S=1資料分佈投影完全重疊之測試結果 E-1
附表4.2 兩群集大小B/S=1資料分佈投影半部重疊之測試結果 E-1
附表4.3 兩群集大小B/S=1資料分佈投影邊緣重疊之測試結果 E-2
附表4.4 兩群集大小B/S=5資料分佈投影完全重疊之測試結果 E-2
附表4.5 兩群集大小B/S=5資料分佈投影半部重疊之測試結果 E-3
附表4.6 兩群集大小B/S=5資料分佈投影邊緣重疊之測試結果 E-3
附表5.1 三維空間兩群集資料之測試結果 F-1
附表5.2 三維空間三群集資料之測試結果 F-1
附表5.3 三維空間四群集資料之測試結果 F-2
附表5.4 三維空間五群集資料之測試結果 F-2
附表6.1 四維空間二群集測試資料集分析結果 G-1
附表6.2 五維空間二群集測試資料集分析結果 G-1
附表6.3 六維空間二群集測試資料集分析結果 G-2
附表7.1 WGCLUS測試資料1分析結果效能評估 H-1
附表7.2 CLIQUE測試資料1分析結果效能評估 H-1
附表7.3 k-means 測試資料1分析結果效能評估 H-2
附表7.4 WGCLUS測試資料2分析結果效能評估 H-2
附表7.5 CLIQUE測試資料2分析結果效能評估 H-3
附表7.6 k-means 測試資料2分析結果效能評估 H-3
附表7.7 層級群集法測試資料2分析結果效能評估 H-4
參考文獻
[彭文正, 2001] 彭文正 譯,原著:Michael J.A. Berry, and Gordon Linoff, “資料採礦-顧客關係管理暨電子行銷之應用”,數博網資訊股份有限公司,維科出版社,ISBN 957-8675-75-5,2001.
[Agrawal et al., 1996] R. Agrawal, H. Mannila, R. Strikant, H. Toivonen, and A. I. Verkamo, “Fast Discovery of Association Rules” in Usama M. Fayyad, Gregory Piatetsky-Shapiro, Padhraic Smyth, and Ramasamy Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, chapter 12, p.p. 307-328. AAAI/MIT Press, 1996.
[Agrawal et al., 1998] R. Agrawal, J. Gehrke, D. Gunopulos, and P. Raghavan, “Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications”, Proceedings 1998 ACM-SIGMOD Int. Conf. Management of Data (SIGMOD’98), pp.94-105, Seattle, WA, June 1998.
[Ankerst et al., 1999] M. Ankerst, M. Breunig, H.-P. Kriegel, and J. Snader, “OPTICS: Ordering points to identify clustering structure”, Proceedings of the ACM SIGMOD Conference, pp.49-60, Philadelphia, PA.,1999.
[Barnett and Lewis, 1984] V. Barnett, and T. Lewis, Outliers in Statistical Data (second edition), John Wiely & Sons Ltd., 1984.
[Berkhin, 2002] P. Berkhin, Survey of Clustering Data Mining Techniques, Accrue Software, Inc., 2002.
[Berry and Linoff, 1996] M. J. A. Berry, G. Linoff, Data Mining Techniques For marketing, Sales and Customer Support, John Willey & Sons, Inc, 1996.
[Bickel and Doksum, 1977] P. J. Bickel, and K. A. Doksum, Mathematical Statistics: Basic Ideas and Selected Topics, San Francisco : Holden-Day, 1977.
[Burrus et al., 1998] C. S. Burrus, R. A. Gopinath, and H. Guo, Introduction to Wavelets and Wavelets Transforms A primer, Prentice Hall, Upper Saddle River, New Jersey, 1998.
[Ester et al., 1996] M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise”, 2nd International Conference on Knowledge Discovery and Data Mining (KDD-96) , pp.226-231, 1996.
[Fayyad and Piatetsky-Shapiro, 1996] U. M. Fayyad, G. P.-Shapiro, and P. Smyth, “From Data Mining to Knowledge Discover: An Overview”, in Usama M. Fayyad, Gregory Piatetsky-Shapiro, Padhraic Smyth, and Ramasamy Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, chapter 1, pp.1-34. AAAI/MIT Press, 1996.
[Ganti et al., 1999] V. Ganti, J. Gehrke, and R. Ramakrishnan, “CACTUS-Clustering Categorical Data Using Summaries”, Proceedings of ACM SIGKDD, pp.73-83, 1999.
[Goil et al., 1999] S. Goil, H. Nagesh, and A. Choudhary, “MAFIA: Efficient and scalable subspace clustering for very large data sets.” Technical Report No. CPDC-TR-9906-010, Center for Parallel and Distributed Computing, Northwestern University Technoloical Institute, Evanston, 1999.
[Halkidi et al., 2000] M. Halkidi, M. Vazirgiannis, Y. Batistakis. "Quality scheme assessment in the clustering process", Proceedings of PKDD (Principles and Practice of Knowledge Discovery in Databases ) 2000 Conference, Lyon, France, 2000.
[Halkidi et al., 2001] M. Halkidi, Y. Batistakis, and M. Vazirgiannis, “Clustering algorithms and validity measures” tutorial paper in the Proceeding of the SSDBM 2001 Conference, 2001.
[Halkidi et al., 2001] M. Halkidi, M. Vazirgiannis. "Clustering Validity Assessment: Finding the optimal partitioning of a data set", in the Proceedings of IEEE - Internationa Conference on Data Mining (ICDM) Conference, California, USA, November 2001.
[Han and Kamber, 2001] Jiawei Han, and Micheline Kamber, Data Mining: Concepts and Techniques, Morgan Kaufmann Publishers, 2001.
[Hartigan, 1975] J. Hartigan, Clustering Algorithms, John Wiley & Sons, New York, NY., 1975.
[Hartigan and Wong, 1979] J. Hartigan, and M. Wong, Algorithm AS136: A k-means clustering algorithm, Applied Statistics, 28, pp.100-108., 1979.
[Hinneburg and Keim, 1998] A. Hinneburg, and D. Keim, “An efficient approach to clustering large multimedia databases with noise”, in Proceeding of the 4th ACM SIGKDD, pp.58-65, New York, NY., 1998.
[Jain and Dubes, 1988] A. Jain, and R. Dubes, Algorithms for Clustering Data, Prentice-Hall, Englewood Cliffs, NJ, 1988.
[Johnson and Wichern, 2002] R. A. Johnson, and D. W. Wichern, Applied Multivariate Statistical Analysis (fifth edition), Prentice-Hall, Inc., 2002.
[Kaufman and Rousseeuw, 1990] L. Kaufman, and P. Rousseeuw, Finding Groups in Data: An Introduction to Cluster Analysis, John Wiley and Sons, New York, NY, 1990.
[Knorr and Ng, 1997] EM. Knorr, RT. Ng, “A unified notion of outliers: properties and computation.” in Proceedings of KDD, pp.219-222, 1997.
[Knorr and Ng, 1998] EM. Knorr, RT. Ng, “Algorithms for mining distance-based outliers in large datasets.”, in Proceedings of VLDB, pp.392-403, 1998.
[Mallat, 1989] S.G. Mallat, “A Theory for Multiresolution Signal Decomposition: The Wavelet Representation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 11, No.7, pp. 674-693, 1989.
[Ng and Han, 1994] R. Ng, J. Han, “Efficient and effective clustering methods for spatial data mining”, in Proceedings of the 20th Conference on VLDB, pp. 144-155, Santiago, Chile, 1994.
[Rezaee et al., 1998] R. Rezaee, B.P.F. Lelieveldt, J.H.C. Reiber, “A new cluster validity index for the fuzzy c-mean”, Pattern Recognition Letters, 19, pp.237-246, 1998.
[Sheikoleslami et al., 1998] G. Sheikholeslami, S. Chatterjee, and A. Zhang, “WaveCluster: A multiresolution clustering approach for very large spatial databases”, in Proceedings of 24th Conference on VLDB, pp.429-439, New York, NY, 1998.
[Smyth, 1996] P. Smyth, “Clustering using Monte Carlo Cross-Validation”, KDD 1996, pp. 126-133, 1996.
[Vetterli and Herley, 1992] M. Vetterli, C. Herley, “Wavelets and Filter Banks: Theory and Design”, IEEE Transactions on Processing, Vol. 40, No.9, pp.426-431, 1992.
[Vetterli and Kovacevic, 1995] M. Vetterli, and J. Kovacevic, Wavelets and Subband Coding, Prentice Hall Englewood Cliffs, NJ. 07632, 1995.
[Wang et al., 1997] W. Wang, J. Yang, and R. Muntz, “STING: a statistical information grid approach to spatial data mining”, In Proceedings of the 23th Conference on VLDB, pp. 186-195, Athens, Greece, 1997.
[Yu et al., 1998] D. Yu, S. Chatterjee, G. Sheikholeslami, and A. Zhang, “Efficiently detecting arbitrary shaped clusters in very large datasets with high dimensions”, Technical Report 98-8, State University of New York at Buffalo, Department of Computer Science and Engineering, November, 1998.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
1. 吳俊憲(民89)。建構主義的教學理論與策略及其在九年一貫課程之相關性探討。人文及社會學科教學通訊,11(4),73-88。
2. 邱聯恭(1997)。再談台灣法學教育。月旦法學雜誌,25,33-36。
3. 何秋香、王興芳(民87)。合作學習教學法的應用∼五年制商專成本會計學為例。87年11月。僑光學報。
4. 朱則剛(民85)。建構主義對教學設計的意義。教學科技與媒體,56,3-12。
5. 方吉正(民87)。情境學習理論之主要觀點剖析。教育資料文摘,卷42(2)。
6. 尤菊芳(民88)。合作學習之理論篇,敦煌英語教學雜誌,22,11-15。
7. 尹玫君(民81)。難以抗拒的潮流-淺談電腦素養。國教之友,44(1),5-14。
8. 王麗雲(1999)。 個案教學法之理論與實施。課程與教學季刊,1999,2(3),117-134。
9. 于富雲(民90)。從理論基礎探究合作學習的教學效益。教育資料與研究,〈38〉,22-28。
10. 吳宗立(民89)。情境學習論在教學上的應用。人文及社會學科教學通訊,11(3)。
11. 林生傳(民87)。建構主義的教學評析。課程與教學季刊,1(3),1-14。
12. 林奇賢(民86)。全球資訊網輔助學習系統-網際網路與國小教育。資訊與教育,58,2-11。
13. 高熏芳、蔡宜君(民90)。案例教學法在師資培育之發展與運用。淡江人文社會學刊,7,90年5月。
14. 高熏芳(民85)。情境學習中教師角色之探討:共同調節師生關係模式之應用。教學科技與媒體,29。
15. 徐新逸(民87)。情境學習對教學革新之回應。研習資訊,15(1)。