跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.75) 您好!臺灣時間:2026/08/19 08:29
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:蔡明原
研究生(外文):Ming-yung Tsai
論文名稱:基於語義概念與網頁特徵之相關網頁擷取
論文名稱(外文):Related Web Page Retrieval Based on Semantic Concepts and Features of Web Pages
指導教授:陳榮靜陳榮靜引用關係
指導教授(外文):Rung-ching Chen
學位類別:碩士
校院名稱:朝陽科技大學
系所名稱:資訊管理系碩士班
學門:電算機學門
學類:電算機一般學類
論文種類:學術論文
論文出版年:2005
畢業學年度:93
語文別:英文
論文頁數:45
中文關鍵詞:語意搜尋潛在語義分析資源描述架構本體論
外文關鍵詞:RDFinformation retrievalontologysemantic search
相關次數:
  • 被引用被引用:1
  • 點閱點閱:689
  • 評分評分:
  • 下載下載:61
  • 收藏至我的研究室書目清單書目收藏:6
網際網路(Internet)快速的發展,使得網路上的資料越來越龐大且複雜,利用搜尋引擎去找到相關的網頁變的越來越重要。傳統搜尋相關網頁的方法,利用關鍵字與機率理論,無法找出潛在語意的網頁。再且,利用本體論(ontology)延伸使用者所輸入的關鍵字來猜測使用者所要表達的概念,然而卻遺失掉網頁內容可能預表達的真正語意。本研究提出一個不只利用本體論來找尋使用者所需要的網頁,更分析網頁所預表達的語意概念。首先,我們利用嵌入本體論的擷取網頁的方式收集網頁,並且分析其可能出現的語意概念,並且同時考慮該概念在網頁中的強度、出現的位置(網頁標籤)及概念在本體論中的關聯來找出其網頁間的相似度來對網頁分群,再且利用潛在語意分析(LSA)演算法找尋群組內可能相關的字詞以輔助使用者於搜尋時所輸入的非概念關鍵字,最後利用資源描述架構(RDF)來描述概念、字詞與網頁群組的關聯,透過這種方式,能夠有效的找尋相關的網頁,且具有語意的找出相關的網頁。
Using search engines to find information on the Internet often fails to satisfy user requirements. Previous search methodologies extended the domain of query keywords by the corresponding domain ontology to find related web pages, but typically they omitted the semantic content of the web pages, resulting in ineffective searches. In this paper, we present a related web page retrieval method that not only considers the corresponding domain ontology but also analyzes the semantic content of web pages. First, the method embeds the corresponding domain ontology of search keyword in order to find web pages from the Internet. Next, the method considers the location of the concept in the web pages, and relationships between concepts in the domain ontology when clustering the web page. Finally, an RDF structure is used to describe the relationships between keywords and web pages. We also used a Latent semantic analysis (LSA) algorithm to find relevant words in order to extend the information in the RDF. Experimental results prove that our method makes queries more effectively.
Table of Contents
Abstract I
中文摘要 II
誌謝 III
Table of Contents IV
List of Figures VI
List of Tables VII
1. Introduction 1
1-1. Background and Motivation 1
1-2. Object 2
2. Traditional method 4
2-1. Traditional information retrieval method 4
2-1-1. Boolean Searching 4
2-1-2. Vector space model 4
2-1-3. Probabilistic model 4
2-2. Using Hyperlink to find related web pages 5
2-3. Using Ontology to find related web pages 8
2-4. The Evaluation 11
2-5. Related technologies 12
2-5-1. Ontology 12
2-5-2. Content Features of web pages 12
2-5-3. RDF 13
2-5-4. Latent semantic analyses 14
3. The Related Web Pages Retrieval Method 16
3-1. The workflow of SCFP (Semantic Concepts and Features of Web pages) 16
3-2. Details of the SCFP method 17
3-2-1. Web pages analysis 18
3-2-2. The Similarity Estimation 20
3-2-3. Web Page Clustering 26
3-2-3. Construct RDF 26
3-2-4. The query 28
4. Experiments 29
4-1. Experiment environment 29
4-2. Ontology 29
4-3. Collected web pages 29
4.4 Initial experiment 31
4.5 Construction of the RDF 36
5. Conclusion and Future work 39
References 40

List of Figures
Figure 1. Hubs and Authorities 6
Figure 2. The flow of Link semantics 10
Figure 3. A simple RDF structure 13
Figure 4. The operation of an SVD 15
Figure 5. The workflow of SCFP method 17
Figure 6. The relationships between concepts 20
Figure 7. The similarity has the same concept 21
Figure 8. The parent and child concepts 22
Figure 9. The similarity of child to parent concepts 23
Figure 10. A domain ontology 24
Figure 11. The bag of the RDF 27
Figure 12. The LSA process of constructing an RDF 28
Figure 13. F-value of K-means 33
Figure 14. F-value of FarthestFirst 35
Figure 15. Extracted words 36
Figure 16. The results of the RDF 38

List of Tables
Table 1. Precision and Recall 11
Table 2. The category of data set of instruments 30
Table 3. The weight of Tag 31
Table 4. Cluster by COBWEB 32
Table 5. Cluster by K-means 33
Table 6. Cluster by FarthestFirst 34
Table 7. All results of cluster for set3 35
Table 8. The related words 37
References
[1] 林智揚、陽豐兆(2004),以知識本體為基礎的多代理人資訊系統之研究-以天氣查詢為例,碩士論文, 大葉大學大學資訊管理學系。
[2] 洪奕璿、吳秀陽(2004),採用本體論推理之使用者意向萃取、查詢擴充與概念式擷取,碩士論文,國立東華大學資訊工程學系,花蓮。
[3] 鐘正男、陽豐兆(2004),以知識本體為基礎的語意查詢系統之研究-以圖書館為例,碩士論文,大葉大學資訊管理學系,彰化。
[4] 台灣網路資訊中心(TWNIC), “2004台灣寬頻網路使用調查”, 2004.
[5] B. Larsen and C. Aone(1999), “Fast and Effective Text Mining Using Linear-time Document Clustering”, Proceedings of the 5th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 16-22.
[6] B.R. Yates(2003), “Information Retrieval in the Web beyond: beyond current search engines,” Approximate reasoning, vol. 34, pp. 97-104.
[7] B.Y. kang, D.W. kim and S.J. lee(2005), “Exploiting Concept Clusters for Content-based Information Retrieval”, Information Sciences, vol. 170, pp. 443-462.
[8] C. Jenkins, M. Jackson, P. Burden, and J. Wallis(1999), “Automatic RDF Metadata Generation for Resource Discovery”, Computer Networks, Vol. 31, pp. 1305-1320.
[9] D. Rafiei and A.O. Mendelzon(2000), “What is this page known for? Computing Web page Reputations,” Computer Networks, vol. 33, pp. 823-835.
[10] D. Fisher(1987), “Knowledge Acquisition Via Incremental Conceptual Clustering”, Machine Learning, Vol. 2, pp. 139-172.
[11] D. Hochbaum and D. Shmoys(1985), “A Best Possible Heuristic for the K-center Problem,” Mathematics of Operations Research, Vol. 10, pp. 180-184.
[12] E. Atlam, M. Fuketa, K. Morita, and J. Aoe(2003), “Documents Similarity Measurement Using Field Association Terms”, Information Processing and Management, Vol. 39, pp. 809-824.
[13] F. Meziane and Y. Rezgui(2004), “A Document Management Methodology based on Similarity Contents,” Information Sciences, Vol. 158, pp. 15-36.
[14] F. Norbert(1999), “Towards Data Abstraction in Networked Information Retrieval Systems,” Information Processing & Management, Vol. 35, pp. 101-119.
[15] H.L. Roger and C.E.H. Chua, and V. C. Storey(2001), “A Smart Web Query Method for Semantic Retrieval of Web Data,” Data & Knowledge Engineering, Vol. 38, pp. 63-84.
[16] H. kaindl, S. Kramer, and L. M. Afonso(2002), “Combining Structure Search and Content Search for the World-Wide Web,” Proceedings of the 13th International Workshop on Database and Expert Systems Applications, pp. 366-372.
[17] H. Tirri(2003), “Search in Vain: Challenges for Internet search,” IEEE Computer, Vol. 36, No. 1, pp. 115-116.
[18] J. B. MacQueen(1967), "Some Methods for Classification and Analysis of Multivariate Observations,” Proceedings of 5-th Berkeley Symposium on Mathematical Statistics and Probability", Vol. 1, pp. 281-297
[19] J. Charlotte, J. Mike and B. Peter(1999), “Automatic RDF metadata generation for resource discovery,” Computer Networks, Vol. 31, pp. 1305-1320.
[20] J. Dean and M. R. Henzinger(1999), “Finding Related Pages in the World Wide Web,” Computer Networks, Vol. 31, pp. 1467-1479.
[21] J. Hodgson(2001), “Do HTML Tags Flag Semantic Content?” IEEE Internet Computing, Vol. 5, pp. 20-25.
[22] J. Hou and Y. Zhang(2003), “Effectively Finding Relevant Web pages from Linkage Information,” IEEE transactions on knowledge and data engineering, Vol. 15, No. 4, pp. 940-951.
[23] J.M. Abasolo and M. Gomez(2000), “MELISA, An Ontology-based Agent for Information Retrieval in Medicine,” Proceedings of the First International Workshop on the Semantic Web, pp.73-82.
[24] J. M. Kleinberg(1998), “Authoritative Sources in a Hyperlinked Environment”, In Proceedings of the 9th Annual ACM-SIAM Symposium on Discrete Algorithms, pp. 668–677.
[25] K.J. Wu, M.C. Chen, and Y. Sun(2004), “Automatic Topics Discovery from Hyperlinked Documents,” Information Processing and Management, Vol. 40, pp. 239-255.
[26] K.M. Risvik and R. Michelsen(2002), “Search Engines and Web Dynamics,” Computer Networks, vol. 39, pp.289-302.
[27] L. Varlamis, M. Vazirgiannis and M. Halkidi(2004), “THESUS, A Closer View on Web Content Management Enhanced with Link Semantics,” IEEE Transactions on knowledge and data engineering, Vol. 16, No. 6, pp. 685-700.
[28] L. Vaughan(2004), “New Measurements for Search Engine Evaluation Proposed and Tested,” Information processing and Management, Vol. 40, pp. 677-691.
[29] M. klein(2001), “XML, RDF, and Relatives,” IEEE Intelligence Systems, Vol. 16, pp.26-28.
[30] M.M. Sufyan(2005), “User Feedback Based Enhancement in Web Search Quality,” Information Sciences, Information Sciences, vol. 170, pp. 153-172.
[31] M.R. Henzinger(2001), “Hyperlink Analysis for The Web,” IEEE Internet computing, Vol. 5, No. 1, pp. 45-50.
[32] M.S. Khan and S.W. Khor(2004), “Web Document Clustering Using a Hybrid Neural Network,” Applied Soft Computing, Vol. 4, pp. 423-432.
[33] P. Y. Lee, and S. C. Hui(2002), “Neural Networks for Web Content Filtering,” IEEE Intelligent Systems, Vol. 17, No. 5, pp. 48-57.
[34] R. Klapsing, G. Neumann, and W. Conen(2001), “Semantics in Web Engineering: Applying the Resource Description Framweork,” IEEE Multimedia, Vol. 8, No. 2, pp. 62-68.
[35] S. Brin and L. Page(1998),“The Anatomy of a Large-scale Hypertextual Web Search Engine,” Computer Networks and ISDN Systems, Vol.30 , pp.107-117.
[36] S. Chakrabarti, M. Joshi and V. Tawde(2001), “Enhanced Topic Distillation Using Text, Markup Tags, and Hyperlinks,” Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval New Orleans, Louisiana, pp 208-216.
[37] S. Oyama, T. Kokubo, and T. Ishida(2004), “Domain-specific Web Search with Keyword Spices,” IEEE Transactions on knowledge and data engineering, Vol. 16, No. 1, pp. 17-27.
[38] T.andreasen, P. Anker and J. Fischer(2004), “Content-based Text Querying with Ontological Descriptors”, Data & Knowledge Engineering, vol. 48, pp. 199-219.
[39] T.R. Gruber(1993), “Translation Approach to Portable Ontology Specifications, Knowledge Acquisition,” Vol. 5, pp. 199-220.
[40] V. Bhat, T. Oates, V. Shanbhag, and C. Nicholas(2004), “Finding Aliases on the Web Using Latent Semantic Analysis,” Data & Knowledge Engineering, Vol. 49, pp. 129-143.
[41] Y. Labrou, T. Finin(1999), “Yahoo ! as An Ontology- Ssing Yahoo! Categories to Describe Documents”, in: Proceedings of the Eighth International Conference on Information and Knowledge Management, pp. 180-187.
[42]“Yahoo” ,http://tw.yahoo.com/
[43]“Google”, http://www.google.com/
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top