跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.75) 您好!臺灣時間:2026/08/19 08:31
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:陳昭男
研究生(外文):Chao-Nan Chean
論文名稱:網頁之型態、時間、與關鍵詞偵測
論文名稱(外文):Detection of Page Type, Time, and Key Terms of Web Pages
指導教授:吳昇
指導教授(外文):Sun Wu
學位類別:碩士
校院名稱:國立中正大學
系所名稱:資訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2003
畢業學年度:91
語文別:中文
中文關鍵詞:搜尋引擎網頁型態關鍵詞
外文關鍵詞:Search EnginePage TypeKey TermTFIDF
相關次數:
  • 被引用被引用:1
  • 點閱點閱:240
  • 評分評分:
  • 下載下載:22
  • 收藏至我的研究室書目清單書目收藏:2
隨著WWW快速的成長,線上資源的量變得更富豐。現今的搜尋引擎不僅提供了一般的網頁搜尋服務,還提供了領域相關或型態相關的搜尋服務來符合使用者的需求。而要提供特定型態之搜尋服務,必優先建立一個自動化的型態偵測的機制。
藉由網頁的統計分析,我們找出了一些適合用於型態偵測的特徵,我們也提出了一個評分方法,來評估網頁屬於哪一個類型。有時網頁內容所描述的時間資訊和網頁的最後修改時間不一樣,我們定義了一些規則來偵測時間資訊。
在擷取關鍵詞時,網頁裡的每個詞有三項特徵要被計算,它們是:位置,也就是詞第一次出現的地方;強調性標籤,詞是否被某幾個HTML標籤所強調;TFIDF,網頁裡詞的普遍性的評量方法。
With the rapid growth of WWW, the amount of online resources is getting richer. Modern search engines not only provide general search service for web pages, but domain-specified or type-specified search service to meet users'' need. To be able to provide type-specified search service, one needs to build up an automatic mechanism for type detection.
By statistical analysis of the web pages, we find out some features which are appropriate for type detection. We also propose a scoring method to evaluate which type the web page belongs. Sometimes, the time information described in the content of the web page may be different from the last modified time of the web page. We define some rules to detect the time information from the web page.
When extracting key terms, three features are calculated for each term in the web page. They are: location, which is the term''s first appearance; emphatic tag, whether the term is emphasized by some kinds of HTML tag or not; and TFIDF, a generality measure of a term''s frequency in a web page.
1. 緒論
1.1. 簡介
1.2. 研究動機
1.3. 章節結構
2. 相關研究
2.1. 文件分類
2.2. 關鍵詞擷取
2.3. 網站型態分類
3. 前置處理
3.1. 資料儲存格式
3.2. 群組同一個網站的網頁
4. 網頁型態偵測
4.1. 型態定義
4.1.1. BBS
4.1.2. E-Commerce
4.1.3. News
4.2. TRAINING DATA
4.3. URL分析
4.3.1. Site Name
4.3.2. CGI程式名稱
4.3.3. CGI參數名稱
4.4. 常見樣式
4.5. RECORD欄位分析
4.6. 評分
4.6.1. Positive Scoring
4.6.2. Negative Scoring
4.6.3. Neighborhood Information
5. 網頁時間偵測
6. TITLE REFINEMENT
7. 關鍵詞偵測
7.1. CANDIDATE TERMS
7.1.1. 位置
7.1.2. 強調性的HTML標籤
7.1.3. TFIDF
7.2. 評分
8. 實驗
8.1. 實驗一:網頁型態偵測
8.2. 實驗二:網頁時間偵測
8.3. 實驗三:TITLE REFINEMENT
8.4. 關鍵詞偵測
9. 結論
9.1. 結論
9.2. FUTURE WORK
9.2.1. 改善Recall
9.2.2. 其他語系
9.2.3. 文件結構分析
10. 參考文獻
[1] Edmundson H.P. , “New Methods in Automatic Extracting”, University of Maryland, 1969.
[2] Salton. G. and Buckley C., “Term Weighting Approaches in Automatic Text Retrieval.”, Information Processing and Management, 24(5):513-523.
[3] Witten, Ian H., Gordon W. Paynter, Eibe Frank, Carl Gutwin & Craig G. Nevill-Manning, “KEA: Practical Automatic Keyphrase Extraction.” Proceedings of the Fourth ACM Conference on Digital Libraries, 1999.
[4] Turney P., “Learning to Extract Keyphrases form Text”, National Research Council of Canada, 1999
[5] Sung-Ting Tsai, “Some Issues in Large Scale Data Gathering”, Department of Computer Science, National Chung-Cheng University, 2002.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top