跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.146) 您好!臺灣時間:2026/09/29 11:05
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:林桀宏
研究生(外文):Chieh-Hung Lin
論文名稱:藉由潛在狄式分佈產生論文主題結構和專家推薦名單
論文名稱(外文):Survey Topic Modeling and Expert Finding Based on Latent Dirichlet Allocation
指導教授:李漢銘李漢銘引用關係
指導教授(外文):Hahn-Ming Lee
口試委員:李漢銘
口試日期:2011-07-27
學位類別:碩士
校院名稱:國立臺灣科技大學
系所名稱:資訊工程系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2011
畢業學年度:99
語文別:英文
論文頁數:90
中文關鍵詞:主題模組、文件分類、專家推薦、潛在狄式分佈
外文關鍵詞:Topic Modeling、Document Clustering、Expert Finding、Latent Dirichlet Allocation
相關次數:
  • 被引用被引用:0
  • 點閱點閱:362
  • 評分評分:
  • 下載下載:37
  • 收藏至我的研究室書目清單書目收藏:0
當一個研究人員想要進入一個新的領域的時候,去讀調查文章是一個快速的
捷徑。調查文章提供許多好處,它介紹並且整理了這個研究領域相關重要性的方
法。使用者可以簡單的瞭解相關領域,並且快速的找到相關的文章。然而,並不
是每一個領域都有調查文章,或者這一個領域的調查文章不夠新。為了處理現實
會遇到的這種狀況,傳統的方法使用引用文字去產生調查文章。不過,引用式的
方法會在成效上有所限制。在本論文中,我們提出了一個名為論文主題結構的方
法,這個方法運用了潛在狄式分佈去建構主題模組或稱論文架構。提出的論文主
題結構使用某些學術關鍵字,並對使用者提供了兩個功能:(1)從線上圖書館
蒐集重要的學術文章;(2)分類收集的論文具有一定的結構方式。在我們方法
論中,對重要文章的特徵選取和對調查文章用潛在狄式分佈式的聚類被提出來。
我們用CSUR蒐集來的調查文章來評估所提出的機制。實驗結果顯示,我們所使用
潛在狄式分佈法有著比其他演算法來的進步。我們同時使用潛在狄式分佈法來解
決專家推薦問題,並且我們的系統可以對每一篇計畫提供公平且合適的專家給審
查委員。
To get into a new research topic for a researcher, it is a shortcut to study survey articles that provide some benefits for readers. Survey articles introduce and summarize significant approaches from those important articles of a certain research topic. Users get these to understand corresponding domains easily, and find related papers quickly.
However, it is not easy to find survey articles in every research domain, or there are no recent survey articles in a specific domain. To deal with this in actual condition, traditional approaches use citing texts to generate surveys. Nevertheless, citation-based approaches might limit the performance. In this thesis, we propose an approach, namely Survey Topic Model (STM), which applies Latent Dirichlet Allocation model (LDA) to facilitate the processes of building Topic Modeling or Survey Structure. The proposed STM provides two functions for readers by using certain academic keywords:
(1) Collecting important research articles from on-line digital libraries;
(2) Categorizing the collected papers with a certain structure manner.
In the proposed methodology, feature selection for important articles and LDA-based clustering for survey articles are proposed. We evaluate the proposed mechanism by survey as a dataset collected from CSUR articles. The experimental results show that LDA-based clustering leads to significant improvement. We also solve Expert Finding problem based on LDA, and our system provides fair and relevant expert lists for each proposal.
ABSTRACT i
ACKNOWLEDGEMENTS iii
1 Introduction 1
1.1 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.2 Challenges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.3 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.4 Proposed Concept . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.5 Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.6 Outline of the Thesis . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2 Background 10
2.1 Survey Articles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
2.2 Topic-Based Clustering . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.3 Other Applications Using LDA . . . . . . . . . . . . . . . . . . . . . 15
3 Survey Topic Modeling 19
3.1 Notation Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.2 System Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
3.2.1 Paper Finding Stage . . . . . . . . . . . . . . . . . . . . . . 21
3.2.2 Pre-Processing Stage . . . . . . . . . . . . . . . . . . . . . . 23
3.2.3 Topic Modeling Stage . . . . . . . . . . . . . . . . . . . . . 25
3.3 Overview of the Proposed System . . . . . . . . . . . . . . . . . . . 28
3.4 Using Latent Dirichlet Allocation in Expert Finding System . . . . . 30
4 Experiments 33
4.1 Experimental Setup/Tools . . . . . . . . . . . . . . . . . . . . . . . . 33
4.2 Dataset Collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
4.3 Metric . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
4.4 Experimental Design . . . . . . . . . . . . . . . . . . . . . . . . . . 37
4.4.1 Experiment 1 - LDA-Based Technique for Survey . . . . . . . 38
4.4.2 Experiment 2 - Number of Papers in Each Section . . . . . . 39
4.4.3 Experiment 3 - General and Specific Terms by TF-IDF . . . . 44
4.4.4 Experiment 4 - Using LDA in Expert Finding System . . . . . 45
4.5 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
5 Conclusions and Further Work 55
5.1 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
5.2 Further Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
APPENDIX 67
[1] Computing Surveys (CSUR), Std. [Online]. Available: http://portal.acm.org/csur/
[2] M. Alavi and D. Leidner, “Review: Knowledge management and knowledge
management systems: Conceptual foundations and research issues,” Manage-
ment Information Systems Quarterly, pp. 107–136, 2001.
[3] S. Bethard, S. Ghosh, J. Martin, and T. Sumner, “Topic model methods for automatically identifying out-of-scope resources,” in Proceedings of the 9th ACM/IEEE-CS Joint Conference on Digital Libraries, 2009, pp. 19–28.
[4] I. B’ır’o, D. Sikl’osi, J. Szab’o, and A. Bencz’ur, “Linked latent dirichlet allocation in web spam filtering,” in Proceedings of the 5th International Workshop on Adversarial Information Retrieval on the Web, 2009, pp. 37–40.
[5] I. B’ır’o, J. Szab’o, and A. Bencz’ur, “Latent dirichlet allocation in web spam filtering,” in Proceedings of the 4th International Workshop on Adversarial Information Retrieval on the Web, 2008, pp. 29–32.
[6] D. Blei and J. Lafferty, “A correlated topic model of science,” In Advances in Neural Information Processing Systems, vol. 1, pp. 17–35, 2007.
[7] D. Blei and J. McAuliffe, “Supervised topic models,” In Advances in Neural
Information Processing Systems, vol. 20, pp. 121–128, 2008.
[8] D. Blei, A. Ng, and M. Jordan, “Latent dirichlet allocation,” The Journal of Machine Learning Research, vol. 3, pp. 993–1022, 2003.
[9] L. Bolelli, S﹐. Ertekin, and C. Giles, “Topic and trend detection in text collections using latent dirichlet allocation,” Advances in Information Retrieval, pp. 776–780, 2009.
[10] J. Boyd-Graber, J. Chang, S. Gerrish, C. Wang, and D. Blei, “Reading tea leaves: How humans interpret topic models,” In Advances in Neural Information Processing Systems, vol. 31, 2009.
[11] D. Cai, Q. Mei, J. Han, and C. Zhai, “Modeling hidden topics on document manifold,” in Proceeding of the 17th ACM Conference on Information and Knowledge
Management, 2008, pp. 911–920.
[12] D. Cai, X. Wang, and X. He, “Probabilistic dyadic data analysis with local and global consistency,” in Proceedings of the 26th Annual International Conference on Machine Learning, 2009, pp. 105–112.
[13] J. Chen-Burger, “Notes on how to write a survey,” July 2009. [On-line]. Available: http://www.aiai.ed.ac.uk/∼jessicac/project/student-study-note/
Writing%20Advices/Write-a-survey.doc
[14] J. Chien and C. Chueh, “Latent dirichlet language model for speech recognition,” in IEEE Spoken Language Technology Workshop. IEEE, 2009, pp. 201–204.
[15] K. Christidis and G. Mentzas, “Using probabilistic topic models in enterprise
social software,” in Business Information Systems, 2010, pp. 23–34.
[16] A. Dempster, N. Laird, and D. Rubin, “Maximum likelihood from incomplete
data via the em algorithm,” Journal of the Royal Statistical Society. Series B
(Methodological), vol. 39, no. 1, pp. 1–38, 1977.
[17] P. Elango and K. Jayaraman, “Clustering images using the latent dirichlet
allocation model,” 2005. [Online]. Available: http://pages.cs.wisc.edu/∼dyer/
cs766/projects/elango05.pdf
[18] A. Elkiss, S. Shen, A. Fader, G. Erkan, D. States, and D. Radev, “Blind men and
elephants: What do citation summaries tell us about a research article?” Journal
of the American Society for Information Science and Technology, vol. 59, no. 1,
pp. 51–62, 2008.
[19] A. Fern’andez and S. G’omez, “Solving non-uniqueness in agglomerative hier-
archical clustering using multidendrograms,” Journal of Classification, vol. 25,
no. 1, pp. 43–65, 2008.
[20] X. Geng and J. Wang, “Toward theme development analysis with topic cluster-
ing,” in International Conference on 2008 Advanced Computer Theory and En-
gineering, 2008, pp. 628–632.
[21] T. Griffiths and M. Steyvers, “Finding scientific topics,” Proceedings of the Na-
tional Academy of Sciences of the United States of America, vol. 101, no. 1, pp.
5228–5235, 2004.
[22] T. Griffiths, M. Steyvers, and J. Tenenbaum, “Topics in semantic representation,”
Psychological Review, vol. 114, no. 2, pp. 211–244, 2007.
[23] D. Hall, D. Jurafsky, and C. Manning, “Studying the history of ideas using topic
models,” in Proceedings of the Conference on Empirical Methods in Natural Lan-
guage Processing, 2008, pp. 363–371.
[24] M. Hall, E. Frank, G. Holmes, B. Pfahringer, P. Reutemann, and I. Witten, “The
WEKA data mining software: An update,” ACM SIGKDD Explorations Newslet-
ter, vol. 11, no. 1, pp. 10–18, 2009.
[25] A. Harzing, Publish or Perish, Std., Rev. 3.0.3813, June 2010. [Online].
Available: http://www.harzing.com/pop.htm
[26] M. Hearst, “TextTiling: segmenting text into multi-paragraph subtopic passages,”
Computational linguistics, vol. 23, no. 1, pp. 33–64, 1997.
[27] J. Hirsch, “An index to quantify an individual’s scientific research output,” Pro-
ceedings of the National Academy of Sciences, vol. 102, no. 46, pp. 16 569–
16 572, 2005.
[28] D. Hochbaum and D. Shmoys, “A best possible heuristic for the k-center prob-
lem,” Mathematics of operations research, pp. 180–184, 1985.
[29] M. Hoffman, D. Blei, F. Bach, and I. Sup’erieure, “Online learning for latent
dirichlet allocation,” in Neural Information Processing Systems, 2010.
[30] T. Hofmann, “Probabilistic latent semantic indexing,” in Proceedings of the 22nd
Annual International ACM SIGIR Conference on Research and Development in
Information Retrieval, 1999, pp. 50–57.
[31] S. Keshav, “How to read a paper,” ACM SIGCOMM Computer Communication
Review, vol. 37, no. 3, pp. 83–84, 2007.
[32] S. Kim, S. Narayanan, and S. Sundaram, “Acoustic topic model for audio in-
formation retrieval,” in IEEE Workshop on Applications of Signal Processing to
Audio and Acoustics, 2009, pp. 37–40.
[33] T. Kohonen and H. Xing, “Contextually self-organized maps of chinese words,”
Advances in Self-Organizing Maps, pp. 16–29, 2011.
[34] T. Kuo, K. Yang, H. Lee, and J. Ho, “A reviewer recommendation system based
on collaborative intelligence,” in in Proceedings of the 2009 IEEE/WIC/ACM
International Conference on Web Intelligence, 2009.
[35] M. Lienou, H. Maitre, and M. Datcu, “Semantic annotation of satellite images
using latent dirichlet allocation,” Geoscience and Remote Sensing Letters, vol. 7,
no. 1, pp. 28–32, 2010.
[36] E. Linstead, P. Rigor, S. Bajracharya, C. Lopes, and P. Baldi, “Mining eclipse
developer contributions via author-topic models,” in 4th International Workshop
on Mining Software Repositories, 2007, pp. 30–34.
[37] S. Lukin, N. Kraft, and L. Etzkorn, “Bug localization using latent dirichlet allo-
cation,” Information and Software Technology, pp. 972–990, 2010.
[38] J. MacQueen, “Some methods for classification and analysis of multivariate ob-
servations,” in Proceedings of the 5th Berkeley Symposium on Mathematical
Statistics and Probabilitym, vol. 1. California, USA, 1967, pp. 281–297.
[39] G. Maskeri, S. Sarkar, and K. Heafield, “Mining business topics in source code
using latent dirichlet allocation,” in Proceedings of the 1st Conference on India
Software Engineering Conference, 2008, pp. 113–120.
[40] S. Mohammad, B. Dorr, M. Egan, A. Hassan, P. Muthukrishan, V. Qazvinian,
D. Radev, and D. Zajic, “Using citations to generate surveys of scientific
paradigms,” in Proceedings of Human Language Technologies: The 2009 Annual
Conference of the North American Chapter of the Association for Computational
Linguistics, 2009, pp. 584–592.
[41] R. Nallapati, A. Ahmed, E. Xing, and W. Cohen, “Joint latent topic models for
text and citations,” in Proceeding of the 14th ACM SIGKDD International Con-
ference on Knowledge Discovery and Data Mining, 2008, pp. 542–550.
[42] H. Nanba, N. Kando, and M. Okumura, “Classification of research papers using
citation links and citation types: Towards automatic review article generation,”
In American Society for Information Science SIG Classification Research Work-
shop: Classification for User Support and Learning, pp. 117–134, 2000.
[43] H. Nanba and M. Okumura, “Automatic detection of survey articles,” Research
and Advanced Technology for Digital Libraries, pp. 391–401, 2005.
[44] D. Newman, J. Lau, K. Grieser, and T. Baldwin, “Automatic evaluation of topic
coherence,” in Human Language Technologies: The 2010 Annual Conference of
the North American Chapter of the Association for Computational Linguistics.
Association for Computational Linguistics, 2010, pp. 100–108.
[45] D. Newman, Y. Noh, E. Talley, S. Karimi, and T. Baldwin, “Evaluating topic
models for digital libraries,” in Proceedings of the 10th Annual Joint Conference
on Digital Libraries, 2010, pp. 215–224.
[46] D. Pelleg and A. Moore, “X-means: Extending k-means with efficient estimation
of the number of clusters,” in Proceedings of the 17th International Conference
on Machine Learning, 2000, pp. 727–734.
[47] M. Porter, “An algorithm for suffix stripping,” Program, pp. 130–137, 1997.
[48] V. Qazvinian and D. Radev, “Scientific paper summarization using citation sum-
mary networks,” in Proceedings of the 22nd International Conference on Com-
putational Linguistics, vol. 1, 2008, pp. 689–696.
[49] V. Qazvinian and D. Radev, “Identifying non-explicit citing sentences for
citation-based summarization,” in Proceedings of the 48th Annual Meeting of
the Association for Computational Linguistics. Association for Computational
Linguistics, 2010, pp. 555–564.
[50] X. Qi and B. Davison, “Web page classification: Features and algorithms,” ACM
Computing Surveys (CSUR), vol. 41, no. 2, pp. 1–31, 2009.
[51] D. Ramage and E. Rosen, Stanford Topic Modeling Toolbox, Std., Rev. 1,
September 2009. [Online]. Available: http://nlp.stanford.edu/software/tmt/tmt-0.
2/
[52] M. Rosen-Zvi, C. Chemudugunta, T. Griffiths, P. Smyth, and M. Steyvers,
“Learning author-topic models from text corpora,” ACM Transactions on Infor-
mation Systems, vol. 28, no. 1, pp. 1–38, 2010.
[53] M. Rosen-Zvi, T. Griffiths, M. Steyvers, and P. Smyth, “The author-topic model
for authors and documents,” in Proceedings of the 20th Conference on Uncer-
tainty in Artificial Intelligence, vol. 70, 2004, pp. 487–494.
[54] N. Slonim, N. Friedman, and N. Tishby, “Unsupervised document classification
using sequential information maximization,” in Proceedings of the 25th annual
international ACM SIGIR conference on Research and development in informa-
tion retrieval, 2002, pp. 129–136.
[55] M. Steyvers and T. Griffiths, “Probabilistic topic models,” Handbook of Latent
Semantic Analysis, pp. 424–440, 2007.
[56] Y. Sun, J. Han, J. Gao, and Y. Yu, “Itopicmodel: Information network-integrated
topic modeling,” in the 9th IEEE International Conference on Data Mining, 2009,
pp. 493–502.
[57] J. Tang, J. Zhang, J. Yu, Z. Yang, K. Cai, R. Ma, L. Zhang, and Z. Su, “Topic
distributions over links on web,” in the 9th IEEE International Conference on
Data Mining, 2009, pp. 1010–1015.
[58] S. Teufel, A. Siddharthan, and D. Tidhar, “Automatic classification of citation
function,” in Proceedings of the 2006 Conference on Empirical Methods in Nat-
ural Language Processing, 2006, pp. 103–110.
[59] S. Teufel, A. Siddharthan, and D. Tidhar, “An annotation scheme for citation
function,” in Proceedings of the 7th SIGdial Workshop on Discourse and Dia-
logue, 2009, pp. 80–87.
[60] K. Tian, M. Revelle, and D. Poshyvanyk, “Using latent dirichlet allocation for au-
tomatic categorization of software,” in the 6th IEEE International Working Con-
ference on Mining Software Repositories. IEEE, 2009, pp. 163–166.
[61] K. Toutanova, D. Klein, C. Manning, and Y. Singer, “Feature-rich part-of-speech
tagging with a cyclic dependency network,” in Proceedings of the 2003 Con-
ference of the North American Chapter of the Association for Computational
Linguistics on Human Language Technology, vol. 1, 2003, pp. 173–180.
[62] H. Wallach, I. Murray, R. Salakhutdinov, and D. Mimno, “Evaluation methods
for topic models,” in Proceedings of the 26th Annual International Conference
on Machine Learning, 2009, pp. 1105–1112.
[63] S. Wan, C. Paris, and R. Dale, “Whetting the appetite of scientists: Producing
summaries tailored to the citation context,” in Proceedings of the 9th ACM/IEEE-
CS Joint Conference on Digital Libraries. ACM, 2009, pp. 59–68.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top