跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.21) 您好!臺灣時間:2026/09/11 08:17
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:黃安婷
研究生(外文):Nancy Huang
論文名稱:微生物源資料之辨識模型探勘
論文名稱(外文):Discriminative Pattern Mining in Microbiomic Data
指導教授:歐陽彥正歐陽彥正引用關係
口試委員:阮雪芬趙坤茂賴飛羆陳倩瑜曾宇鳳
口試日期:2016-07-27
學位類別:博士
校院名稱:國立臺灣大學
系所名稱:資訊工程學研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2016
畢業學年度:104
語文別:英文
論文頁數:70
中文關鍵詞:辨識模型探勘辨識模型相關性辨識模型冗餘性辨識模型選取微生物源資料
外文關鍵詞:discriminative patternspattern miningpattern relevancypattern redundancypattern selectionmicrobiomic data
相關次數:
  • 被引用被引用:0
  • 點閱點閱:1711
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
Machine learning classifiers have long been used to solve biological problems by predicting the target class (e.g. disease state, bacterial taxonomy, etc.) of unseen samples. A favorable and important byproduct of a special type of classifier is “interpretability” (also known as “comprehensibility”), which could be utilized to offer explanations as to why and how a sample is assigned to the predicted class. Interpretable classifiers produce “discriminative patterns” that lead to different prediction results, and provide insights to critical properties of the biological problem by capturing a greater extent of underlying semantics than single features. Discriminative patterns can be directly utilized by pattern-based classifiers to predict unseen samples by a majority voting or aggregation mechanism. In this case, we are concerned with not only finding useful individual patterns, but also the effectiveness of the pattern set as a whole. Thus, it is imperative to ensure the relevancy and non-redundancy of the discriminating patterns. Few studies have evaluated pattern redundancy via examining samples covered by the patterns; and in those that do, the focus has been mostly on the proportion of overlapping samples, suggesting that a great deal of information on non-overlapping samples were overlooked. In addition, traditional pattern mining approaches often require the generation of a complete set of initial patterns and a global discretization of continuous attributes, both of which are impractical for high-dimensional biological datasets of complex nature. We address the above issues by presenting a novel pattern selection algorithm that estimates pattern redundancy by not only the proportion of overlapping samples, but also the resemblance of non-overlapping samples. The proposed method was applied on three real microbiomic datasets, with the aim of providing new insights on the interactions between microbial factors and their effects on the host. When compared with other robust classifiers and feature selection heuristics, our pattern selection algorithm led to diverse and compact sets of final patterns that demonstrated comparable or even superior predictive capabilities.


摘要.....i
ABSTRACT.....ii
Chapter 1. Background and Significance.....1
1.1 Microbiomic Studies.....1
1.1.1 Obesity and the Gut Microbiota.....2
1.2 Discriminative Pattern Mining.....5
1.2.1 Terminologies.....5
1.2.2 Current Approaches and Their Limitations.....5
1.3 Research Significance.....7
Chapter 2. The Phylogeny-based Pattern Selection Algorithm (PBPS).....9
2.1 Pattern Generation and Initial Selection.....9
2.2 Redundancy Estimation and Pattern Pruning.....13
2.3 Correcting for Multiple Hypothesis Testing.....18
Chapter 3. Experimental Framework and Results.....20
3.1 Datasets and Features.....20
3.2 Experimental Framework.....22
3.3 Results.....26
Chapter 4. Discussion.....34
4.1 The Combined Force of Association and Classification.....34
4.2 Characteristics of PBPS.....36
4.3 Microbial Abundance Patterns.....39
4.4 Limitation.....42
Chapter 5. Conclusion and Future Work.....44
Bibliography.....45


[1]J.C. Wooley, A. Godzik, and I. Friedberg “A primer on metagenomics,” PLoS Comput. Biol., 2010. doi:10.1371/journal.pcbi.1000667.
[2]J.C. Venter et al., “Environmental genome shotgun sequencing of the Sargasso Sea,” Science, vol. 304, pp. 66-74, 2004.
[3]G. Srinivas et al., “Genome-wide mapping of gene-microbiota interactions in susceptibility to autoimmune skin blistering,” Nature Communications, 2013. doi:10.1038/ncomms3462.
[4]M.L. Zupancic et al., “Analysis of the gut microbiota in the Old Order Amish and its relation to the metabolic syndrome,” PLoS ONE, 2012. doi:10.1371/journal.pone.0043052.
[5]W.A. de Steenhuijsen Piters et al., “Dysbiosis of upper respiratory tract microbiota in elderly pneumonia patients,” The ISME Journal, vol. 10, pp. 97-108, 2015.
[6]AMA News Room,
http://www.ama-assn.org/ama/pub/news/news/2013/2013-06-18-new-
ama-policies-annual-meeting.page, 2013.
[7]NHLBI Obesity Education Initiative, “Clinical guidelines on the identification, evaluation, and treatment of overweight and obesity in adults: the evidence report,” Obesity Research, vol. 6, pp. 51S-209S, 1998.
[8]OECD Health Division, “OECD obesity update 2014,”
http://www.oecd.org/els/health-systems/Obesity-Update-2014.pdf, 2014.
[9]WHO Media Center, “Obesity and overweight,”
http://www.who.int/mediacentre/factsheets/fs311/en/, 2014.
[10]R.E. Ley, P.J. Turnbaugh, S. Klein, and J.I. Gordon, “Microbial ecology: human gut microbes associated with obesity,” Nature, vol. 444, pp. 1022-1023, 2006.
[11]P.J. Turnbaugh et al., “A core gut microbiome in obese and lean twins,” Nature, vol. 457, pp. 480-484, 2009.
[12]Y. Sanz, R. Rastmanesh, and C. Agostonic, “Understanding the role of gut microbes and probiotics in obesity: how far are we?” Pharmacological Research, vol. 69, pp. 144-155, 2013.
[13]J.P. Furet et al., “Differential adaptation of human gut microbiota to bariatric surgery-induced weight loss: links with metabolic and low-grade inflammation markers,” Diabetes, vol. 59, pp. 3049-3057, 2010.
[14]F. Armougom, M. Henry, B. Vialettes, D. Raccah, and D. Raoult, “Monitoring bacterial community of human gut microbiota reveals an increase in Lactobacillus in obese patients and methanogens in anorexic patients,” PLoS ONE, 2009. doi:10.1371/journal.pone.0007125.
[15]E. Munukka et al., “Women with and without metabolic disorder differ in their gut microbiota composition,” Obesity, vol. 20, pp. 1082-1087, 2012.
[16]R. Jumpertz et al., “Energy-balance studies reveal associations between gut microbes, calorie load, and nutrient absorption in humans,” American Journal of Clinical Nutrition, vol. 94, pp. 58-65, 2011.
[17]V. Mai, Q.M. McCrary, R. Sinha, and M. Glei, “Associations between dietary habits and body mass index with gut microbiota composition and fecal water genotoxicity: an observational study in African American and Caucasian American volunteers,” Nutrition Journal, 2009. doi:10.1186/1475-2891-8-49.
[18]M.C. Collado, E. Isolauri, K. Laitinen, and S. Salminen, “Distinct composition of gut microbiota during pregnancy in overweight and normal-weight women,” American Journal of Clinical Nutrition, vol. 88, pp. 894-899, 2008.
[19]Y. Sun et al., “Advanced computational algorithms for microbial community analysis using massive 16S rRNA sequence data,” Nucleic Acids Research, 2010. doi:10.1093/nar/gkq872.
[20]L. Bervoets et al., “Differences in gut microbiota composition between obese and lean children: a cross-sectional study,” Gut Pathogens, 2013. doi:10.1186/1757-4749-5-10.
[21]D. Knights, E. K. Costello, and R. Knight, “Supervised classification of human microbiota,” FEMS microbiology reviews, vol. 35, pp. 343-359, 2011.
[22]T. Yatsunenko et al., “Human gut microbiome viewed across age and geography,” Nature, 2012. doi:10.1038/nature11053.
[23]C.A. Lozupone et al., “Alterations in the gut microbiota associated with HIV-1 infection,” Cell Host & Microbe, vol. 14, pp. 329-339, 2013.
[24]Hummelen et al., “Deep sequencing of the vaginal microbiota of women with HIV,” PLoS ONE, 2010. doi:10.1371/journal.pone.0012078.
[25]Skraban et al., “Gut microbiota patterns associated with colonization of different Clostridium difficile ribotypes,” PLoS ONE, 2013. doi:10.1371/journal.pone.0058005.
[26]A.K. Bartram, M.D.J. Lynch, J.C. Stearns, G. Moreno-Hagelsieb, and J.D. Neufeld, “Generation of multimillion-sequence 16S rRNA gene libraries from complex microbial communities by assembling paired-end Illumina reads,” Appl. Environ. Microbiol., vol. 77, pp. 3846-3852, 2011.
[27]M. García-Borroto, J.F. Martínez-Trinidad, and J.A. Carrasco-Ochoa, “A survey of emerging patterns for supervised classification,” Artificial Intelligence Review, vol. 42, pp. 705-721, 2014.
[28]P.K. Novak, N. Lavrač, and G.I. Webb, “Supervised descriptive rule discovery: a unifying survey of contrast set, emerging pattern and subgroup mining,” The Journal of Machine Learning Research, vol. 10, pp. 377-403, 2009.
[29]X. Liu, J. Wu, F. Gu, J. Wang, and Z. He, “Discriminative pattern mining and its applications in bioinformatics,” Briefings in Bioinformatics, vol. 16, pp. 884-900, 2015.
[30]G. Dong and J. Bailey, Contrast Data Mining: Concepts, Algorithms, and Applications. Boca Raton: CRC Press, 2013.
[31]G. Dong and N. Fore, “Discovering dynamic logical blog communities based on their distinct interest profiles,” In The International Conference on Social Eco-Informatics (SOTICS), pp. 24-30, 2011.
[32]L. Kobylinski and K. Walczak, “Jumping emerging patterns with occurrence count in image classification,” In Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp. 904-909, 2008.
[33]J. Li, H. Liu, J.R. Downing, A.E.J. Yeoh, and L. Wong, “Simple rules underlying gene expression profiles of more than six subtypes of acute lymphoblastic leukemia (ALL) patients,” Bioinformatics, vol. 19, pp. 71-78, 2003.
[34]A.L. Boulesteix, G. Tutz, and K. Strimmer, “A CART-based approach to discover emerging patterns in microarray data,” Bioinformatics, vol. 19, pp. 2465-2472, 2003.
[35]A. Savasere, E.R. Omiecinski, and S.B. Navathe, “An efficient algorithm for mining association rules in large databases,” In The International Conference on Very Large Data Bases (VLDB), pp. 432-443, 1995.
[36]B. Liu, W. Hsu, and Y. Ma, “Integrating classification and association rule mining,” In ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pp. 80-86, 1998.
[37]J. Li, G. Liu, and L. Wong, “Mining statistically important equivalence classes and delta-discriminative emerging patterns,” In ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pp. 430-439, 2007.
[38]X. Yin and J. Han, “CPAR: classification based on predictive association rules,” In SIAM International Conference on Data Mining (SDM), pp. 331-335, 2003.
[39]G. Grahne and J. Zhu, “Fast algorithms for frequent itemset mining using FP-trees,” IEEE Transactions on Knowledge and Data Engineering, vol. 17, pp. 1347-1362, 2005.
[40]M. García-Borroto, J.F. Martínez-Trinidad, and J.A. Carrasco-Ochoa, “A new emerging pattern mining algorithm and its application in supervised classification,” In The Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp. 150-157, 2010.
[41]G. Dong, X. Zhang, L. Wong, and J. Li, “CAEP: classification by aggregating emerging patterns,” In International Conference on Discovery Science (DS), pp. 30-42, 1999.
[42]X. Zhang, G. Dong, and K. Ramamohanarao, “Information-based classification by aggregating emerging patterns,” In International Conference on Intelligent Data Engineering and Automated Learning (IDEAL), pp. 48-53, 2000.
[43]M. García-Borroto, J.F. Martínez-Trinidad, J.A. Carrasco-Ochoa, M.A. Medina-Pérez, and J. Ruiz-Shulcloper, “LCMine: an efficient algorithm for mining discriminative regularities and its application in supervised classification,” Pattern Recognition, vol. 43, pp. 3025-3034, 2010.
[44]H. Cheng, X. Yan, J. Han, and C.W. Hsu, “Discriminative frequent pattern analysis for effective classification,” In IEEE International Conference on Data Engineering (ICDE), pp. 716-725, 2007.
[45]J. Li, C. Wang, L. Cao, and P.S. Yu, “Efficient selection of globally optimal rules on large imbalanced data based on rule coverage relationship analysis,” In SIAM International Conference on Data Mining (SDM), pp. 216-224, 2013.
[46]A.P. Martin, “Phylogenetic approaches for describing and comparing the diversity of microbial communities,” Appl. Environ. Microbiol., vol. 68, pp. 3673-3682, 2002.
[47]N. Huang and Y.J. Oyang, “Microbial abundance patterns of host obesity inferred by the structural incorporation of association measures into interpretable classifiers,” In IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 315-319, 2014.
[48]L. Yu and H. Liu, “Efficient feature selection via analysis of relevance and redundancy,” The Journal of Machine Learning Research, vol. 5, pp. 1205-1224, 2004.
[49]L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5-32, 2001.
[50]M. García-Borroto, J.F. Martínez-Trinidad, and J.A. Carrasco-Ochoa, “Finding the best diversity generation procedures for mining contrast patterns,” Expert Systems with Applications, vol. 42, pp. 4859-4866, 2015.
[51]W.L. Holcomb, T. Chaiworapongsa, D.A. Luke, and K.D. Burgdorf, “An odd measure of risk: use and misuse of the odds ratio,” Obstetrics & Gynecology, vol. 98, pp. 685-688, 2001.
[52]A.J. Viera, “Odds ratios and risk ratios: what’s the difference and why does it matter?” Southern Medical Journal, 2008. doi:10.1097/SMJ.0b013e31817a7ee4.
[53]W. Wang and Z. Zhou, “A review of associative classification approaches,” Transaction on IoT and Cloud Computing, vol. 2, pp. 68-77, 2014.
[54]J. Han, J. Pei, and Y. Yin, “Mining frequent patterns without candidate generation,” In ACM SIGMOD Conference (SIGMOD), pp. 1-12, 2000.
[55]L. Rokach and O. Maimon, “Top-down induction of decision tree classifiers – a survey,” IEEE Transactions on Systems, Man, and Cybernetics – Part C: Applications and Reviews, vol. 35, pp. 476-487, 2005.
[56]M. Slatkin and W.P. Maddison, “A cladistic measure of gene flow inferred from the phylogenies of alleles,” Genetics, vol. 123, pp. 603-613, 1989.
[57]W.M. Fitch, “Toward defining the course of evolution: minimum change for a specific tree topology,” Systematic Biology, vol. 20, pp. 406-416, 1971.
[58]W.P. Maddison and M. Slatkin, “Null models for the number of evolutionary steps in a character on a phylogenetic tree,” Evolution, vol. 45, pp. 1184-1197, 1991.
[59]Y. Benjamini and Y. Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” Journal of the Royal Statistical Society. Series B (Methodological), vol. 57, pp. 289-300, 1995.
[60]P.D. Schloss and J. Handelsman, “Introducing TreeClimber, a test to compare microbial community structures,” Appl. Environ. Microbiol., vol. 72, pp. 2379-2384, 2006.
[61]O.J. Dunn, “Multiple comparisons among means,” Journal of the American Statistical Association, vol. 56, pp. 52-64, 1961.
[62]I. Batal and M. Hauskrecht, “A concise representation of association rules using minimal predictive rules,” In The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery (ECML-PKDD), pp. 87-102, 2010.
[63]L. Fu, B. Niu, Z. Zhu, S. Wu, and W. Li, “CD-HIT: accelerated for clustering the next-generation sequencing data,” Bioinformatics, vol. 28, pp. 3150-3152, 2012.
[64]P.D. Schloss et al., “Introducing mothur: open-source, platform-independent, community-supported software for describing and comparing microbial communities,” Appl. Environ. Microbiol., vol. 75, pp. 7537-7541, 2009.
[65]T.Z. DeSantis et al., “Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB,” Appl. Environ. Microbiol., vol. 72, pp. 5069-5072, 2006.
[66]World Health Organization, “Global Database on Body Mass Index,”
http://apps.who.int/bmi/index.jsp?introPage=intro_3.html, 2013.
[67]J.D. Thompson, T.J. Gibson, and D.G. Higgins, “Multiple sequence alignment using ClustalW and ClustalX,” Current Protocols in Bioinformatics, 2002. doi:10.1002/0471250953.bi0203s00.
[68]A. Stamatakis, “RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies,” Bioinformatics, vol. 30, pp. 1312-1313, 2014.
[69]P.J. McMurdie and S. Holmes, “phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data,” PLoS ONE, 2013. doi:10.1371/journal.pone.0061217.
[70]C. Lozupone and R. Knight, “UniFrac: a new phylogenetic method for comparing microbial communities,” Appl. Environ. Microbiol., vol. 71, pp. 8228-8235, 2005.
[71]C.A. Lozupone, M. Hamady, S.T. Kelley, and R. Knight, “Quantitative and qualitative β diversity measures lead to different insights into factors that structure microbial communities,” Appl. Environ. Microbiol., vol. 73, pp. 1576-1585, 2007.
[72]A. Statnikov et al., “A comprehensive evaluation of multicategory classification methods for microbiomic data,” Microbiome, 2013. doi:10.1186/2049-2618-1-11.
[73]C.C. Chang and C.J. Lin, “LIBSVM: a library for support vector machines,” ACM Transactions on Intelligent Systems and Technology (TIST), 2011. doi:10.1145/1961189.1961199.
[74]J. Carbonell and J. Coldstein, “The use of mmr, diversity-based reranking for reordering documents and producing summaries,” In ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), pp. 335-336, 1998.
[75]T. Fawcett, “ROC graphs: notes and practical considerations for researchers,” Technical Report HPL-2003-4, HP Labs, 2004.
[76]X. Robin et al., “pROC: an open-source package for R and S+ to analyze and compare ROC curves,” BMC Bioinformatics, 2011. doi:10.1186/1471-2105-12-77.
[77]C.G. Whitney et al., “Increasing prevalence of multidrug-resistant Streptococcus pneumoniae in the United States,” The New England Journal of Medicine, vol. 343, pp. 1917-1924, 2000.
[78]D.M. Dziuda, Data Mining for Genomics and Proteomics: Analysis of Gene and Protein Expression Data. New York: Wiley, 2010.
[79]J.U. Scher et al., “Expansion of intestinal Prevotella copri correlates with enhanced susceptibility to arthritis,” eLife, 2013. doi:http://dx.doi.org/10.7554/eLife.01202.
[80]A.L. Perry and P.A. Lambert, “Propionibacterium acnes,” Letters in Applied Microbiology, vol. 42, pp. 185-188, 2006.
[81]D. Marriott, D. Stark, and J. Harkness, “Veillonella parvula discitis and secondary bacteremia: a rare infection complicating endoscopy and colonoscopy?” Journal of Clinical Microbiology, vol. 45, pp. 672-674, 2007.
[82]C.J. Adler et al., “Sequencing ancient calcified dental plaque shows changes in oral microbiota with dietary shifts of the Neolithic and Industrial revolutions,” Nature Genetics, vol. 45, pp. 450-456, 2013.


QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top