跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.64) 您好!臺灣時間:2026/08/26 09:08
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:黃盈禎
研究生(外文):Ying-Chen Huang
論文名稱:以刪減叢集法找尋蛋白質磷酸化鄰近區域之特定樣本
論文名稱(外文):Q-Motif: Phosphorylation Motif Finding by Subtractive Clustering
指導教授:鍾翊方
指導教授(外文):I-Fang Chung
學位類別:碩士
校院名稱:國立陽明大學
系所名稱:生物醫學資訊研究所
學門:工程學門
學類:生醫工程學類
論文種類:學術論文
論文出版年:2011
畢業學年度:99
語文別:中文
論文頁數:82
中文關鍵詞:激酶激酶激酶激酶激酶
外文關鍵詞:phosphorylation motifsubtractive clusteringkinasemotif findingdensity
相關次數:
  • 被引用被引用:0
  • 點閱點閱:214
  • 評分評分:
  • 下載下載:10
  • 收藏至我的研究室書目清單書目收藏:0
隨著生物定序技術的進步,已完成定序的蛋白質序列數目快速地大量增加,發展一套計算演算法找尋蛋白質磷酸化鄰近區域之特定樣本(以下均稱motif),應用於生物上成為相當重要的議題。本研究中,我們提出一個新的非監督式方法(unsupervised)從大量資料中萃取出磷酸化motif,我們稱作為motif快速探測器(Quick Motif Finder,簡稱Q-Motif)。Q-Motif採用刪減叢集法,從大量磷酸化資料中發現隱藏其中有意義的統計資訊,以及找出motif。透過刪減叢集法能夠將相似的磷酸化資料群聚成一個小組,並分析和辨識可能磷酸化的motif;而這些可能磷酸化的motif會藉由算分機制過濾出可能性較顯著的磷酸化motif。
我們運用Q-Motif測試幾組既有和新的資料,跟一個有名且常被使用的方法Motif-X (Schwartz and Gygi, 2005)並應用其算分方法將結果做比較。大約有80%的案例能夠辨識和Motif-X完全相同且統計上顯著的motif,並且Q-Motif還能額外發掘新的motif;而這些額外的新發現,我們幾乎都能夠從文獻中驗證它們確實是磷酸化motif。由於我們認為motif-x的單一位置給分和資料移除方式,無法完整確認整個motif在磷酸化資料中有夠強烈的表現,因此我們也針對計算分數方面重新改良。比較兩種算分結果,某些原本給分較高的motif,透過新的計分法計算分數變低;可以發現資料移除再做分數計算的概念,很明顯影響最終motif的呈現。從結果而言,我們認為此研究清楚地建立和改良一個相當不錯尋找motif的新方法。
我們的研究中提出採用疊代運算方法,對磷酸化數據資料做探索式數據分析。透過幾組實際磷酸化資料和二組人工合成資料測試,能夠證實Q-Motif找尋motif的效能。此外,此方法相當具有普遍性,因此還能夠應用在發掘除了磷酸化外,其他不同型式的motif。

Development of computational algorithms to discover biologically relevant phosphorylation motifs is becoming more and more important with the rapid increase in the proteomic sequences. Here we present a novel unsupervised method, called Quick Motif Finder, (in short, Q-Motif) to extract phosphorylation motifs. Q-Motif adopts the subtractive clustering algorithm to discover motifs exploiting statistical information hidden in phosphorylated data. The phosphorylated data is clustered into homogeneous groups, which are analyzed to identify candidate motifs. These candidate motifs are then filtered to find actual motifs with statistically significant motif scores.
We have applied Q-Motif on several new and existing data sets and compared its performance with a well known state-of-the-art method, Motif-X (Schwartz, 2005). In 80% cases Q-Motif could identify all statistically significant motifs extracted by the state-of-the-art method. In addition to this, Q-Motif uncovers several novel motifs. For most of these additional motifs, we have verified their existence using evidences from the literature. Because motif scores were calculated by Motif-X scoring method, we consider it could not represent the statistical significant of every motif extracted in original phosphorylated data. Therefore, we re-define the definition of scoring method and compare the results with different scoring. These clearly establish the excellent motif discovery ability of our algorithm.
An iterative algorithm proposed here uses exploratory data analysis to discover motifs from phosphorylated data. The effectiveness of Q-Motif has been demonstrated using several real data sets as well as using a synthetic data set. The method is quite general in nature and can be used to find other type of motifs also.

口試委員會審定書 #
中文摘要 i
ABSTRACT ii
CONTENTS iii
LIST OF FIGURES v
LIST OF TABLES vi
Chapter 1 緒論 1
1.1 研究背景…………………. 1
1.2 相關研究 3
1.3 研究動機 6
Chapter 2 材料與方法 9
2.1 資料來源………………………. 9
2.2 前景資料 (Foreground data) 11
2.2.1 F-Human centered on serine (“〖FH〗^S”) 11
2.2.2 F-Human centered on serine/threonine (“〖FH〗^ST”) 11
2.2.3 F-motif-x offered (“FM”) 11
2.2.4 F-Synthetic (“FS”) 12
2.2.5 F-Mass Spectrometry (“FMS”) 12
2.3 背景資料 (Background data) 13
2.3.1 B-Human centered on serine (“〖BH〗^S”) 13
2.3.2 B-Human centered on serine/threonine (“〖BH〗^ST”) 13
2.3.3 B-Mouse (“BM”) 13
2.4 Q-Motif流程 15
2.4.1 步驟一:建造分數矩陣將序列資料編碼 18
2.4.2 步驟二:藉由群聚尋找有潛力的motif並決定候選motif清單 19
2.4.3 步驟三:透過給分評估最後的motif清單 26
Chapter 3 研究結果 29
3.1 實驗一:分析特定激酶的人類前景資料和背景資料 (〖FH〗^S 和 〖BH〗^S) 30
3.1.1 以Serine為中心,個別分析PKA 、PKC、CK2、CDK激酶(〖FH〗_PKA^S 、〖FH〗_PKC^S 、〖FH〗_CK2^S 、〖FH〗_CDK^S 和 〖BH〗^S) 31
3.1.2 以Serine/Threonine為中心,個別分析PKA 、PKC、CK2、CDK激酶 (〖FH〗_PKA^ST 、〖FH〗_PKC^ST 、〖FH〗_CK2^ST 、〖FH〗_CDK^ST 和 〖BH〗^ST) 32
3.2 實驗二:分析Motif-X文獻中測試的實驗資料 (FM 和 〖BH〗^S) 36
3.3 實驗三:分析老鼠蛋白質質譜儀資料 (FMS 和 BM) 38
3.4 實驗四:分析Motif-X文獻中提供的合成資料 (FS 和 〖BH〗^S) 43
Chapter 4 討論 47
4.1 Q-Motif和F-Motif結果比較 47
4.2 利用文獻註解驗證磷酸化Motif 53
4.3 分析Motif在生物體內的意義 55
4.4 分析Motif在生物功能和註解的專一性 65
Chapter 5 結論 68
參考文獻 69
附錄 72


Amanchy, R., Periaswamy, B., Mathivanan, S., Reddy, R., Tattikota, S.G., and Pandey, A. (2007). A curated compendium of phosphorylation motifs. Nat Biotechnol 25, 285-286.

Bailey, T.L., and Elkan, C. (1995). The value of prior knowledge in discovering motifs with MEME. Proc Int Conf Intell Syst Mol Biol 3, 21-29.

Claverie, J.M., and Audic, S. (1996). The statistical significance of nucleotide position-weight matrix matches. Comput Appl Biosci 12, 431-439.

Dinkel, H., Chica, C., Via, A., Gould, C.M., Jensen, L.J., Gibson, T.J., and Diella, F. (2011). Phospho.ELM: a database of phosphorylation sites--update 2011. Nucleic Acids Res 39, D261-267.

Durbin, R.S.R.E., Anders Krogh, Graeme Mitchison (1998). Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids.

Finn, R.D., Mistry, J., Tate, J., Coggill, P., Heger, A., Pollington, J.E., Gavin, O.L., Gunasekaran, P., Ceric, G., Forslund, K., et al. (2010). The Pfam protein families database. Nucleic Acids Res 38, D211-222.

Han, S.J., Vaccari, S., Nedachi, T., Andersen, C.B., Kovacina, K.S., Roth, R.A., and Conti, M. (2006). Protein kinase B/Akt phosphorylation of PDE3A and its role in mammalian oocyte maturation. EMBO J 25, 5716-5725.

He, Z., Yang, C., Guo, G., Li, N., and Yu, W. (2011). Motif-All: discovering all phosphorylation motifs. BMC Bioinformatics 12 Suppl 1, S22.

Khaled Hammouda, F.K. (2000). A Comparative Study of Data Clustering Techniques. Computer and Information Science 625, 1-21.

Kitagawa, M., Higashi, H., Jung, H.K., Suzuki-Takahashi, I., Ikeda, M., Tamai, K., Kato, J., Segawa, K., Yoshida, E., Nishimura, S., et al. (1996). The consensus motif for phosphorylation by cyclin D1-Cdk4 is different from that for phosphorylation by cyclin A/E-Cdk2. EMBO J 15, 7060-7069.

Krogh, A., Larsson, B., von Heijne, G., and Sonnhammer, E.L. (2001). Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes. J Mol Biol 305, 567-580.

Malbon, C.C. (2005). G proteins in development. Nat Rev Mol Cell Biol 6, 689-701.
Manning, G., Whyte, D.B., Martinez, R., Hunter, T., and Sudarsanam, S. (2002). The protein kinase complement of the human genome. Science 298, 1912-1934.

Munoz, E., Zubiaga, A.M., and Huber, B.T. (1991). Tyrosine protein phosphorylation is required for protein kinase C-mediated proliferation in T cells. FEBS letters 279, 319-322.

Nishikawa, K., Toker, A., Johannes, F.J., Songyang, Z., and Cantley, L.C. (1997). Determination of the specific substrate sequence motifs of protein kinase C isozymes. The Journal of biological chemistry 272, 952-960.

Pearson, R.B., and Kemp, B.E. (1991). Protein kinase phosphorylation site sequences and consensus specificity motifs: tabulations. Methods Enzymol 200, 62-81.

Pessin, J.E., and Saltiel, A.R. (2000). Signaling pathways in insulin action: molecular targets of insulin resistance. The Journal of Clinical Investigation 106, 165-169.

Ramars Amanchy, K.K., Suresh Mathivanan, Balamurugan Periaswamy, Raghunath Reddy, Wan-Hee Yoon, Jos, and Joore, M.A.B., Leslie Cope and Akhilesh Pandey (2011). Identification of Novel Phosphorylation Motifs Through an Integrative Computational and Experimental Analysis of the Human Phosphoproteome. Proteomics & Bioinformatics 022-035 (2011) - 022.

Rascon, A., Degerman, E., Taira, M., Meacci, E., Smith, C.J., Manganiello, V., Belfrage, P., and Tornqvist, H. (1994). Identification of the phosphorylation site in vitro for cAMP-dependent protein kinase on the rat adipocyte cGMP-inhibited cAMP phosphodiesterase. J Biol Chem 269, 11962-11966.

Ren, S., Yang, G., He, Y., Wang, Y., Li, Y., and Chen, Z. (2008). The conservation pattern of short linear motifs is highly correlated with the function of interacting protein domains. BMC Genomics 9, 452.

Rigoutsos, I., and Floratos, A. (1998). Combinatorial pattern discovery in biological sequences: The TEIRESIAS algorithm. Bioinformatics 14, 55-67.
Ritz, A., Shakhnarovich, G., Salomon, A.R., and Raphael, B.J. (2009). Discovery of phosphorylation motif mixtures in phosphoproteomics data. Bioinformatics 25, 14-21.
Rodriguez, M., Li, S.S., Harper, J.W., and Songyang, Z. (2004). An oriented peptide array library (OPAL) strategy to study protein-protein interactions. J Biol Chem 279, 8802-8807.

Schwartz, D., and Gygi, S.P. (2005). An iterative statistical approach to the identification of protein phosphorylation motifs from large-scale data sets. Nat Biotechnol 23, 1391-1398.

Shah, O.J., Ghosh, S., and Hunter, T. (2003). Mitotic regulation of ribosomal S6 kinase 1 involves Ser/Thr, Pro phosphorylation of consensus and non-consensus sites by Cdc2. The Journal of biological chemistry 278, 16433-16442.

Sharma, P., Veeranna, Sharma, M., Amin, N.D., Sihag, R.K., Grant, P., Ahn, N., Kulkarni, A.B., and Pant, H.C. (2002). Phosphorylation of MEK1 by cdk5/p35 down-regulates the mitogen-activated protein kinase pathway. The Journal of biological chemistry 277, 528-534.

Songyang, Z., Lu, K.P., Kwon, Y.T., Tsai, L.H., Filhol, O., Cochet, C., Brickey, D.A., Soderling, T.R., Bartleson, C., Graves, D.J., et al. (1996). A structural basis for substrate specificities of protein Ser/Thr kinases: primary sequence preference of casein kinases I and II, NIMA, phosphorylase kinase, calmodulin-dependent kinase II, CDK5, and Erk1. Mol Cell Biol 16, 6486-6493.

Steuer, R., Kurths, J., Daub, C.O., Weise, J., and Selbig, J. (2002). The mutual information: detecting and evaluating dependencies between variables. Bioinformatics 18 Suppl 2, S231-240.

Yi-Cheng Chen, K.A., Chu-Wen Yang, Yao-Tsung Wang, Nikhil R. Pal, and I-Fang Chung (2011). Discovery of Protein Phosphorylation Motifs Through Exploratory Data Analysis. PLoS One.

Zanivan, S., Gnad, F., Wickstrom, S.A., Geiger, T., Macek, B., Cox, J., Fassler, R., and Mann, M. (2008). Solid tumor proteome and phosphoproteome analysis by high resolution mass spectrometry. J Proteome Res 7, 5314-5326.

連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關論文
 
無相關期刊