跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.126) 您好!臺灣時間:2026/08/28 02:52
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:應效軒
研究生(外文):Keith Ying
論文名稱:學生標註行為分類方法之研究設計
論文名稱(外文):Designing Chromosome-based Comparison Algorithms for Clustering Students' Annotation Behaviors
指導教授:賀嘉生賀嘉生引用關係
指導教授(外文):Jia-Sheng Heh
學位類別:碩士
校院名稱:中原大學
系所名稱:資訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2011
畢業學年度:99
語文別:英文
論文頁數:69
中文關鍵詞:註記閱讀學習註記行為基於染色體的分群
外文關鍵詞:F-measureChromosome-basedClusteringAnnotation BehaviorLearningReadingAnnotation
相關次數:
  • 被引用被引用:0
  • 點閱點閱:257
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
從以前到現在,即使有愈來愈多的工具可以輔助使用者進行閱讀或是學習,對於使用者來說標註仍然是最重要的工具。標註屬於私人的註記,如果有人的註記相同或是類似的話,也許在閱讀或是學習上有著相同的想法。因此,我們把使用者的註記當作是一個人,就像是每個人會有不同的染色體(Chromosome),藉由用使用者的註記來分類使用者進而分析使用者的註記行為。

在本論文中,首先我們將使用者的資料,藉由兩個門檻值,來降低使用者註記資料的大小,因為使用者的註記資料其實很龐大,不利於分析。第二,介紹4個基於染色體的方法,這些方法將會把使用者的註記資料,轉成2維的平面資料,在這2維的平面上面可以觀察使用者的差異。第三,用分群的方法將2維的平面資料做分群。最後,將分群的結果與專家以及MATLAB所做的分群的結果用用F-measure進行比較。

結果顯示,我們所做的4個基於染色體的方法,在準確率來說,有些的確比單純使用MATLAB分群還要好,這意味著將資料先轉換成2維的平面資料進行分群,結果比不轉換來的好。在整體所花的時間來說,4個基於染色體的方法都比MATLAB來的快,這意味著4個基於染色體的方法需要轉換的時間,轉換後再分群,仍然比MATLAB不轉換還來的快。


From the past to now, even though we have many tools for enhance our reading or learning, the annotation is still important to users. The annotation is a very private note, which is made by user for enhance his/her reading or learning, so that if users have similar annotation may have same thought on reading or learning. Therefore, we treat the annotation as a presentation of a person such as the chromosome of a human, and then classify the users by their annotation data for analysis the users’ annotation behaviors.

In this research, first we created two thresholds to reduce the original annotation data. The reduction is because the original annotation data is a long data sequence and hard to analysis. Second, we developed four chromosome-based approaches for the reduced annotation data and generate 2-dimensional axis data to present the difference between users. Third, we used the clustering methods to clustering the 2-dimensional axis data. Finally, we take the results of clusters, which are made by expert and MATLAB, to compare with the results with our approaches, and use the F-measure to evaluate the performance of these methods.

Results show some of our approaches did better than MATLAB, and that is means the accuracy of clustering with transformed data is better than clustering without transformation. And another finding is that even though the clustering with transformed data need the transformation time, but the clustering with transformed data is also quick than the clustering with original data.


Index
摘要 I
Abstract II
致謝 III
Index IV
Figure index VI
Table index VIII
1. Introduction 1
1.1 Motivation 1
1.2 Goals and Contributions 1
1.3 Chapter Descriptions 1
2. Research Background 3
2.1 Annotation 3
2.2 Bio-inspired Computing 4
2.3 Clustering 4
2.3.1 Types of Clustering 5
2.3.1.1 Hierarchical Clustering 5
2.3.1.2 Partitioning Clustering 6
2.3.2 Similarity 7
2.3.3 Linkage Criterion 8
2.4 Performance Measures 9
3. Chromosome-based Approaches 11
3.1 Annotation Pattern 11
3.2 Chromosome-based Annotation Behavior Comparison 13
3.2.1 Chromosome-based Standard Approach 13
3.2.2 Chromosome-based Quantitative Approach 14
3.2.3 Chromosome-based Cosine Approach 14
3.2.4 Chromosome-based Diffusion Approach 15
4. Algorithm 17
4.1 Annotation Pattern Transformation 17
4.2 Chromosome-based Approaches 18
4.2.1 Standard and Quantitative Approaches 18
4.2.2 Cosine Approach 20
4.2.3 Diffusion Approach 21
4.3 Clustering Algorithm 22
4.4 Performance Evaluation 25
5. Results 27
5.1 User Data 27
5.2 Comparisons 27
5.2.1 Algorithms Cost Time Comparisons 27
5.2.2 Performance Comparisons 29
5.2.2.1 Users Clustering 29
5.2.2.2 Cluster Relations 35
5.3 Discussions 39
6. Conclusions 41
6.1 Summary 41
6.2 Future Works 41
References 43
Appendix 47
Appendix A: F(0.5) of all approaches in each threshold of users clustering 47
Appendix B: F(2) of all approaches in each threshold of users clustering 50
Appendix C: Recall of all approaches in each threshold of cluster relations 53
Appendix D: F(0.5) of all approaches in each threshold of cluster relations 56
Appendix E: F(2) of all approaches in each threshold of cluster relations 59


Figure index
Figure 1 Data set 5
Figure 2 Hierarchical tree 5
Figure 3 Data set 6
Figure 4 Clustered data 7
Figure 5 Result of the clustering 7
Figure 6 Original data 11
Figure 7 Annotated data 11
Figure 8 Transformed annotation data 11
Figure 9 Transformed data 12
Figure 10 Example of sentence threshold 12
Figure 11 Paragraph annotation data example 12
Figure 12 Example of paragraph threshold 13
Figure 13 Transformed annotation data of User 1 and User 2 13
Figure 14 Transformed annotation data of User 1 and User 3 13
Figure 15 Diffusion of water droplets 15
Figure 16 Annotation data of user 1 and user 2 15
Figure 17 Chromosome of user 1 and user 2 16
Figure 18 Pseudo code of annotation data transformation 18
Figure 19 Transformed annotation data of user 1 and user 2 18
Figure 20 Transformed annotation data of user 1 and user 3 19
Figure 21 Transformed annotation data of 4 users 22
Figure 22 Hierarchical tree of standard approach 23
Figure 23 Hierarchical tree of quantitative approach 23
Figure 24 Hierarchical tree of cosine approach 24
Figure 25 Hierarchical tree of diffusion approach 24
Figure 26 Precision of MATLAB approach in each threshold of users clustering 30
Figure 27 2-dimension Precision of MATLAB approach in each threshold of users clustering 30
Figure 28 Precision of standard approach in each threshold of users clustering 31
Figure 29 Precision of quantitative approach in each threshold of users clustering 31
Figure 30 Precision of cosine approach in each threshold of users clustering 32
Figure 31 Precision of diffusion approach in each threshold of users clustering 32
Figure 32 Recall of MATLAB approach in each threshold of users clustering 33
Figure 33 Recall of standard approach in each threshold of users clustering 33
Figure 34 Recall of quantitative approach in each threshold of users clustering 34
Figure 35 Recall of cosine approach in each threshold of users clustering 34
Figure 36 Recall of diffusion approach in each threshold of users clustering 35
Figure 37 Precision of MATLAB approach in each threshold of cluster relations 36
Figure 38 2-dimansion precision of MATLAB approach in each threshold of cluster relations 37
Figure 39 Precision of standard approach in each threshold of cluster relations 37
Figure 40 Precision of quantitative approach in each threshold of cluster relations 38
Figure 41 Precision of cosine approach in each threshold of cluster relations 38
Figure 42 Precision of diffusion approach in each threshold of cluster relations 39
Figure 43 F(0.5) of MATLAB approach in each threshold of users clustering 47
Figure 44 F(0.5) of standard approach in each threshold of users clustering 47
Figure 45 F(0.5) of quantitative approach in each threshold of users clustering 48
Figure 46 F(0.5) of cosine approach in each threshold of users clustering 48
Figure 47 F(0.5) of diffusion approach in each threshold of users clustering 49
Figure 48 F(2) of MATLAB approach in each threshold of users clustering 50
Figure 49 F(2) of standard approach in each threshold of users clustering 50
Figure 50 F(2) of quantitative approach in each threshold of users clustering 51
Figure 51 F(2) of Cosine approach in each threshold of users clustering 51
Figure 52 F(2) of diffusion approach in each threshold of users clustering 52
Figure 53 recall of MATLAB approach in each threshold of cluster relations 53
Figure 54 recall of standard approach in each threshold of cluster relations 53
Figure 55 recall of quantitative approach in each threshold of cluster relations 54
Figure 56 recall of cosine approach in each threshold of cluster relations 54
Figure 57 recall of diffusion approach in each threshold of cluster relations 55
Figure 58 F(0.5) of MATLAB approach in each threshold of cluster relations 56
Figure 59 F(0.5) of standard approach in each threshold of cluster relations 56
Figure 60 F(0.5) of quantitative approach in each threshold of cluster relations 57
Figure 61 F(0.5) of cosine approach in each threshold of cluster relations 57
Figure 62 F(0.5) of diffusion approach in each threshold of cluster relations 58
Figure 63 F(2) of MATLAB approach in each threshold of cluster relations 59
Figure 64 F(2) of standard approach in each threshold of cluster relations 59
Figure 65 F(2) of quantitative approach in each threshold of cluster relations 60
Figure 66 F(2) of cosine approach in each threshold of cluster relations 60
Figure 67 F(2) of diffusion approach in each threshold of cluster relations 61


Table index
Table 1 Original user annotation data 27
Table 2 Average algorithms cost time 27
Table 3 Average clustering cost time 28
Table 4 Total cost time 28
Table 5 Average precision-recall of users clustering 29
Table 6 Average F-measure of users clustering 29
Table 7 Average precision-recall of cluster relations 35
Table 8 Average F-measure of cluster relations 36
Adler, A., Gujar, A., Harrison, B.L., O'Hara, K., Sellen, A., (1998), "A diary study of work-related reading: design implications for digital reading devices", Proceedings of the SIGCHI conference on Human factors in computing systems, pp. 241-248, April 18-23, Los Angeles, California, United States
Abdeslam, D. O., Wira, P., Merckle, J., Flieller, D., Chapuis, Y.-A., (2007), "A Unified Artificial Neural Network Architecture for Active Power Filters", IEEE Transactions on Industrial Electronics, 54(1), pp. 61-76
Ball, E., Franks, H., Jenkins, J., McGrath, M., Leigh, J., (2009), "Annotation is a valuable tool to enhance learning and assessment in student essays", Nurse Education Today, 29(3), pp. 284-291
Billsus, D., Pazzani, M. J., (1998), "Learning Collaborative Information Filters", In the Proceedings of the Fifteenth International Conference on Machine Learning(ICML'98), pp. 46-54, July 24-27, 1998, Madison, Wisconsin, USA
Bien, J., Tibshirani, R., (2011), "Hierarchical Clustering With Prototypes via Minimax Linkage", Journal of the American Statistical Association, Accepted
Chang, C. K., Chen, G. D., & Chen, C. K., (2006), "Using computer-based annotation to promote the summary of reading comprehension", Communication of IICM, 9(1), pp. 41-58
Chao, P.-Y., Chen, G.-D., Chang, C.-W., (2010), "Developing a Cross-media System to Facilitate Question-Driven Digital Annotations on Paper Textbooks", Journal of Educational Technology & Society, 13(4), pp. 38-49
Cui, X., Gao, J., Potok, T. E., (2006), "A flocking based algorithm for document clustering analysis", Journal of Systems Architecture, 52(8-9), pp. 505-515
Ding, C., He, X., (2004), "K-means clustering via principal component analysis", In Proceedings of the twenty-first international conference on Machine learning (ICML '04), July 4-8, 2004, Banff, Alberta, Canada
Ding, C., Li, T., (2007), "Adaptive dimension reduction using discriminant analysis and K-means clustering", In Proceedings of the 24th international conference on Machine learning (ICML '07), June 20-24, 2007, Corvallis, Oregon
Hripcsak, G., Rothschild, A. S., (2005), "Agreement, the F-Measure, and Reliability in Information Retrieval", Journal of the American Medical Informatics Association, 12(3), pp. 296-298
Hwang, W.-Y., Hsu, G.-L., (2011), "The effects of pre-reading and sharing mechanism with annotations on learning", The Turkish Online Journal of Educational Technology(TOJET), 10(2), pp. 234-249
Hwang, W.-Y., Wang, C.-Y., Sharples, M., (2007), "A study of multimedia annotation of Web-based materials", Computers & Education, 48(4), pp. 680-699
Jain, A. K., (2010), "Data clustering: 50 years beyond K-means", Pattern Recognition Letters, 31(8), pp. 651-666
Jeon, J., Lavrenko, V., Manmatha, R., (2003), "Automatic image annotation and retrieval using cross-media relevance models", In Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval (SIGIR '03), July 28-Auguest 01, 2003, Toronto, Canada
Khan, L., Awad, M., Thuraisingham, B., (2007), "A new intrusion detection system using support vector machines and hierarchical clustering", The International Journal on Very Large Data Bases, 16(4), pp. 507-521
Kao, Y.-T., Zahara, E., (2008), "A hybrid genetic algorithm and particle swarm optimization for multimodal functions", Applied Soft Computing, 8(2), pp. 849-857
Ke, X., Li, S., Cao, D., (2011), "K-nearest Neighbors Relevance Annotation Model for Distance Education", International Journal of Distance Education Technologies, 9(1), pp. 86-100
Kurhila, J., Miettinen, M., Nokelainen, P., Floréen, P., Tirri, H., (2003), "Peer-to-Peer Learning with Open-Ended Writable Web", Proceedings of the 8th annual conference on Innovation and technology in computer science education, 35(3), pp. 173-177
Lewis, D. D. , Gale, W. A., (1994), "A sequential algorithm for training text classifiers", Proceedings of the 17th annual international ACM SIGIR conference on Research and development in information retrieval(SIGIR '94), p.3-12, July 03-06, 1994, Dublin, Ireland
Lavrenko, V., Manmatha, R., Jeon, J., (2003), "A Model for Learning the Semantics of Pictures", In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS '03), December 8-13, 2003, Whistler, British Columbia, Canada
Li, M.J., Ng, M.K., Cheung, Y.-m., Huang, J.Z., (2008), "Agglomerative Fuzzy K-Means Clustering Algorithm with Selection of Number of Clusters", IEEE Transactions on Knowledge and Data Engineering, 20(11), pp. 1519-1534
Langfelder, P., Zhang, B., Horvath, S., (2008), "Defining clusters from a hierarchical cluster tree: the Dynamic Tree Cut package for R", Bioinformatics, 24(5), pp. 719-720
Marshall, C.C., (1997), "Annotation: from paper books to the digital library", Proceedings of the second ACM international conference on Digital libraries, pp. 131-140
Morzy, T., Wojciechowski, M., Zakrzewicz, M., (1999), "Pattern-Oriented Hierarchical Clustering", Proceedings of the third East-European Symposium on Advances in Databases and Information Systems(ADBIS'99), pp. 179-190
Nahapetian, N., Analoui, M., Motlagh, M. R. J., (2009), "Training set generation using fuzzy logic and dynamic chromosome based Genetic Algorithms for plant identifiers", IEEE Symposium on Computational Intelligence in Control and Automation, pp.49-56, March 30-April 2, 2009, Nashville, TN, USA
Neisser, U., (1997), "Rising scores on intelligence tests", American Scientist, 85(5), pp. 440-447
Nokelainen, P., Miettinen, M., Kurhila, J., Floréen, P., Tirri, H., (2005), "A Shared Document-Based Annotation Tool to Support Learner-Centred Collaborative Learning", British Journal of Educational Technology, 36(5), pp. 757-770
O'Hara, K., Sellen, A., (1997), "A Comparison of Reading Paper and On-Line Documents", Proceedings of the SIGCHI conference on Human factors in computing systems, pp. 335-342, March 22-27, 1997, Atlanta, Georgia, United States
Pezzella, F., Morganti, G., Ciaschetti, G., (2008), "A genetic algorithm for the Flexible Job-shop Scheduling Problem ", Computers & Qperations Research, 35(10), pp. 3202-3212
Rasmussen, M., Karypis, G., (2004), "gcluto: An interactive clustering, visualization, and analysis system", Technical Report CSE/UMN TR 04-021, Univ. of Minnesota, Dep. of Computer Science and Engineering
Rau, P.-L. P., Chen, S.-H., Chin, Y.-T., (2004), "Developing web annotation tools for learners and instructors", Interacting with Computers, 16(2), pp. 163-181
Ritter, G., Gallegos, M. T., Gaggermeier, K., (1995), "Automatic context-sensitive karyotyping of human chromosomes basednext term on elliptically symmetric statistical distributions", Pattern Recognition, 28(6), pp. 823-831
Shepitsen, A., Gemmell, J., Mobasher, B., Burke, R., (2008), "Personalized recommendation in social tagging systems using hierarchical clustering", In Proceedings of the 2008 ACM conference on Recommender systems (RecSys '08), pp. 259-266, October 23-25, 2008, Lausanne, Switzerland
Steinbach, M.,Karypis, G., Kumar, V., (2000), "A Comparison of Document Clustering Techniques", In KDD-2000 Workshop on Text Mining, August 20, 2000, Boston, MA, USA
Stafford, R., (2010), "Constraints of Biological Neural Networks and Their Consideration in AI Applications", Advances in Artificial Intelligence, 2010, Article ID 845723, pp. 1-6
Terwilliger, J. D., Speer, M., Ott, J., (1993), "Chromosome-Based Method for Rapid Computer Simulation in Human Genetic Linkage Analysis", Genetic Epidemiology, 10(4), pp. 217-224
Wu, Z., Dong, H., Liang, Y., McKay, R.I., (2003), "A chromosome-based evaluation model for computer defense immune systems", The 2003 Congress on Evolutionary Computation - CEC 2003, 2, pp. 1363- 1369, Dec. 8-12, Canberra, Australia
Xu, R., Wunsch, D., II, (2005), "Survey of clustering algorithms", IEEE Transactions on Neural Networks, 16(3), pp. 645-678
Yeh, S.-W., Lo, J.-J., (2009), "Using online annotations to support error correction and corrective feedback", Computer & Education, 52(4), pp. 882-892

電子全文 電子全文(本篇電子全文限研究生所屬學校校內系統及IP範圍內開放)
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top