跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.43) 您好!臺灣時間:2026/09/01 16:48
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:林季緯
研究生(外文):LIN, JI-WEI
論文名稱:基於多任務學習的多目標分類問題 - 以自然語言處理問題為例
論文名稱(外文):Multi-label Classification with Multi-task learning - A case study of Natural Language Processing
指導教授:劉建良劉建良引用關係
指導教授(外文):Liu, Chien-Liang
口試委員:劉建良巫木誠陳勝一
口試委員(外文):Liu, Chien-Liang
學位類別:碩士
校院名稱:國立交通大學
系所名稱:工業工程與管理系所
學門:工程學門
學類:工業工程學類
論文種類:學術論文
論文出版年:2019
畢業學年度:108
語文別:英文
論文頁數:38
中文關鍵詞:深度學習機器學習多任務學習⾃然語⾔處理
外文關鍵詞:Deep learningMachine learningMulti-task learningNatural language processing
相關次數:
  • 被引用被引用:0
  • 點閱點閱:744
  • 評分評分:
  • 下載下載:12
  • 收藏至我的研究室書目清單書目收藏:0
⾃然語⾔處理問題⼀直以來都是許多學者研究的議題,例如: 偵測垃圾郵件、⽂本
分類、影評分級和惡意評論偵測等等。近年來隨著⼈⼯智慧的發展,以及電腦設備的
提升,應⽤⼈⼯智慧來處理⾃然語⾔處理問題的⽅法也越來越多。像是深度學習模型
利⽤⼤量的訓練資料去⾃動抽取特徵,並擷取出重要特徵進⾏分類。⽬前已被廣泛應
⽤在語⾳辨識、物件偵測以及圖像辨識等領域,並且獲得⾮常好的效果。多⽬標分類
的問題是指⼀筆資料可能會有多個標籤答案,像是在⾃然語⾔處理問題中惡意評論的
偵測,可能這則評論同時包含了歧視和威脅,另則句⼦則是包含了威脅、猥褻與髒話
字眼。這種問題稱之為多⽬標分類。多任務學習則是將單⼀標籤視為單⼀的任務,透
過多個相關的任務⼀起學習,透過共享的參數讓他們在學習中互相分享他們學習到的
資訊,來讓模型的各個任務準確度提升,並縮短模型訓練時間。本研究將針對多⽬標
分類的問題(Multi-label classification) 使⽤多任務學習(Multi-task learning) 來解決。
本論⽂使⽤了⼀個惡意評論偵測分類的資料集,這個資料集中每筆資料可能有: 惡毒、
極度惡毒、猥褻、威脅、歧視和侮辱等六種標籤,以及另外兩個類似的y然語⾔處理
資料集。並使⽤多任務學習的架構來解決此問題。由於⾃然語⾔處理問題屬於時間序
列的問題,⽽在深度學習模型中Recurrent Neural Networks (RNN) 的架構較能處理時
間序列的問題,所以我們透過bidirectional GRU 來建構我們的模型,並在bidirectional
GRU 後加⼊⼀層CNN 做更細部的特徵萃取。本論⽂將多任務學習的架構應⽤於三個
⾃然語⾔處理的資料集,對⽐其他深度學習模型,我們的模型獲得⾮常好的效果。
Natural language processing (NLP) is always an important research topic. Many
NLP problems such as spam filtering, text classification, movie rating, and toxic comment
detection, are present in our daily life. In recent years, with the development of
deep learning and the improvement of computing power, many studies have applied deep
learning to deal with NLP problems. Deep learning model uses a large amount of training
data to automatically extract features, giving a base to extract important and discriminative
features from data, and achieve promising results. Therefore, deep learning has
been widely used in speech recognition, object detection and image recognition, and has
achieved good results. Multi-label classification problem refers to the problem that each
data sample may have multiple labels. In the toxic comments classification problem, a
comment may consists of both identity hates and threats, while the other comments may
contain threats, obscene and toxic. This thesis proposes to use multi-task learning to deal
with multi-label classification problems, in which we treat each single label as a single task.
For each task, it is possible to improve the performance by sharing representation between
tasks. This thesis uses three different NLP classification datasets to conduct experiments.
We use our proposed multi-task learning architecture to deal with this problem. Since
the natural language processing problem is a sequence problem, and the Recurrent Neural
Networks (RNN) architecture is more capable of processing sequence problems in the
deep learning model. Therefore, we construct our model through bidirectional GRU, and
then apply CNN for further feature learning. This thesis conducts experiments on three
datasets. The experimental results indicate that our proposed method outperforms the
other alternatives.
Contents
1 Introduction 1
1.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1
1.2 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.3 Objective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2 Related Work 5
2.1 Single-task learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
2.2 Multi-task learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2.1 Regularization based multi-task learning . . . . . . . . . . . . . . . 9
2.2.2 Deep Multi-task Learning . . . . . . . . . . . . . . . . . . . . . . . 9
2.3 Transfer learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
2.3.1 Instance-based transfer learning . . . . . . . . . . . . . . . . . . . . 11
2.3.2 Feature-representation transfer learning . . . . . . . . . . . . . . . 11
2.3.3 Parameter-transfer learning . . . . . . . . . . . . . . . . . . . . . . 12
2.4 Text mining . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
2.5 Convolution Neural Network . . . . . . . . . . . . . . . . . . . . . . . . . . 13
2.6 Recurrent Neural Network . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
3 Proposed Method 15
3.1 Data Pre-processing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
3.2 Multi-label Text Classification . . . . . . . . . . . . . . . . . . . . . . . . . 18
3.2.1 Word embedding . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
3.2.2 Multi-task learning . . . . . . . . . . . . . . . . . . . . . . . . . . . 18
3.2.3 Loss function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
4 Experimental Results 21
4.1 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21
4.2 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
4.3 Experimental settings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
4.4 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25
5 Discussion 28
5.1 Word Embedding pre-trained model . . . . . . . . . . . . . . . . . . . . . . 28
5.2 Deep Learning Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . 29
5.3 Loss Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
6 Conclusions and Future Work 32
References 34
References
[1] Monika Hengstler, Ellen Enkel, and Selina Duelli. “Applied artificial intelligence
and trust—The case of autonomous vehicles and medical assistance devices”. In:
Technological Forecasting and Social Change 105 (2016), pp. 105–120.
[2] Takeshi Nakazawa and Deepak V Kulkarni. “Wafer map defect pattern classification
and image retrieval using convolutional neural network”. In: IEEE Transactions on
Semiconductor Manufacturing 31.2 (2018), pp. 309–314.
[3] Murtaza Roondiwala, Harshal Patel, and Shraddha Varma. “Predicting stock prices
using LSTM”. In: International Journal of Science and Research (IJSR) 6.4 (2017),
pp. 1754–1756.
[4] Iulian V Serban et al. “A deep reinforcement learning chatbot”. In: arXiv preprint
arXiv:1709.02349 (2017).
[5] Lorien Pratt and Sebastian Thrun. Second special issue on inductive transfer-Guest
editors’ introduction. 1997.
[6] Rich Caruana. “Multitask learning”. In: Machine learning 28.1 (1997), pp. 41–75.
[7] Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. “Deep learning”. In: nature
521.7553 (2015), p. 436.
[8] Yann LeCun, Yoshua Bengio, et al. “Convolutional networks for images, speech, and
time series”. In: The handbook of brain theory and neural networks 3361.10 (1995),
p. 1995.
[9] Kilian Weinberger et al. “Feature hashing for large scale multitask learning”. In:
arXiv preprint arXiv:0902.2206 (2009).
[10] Ronan Collobert and Jason Weston. “A unified architecture for natural language
processing: Deep neural networks with multitask learning”. In: Proceedings of the
25th international conference on Machine learning. ACM. 2008, pp. 160–167.
[11] Muhammad Ghifary et al. “Domain generalization for object recognition with multitask
autoencoders”. In: Proceedings of the IEEE international conference on computer
vision. 2015, pp. 2551–2559.
[12] Jian Zhang, Zoubin Ghahramani, and Yiming Yang. “Learning multiple related
tasks using latent independent component analysis”. In: Advances in neural information
processing systems. 2006, pp. 1585–1592.
[13] Rie Kubota Ando and Tong Zhang. “A framework for learning predictive structures
from multiple tasks and unlabeled data”. In: Journal of Machine Learning Research
6.Nov (2005), pp. 1817–1853.
[14] Daxiang Dong et al. “Multi-task learning for multiple language translation”. In:
Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics
and the 7th International Joint Conference on Natural Language Processing
(Volume 1: Long Papers). Vol. 1. 2015, pp. 1723–1732.
[15] Hua Wang et al. “High-order multi-task feature learning to identify longitudinal
phenotypic markers for alzheimer’s disease progression prediction”. In: Advances in
neural information processing systems. 2012, pp. 1277–1285.
[16] Zhizheng Wu et al. “Deep neural networks employing multi-task learning and stacked
bottleneck features for speech synthesis”. In: 2015 IEEE international conference
on acoustics, speech and signal processing (ICASSP). IEEE. 2015, pp. 4460–4464.
[17] Theodoros Evgeniou and Massimiliano Pontil. “Regularized multi–task learning”.
In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge
discovery and data mining. ACM. 2004, pp. 109–117.
[18] Guillaume Obozinski, Ben Taskar, and Michael I Jordan. “Joint covariate selection
and joint subspace selection for multiple classification problems”. In: Statistics and
Computing 20.2 (2010), pp. 231–252.
[19] Yongxin Yang and Timothy M Hospedales. “Trace norm regularised deep multi-task
learning”. In: arXiv preprint arXiv:1606.04038 (2016).
[20] Daniel Dahlmeier and Hwee Tou Ng. “Grammatical error correction with alternating
structure optimization”. In: Proceedings of the 49th Annual Meeting of the
Association for Computational Linguistics: Human Language Technologies-Volume
1. Association for Computational Linguistics. 2011, pp. 915–923.
[21] Rie Kubota Ando. “Applying alternating structure optimization to word sense disambiguation”.
In: Proceedings of the Tenth Conference on Computational Natural
Language Learning. Association for Computational Linguistics. 2006, pp. 77–84.
[22] Jiayu Zhou, Jianhui Chen, and Jieping Ye. “Clustered multi-task learning via alternating
structure optimization”. In: Advances in neural information processing
systems. 2011, pp. 702–710.
[23] Lei Han and Yu Zhang. “Learning tree structure in multi-task learning”. In: Proceedings
of the 21th ACM SIGKDD International Conference on Knowledge Discovery
and Data Mining. ACM. 2015, pp. 397–406.
[24] Jonathan Baxter. “A Bayesian/information theoretic model of learning to learn via
multiple task sampling”. In: Machine learning 28.1 (1997), pp. 7–39.
[25] Long Duong et al. “Low resource dependency parsing: Cross-lingual parameter sharing
in a neural network parser”. In: Proceedings of the 53rd Annual Meeting of the
Association for Computational Linguistics and the 7th International Joint Conference
on Natural Language Processing (Volume 2: Short Papers). Vol. 2. 2015,
pp. 845–850.
[26] Sinno Jialin Pan and Qiang Yang. “A survey on transfer learning”. In: IEEE Transactions
on knowledge and data engineering 22.10 (2010), pp. 1345–1359.
[27] Omer Levy and Yoav Goldberg. “Dependency-based word embeddings”. In: Proceedings
of the 52nd Annual Meeting of the Association for Computational Linguistics
(Volume 2: Short Papers). Vol. 2. 2014, pp. 302–308.
[28] Tomas Mikolov et al. “Distributed representations of words and phrases and their
compositionality”. In: Advances in neural information processing systems. 2013,
pp. 3111–3119.
[29] Jeffrey Pennington, Richard Socher, and Christopher Manning. “Glove: Global vectors
for word representation”. In: Proceedings of the 2014 conference on empirical
methods in natural language processing (EMNLP). 2014, pp. 1532–1543.
[30] Tomas Mikolov et al. “Advances in pre-training distributed word representations”.
In: arXiv preprint arXiv:1712.09405 (2017).
[31] Wenpeng Yin et al. “Comparative study of CNN and RNN for natural language
processing”. In: arXiv preprint arXiv:1702.01923 (2017).
[32] Tomáš Mikolov et al. “Recurrent neural network based language model”. In: Eleventh
annual conference of the international speech communication association. 2010.
[33] Sepp Hochreiter and Jürgen Schmidhuber. “Long short-term memory”. In: Neural
computation 9.8 (1997), pp. 1735–1780.
[34] Xiang Zhang, Junbo Zhao, and Yann LeCun. “Character-level convolutional networks
for text classification”. In: Advances in neural information processing systems.
2015, pp. 649–657.
[35] Kyunghyun Cho et al. “Learning phrase representations using RNN encoder-decoder
for statistical machine translation”. In: arXiv preprint arXiv:1406.1078 (2014).
[36] Qingyu Zhou et al. “Selective encoding for abstractive sentence summarization”. In:
arXiv preprint arXiv:1704.07073 (2017).
[37] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. “Imagenet classification
with deep convolutional neural networks”. In: Advances in neural information processing
systems. 2012, pp. 1097–1105.
[38] Kaiming He et al. “Deep residual learning for image recognition”. In: Proceedings of
the IEEE conference on computer vision and pattern recognition. 2016, pp. 770–778.
[39] Nitesh V Chawla. “Data mining for imbalanced datasets: An overview”. In: Data
mining and knowledge discovery handbook. Springer, 2005, pp. 853–867.
[40] Ahmed Elnaggar et al. “Stop Illegal Comments: A Multi-Task Deep Learning Approach”.
In: Proceedings of the 2018 Artificial Intelligence and Cloud Computing
Conference. ACM. 2018, pp. 41–47.
[41] Betty van Aken et al. “Challenges for toxic comment classification: An in-depth
error analysis”. In: arXiv preprint arXiv:1809.07572 (2018).
[42] Alon Rozental and Daniel Fleischer. “Amobee at semeval-2018 task 1: GRU neural
network with a CNN attention mechanism for sentiment classification”. In: arXiv
preprint arXiv:1804.04380 (2018).
[43] Saurabh Srivastava, Prerna Khurana, and Vartika Tewari. “Identifying aggression
and toxicity in comments using capsule network”. In: Proceedings of the First Workshop
on Trolling, Aggression and Cyberbullying (TRAC-2018). 2018, pp. 98–105.
[44] Isuru Gunasekara and Isar Nejadgholi. “A Review of Standard Text Classification
Practices for Multi-label Toxicity Identification of Online Content”. In: Proceedings
of the 2nd Workshop on Abusive Language Online (ALW2). 2018, pp. 21–25.
[45] Siyuan Li. “Application of recurrent neural networks in toxic comment classification”.
PhD thesis. UCLA, 2018.
[46] Yanghoon Kim, Hwanhee Lee, and Kyomin Jung. “AttnConvnet at SemEval-2018
Task 1: attention-based convolutional neural networks for multi-label emotion classification”.
In: arXiv preprint arXiv:1804.00831 (2018).
[47] Hardik Meisheri and Lipika Dey. “TCS Research at SemEval-2018 Task 1: Learning
Robust Representations using Multi-Attention Architecture”. In: Proceedings of The
12th International Workshop on Semantic Evaluation. 2018, pp. 291–299.
[48] Ji Ho Park, Peng Xu, and Pascale Fung. “PlusEmo2Vec at SemEval-2018 Task
1: Exploiting emotion knowledge from emoji and# hashtags”. In: arXiv preprint
arXiv:1804.08280 (2018).
[49] Tomas Mikolov et al. “Efficient estimation of word representations in vector space”.
In: arXiv preprint arXiv:1301.3781 (2013).
[50] Armand Joulin et al. “FastText.zip: Compressing text classification models”. In:
arXiv preprint arXiv:1612.03651 (2016).
[51] Tsung-Yi Lin et al. “Focal loss for dense object detection”. In: Proceedings of the
IEEE international conference on computer vision. 2017, pp. 2980–2988.
連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊