跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.60) 您好!臺灣時間:2026/08/02 02:49
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:蔡長宏
研究生(外文):Tsai, Chang-Hung
論文名稱:應用於類神經網路與機器學習之侷限型波茲曼模型處理器設計
論文名稱(外文):Restricted Boltzmann Machine (RBM) Processor Design for Neural Network and Machine Learning Applications
指導教授:李鎮宜
指導教授(外文):Lee, Chen-Yi
口試日期:2016-10-05
學位類別:博士
校院名稱:國立交通大學
系所名稱:電子研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2016
畢業學年度:105
語文別:英文
論文頁數:113
中文關鍵詞:侷限型波茲曼模型類神經網路機器學習
外文關鍵詞:Restricted Boltzmann MachineNeural NetworkMachine Learning
相關次數:
  • 被引用被引用:0
  • 點閱點閱:609
  • 評分評分:
  • 下載下載:23
  • 收藏至我的研究室書目清單書目收藏:0
近年來,機器學習已被廣泛地應用於訊號處理系統之中以提供智慧處理能力,例如AdaBoost、K-NN、mean-shift和SVM應用於資料辨識,以及HOG和SIFT應用於影像特徵擷取。而在過去幾十年之中,類神經網路在多項應用之中被視為最佳解決方案之一,有別於傳統方法,類神經網路將資料特徵擷取與資料辨識整合並串聯於架構之中,並透過大量資料的訓練,得到一個強大且能夠準確進行資料辨識的類神經網路模型。然而,雖然更多層與更複雜的類神經網路模型可以獲得更準確的辨識效能,但傳統的順向計算與錯誤回饋之類神經網路模型訓練演算法無法有效地針對多層之類神經網路進行訓練。另一方面,將訓練資料庫中大量資料進行標示所需的代價,以及如何在沒有相關背景知識下進行類神經網路的初始化都將是類神經網路模型訓練中的一大問題。
在本論文中,我們將設計並實現一侷限型波茲曼模型(RBM)處理器。在此處理器中,整合了32個本論文所提出之侷限型波茲曼模型(RBM)運算單元進行平行處理,並支援單層最多4096個類神經節點(neuron)和單筆測試資料最多128種類別的資料辨識能力。當處理器操作在類神經網路模型學習模式下,可達到多筆訓練資料(batch-level)之間的平行運算,並同時支援監督式(supervised)與非監督式(unsupervised)之侷限型波茲曼模型(RBM)模型訓練。當操作在類神經網路資料辨識模式下,可達到多筆測試資料(sample-level)之間的平行運算。此外,本論文中亦提出多項技術用於改善運算效能、硬體設計複雜度、外部記憶體頻寬、和功率消耗。
在本論文中,我們透過兩種實作方式實現所提出之侷限型波茲曼模型(RBM)處理器。第一種是使用Xilinx Virtex-7系列之現場可程式化閘陣列(FPGA)實現,當處理器操作在125 MHz工作頻率下,本處理器使用了114.0k個查找表(LUT)、107.1k個正反器(flip-flop)和80個記憶體區塊(block memory block)。第二種是使用UMC 65奈米製程實現,本侷限型波茲曼模型(RBM)處理器晶片在8.8 mm2面積內包含了2.2M個邏輯閘(gate)和128kB內存記憶體(SRAM),並將32個侷限型波茲曼模型(RBM)運算單元整合於2個運算叢集(cluster)之中。當操作在1.2V供電電壓下,本晶片可操作在最高210 MHz工作頻率進行模型學習與資料辨識。
依據量測結果,基於現場可程式化閘陣列(FPGA)之系統雛型平台可分別達到4.60G neuron weights/s (NWPS)之模型學習效能和3.87G NWPS之資料辨識效能。而本侷限型波茲曼模型(RBM)處理器晶片操作在210 MHz工作頻率下,可分別達到4.61G NWPS與69.50 pJ/NW之模型學習效能以及3.86G NWPS與81.20 pJ/NW之資料辨識效能。
和使用一般處理器(CPU)與多核心處理器(multi-core processor)相比,本論文所提出之侷限型波茲曼模型(RBM)處理器能以更快的運算效能與更高的能源效益進行侷限型波茲曼模型(RBM)之模型訓練和資料辨識。因此,針對物聯網(IoT)與手持裝置這類有能源約束考量的裝置,本論文所提出之侷限型波茲曼模型(RBM)處理器將可提供一高效能與高能源效益解決方案,以提供裝置具有智慧處理能力並達到即時模型訓練和即時資料辨識功能。
Recently, machine learning techniques have been widely applied to signal processing systems to support intelligent capabilities, such as AdaBoost, K-NN, mean-shift, and SVM for data classification, and HOG and SIFT for feature extraction in multimedia applications. In the past decades, the neural network (NN) algorithms are considered one of the state-of-the-art solutions in many applications, and both feature extraction and data classification are integrated and cascaded in neural networks. In the big data era, the huge dataset benefits neural network learning algorithms to train a powerful and accurate model for machine learning applications. Since the network structure becomes deeper and deeper to achieve more accurate performance for applications, the traditional neural network learning algorithm with feedforwarding and error backpropagation is inefficient to train multi-layer neural networks. Moreover, the data labeling is very expensive especially for big dataset, and how to initialize a neural network without any domain knowledge is also a crucial issue for model training.
In this dissertation, a restricted Boltzmann machine (RBM) processor is designed and implemented. In the proposed RBM processor, 32 proposed RBM cores are integrated for parallel computing with the neural network structure of maximal 4k neurons per layer and 128 candidates per sample for inference. Operated in the learning mode, the batch-level parallelism is achieved for RBM model training with supervised and unsupervised learning. And the sample-level parallelism is achieved for data classification operated in the inference mode. Moreover, several features are proposed and implemented in the proposed RBM processor to save computation time, hardware cost, external memory bandwidth, and power consumption.
To realize the proposed RBM processor, two implementations are designed in this dissertation. Implemented in Xilinx Virtex-7 FPGA, the proposed RBM processor is operated at 125 MHz and occupies 114.0k LUTs, 107.1k flip-flops, and 80 block memory blocks. Implemented in UMC 65nm LL RVT CMOS technology, the proposed RBM processor chip costs 2.2M gates and 128kB internal SRAM with 8.8 mm2 area to integrate 32 proposed RBM cores in 2 clusters, and the maximal operating frequency of this chip achieves 210 MHz in both learning and inference modes operated at 1.2V supply voltage.
According to the measurement results, the proposed FPGA-based system prototype platform achieves 4.60G neuron weights/s (NWPS) learning performance and 3.87G NWPS inference performance for RBM model training and data classification, respectively. And the proposed RBM processor chip operated at 210MHz to achieve 4.61G NWPS and 3.86G NWPS performance with 69.50 pJ/NW and 81.20 pJ/NW energy efficiency in learning and inference modes, respectively.
Compared to the software solution implemented on CPU and powerful multi-core processors, the proposed RBM processor achieves faster processing time and higher energy efficiency in both RBM model learning and data inference, respectively. Since the battery life is a crucial issue in IoT and handheld devices, our proposal achieves an energy-efficient solution to integrate the proposed RBM processor chip into the emerging energy-constrained devices to support intelligent capabilities with learning and inference for in-time model training and real-time decision making.
Chapter 1 Introduction 1
1.1 ML-based Data Classification Trends 1
1.2 Motivation and Design Challenge 4
1.3 Dissertation Organization 7
Chapter 2 Restricted Boltzmann Machine Algorithm 9
2.1 Background of Restricted Boltzmann Machine 9
2.2 Related Works 13
2.2.1 Acceleration by General Purpose Processor 13
2.2.2 Acceleration by Dedicated Hardware 14
2.3 Summary 16
Chapter 3 Restricted Boltzmann Machine Processor 18
3.1 Neuron Computation Unit (NCU) 29
3.2 Low Power Neuron Binarizer (LPNB) 32
3.3 User-Defined Connection Map (UDCM) 38
3.4 Early Stopping (ES) Mechanism 42
3.5 Summary 47
Chapter 4 System Implementation and Integration 48
4.1 System Prototype Platform Design 50
4.1.1 RBM Processor 53
4.1.2 Communication Interface 58
4.1.3 Data Storage Interface 63
4.1.4 System Implementation Result 66
4.2 RBM Processor Chip Implementation 69
4.2.1 Chip Implementation Result 70
4.2.2 RBM Processor Chip Interface 73
4.2.3 DUT Board Design 77
Chapter 5 Verification and Performance Evaluation 79
5.1 Verification Environment 80
5.1.1 Functionality Verification 81
5.1.2 FPGA-based Hardware System Verification 83
5.2 Performance Evaluation 85
5.3 Applications 98
Chapter 6 Conclusion & Future Research 107
6.1 Conclusion 107
6.2 Recommendations for Future Research 110
[1] Hirotsugu Shikano, Kiyoto Ito, Kazuhide Fujita, and Tadashi Shibata, "A Real-Time Learning Processor Based on K-means Algorithm with Automatic Seeds Generation," International Symposium on System-on-Chip, pp. 1-4, Nov. 2007.
[2] Hanaa Hussain, Khaled Benkrid, Chuan Hong, and Huseyin Seker, "An adaptive FPGA implementation of multi-core K-nearest neighbour ensemble classifier using dynamic partial reconfiguration," International Conference on Field Programmable Logic and Applications, pp. 627-630, Aug. 2012.
[3] Muhammad Awais Bin Altaf and Jerald Yoo, "A 1.52 uJ/classification patient-specific seizure classification processor using Linear SVM," IEEE International Symposium on Circuits and Systems, pp. 849-852, May 2013.
[4] Chang-Hung Tsai, Hui-Hsuan Lee, Wan-Ju Yu, and Chen-Yi Lee, "A 2 GOPS quad-mean shift processor with early termination for machine learning applications," IEEE International Symposium on Circuits and Systems, pp. 157-160, Jun. 2014.
[5] Jung Kuk Kim, Phil Knag, Thomas Chen, and Zhengya Zhang, "A 6.67mW sparse coding ASIC enabling on-chip learning and inference," Symposium on VLSI Circuits, pp. 1-2, Jun. 2014.
[6] Jung Kuk Kim, Phil Knag, Thomas Chen, and Zhengya Zhang, "A 640M pixel/s 3.65mW sparse event-driven neuromorphic object recognition processor with on-chip learning," Symposium on VLSI Circuits, pp. 50-51, Jun. 2015.
[7] Seongwook Park, Kyeongryeol Bong, Dongjoo Shin, Jinmook Lee, Sungpill Choi, and Hoi-Jun Yoo, "A 1.93TOPS/W scalable deep learning/inference processor with tetra-parallel MIMD architecture for big-data applications," IEEE International Solid- State Circuits Conference, pp. 80-81, Feb. 2015.
[8] Henry A. Rowley, Shumeet Baluja, and Takeo Kanade, "Neural network-based face detection," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 20, no. 1, pp. 23-38, Jan. 1998.
[9] Ruslan Salakhutdinov and Geoffrey Hinton, "Deep Boltzmann machines," International Conference on Artificial Intelligence and Statistics, vol. 5, pp. 448-455, Apr. 2009.
[10] Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, "Imagenet classification with deep convolutional neural networks," Neural Information Processing Systems, pp. 1097-1105, Dec. 2012.
[11] Vinayak Gokhale, Jonghoon Jin, Aysegul Dundar, Berin Martini, and Eugenio Culurciello, "A 240 G-ops/s Mobile Coprocessor for Deep Neural Networks," IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 696-701, Jun. 2014.
[12] Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, and Brian Kingsbury, "Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups," IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82-97, Nov. 2012.
[13] Abdel-rahman Mohamed, George E. Dahl, Geoffrey Hinton, "Acoustic Modeling Using Deep Belief Networks," IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, no. 1, pp. 14-22, Jan. 2012.
[14] Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu, "Convolutional Neural Networks for Speech Recognition," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, no. 10, pp. 1533-1545, Oct. 2014.
[15] Yuanfang Ren and Yan Wu, "Convolutional deep belief networks for feature extraction of EEG signal," International Joint Conference on Neural Networks, pp. 2850-2853, Jul. 2014.
[16] Robert Hecht-Nielsen, "Theory of the backpropagation neural network," International Joint Conference on Neural Networks, vol. 1, pp. 593-605, Jun. 1989.
[17] Christopher M. Bishop, Pattern Recognition and Machine Learning, Springer, 2006.
[18] Geoffrey Hinton, Simon Osindero, and Yee-Whye Teh, "A Fast Learning Algorithm for Deep Belief Nets," Neural Computation, vol. 18, no. 7, pp. 1527-1554, Jul. 2006.
[19] Asja Fischer and Christian Igel, "Training restricted Boltzmann machines: an introduction," Pattern Recognition, vol. 47, no. 1, pp. 25-39, 2014.
[20] Jian Ouyang, Shiding Lin, Wei Qi, Yong Wang, Bo Yu, and Song Jiang, "SDA: Software-Defined Accelerator for Large-Scale DNN Systems," HotChips, Aug. 2014.
[21] Daniel Ly and Paul Chow, "High-Performance Reconfigurable Hardware Architecture for Restricted Boltzmann Machines," IEEE Transactions on Neural Networks, vol. 21, no. 11, pp. 1780-1792, Nov. 2010.
[22] Mahmoud Al-Nsour and Hoda S. Abdel-Aty-Zohdy, "Implementation of programmable digital sigmoid function circuit for neuro-computing," IEEE Midwest Symposium on Circuits and Systems, pp. 571-574, Aug. 1998.
[23] Chang-Hung Tsai, Yu-Ting Chih, Wing H. Wong, and Chen-Yi Lee, "A Hardware-Efficient Sigmoid Function with Adjustable Precision for Neural Network System," IEEE Transactions on Circuits and Systems II, vol. 62, no. 11, pp. 1073-1077, Nov. 2015.
[24] Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton. "Imagenet classification with deep convolutional neural networks," Neural Information Processing Systems, pp. 1097-1105, Dec. 2012.
[25] Geoffrey Hinton, "A Practical Guide to Training Restricted Boltzmann Machines," Technical Report UTML TR 2010003, Department of Computer Science, University of Toronto, 2010.
[26] Yoav Freund and Robert E. Schapire, "A decision-theoretic generalization of on-line learning and an application to boosting," European Conference on Computational Learning Theory, pp. 23-37, Mar. 1995.
[27] Corinna Cortex and Vladimir Vapnik, "Support-Vector Networks," Machine Learning, vol. 20, no. 3, pp. 273-297, Sep. 1995.
[28] David G. Lowe, "Object recognition from local scale-invariant features," IEEE International Conference on Computer Vision, vol. 2, pp. 1150-1157, Sep. 1999.
[29] James MacQueen, "Some methods for classification and analysis of multivariate observations," Berkeley symposium on mathematical statistics and probability, vol. 1, no. 14, pp. 281-297, Jue. 1965.
[30] Thomas M. Cover and Peter E. Hart, "Nearest neighbor pattern classification," IEEE Transactions on Information Theory, vol. 13, no. 1, pp. 21-27, Jan. 1967.
[31] Keinosuke Fukunaga and Larry Hostetler, "The estimation of the gradient of a density function, with applications in pattern recognition," IEEE Transactions on Information Theory, vol. 21, no. 1, pp. 32-40, Jan. 1975.
[32] David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams, "Learning representations by back-propagation errors," Nature, no. 323, pp. 533-536, Oct. 1986.
[33] Kyoung-Su Oh and Keechul Jung, "GPU implementation of neural networks," Pattern Recognition, no. 41, no. 8, pp. 2684-2692, Aug. 2008.
[34] Dave Steinkraus, Ian Buck, and Patrice Y. Simard, "Using GPUs for machine learning algorithms," International Conference on Document Analysis and Recognition, vol. 2, pp. 1115-1120, Aug. 2005.
[35] Yangqing Jia, Caffe: An open source convolutional architecture for fast feature embedding. http://caffe.berkeleyvision.org/
[36] Daniel L. Ly and Paul Chow, "A multi-FPGA architecture for stochastic Restricted Boltzmann Machines," International Conference on Field Programmable Logic and Applications, pp. 168-173, Aug. 2009.
[37] Sang Kyun Kim, Lawrence C. McAfee, Peter L. McMahon, and Kunle Olukotun, "A highly scalable Restricted Boltzmann Machine FPGA implementation," International Conference on Field Programmable Logic and Applications, pp. 367-372, Aug. 2009.
[38] Sang Kyun Kim, Peter L. McMahon, and Kunle Olukotun, "A Large-Scale Architecture for Restricted Boltzmann Machines," IEEE Annual International Symposium on Field-Programmable Custom Computing Machines, pp. 201-208, May 2010.
[39] Jordan L. Holt and Thomas E. Baker, "Back propagation simulations using limited precision calculations," International Joint Conference on Neural Networks, vol. 2, pp. 121-126, Jul. 1991.
[40] Yann LeCun, Corinna Cortes, and Christopher Burges, "The MNIST database of handwritten digits," available at http://yann.lecun.com/exdb/mnist/
[41] Shu-Yu Hsu, Ying-Chieh Ho, Po-Yao Chang, Chau-Chin Su, and Chen-Yi Lee, "A 48.6-to-105.2 μW Machine Learning Assisted Cardiac Sensor SoC for Mobile Healthcare Applications," IEEE Journal of Solid-State Circuits, vol. 49, no. 4, pp. 801-811, Apr. 2014.
連結至畢業學校之論文網頁點我開啟連結
註: 此連結為研究生畢業學校所提供,不一定有電子全文可供下載,若連結有誤,請點選上方之〝勘誤回報〞功能,我們會盡快修正,謝謝!
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊