|
[1]D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," Nature, vol. 323, no. 6088, pp. 533-536, 1986/10/01 1986. [2]Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, "Gradient-based learning applied to document recognition," Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, 1998. [3]A. Krizhevsky, I. Sutskever, and G. E. Hinton, "ImageNet Classification with Deep Convolutional Neural Networks," NIPS, 2012. [4]C. Szegedy et al., "Going deeper with convolutions," in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 1-9. [5]K. He, X. Zhang, S. Ren, and J. Sun, "Deep Residual Learning for Image Recognition," in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. [6]K. Ovtcharov, O. Ruwase, J.-Y. Kim, J. Fowers, K. Strauss, and E. S. Chung, "Accelerating Deep Convolutional Neural Networks Using Specialized Hardware," 2015. [7]S. Mittal, A Survey of FPGA-based Accelerators for Convolutional Neural Networks. 2018. [8]N. P. Jouppi et al., "In-datacenter performance analysis of a tensor processing unit," pp. 1-12, 2017. [9]U. Köster et al., Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks. 2017. [10]S. Han et al., "EIE: Efficient Inference Engine on Compressed Deep Neural Network," in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 243-254. [11]Y. Chen, T. Krishna, J. S. Emer, and V. Sze, "Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks," IEEE Journal of Solid-State Circuits, vol. 52, no. 1, pp. 127-138, 2017. [12]Z. Yuan et al., "Sticker: A 0.41-62.1 TOPS/W 8Bit Neural Network Processor with Multi-Sparsity Compatible Convolution Arrays and Online Tuning Acceleration for Fully Connected Layers," in 2018 IEEE Symposium on VLSI Circuits, 2018, pp. 33-34. [13]J. Song et al., "7.1 An 11.5TOPS/W 1024-MAC Butterfly Structure Dual-Core Sparsity-Aware Neural Processing Unit in 8nm Flagship Mobile SoC," in 2019 IEEE International Solid- State Circuits Conference - (ISSCC), 2019, pp. 130-132. [14]J. Lee, C. Kim, S. Kang, D. Shin, S. Kim, and H. Yoo, "UNPU: An Energy-Efficient Deep Neural Network Accelerator With Fully Variable Weight Bit Precision," IEEE Journal of Solid-State Circuits, vol. 54, no. 1, pp. 173-185, 2019. [15]S. Yin et al., "An Ultra-High Energy-Efficient Reconfigurable Processor for Deep Neural Networks with Binary/Ternary Weights in 28NM CMOS," in 2018 IEEE Symposium on VLSI Circuits, 2018, pp. 37-38. [16]M. Anders et al., "2.9TOPS/W Reconfigurable Dense/Sparse Matrix-Multiply Accelerator with Unified INT8/INTI6/FP16 Datapath in 14NM Tri-Gate CMOS," in 2018 IEEE Symposium on VLSI Circuits, 2018, pp. 39-40. [17]S. Zhang et al., "Cambricon-X: An accelerator for sparse neural networks," in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 1-12. [18]D. Shin, J. Lee, J. Lee, and H. Yoo, "14.2 DNPU: An 8.1TOPS/W reconfigurable CNN-RNN processor for general-purpose deep neural networks," in 2017 IEEE International Solid-State Circuits Conference (ISSCC), 2017, pp. 240-241. [19]Z. Yuan, Y. Liu, J. Yue, J. Li, and H. Yang, "CORAL: Coarse-grained reconfigurable architecture for Convolutional Neural Networks," in 2017 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 2017, pp. 1-6. [20]G. Venkatesh, E. Nurvitadhi, and D. Marr, "Accelerating Deep Convolutional Networks using low-precision and sparsity," in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 2861-2865. [21]A. Parashar et al., "SCNN: An accelerator for compressed-sparse convolutional neural networks," in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), 2017, pp. 27-40. [22]G. Desoli et al., "14.1 A 2.9TOPS/W deep convolutional neural network SoC in FD-SOI 28nm for intelligent embedded systems," in 2017 IEEE International Solid-State Circuits Conference (ISSCC), 2017, pp. 238-239. [23]S. Wang, D. Zhou, X. Han, and T. Yoshimura, "Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks," in Design, Automation & Test in Europe Conference & Exhibition (DATE), 2017, 2017, pp. 1032-1037. [24]Y. Shen, M. Ferdman, and P. Milder, "Escher: A CNN Accelerator with Flexible Buffering to Minimize Off-Chip Transfer," in 2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2017, pp. 93-100. [25]C. Wang, L. Gong, Q. Yu, X. Li, Y. Xie, and X. Zhou, "DLAU: A Scalable Deep Learning Accelerator Unit on FPGA," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 36, no. 3, pp. 513-517, 2017. [26]Y. Umuroglu et al., "FINN: A Framework for Fast, Scalable Binarized Neural Network Inference," arXiv:1612.07119, 2016. [27]C. Zhang, F. Zhenman, Z. Peipei, P. Peichen, and C. Jason, "Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks," in 2016 IEEE/ACM International Conference on Computer-Aided Design (ICCAD), 2016, pp. 1-8. [28]L. Huimin, F. Xitian, J. Li, C. Wei, Z. Xuegong, and W. Lingli, "A high performance FPGA-based accelerator for large-scale convolutional neural networks," in 2016 26th International Conference on Field Programmable Logic and Applications (FPL), 2016, pp. 1-9. [29]S. Wijeratne, S. Jayaweera, M. Dananjaya, and A. Pasqual, "Reconfigurable co-processor architecture with limited numerical precision to accelerate deep convolutional neural networks," in 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2018, pp. 1-7. [30]J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, "Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing," in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 1-13. [31]P. Judd, A. Delmas, S. Sharify, and A. Moshovos, "Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing," arXiv:1705.00125, 2017. [32]P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, "Stripes: Bit-serial deep neural network computing," in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 1-12. [33]A. Erdem, C. Silvano, T. Boesch, A. Ornstein, S. Singh, and G. Desoli, "Design Space Exploration for Orlando Ultra Low-Power Convolutional Neural Network SoC," in 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2018, pp. 1-7. [34]Y. Chen et al., "DaDianNao: A Machine-Learning Supercomputer," in 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, 2014, pp. 609-622. [35]S. Venkataramani et al., "SCALEDEEP: A scalable compute architecture for learning and evaluating deep networks," in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), 2017, pp. 13-26. [36]J. Lee, J. Lee, D. Han, J. Lee, G. Park, and H. Yoo, "7.7 LNPU: A 25.3TFLOPS/W Sparse Deep-Neural-Network Learning Processor with Fine-Grained Mixed Precision of FP8-FP16," in 2019 IEEE International Solid- State Circuits Conference - (ISSCC), 2019, pp. 142-144. [37]B. Fleischer et al., "A Scalable Multi- TeraOPS Deep Learning Processor Core for AI Training and Inference," in 2018 IEEE Symposium on VLSI Circuits, 2018, pp. 35-36. [38]Z. Wenlai et al., "F-CNN: An FPGA-based framework for training Convolutional Neural Networks," in 2016 IEEE 27th International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2016, pp. 107-114. [39]X. Han, D. Zhou, S. Wang, and S. Kimura, "CNN-MERP: An FPGA-based memory-efficient reconfigurable processor for forward and backward propagation of convolutional neural networks," in 2016 IEEE 34th International Conference on Computer Design (ICCD), 2016, pp. 320-327. [40]V. Sze, Y. Chen, T. Yang, and J. S. Emer, "Efficient Processing of Deep Neural Networks: A Tutorial and Survey," Proceedings of the IEEE, vol. 105, no. 12, pp. 2295-2329, 2017. [41]P. Lin, M. Sun, C. Kung, and T. Chiueh, "FloatSD: A New Weight Representation and Associated Update Method for Efficient Convolutional Neural Network Training," IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 9, no. 2, pp. 267-279, 2019. [42]N. Wang, J. Choi, D. Brand, C.-Y. Chen, and K. Gopalakrishnan, "Training Deep Neural Networks with 8-bit Floating Point Numbers," Proc. Adv. Neural Inf. Process. Syst., pp. 7685-7694, 2018. [43]P. Lin, "Low-complexity Convolutional Neural Network Training and Low Power Circuit Design of its Processing Element," M.A. Thesis, National Taiwan University, 2017. [44]F. Rosenblatt, "The Perceptron--a perceiving and recognizing automaton," Report 85-460-1, Cornell Aeronautical Laboratory, 1957. [45]M. A. Nielsen, "Neural Networks and Deep Learning," Determination Press, 2015. [46]T. H. Juang, "Energy-Efficient Accelerator Architecture for Neural Network Training and its Circuit Design," Master Thesis, National Taiwan University, 2018. [47]Y. Jia et al., "Caffe: Convolutional Architecture for Fast Feature Embedding," Proc. ACM Int. Conf. Multimedia, pp. 675-678, 2014. [48]TSRI. (7-Jul-2019). TSRI AI SoC設計平台簡介. Available: https://www.tsri.org.tw/aisoc/aisoc.jsp [49]ARM, AMBA AXI and ACE Protocol Specification AXI3, AXI4, and AXI4-Lite and ACE-Lite. 2011. [50]P. Elenius and L. Levine. (2000) Comparing Flip-Chip and Wire-Bond Interconnection Technologies. Chip Scale Review. [51]I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning MIT Press, 2016. [52]P. Bright. (2017). Google brings 45 teraflops tensor flow processors to its compute cloud. Available: https://arstechnica.com/information-technology/2017/05/google-brings-45-teraflops-tensor-flow-processors-to-its-compute-cloud/ [53]P. Clark. (2016). AI Chip Startup Shares Insights. Available: https://www.eetimes.com/document.asp?doc_id=1330739 [54](2017). Graphcore at NIPS 2017 Presentations. Available: https://www.graphcore.ai/nips2017_presentations
|