跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.73) 您好!臺灣時間:2026/07/22 11:51
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:蕭丕承
研究生(外文):Pi-Chen Hsiao
論文名稱:管線化及叢集化之超長指令集數位信號處理器之高效能資料路徑設計
論文名稱(外文):Efficient Datapath Design for Clustered & Pipelined VLIW DSP Processors
指導教授:劉志尉
指導教授(外文):Chih-Wei Liu
學位類別:碩士
校院名稱:國立交通大學
系所名稱:電子工程系所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2005
畢業學年度:94
語文別:英文
論文頁數:71
中文關鍵詞:前饋叢集化暫存器組超長指令集數位信號處理器管線化
外文關鍵詞:forwardingclusteringregister filevery long instruction word (VLIW)digital signal processor (DSP)pipelining
相關次數:
  • 被引用被引用:0
  • 點閱點閱:472
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
大部分的數位訊號處理應用程式都具有高度資料階層以及指令階層平行度的特性,因此可以藉由叢集化以及加深管線的方式來增加資料路徑的效率。然而,複雜的前饋(forwarding)網路以及叢集間連結(inter-cluster communication)網路抵銷了叢集化及加深管線所提升的效能。這篇論文以標準單元設計為基礎利用正反器和多工器來分析前饋單元與叢集間連結機制的複雜度。藉由這些分析,我們提出複雜度感知的前饋單元架構和以記憶體讀取/寫入(load/store)指令為基礎的叢集間連結機制。除此之外,我們還提出了分散式乒乓暫存器組架構來進一步降低叢集內暫存器組的複雜度。在實作的部分,我們使用UMC 0.13um 1P8M CMOS製程來實現我們設計。實驗的結果顯示,我們提出了前饋單元架構可以增加13.2%的運作時脈,而分散式乒乓暫存器組搭配我們提出的叢集間連結機制則可以減少76.8%的面積和46.9%的暫存器存取時間。對於可攜帶型裝置的應用方面,我們另外提出了與原本應用程式完全相容的折疊式資料路徑架構。比起原本的設計,這種架構可以節省55.33%的面積和增加26.3%的運作速度。最後,我們利用前述的前饋單元和叢集架構設計並實現了一個完整的4-way 超長指令集(VLIW)數位訊號處理器。實作與模擬的結果顯示在UMC 0.13um 1P8M CMOS的製程下,其最高工作頻率為333MHz,且具有近似於現在市面上數位訊號處理器的運算能力。
Most DSP applications feature a high degree of data-level and instruction-level parallelism, which enables efficient datapath design with clustering and deep pipelining. However, the ad-hoc data forwarding and inter-cluster communications in most processors significantly compensate the advantages. This thesis presents analytical formulae which are based on cell-based implementation with flip-flops and multiplexers to analyze the complexity of forwarding unit and inter-cluster communication mechanisms. We also propose a complexity-aware data forwarding architecture and a simple inter-cluster communication mechanism based on load/store instruction pairs. Moreover, we introduce the distributed & ping-pong register file to further reduce the complexity of register file inside clusters. In the experiments with UMC 0.13um 1P8M CMOS technology, our proposed forwarding architecture can improve cycle time by 13.2%, while the distributed ping-pong register file collocated with proposed inter-cluster communication mechanism can reduce the area and access time of register file by 76.8% and 46.9%. For portable applications, we bring up the folded datapath with binary compatibility which saves 55.33% area and increases the clock speed by 26.3%. Finally, we implement the proposed forwarding unit and the proposed inter-cluster communication mechanism with distributed & ping-pong register file organization in a complete 4-way VLIW DSP processor which can operate at 333MHz and shows comparable performance with state-of-the-art DSPs.
ABSTRACT (CHINESE)...................................................................................................................................... i
ABSTRACT (ENGLISH) ...................................................................................................................................iii
ACKNOWLEDGEMENT................................................................................................................................... v
CONTENTS........................................................................................................................................................vii
LIST OF TABLES.............................................................................................................................................viii
LIST OF FIGURES............................................................................................................................................. ix
1 INTRODUCTION............................................................................................................................................. 1
1.1 DSP PROCESSORS ........................................................................................................... 2
1.2 DESIGN PROBLEMS ............................................................................................................. 4
1.3 THESIS ORGANIZATION ........................................................................................................ 6
2 PIPELINED DATAPATHS............................................................................................................. 7
2.1 PIPELINE DESIGN ............................................................................................................... 8
2.2 DATA HAZARDS ............................................................................................................. 14
2.3 COMPLEXITY-AWARE DATA FORWARDING .............................................. 16
2.4 FORWARDING THROUGH REGISTER FILE .................................................... 19
3 REGISTER FILE ORGANIZATION ............................................ 21
3.1 COMPLEXITY OF SYNTHESIZED REGISTER FILE..................................................... 22
3.2 CLUSTERING .......................................................................................................... 25
3.2.1 Inter-cluster communication mechanisms............................................... 26
3.2.2 Proposed inter-cluster communication mechanism ............................................... 30
3.3 DISTRIBUTED & PING-PONG REGISTER ORGANIZATION ............................................ 34
4 SCALED DATAPATH WITH BINARY COMPATIBILITY............................................ 43
4.1 FOLDED DATAPATH ................................................ 44
4.2 MEMORY SUBSYSTEM & INTER-CLUSTER COMMUNICATION ISSUES ................................................ 47
5 SIMULATION & IMPLEMENTATION RESULTS................... 51
5.1 PICADSP ............................................ 52
5.2 DESIGN FLOW..................................................... 55
5.3 EXPERIMENT RESULTS ................................. 59
6 CONCLUSIONS & FUTURE WORK.............................. 67
REFERENCE ............................................... 69
[1] Y. H. Hu, Programmable Digital Signal Processors – Architecture, Programming, and Applications, Marcel Dekker Inc., 2002
[2] J. L Hennessy, and D. A. Patterson, Computer Architecture – A Quantitative Approach, 3rd Edition, Morgan Kaufmann, 2002
[3] J. A. Fisher, P. Faraboschi, and C. Young, Embedded Computing – A VLIW Approach to Architecture, Compiler, and Tools, Morgan Kaufmann, 2005
[4] S. Rixner, et al., “Register organization for media processing,” in Proc. HPCA-6, pp.375-386, 2000
[5] T. J. Lin, et al., “A unified processor architecture for RISC & VLIW DSP,” in Proc. GLSVLSI, 2005
[6] D. A. Patterson and J. L. Hennessy, Computer Organization & Design – the Hardware/Software Interface, 2nd Edition, Morgan Kaufmann, 1998.
[7] A. Terechko, M. Garg, and H. Corporaal, “Evaluation of speed and area of clustered VLIW Processors,” in Proc. VLSID, pp.557-563, 2005
[8] A. Terechko, E. L. Thenaff, M. Garg, J. Eijndhoven, and H. Corporaal, “Inter-cluster communication models for clustered VLIW processors,” in Proc. HPCA-9, pp.354-364, 2003
[9] P. Faraboschi, G. Brown, J. A. Fisher, G. Desoll, and F. M. O. Homewood, “Lx: a technology platform for customizable VLIW embedded processing,” in Proc. ISCA, pp.203-213, 2000
[10] G. G. Pechanek and S. Vassiliadis, “The ManArray embedded processor architecture,” in Proc. Euromicro Conf., pp.348-355, 2000
[11] E. F. Barry, G. G. Pechanek, and P. R. Marchand, “Register file indexing methods and apparatus for providing indirect control of register file addressing in a VLIW processor,” International Application Published under the Patent Cooperation Treaty (PCT), WO 00/54144, Mar. 9 2000
[12] TMS320C6000 CPU and Instruction Set Reference Guide, Texas Instruments Inc., 2000
[13] K. Arora, H. Sharangpani, and R. Gupta, “Copied register files for data processors having many execution units” U.S. Patent 6629 232, Sep. 30, 2003
[14] TMS320C64x DSP Library Programmer's Reference, Texas Instruments Inc., Apr 2002
[15] P.C. Hsiao, T. J. Lin, C. W. Liu, and C. W. Jen, “Efficient datapath design for clustered & pipelined digital signal processors,” in Proc. VLSI design/CAD, Aug. 2005
[16] J. H. Tseng and K. Asanovic, “Banked multiported register files for high-frequency superscalar microprocessors,” in Proc. ISCA, pp.62-71, 2003
[17] A. V. Oppenheim, R. W. Schafer, and J. R. Buck, Discrete-Time Signal Processing, 2nd Edition, Prentice Hall, 1999
[18] W. B. Pennebaker and J. L. Mitchell, JPEG – Still Image Data Compression Standard, Van Nostrand Reinhold, 1993
[19] T. J. Lin, C. C. Lee, C. W. Liu, and C. W. Jen, “A novel register organization for VLIW digital signal processors,” in Proc. VLSI-TSA-DAT, April 2005
[20] T. Kumura, M. Ikekawa, M. Yoshida, and I. Kuroda, “VLIW DSP for mobile applications,” IEEE Signal Processing Mag., pp.10-21, July 2002
[21] H. Pan and K. Asanovic, “Heads and tails: a variable-length instruction format supporting parallel fetch and decode,” in Proc. CASES, 2001
[22] A. M. Tekalp, Digital Video Processing, Prentice Hall, 1995
[23] V. Zyuban and P. Kogge, “The energy complexity of register files,” in Proc. ISLPED, pp.305-310, 1998
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top