跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.73) 您好!臺灣時間:2026/07/22 11:51
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:林聖勳
研究生(外文):Sheng-Hsun Lin
論文名稱:可被窄寬度運算共用之單一管線資料路徑設計
論文名稱(外文):A Single Pipeline Datapath Design for Joinable Narrow-operand Operations
指導教授:鍾崇斌
指導教授(外文):Chung-Ping Chung
學位類別:碩士
校院名稱:國立交通大學
系所名稱:資訊科學與工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2006
畢業學年度:94
語文別:英文
論文頁數:65
中文關鍵詞:資料路徑共用窄寬度運算
外文關鍵詞:datapath sharingnarrow-operand operations
相關次數:
  • 被引用被引用:0
  • 點閱點閱:195
  • 評分評分:
  • 下載下載:8
  • 收藏至我的研究室書目清單書目收藏:0
現今大多數之處理器之架構都採32位元或者更高之位元數. 然而整數運算大多數時間都不會用到資料路徑中完整的寬度. 若將Source Operand Bus及ALU分割為多組block, 則此資料路徑便具備同時處理一道以上運算之能力.本論文提出讓單一資料路徑被兩道具備窄寬度特性之指令合用之設計. 讓其中一道運算之Operand Block以反轉(Turnaround)之次序, 來達到極有效率之資料路徑合用機制.相較於傳統以直接Shift方式, 本設計不論就面積與電路延遲上, 都有較佳之表現. 此外, 本論文亦提出一縮短ALU因分割為多組ALU Block所造成延遲之設計. 最後, 本論文提出將此機制整合至一典型五階MIPS管線之方式.
Most general-purpose processors and embedded processors have 32-bit word widths or wider. However, integer operations rarely need the full 32-bit dynamic range of the datapath. If we partition the operand bus, result bus, and ALU into several blocks, the datapath could perform more than one operation in parallel.
In this thesis, mechanisms to join two narrow-operand operations together to share a single datapath are proposed. We proposed one novel ALU-sharing scheme by turning around operation-block ordering. Efficient designs to merge operands to share buses and ALU based on the technique are proposed and discussed. Compared with traditional “shift” approach, the turnaround approach has many advantages on area and delay. Besides, a technique to mitigate the delay overhead of the partitioned ALU by swapping operands is proposed. We also made performance simulation to help decide how to partition the datapath. Finally, how to integrate such datapath into a MIPS five-stage pipeline and required modifications are discussed.
摘要 i
Abstract ii
誌謝 iii
Table of Contents iv
List of Figures vi
List of Tables viii
Chapter 1. Introduction 1
1.1. Narrow-Operand ALU Operations 1
1.2. Significant Bit-width of an ALU Operation 2
1.3. Organization of this Thesis 4
Chapter 2. Background and Related Work 5
2.1. Distribution of Significant Bit-widths of ALU Operations 5
2.2. Datapath in Multi-Bitwidth Pipeline 8
2.3. More Flexible ALU-sharing Mechanism 10
2.4. Motivation 12
2.5. Objective 13
Chapter 3. Design 15
3.1. Overview of Datapath-Sharing 15
3.1.1. Constraints for Joining Two Instructions 15
3.1.1.1. Structural Hazards 15
3.1.1.2. Data Hazards 16
3.1.2. How ALU is Shared by Two Operations 17
3.1.3. Possible Modifications in a Five-stage MIPS-like Pipeline 18
3.2. Modifications in Instruction-Decode Stage 20
3.2.1. Type-Check and Data-dependency Check 20
3.2.2. Width-Check 20
3.2.2.1. Width-Determination Logic 20
3.2.2.2. Operand-boundary Signals 25
3.2.2.3. Width-Check Logic 26
3.3. Modifications in Execution Stage 30
3.3.1. Deciding the Block Width 30
3.3.2. Merging Operands from Two Operations 31
3.3.3. ALU in the Turnaround Approach 34
3.3.3.1. Operation Swapping 39
3.3.3.2. Widening the ALU 42
3.4. Modifications in Memory Access and Write-back Stage 42
3.4.1. Sign-extending the Joined Results 43
3.5. Integrating the Design into a Five-stage MIPS-like Pipeline Datapath 45
Chapter 4. Experiments 49
4.1. Goals of Our Experiments 49
4.2. Simulation Environment 49
4.3. Comparison between Different Operand-Merging Schemes 50
4.4. Hardware Cost of Operation-Swapping Mechanism 52
4.5. Choosing ALU Block Width 54
4.6. Choosing ALU Width 57
4.7. Final Proposal of the ALU Design 58
Chapter 5. Conclusion and Future Work 61
5.1. The Turnaround-based Sharing Mechanism 61
5.2. Reducing Additional Delay by Avoiding Unnecessary Partitioning 61
5.3. Applying the Design to Architectures with Shifter Concatenated with ALU 62
5.4. Clock Gating to Freeze Unused Blocks 63
5.5. Future Work 63
References 65
1. Gabriel H. Loh – Exploiting Data-Width Locality to Increase Superscalar Execution Bandwidth, Proceedings of 35th Annual IEEE/ACM International Symposium on Microarchitecture, 2002. (MICRO-35)
2. R Sheen, S Wang, OTC Chen, Ruey-Liang Ma – Power consumption of a 2's complement adder minimized by effective dynamic data ranges, Proceedings of the 1999 IEEE International Symposium on Circuits and Systems, 1999. ISCAS '99
3. D Brooks, M Martonosi – Dynamically Exploiting Narrow Width Operands to Improve Processor Power and Performance, Proceedings. Fifth International Symposium On High-Performance Computer Architecture, 1999
4. Steve Furber, ARM System-on-Chip Architecture, 2nd Edition
5. Wayne Wolf, Modern VLSI Design – System-on-Chip Design
6. John L. Hennessy and David A. Patterson, Computer Architecture – A Quantitative Approach 3rd Edition
7. LEON2 Processor, http://www.gaisler.com/products/leon2/leon.html
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top