資料載入處理中...
跳到主要內容
臺灣博碩士論文加值系統
:::
網站導覽
|
首頁
|
關於本站
|
聯絡我們
|
國圖首頁
|
常見問題
|
操作說明
English
|
FB 專頁
|
Mobile
免費會員
登入
|
註冊
切換版面粉紅色
切換版面綠色
切換版面橘色
切換版面淡藍色
切換版面黃色
切換版面藍色
功能切換導覽列
(216.73.216.73) 您好!臺灣時間:2026/07/22 11:51
字體大小:
字級大小SCRIPT,如您的瀏覽器不支援,IE6請利用鍵盤按住ALT鍵 + V → X → (G)最大(L)較大(M)中(S)較小(A)小,來選擇適合您的文字大小,如為IE7或Firefoxy瀏覽器則可利用鍵盤 Ctrl + (+)放大 (-)縮小來改變字型大小。
字體大小變更功能,需開啟瀏覽器的JAVASCRIPT功能
:::
詳目顯示
recordfocus
第 1 筆 / 共 1 筆
/1
頁
論文基本資料
摘要
外文摘要
目次
參考文獻
電子全文
紙本論文
QR Code
本論文永久網址
:
複製永久網址
Twitter
研究生:
林聖勳
研究生(外文):
Sheng-Hsun Lin
論文名稱:
可被窄寬度運算共用之單一管線資料路徑設計
論文名稱(外文):
A Single Pipeline Datapath Design for Joinable Narrow-operand Operations
指導教授:
鍾崇斌
指導教授(外文):
Chung-Ping Chung
學位類別:
碩士
校院名稱:
國立交通大學
系所名稱:
資訊科學與工程研究所
學門:
工程學門
學類:
電資工程學類
論文種類:
學術論文
論文出版年:
2006
畢業學年度:
94
語文別:
英文
論文頁數:
65
中文關鍵詞:
資料路徑共用
、
窄寬度運算
外文關鍵詞:
datapath sharing
、
narrow-operand operations
相關次數:
被引用:0
點閱:195
評分:
下載:8
書目收藏:0
現今大多數之處理器之架構都採32位元或者更高之位元數. 然而整數運算大多數時間都不會用到資料路徑中完整的寬度. 若將Source Operand Bus及ALU分割為多組block, 則此資料路徑便具備同時處理一道以上運算之能力.本論文提出讓單一資料路徑被兩道具備窄寬度特性之指令合用之設計. 讓其中一道運算之Operand Block以反轉(Turnaround)之次序, 來達到極有效率之資料路徑合用機制.相較於傳統以直接Shift方式, 本設計不論就面積與電路延遲上, 都有較佳之表現. 此外, 本論文亦提出一縮短ALU因分割為多組ALU Block所造成延遲之設計. 最後, 本論文提出將此機制整合至一典型五階MIPS管線之方式.
Most general-purpose processors and embedded processors have 32-bit word widths or wider. However, integer operations rarely need the full 32-bit dynamic range of the datapath. If we partition the operand bus, result bus, and ALU into several blocks, the datapath could perform more than one operation in parallel.
In this thesis, mechanisms to join two narrow-operand operations together to share a single datapath are proposed. We proposed one novel ALU-sharing scheme by turning around operation-block ordering. Efficient designs to merge operands to share buses and ALU based on the technique are proposed and discussed. Compared with traditional “shift” approach, the turnaround approach has many advantages on area and delay. Besides, a technique to mitigate the delay overhead of the partitioned ALU by swapping operands is proposed. We also made performance simulation to help decide how to partition the datapath. Finally, how to integrate such datapath into a MIPS five-stage pipeline and required modifications are discussed.
摘要 i
Abstract ii
誌謝 iii
Table of Contents iv
List of Figures vi
List of Tables viii
Chapter 1. Introduction 1
1.1. Narrow-Operand ALU Operations 1
1.2. Significant Bit-width of an ALU Operation 2
1.3. Organization of this Thesis 4
Chapter 2. Background and Related Work 5
2.1. Distribution of Significant Bit-widths of ALU Operations 5
2.2. Datapath in Multi-Bitwidth Pipeline 8
2.3. More Flexible ALU-sharing Mechanism 10
2.4. Motivation 12
2.5. Objective 13
Chapter 3. Design 15
3.1. Overview of Datapath-Sharing 15
3.1.1. Constraints for Joining Two Instructions 15
3.1.1.1. Structural Hazards 15
3.1.1.2. Data Hazards 16
3.1.2. How ALU is Shared by Two Operations 17
3.1.3. Possible Modifications in a Five-stage MIPS-like Pipeline 18
3.2. Modifications in Instruction-Decode Stage 20
3.2.1. Type-Check and Data-dependency Check 20
3.2.2. Width-Check 20
3.2.2.1. Width-Determination Logic 20
3.2.2.2. Operand-boundary Signals 25
3.2.2.3. Width-Check Logic 26
3.3. Modifications in Execution Stage 30
3.3.1. Deciding the Block Width 30
3.3.2. Merging Operands from Two Operations 31
3.3.3. ALU in the Turnaround Approach 34
3.3.3.1. Operation Swapping 39
3.3.3.2. Widening the ALU 42
3.4. Modifications in Memory Access and Write-back Stage 42
3.4.1. Sign-extending the Joined Results 43
3.5. Integrating the Design into a Five-stage MIPS-like Pipeline Datapath 45
Chapter 4. Experiments 49
4.1. Goals of Our Experiments 49
4.2. Simulation Environment 49
4.3. Comparison between Different Operand-Merging Schemes 50
4.4. Hardware Cost of Operation-Swapping Mechanism 52
4.5. Choosing ALU Block Width 54
4.6. Choosing ALU Width 57
4.7. Final Proposal of the ALU Design 58
Chapter 5. Conclusion and Future Work 61
5.1. The Turnaround-based Sharing Mechanism 61
5.2. Reducing Additional Delay by Avoiding Unnecessary Partitioning 61
5.3. Applying the Design to Architectures with Shifter Concatenated with ALU 62
5.4. Clock Gating to Freeze Unused Blocks 63
5.5. Future Work 63
References 65
1. Gabriel H. Loh – Exploiting Data-Width Locality to Increase Superscalar Execution Bandwidth, Proceedings of 35th Annual IEEE/ACM International Symposium on Microarchitecture, 2002. (MICRO-35)
2. R Sheen, S Wang, OTC Chen, Ruey-Liang Ma – Power consumption of a 2's complement adder minimized by effective dynamic data ranges, Proceedings of the 1999 IEEE International Symposium on Circuits and Systems, 1999. ISCAS '99
3. D Brooks, M Martonosi – Dynamically Exploiting Narrow Width Operands to Improve Processor Power and Performance, Proceedings. Fifth International Symposium On High-Performance Computer Architecture, 1999
4. Steve Furber, ARM System-on-Chip Architecture, 2nd Edition
5. Wayne Wolf, Modern VLSI Design – System-on-Chip Design
6. John L. Hennessy and David A. Patterson, Computer Architecture – A Quantitative Approach 3rd Edition
7. LEON2 Processor, http://www.gaisler.com/products/leon2/leon.html
電子全文
國圖紙本論文
推文
當script無法執行時可按︰
推文
網路書籤
當script無法執行時可按︰
網路書籤
推薦
當script無法執行時可按︰
推薦
評分
當script無法執行時可按︰
評分
引用網址
當script無法執行時可按︰
引用網址
轉寄
當script無法執行時可按︰
轉寄
top
相關論文
相關期刊
熱門點閱論文
無相關論文
1.
黃鑑水、張憲卿、劉桓吉, 1994,臺灣南部觸口斷層之地質調查與探勘,經濟部中央地質調查所彙刊,第九號,29-50頁。
1.
考量管線時間之延伸指令集
2.
細線化前之區塊深度值測試與其對系統設計之影響
3.
用於區塊繪圖之階層式儲存方式設計
4.
以低複雜度延伸有效的指令窗以容忍資料讀取失誤延遲
5.
指令快取記憶體的電源管理-(對程式流程有感知能力的昏睡指令記憶體)
6.
快取誤失類型辨認及其在動態預測快取誤失位置之用途
7.
繪圖處理器之材質貼圖下有效率之材質記憶體系統設計
8.
可動態重組之處理單元於頂點與像素處理
9.
可動態重組之材質處理單元
10.
在 FPGA 上實現線上遞迴式獨立成分分析算法
11.
混熱擾流力與分子動能模擬之 GPU 加速
12.
低功耗高效能多核心視訊解碼器設計
13.
霍夫轉換與影像中規律排列圓形區域之偵測
14.
晚點較好:延遲和聚集固態硬碟TRIM指令與發送時機管理方法
15.
藉由全動態斷定執行減少難以預測的分支指令對於效能的影響
簡易查詢
|
進階查詢
|
熱門排行
|
我的研究室