跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.66) 您好!臺灣時間:2026/08/15 05:02
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:張添烜
研究生(外文):Tian-Sheuan Chang
論文名稱:位元層次內積運算之超大型積體電路架構設計
論文名稱(外文):VLSI ARCHITECTURE DESIGN FOR BIT-LEVEL INNER PRODUCT
指導教授:任建葳任建葳引用關係
指導教授(外文):Prof. Chein-Wei Jen
學位類別:博士
校院名稱:國立交通大學
系所名稱:電子工程系
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:1999
畢業學年度:88
語文別:中文
論文頁數:123
中文關鍵詞:共享子表示式數位傅立葉轉換數位餘弦轉換迴旋表示式低功率設計可程式化濾波器
外文關鍵詞:common subexpression sharingDFTDCTcyclic formulationlow powerprogrammable filters
相關次數:
  • 被引用被引用:0
  • 點閱點閱:267
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
內積運算為數位訊號處理之應用,例如多媒體處理和通訊系統中的一個重要建構單元。由於應用範圍的廣泛,如何有效率的實現來滿足不同的應用需求成為一重要研究課題。在本論文中,我們對此課題以探討內積運算的位元層次設計空間加以研究,在探討中將包含可程式化和不可程式化的運算元。
對不可程式化的內積運算,我們考慮固定運算元的常數及數值特性,以使所得的乘法器為一硬體連線並具有共享子表示式之特性。因此,我們提出一個新的分散式算術技巧,此技巧將固定的輸入展開為位元層次,且我們利用共享部分乘積之和與固定輸入的稀疏非零位元特性,來減少運算數目。此所提的分散式算術技巧,已成功地運用在一個二維反數位餘弦函數晶片設計、一個處理器核心設計和一個FPGA實現上。此所提的處理器核心設計、可應用於即時H.263壓縮和數位相機之應用。在設計中,我們已將分散式算術技巧運用到極至:只用一個字元加法器和移位器。再者,此設計更進一步結合快速直接二維數位餘弦函數演算法可減少運算週期。本設計方法具簡單、規則、且容易擴展至其他更高輸出率的應用特性。對FPGA的實現,由於FPGA本身具位元運算特性,因此所提的分散式算術技巧非常適合於此結構,所得之架構設計和傳統的的分散式算術相比,可減少超過三分之二的硬體代價。
除了上述利用共享子表示式作架構最佳化外,我們也考慮演算法的重新列式。演算法的重新列式是將轉換的數學式重組為迴旋表示式,以促成更好的共享子表示式架構。我們提出兩個新的有效率的數位傅立葉轉換架構,此所提的架構結合數位傅立葉轉換係數的對稱性以增加所得的輸出率。對質數長度數位傅立葉轉換設計,可節省80%的閘面積且具有兩倍的輸出率。對二的乘冪長度的數位傅立葉轉換,和現有最佳設計相比,我們的設計具有相當競爭性的面積-時間複雜度。
針對可攜式應用環境,我們也提出了利用差分係數和差分輸入的低功率設計技巧。使用差分技巧比起使用直接輸入和係數,所需要的位元數要少,因此可以減少運算單元的大小而降低功率消耗。我們提出一個改進的演算法以有效的產生差分係數,並使差分係數方法可應用於濾波器的全頻寬而非如先前方法只用於窄頻濾波器。使用固定係數濾波器的模擬結果顯示變化動作減少幅度在全頻寬上可從1%到53%,面積可減少至50%,所得的設計比起先前方法所得的結果在應用廣度、功率消耗和面積上都較為優越。
對可程式化的濾波器設計,我們提出數字串列的架構,此設計在演算法層面用分散式算術技巧加以重組以免除累加的運算,並使用(p, q)壓縮器取代Booth編碼以得到高速運算。所得之架構和先前的設計相比,可節省達17%的硬體面積。
Inner product is an important building block to many DSP applications such as multimedia, wireless and communication systems. Due to the wide range of applications, the study on efficient implementations to meet different application requirements becomes an important research topic. In this dissertation, we study this topic by exploring the bit-level design space of inner product, including both programmable and non-programmable operands.
For non-programmable inner product, we explore its design space by considering the constant and the numerical property of the fixed operands such that the resulting multiplication is a hardwired one with common subexpression sharing. Thus, we propose a new distributed arithmetic (DA) technique that expands the fixed input into bit level so that we can take advantage of shared partial sum-of-products and sparse nonzero bits in the fixed input to reduce the number of computations. The proposed DA has been applied to a 2-D IDCT chip design, a processor core design, and FPGA implementations. The processor core design, which can be used in digital still camera and real time H.263 encoding, explores the sharing properties of the proposed DA to the extreme case: only one word adder and shifter. Furthermore, it may combine the fast direct 2-D DCT algorithm to reduce the computation cycles. The resulting architecture is quite simple, regular and easily scalable to other higher throughput applications. For FPGA implementations, due to its bit level grain size, the design with well-suited proposed DA can offer savings in excess of two-thirds of hardware cost, when comparing with the design by using conventional DA.
Besides architecture optimization with common subexpression sharing, we also consider the algorithm reformulation. The algorithm reformulation formulates transform equations into cyclic convolution form to enable better sharing with common subexpression. We have proposed two efficient DFT designs that also combine the symmetry property of DFT coefficients to increase the resulting throughput. The prime-length DFT design can save 80% of gate area with two-times fast of throughput for length N=61. The power-of-two length DFT design achieves competitive area-time complexity comparing with previous designs.
For portable applications, we also consider low power filter realization by using differential coefficients and inputs instead of using them directly such that fewer bits are required thereby reducing the size of arithmetic units and power dissipation. We present an improved algorithm to effectively generate differential coefficients so that the differential coefficients methods can be applied to full bandwidth of filters instead of only narrow band filters in previous approaches. Simulations with fixed coefficient filters indicate reduction in transition activity ranging from 1% to 53% over the full range of filter bandwidths. Reduction in area can be up to 50% due to less coefficient precision. The resulting design is superior to the one with previous approaches in applicability, power consumption, and area.
For programmable filters, we present a digit-serial architecture that uses DA form in the algorithm level for accumulation-free operations, and (p, q) compressor instead of Booth encoding for high-speed operations. The resulting design can save up to 17% hardware cost comparing with the previous approach.
Cover
摘要
Abstract
Table of Contents
CHAPTER 1 INTRODUCTION
1.1 Brr-LEVEL INNER PRODUCT
1.2 CLASSIFICATION RULE
1.3 WORD-LEVEL DESIGN
1.4 DESIGN WITH BRR-LEVEL INPUTS
1.5 DESIGNS WITH BRT-LEVEL COEFFICIENT
1.6 SUMMARY OF THE DESIGN TECHNIQUES
1.7 DISSERTATION ORGANIZATION AND CONTRIBUTION
CHAPTER 2 DESIGN AND APPLICATIOS OF COMMON SUBEXPRESSION SHARING
2.1 REVEW OF COMMON SUBEXPRESSION SHARING
2.2 ADDER-BASED DA
2-3 2-D IDCT PROCESSOR WITH ADDER-BASED DA
2.4 APROCESSOR-CORE DESIGN FOR DCT/IDCT
2.5 HARDWARE-EFFECIENT IMPLEMENTATIONS FOR TRANSFORMS IN FPGAS
2.6 SUMMARY
CHAPTER 3 TRANSFORM DESIGNS WITH CYCLIC FORMULATION AND COMMON SUBEXPRESSION SHARING
3.1 WHY CYCLIC FORMULATIONS?
3.2 PRIME-LENGTH DFT DESIGNS
3.3 POWER-OF TWO LENGTH DFT
3.4 SUMMARY
CHAPTER 4 LOW POWER DIGITAL FILTER REALIZATION WITH DIFFERENTIAL COEFFICIENTS AND INPUTS
4.1 ALGORITHM FORMULATION
4.2 RESULTS
4.3 SUMMARY
CHAPTER 5 HARDWARE EFFICIENT PROGRAMMABLE FILTER DESIGN
5.1 ALGORITHM REPORMULATION
5.2 ARCHITECTURE DESIGN
5.3 LOGIC LEVEL DESIGN WITH THE DOUBLE EDGE TRIGGERED DFFS
5.4 SUMMARY
CHAPTER 6 CONCLUSIONS AND FUTURE DIRECTIONS
6.1 CONCLUDING REMARKS
6.2 FUTURE DIRECTIONS
APPENDIX
REFERENCES
[1] P. Pirsch, N, Demassieux, and W. Gehrke, “VLSI architectures for video compression — a survey,” Proc. IEEE, vol. 83, no. 2, pp. 220-246, Feb. 1995.
[2] T. H. Meng et al, “Portable video-on-demand in wireless communication,” Proc. IEEE, vol. 83, no. 4, pp. 659-680, Apr. 1994.
[3] E. Bidet et al, ”A fast single-chip implementation of 8192 complex point FFT,” IEEE J. Solid-State Circuits, vol. 30, no. 3, pp. 300-305, Mar. 1995.
[4] K. Hwang, Computer arithmetic: principles, architecture, and design. John Wiley & Sons Inc, 1979.
[5] A. Peled and B. Liu, "A new hardware realization of digital filters," IEEE Trans. on Acoust., Speeach, Signal Processing, vol. ASSP-22, no. 6. pp. 456-462, Dec. 1974.
[6] C. H. Wei and J. S. Lou, "Multimemory block structure for implementing a digital adaptive filter using distributed arithmetic," Proc. Inst. Elec. Eng., pt.G, vo. 133, pp. 19-26, Feb, 1986.
[7] S. A. White, "Applications of distributed arithmetic to digital sequence processing: A tutorial review," IEEE ASSP Mag., vol. 6, no. 3, pp. 5-19, July 1989.
[8] C. Chen, T. Chang and C. Jen, "The IDCT Processor on the Adder-based Distributed Arithmetic," Proc. Symp. VLSI Circ., pp.36-37, 1996.
[9] T. S. Chang, C. W. Jen and C. S. Chen, "A new distributed arithmetic algorithm and its applications to IDCT,” accepted for publication in IEE Proceedings: Circuits, Device and Systems, 1999.
[10] D.R. Bull and D. H. Horrocks, "Primitive operator digital filters," IEE Proc. Circuits Devices Syst., vol. 138, no. 3, pp. 401-412, June 1991
[11] A. G. Dempster and M. D. Macleod, "Constant integer multiplication using minimum adders," IEE Proc. Circuits Devices Syst., vol. 141, no. 5, Oct. 1991.
[12] A. G. Dempster and M. D. Macleod, "Use of minimum-adder multiplier blocks in FIR digital filters," IEEE Trans. Circuits Sys. vol. 42, no. 9, pp. 569-577, Sept. 1995.
[13] M. Potkonjak, M. Srivastava and A. P. Chandrakasan, " Multiple constant multiplications: efficient and versatile framework and algorithms for exploring common subexpression elimination," IEEE Trans. Computer-Aided Design., vol. 15, no. 2, pp. 151-165, Feb. 1996.
[14] M. Mehendale, S. d. Sherlekar, and g. Vekantesh, “Synthesis of multiplierless FIR filters with minimum number of additions,” Proc. ICCAD, pp. 668-671, 1995.
[15] R. Pasko et al, “A new algorithm for elimination for common subexpressions,” IEEE Trans. Computer-Aided Design., vol. 18, no. 1, pp. 58-68, Jan. 1999.
[16] C. C. Ju, “A high-throughput DCT/IDCT architecture and design methodology with application to real-time digital video codec system and associated CAD design,” master thesis, Nat’l Chiao-Tung Univ. 1997.
[17] R. I. Hartley, "Subexpression sharing in filters using canonic signed digit multipliers," IEEE Trans. Circuits Sys. vol. 43, no. 2, pp. 677-688, Oct. 1996.
[18] S. Y. Kung, VLSI Array Processors, New Jersey: Prentice-Hall, 1988
[19] TMS320C3X users guide, Texas Instruments, 1990.
[20] TMS320C80 (MVP) master processor user''s guide, Texas Instruments, 1994.
[21] G. K. Ma and F. J. Taylor, “Multiplier policies for digital signal processing,“ IEEE ASSP Mag., vol. 7, no. 1, pp. 6-20, Jan. 1990.
[22] R. Hartley, and K. K. Parhi, Digit-serial computation, Kluwer Academic Publishers, 1995.
[23] A. V. Oppenheim and R. W. Schafer, Discrete-time signal processing, Prentice-Hall, 1989.
[24] J. I. Guo, C-M. Liu, and C-W Jen, “The efficient memory-based VLSI array designs for DFT and DCT,” IEEE Trans. Circuits Syst. II. vol. 39, pp. 723-733, Oct. 1992.
[25] H-R Lee, C-W Jen and C-M Liu, “A new hardware-efficient architecture for programmable FIR filters,” IEEE Trans. Circuits Syst., vo.43, no.9, pp.637-644, Sep. 1996.
[26] J. R. Choi, L. H. Jang, S. W. Jung and J. H. Choi, “ Structured design of a 288-tap FIR filter by optimized partial product tree compression,” IEEE J. Solid-State Circuits, vol. 32. No. 3, pp. 468-476, Mar. 1997.
[27] D. Reuver and H. Klar, “A configurable convolution chip with programmable coefficients,” IEEE J. Solid-State Circuits, vol.27, pp.1121-1123, July 1992.
[28] D. L. Jones, "Efficient computation of time-varying and adaptive filters," IEEE Trans. on Signal Processing,vol.41, no. 3. pp.1077-1086, March, 1993.
[29] S. Ramprasad, N. R. Shanbhag, and I. N. Hajj, “ Decorrelating (DECOR) transformations for low-power adaptive filters,” Proc. of Intl. Symp. on Low-Power Electronics Design, August 1998, Monterey, CA.
[30] H. Samueli, “An improved search algorithm for the design of multiplierless FIR filters with powers-of-two coefficients,” IEEE Trans. Circuits Syst., vol. 36, pp. 1044-1047. July, 1989.
[31] Y. C. Lim et al, “Signed power-of-two term allocation scheme for the design of digital fitlers,” IEEE Trans. Circuits Syst., vol. 46, pp. 577-584. May, 1999.
[32] M. T. Sun, T. C. Chen and A. M. Gottlieh, “VLSI implementation of a 16x16 discrete cosine transform,” IEEE Trans. Circuits, Systems, vol. CAS-36, pp.610-617, Apr. 1989.
[33] S. Uramoto, Y. Inoue, A. Takabatake, J. Takeda, H. Yamashita, H. Terane, and M. Yoshimoto, ‘A 100Mhz 2-D discrete cosine transform core processor,’ IEEE J. Solid State Circuits, vol. 27, no. 4, pp. 492-499, Apr. 1992.
[34] K. R. Rao and P. Yip, Discrete Cosine Transforms - Algorithms, Advantages, Applications, Boston, MA: Academic, 1990.
[35] K. R. Rao, and J. J. Hwang, Techniques and Standards for image, Video and Audio Coding, New Jersey: Prentice Hall, 1996
[36] K. Nourji and N. Demassieux, "Optimization of real-time VLSI architectures for distributed arithmetic-based algorithms: application to HDTV filters," IEEE ISCAS, vol. 4. pp.223-226,1994.
[37] K. Suzuki, M. Yamashina, J. Goto, T. Inoue, Y. Koseki, T. Horiuchi, N. Hamatake, K. Kumagai, T. Enomoto, and H. Yamada, “A 2.4ns, 16-bit, 0.5um, CMOS arithmetic logic unit for microprogrammable video sequence processor LSIs,” Proc. IEEE CICC., May, 1993, pp.12.4.1-12.4.4.
[38] IEEE Std 1180-1990: ‘IEEE standard specifications for the Implementations of 8×8 inverse discrete cosine transform’.
[39] T. S. Chang, C-S Kung and C.-W. Jen, “A simple processor core design for DCT/IDCT,” accepted for publication in IEEE Trans. Circuits and Systems for Video Technology, 1999.
[40] M. Kovac and N. Ranganathan, "JAGUAR: A VLSI Architecture for JPEG Image Compression Standard," Proceedings of IEEE, vol. 83, no. 2, pp. 247-258, Feb 1995.
[41] A. Madisetti and A. N. Wilson, Jr., “A 100 Mhz 2-D 8×8 DCT/IDCT Processor for HDTV Applications,” IEEE Trans. Video Technology, vol. 5, no. 2, pp. 158-164, Apr. 1995.
[42] C.-Y. Hung, and P. Landman, “Compact inverse discrete cosine transform circuit for MPEG video decoding,” IEEE Workshop on Signal Processing Systems, pp. 364 — 373, 1997.
[43] Y. Katayama, T. Kitsuki, and Y. Ooi, “A block processing unit in a single-chip MPEG-2 video encoder LSI,” IEEE Workshop on Signal Processing Systems, pp. 459-468, 1997.
[44] R. Rambaldi, A. Ugazzoni, and R. Guerrieri, “A 35 W 1.1 V gate array 8×8 IDCT processor for video-telephony,” in Proc. IEEE ICASSP, vol. 5. pp.2993-2996, 1998.
[45] K. Okamoto et al, “ A DSP for DCT-based and wavelet-based video codecs for consumer applications,” IEEE J. Solid-State Circuits, vol. 32, no. 3, pp. 460-467, Mar. 1997.
[46] T. Xanthopoulos and A. Chandrakasan, “ A low-power IDCT macrocell for MPEG2 MP@ML exploiting data distribution properties for minimal activity,” in Proc. Symp. VLSI Circ., pp.38-39, 1998.
[47] H. S. Hou, “A fast recursive algorithm for computing the discrete cosine transform,” IEEE Trans. on Acoust., Speech, Signal Processing, vol. ASSP-35, no. 10. pp. 1455-1461, Oct. 1987.
[48] C. Loeffler, A. Ligtenberg and G. S. Moschytz, “Practical fast 1-D DCT algorithms with 11 multiplications,” in Proc. IEEE ICASSP, vol. 2, pp. 988-991.1989.
[49] N. I. Cho and S. U. Lee, “Fast algorithm and implementations of 2-D DCT,” IEEE Trans. Circuits Sys., vol. 38, no. 3, pp. 297-305, Mar. 1991.
[50] Y-P Lee, T-H Chen, L. G. Chen, M. J. Chen, and C. W Ku, “A cost-effective architecture for 8×8 two-dimensional DCT/IDCT using direct method,” IEEE Trans. Video Technology, vol. 7, no. 3, June 1997.
[51] L.-G. Chen; J.-Y. Jiu, H.-C. Chang, Y.-P. Lee, and C.-W. Ku, “Low power 2D DCT chip design for wireless multimedia terminals,” Proc. ISCAS, vol.4, pp. 41-44, 1998
[52] S.-C. Hsia, B.-D. Liu, J.-F. Yang, and B.-L. Bai, “VLSI implementation of parallel coefficient-by-coefficient two-dimensional IDCT processor,” IEEE Trans. Video Technology, vol. 5, pp. 396-406, Oct. 1995
[53] J. Golston, “Single-chip H.324 videoconferencing, ” IEEE Micro, vol. 16, no. 4, pp. 21-33, Aug. 1996.
[54] W. Houl, “An 8×8 discrete cosine transform implementation on the TMS320C25 or TMS320C30,” Application report: SPRA115, Texas Instruments., 1997.
[55] “TMS320C62x assembly benchmarks”, http://www.ti.com/sc/docs/dsps/products/c6000/62xbench.htm, 1997.
[56] M. Yoshida, H. Ohtomo, I. Kuroda, “A new generation 16-bit general purpose programmable DSP and its video rate application,” IEEE Workshop on VLSI Signal Processing, pp. 93 —101, 1993.
[57] I. Kuroda, “Processor architecture driven algorithm optimization for fast 2-D-DCT,” IEEE Workshop on VLSI Signal Processing, VIII, pp. 481 —490, 1995.
[58] Compass, PASSPORT library, 0.6 micron 5-volt high performance standard cell library, 1996.
[59] T.-S. Chang and C.-W. Jen, “Hardware-efficient implementations for transforms in programmable logic device,” submitted to IEE Proceedings - Computers and Digital Techniques, 1999.
[60] C.H, Dick, “FPGA based systolic array architectures for computing the discrete Fourier transform”, IEEE International Symposium on Circuits and Systems, 1996. vol.2, pp. 465 — 468.
[61] J. Heron, D. Trainor, R. Woods, M. P Fargues, R.D. Hippenstiel, “Image compression algorithms using re-configurable logic”, Thirty-First Asilomar Conference on Signals, Systems & Computers, 1997 vol. 1 , pp. 399 —403.
[62] N. W. Bergmann, and Y. Y. Chung, “Video compression with custom computers”, IEEE Transactions on Consumer Electronics, vol. 43, no. 3, pp. 925-933, Aug. 1997.
[63] D. R. Bull and G. Wacey, “Bit-serial digital filter architecture using RAM-based delay operators”, IEE Proc.-Circuits Devices Syst., vol. 141, no. 5, Oct. 1994.
[64] G. R. Goslin and B. Newgard, ?-Tap, 8-Bit FIR Filter Application Guide," Xilinx Publications, 1994.
[65] Xilinx, The Programmable Logic Data Book, 1994.
[66] T.-S. Chang, J.-I. Guo and C.-W. Jen, “Hardware efficient DFT designs with cyclic convolution and subexpression sharing” submitted to IEEE Trans on Circuits and System, ⅡAnalog and Digital Signal Processing, 1999.
[67] T.-S. Chang, J.-I. Guo and C.-W. Jen, “A new hardware efficient algorithm and architecture for 1-d discrete Fourier transform” submitted to IEEE Trans. on Signal Processing, 1999.
[68] H. T. Kung, “ Why systolic architectures?” Comput. Mag., vol. 15, pp.37-45, Jan. 1982.
[69] J. A. Beraldin, T. Aboulnasr, and W. Steenart, “Efficient one-dimensional systolic array realization of discrete Fourier transform,” IEEE Trans. Acoust. Speech, Signal Processing, vol. 36, pp. 1665-1667, Oct, 1988.
[70] J. A. Beraldin, T. Aboulnasr, and W. Steenaart, “Efficient one-dimensional systolic array realization of the discrete Fourier transform,” IEEE Trans. Circuits Syst., vol. 36, no. 1, Jan. 1989.
[71] W. H. Fang and M. L. Wu, “An efficient unified systolic architecture for the computation of discrete trigonometric transforms,” Proc. ISCAS, vol.3, pp. 2092 —2095, 1997
[72] L. W. Chang and M. Y. Chen, “A new systolic array for discrete Fourier transform,” IEEE Trans. Acoust. Speech, Signal Processing, vol. 36, pp. 1665-1666, Oct. 1988.
[73] N. R. Murthy and M. N. s. Swamy, “On the real-time computation of DFT and DCT through systolic architectures,” IEEE Trans. Signal Processing, vol. 42, no. 4, pp. 988-991, Apr. 1994.
[74] E. Chan and S. Panchanathan, “A VLSI architecture for DFT, ” Proc. the 36th Midwest Symposium on Circuits and Systems, vol. 1, pp. 292-295, 1993
[75] S. G. Sedukhin, “A new systolic architecture for pipeline prime factor DFT algorithm,” Fourth Great Lakes Symposium on VLSI. pp. 40-45, 1994.
[76] S. Gudvangen, and A. Holt, “Computation of prime factor DFT and DHT/DCT algorithms using cyclic and skew-cyclic bit-serial semisystolic IC convolvers,” IEE Proceedings-G, vol. 137, no. 5, pp. 373 — 389, Oct. 1990.
[77] T. S. Chang, and C-W. Jen, “Hardware efficient transform designs with cyclic formulation and subexpression sharing,” Proc. ISCAS, vol. 2. pp. 398-401, 1998.
[78] C.M. Radar, “Discrete Fourier transforms when the number of data samples is prime,” Proc. IEEE, vol. 56, pp. 1107-1108, 1968.
[79] C.S. Burrus and T.W. Parks, DFT/FFT and convolution algorithms, John Wiley & Sons, 1985.
[80] L. R. Rabiner and B. Gold, Theory and Application of Digital Signal Processing. Chap. 10, Prentice-Hall, 1975.
[81] Y.-H. Chan, and W.-C. Siu, “A cyclic correlated structure for the realization of the discrete cosine transform,” IEEE Trans. Circuits Syst. II. vol. 39, pp. 109-113, Feb, 1992
[82] V. Boriakoff, “FFT computation with systolic arrays, a new architecture,” IEEE Trans. Circuits Syst. II. vol. 41, no. 4, pp. 278-284, Apr. 1994.
[83] S.-F. Hsiao and C.-Y. Yen, “New unified VLSI architectures for computing DFT and other transforms,” Proc. ISCAS, 1997.
[84] E. H. Wold and A. M. Despain, “Pipeline and parallel-pipeline FFT processors for VLSI implementations,” IEEE Trans. on computers, vol. C-33, no. 5, pp. 414-426, May, 1984.
[85] M. Vergara et al, “A 195KFFT/s (256-points) high performance FFT/IFFT processor for OFDM applications,” 1998.
[86] L. Jia et al, “A new VLSI-oriented FFT algorithm and implementation”, Proc. IEEE ASIC Conference, pp. 337 — 341, 1998
[87] S. He and M. Torkelson, “Design and implementation of a 1024-point pipeline FFT processor,” IEEE Proc. CICC, pp. 131-134, 1998.
[88] H. K. Garg, Digital signal processing algorithms: number theory, convolution, fast Fourier transforms, and applications, CRC Press, 1998.
[89] C. Joanblanq, F. Rothan, and P. Senn, “A video delay line compiler,” in Proc. ISCAS, May, 1990.
[90] M. Mehendale, S. D. Sherlekar, and G. Venkatesh, “Coefficient optimization for low power realization of FIR filters,” in IEEE Workshop on VLSI Signal Processing, pp. 352-361, 1995.
[91] Y. C. Lim, “Predictive coding for FIR filter wordlength reduction,” IEEE Tran. Circuits Syst, vol. 32, no. 4, pp. 365-372, Apr. 1985.
[92] N Sankarayya, K. Roy, and D. Bhattacharya, “Algorithms for low Power and high speed FIR filter realization using differential coefficients,” IEEE Tran. Circuits Syst. II. vol. 44. No. 6. pp.488-497, June, 1997.
[93] T. S. Chang, and C. W. Jen, “Low power FIR filter realizations with differential coefficients and inputs,” Proc. ICASSP, pp. 3009-3012, 1998. May, Seattle, WA.
[94] T. S. Chang, Y-H. Chu, and C.-W. Jen, “Low power FIR filter realization with differential coefficients and inputs,” revised, IEEE Trans on Circuits and System, ⅡAnalog and Digital Signal Processing, 1999.
[95] N Sankarayya, K. Roy, and D. Bhattacharya, “Optimizing computations in a transposed direct form realization of floating-point LTI-FIR systems,” Proc. ICCAD. pp. 120-125, Nov, 1997.
[96] T.-S. Chang and C.-W. Jen, “High speed pipelined programmable FIR filter design,” submitted to IEE Proceedings: Circuits, Device and Systems, 1999.
[97] B. Edwards, A. Corry, N. Weste, and C. Greenberg, “ A single chip ghost canceller,” in Proc. IEEE 1992 Custom Integrated Circuits Conf., pp.26.5.1-4, May 1992.
[98] D. J. Pearson et al. "Digital FIR filters for high speed PRML disk read channels,", IEEE J. Solid-State Circuits, vo.30, no. 12, pp. 1517-1523, Dec. 1995.
[99] M. Sid-ahmed, “A systolic realization for 2-D digital filters,” IEEE Trans. on Acoust., Speeach, Signal Processing, vol. 37, pp.560-565, Apr. 1989.
[100] O. L.Mac Sorley, “ High speed arithmetic in binary computers,” IRE Proc., vol. 49, pp. 67-91, Jan. 1961.
[101] D. Villeger and V. G. Oklobdzija, “Evaluation of Booth encoding techniques for parallel multiplier implementation.” Electronics Letters, vol. 29, no. 23, pp.2016-2017, Nov. 1993.
[102] P. Song and G. De Michelli, “Circuits and architecture trade-offs for high speed multiplication,” IEEE J. Solid-State Circuits, vol. 26, no. 9, pp.1184-1198, Sept. 1991.
[103] R. Hossain, L. D. Wronski, and A. Albicki “Low power design using double edge flips-flops,” IEEE Trans. VLSI systems, vol. 2. No. 2, pp.261-265, June 1994.
[104] V. G. Oklobdzija, D. Villeger and S. S. Liu, “A method for speed optimized partial product reduction and generation of fast parallel multipliers using an algorithmic approach,” IEEE Trans. Computers, vol. 45, no. 3, pp. 294-306, Mar. 1996.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top
無相關期刊