|
[1] P. Pirsch, N, Demassieux, and W. Gehrke, “VLSI architectures for video compression — a survey,” Proc. IEEE, vol. 83, no. 2, pp. 220-246, Feb. 1995. [2] T. H. Meng et al, “Portable video-on-demand in wireless communication,” Proc. IEEE, vol. 83, no. 4, pp. 659-680, Apr. 1994. [3] E. Bidet et al, ”A fast single-chip implementation of 8192 complex point FFT,” IEEE J. Solid-State Circuits, vol. 30, no. 3, pp. 300-305, Mar. 1995. [4] K. Hwang, Computer arithmetic: principles, architecture, and design. John Wiley & Sons Inc, 1979. [5] A. Peled and B. Liu, "A new hardware realization of digital filters," IEEE Trans. on Acoust., Speeach, Signal Processing, vol. ASSP-22, no. 6. pp. 456-462, Dec. 1974. [6] C. H. Wei and J. S. Lou, "Multimemory block structure for implementing a digital adaptive filter using distributed arithmetic," Proc. Inst. Elec. Eng., pt.G, vo. 133, pp. 19-26, Feb, 1986. [7] S. A. White, "Applications of distributed arithmetic to digital sequence processing: A tutorial review," IEEE ASSP Mag., vol. 6, no. 3, pp. 5-19, July 1989. [8] C. Chen, T. Chang and C. Jen, "The IDCT Processor on the Adder-based Distributed Arithmetic," Proc. Symp. VLSI Circ., pp.36-37, 1996. [9] T. S. Chang, C. W. Jen and C. S. Chen, "A new distributed arithmetic algorithm and its applications to IDCT,” accepted for publication in IEE Proceedings: Circuits, Device and Systems, 1999. [10] D.R. Bull and D. H. Horrocks, "Primitive operator digital filters," IEE Proc. Circuits Devices Syst., vol. 138, no. 3, pp. 401-412, June 1991 [11] A. G. Dempster and M. D. Macleod, "Constant integer multiplication using minimum adders," IEE Proc. Circuits Devices Syst., vol. 141, no. 5, Oct. 1991. [12] A. G. Dempster and M. D. Macleod, "Use of minimum-adder multiplier blocks in FIR digital filters," IEEE Trans. Circuits Sys. vol. 42, no. 9, pp. 569-577, Sept. 1995. [13] M. Potkonjak, M. Srivastava and A. P. Chandrakasan, " Multiple constant multiplications: efficient and versatile framework and algorithms for exploring common subexpression elimination," IEEE Trans. Computer-Aided Design., vol. 15, no. 2, pp. 151-165, Feb. 1996. [14] M. Mehendale, S. d. Sherlekar, and g. Vekantesh, “Synthesis of multiplierless FIR filters with minimum number of additions,” Proc. ICCAD, pp. 668-671, 1995. [15] R. Pasko et al, “A new algorithm for elimination for common subexpressions,” IEEE Trans. Computer-Aided Design., vol. 18, no. 1, pp. 58-68, Jan. 1999. [16] C. C. Ju, “A high-throughput DCT/IDCT architecture and design methodology with application to real-time digital video codec system and associated CAD design,” master thesis, Nat’l Chiao-Tung Univ. 1997. [17] R. I. Hartley, "Subexpression sharing in filters using canonic signed digit multipliers," IEEE Trans. Circuits Sys. vol. 43, no. 2, pp. 677-688, Oct. 1996. [18] S. Y. Kung, VLSI Array Processors, New Jersey: Prentice-Hall, 1988 [19] TMS320C3X users guide, Texas Instruments, 1990. [20] TMS320C80 (MVP) master processor user''s guide, Texas Instruments, 1994. [21] G. K. Ma and F. J. Taylor, “Multiplier policies for digital signal processing,“ IEEE ASSP Mag., vol. 7, no. 1, pp. 6-20, Jan. 1990. [22] R. Hartley, and K. K. Parhi, Digit-serial computation, Kluwer Academic Publishers, 1995. [23] A. V. Oppenheim and R. W. Schafer, Discrete-time signal processing, Prentice-Hall, 1989. [24] J. I. Guo, C-M. Liu, and C-W Jen, “The efficient memory-based VLSI array designs for DFT and DCT,” IEEE Trans. Circuits Syst. II. vol. 39, pp. 723-733, Oct. 1992. [25] H-R Lee, C-W Jen and C-M Liu, “A new hardware-efficient architecture for programmable FIR filters,” IEEE Trans. Circuits Syst., vo.43, no.9, pp.637-644, Sep. 1996. [26] J. R. Choi, L. H. Jang, S. W. Jung and J. H. Choi, “ Structured design of a 288-tap FIR filter by optimized partial product tree compression,” IEEE J. Solid-State Circuits, vol. 32. No. 3, pp. 468-476, Mar. 1997. [27] D. Reuver and H. Klar, “A configurable convolution chip with programmable coefficients,” IEEE J. Solid-State Circuits, vol.27, pp.1121-1123, July 1992. [28] D. L. Jones, "Efficient computation of time-varying and adaptive filters," IEEE Trans. on Signal Processing,vol.41, no. 3. pp.1077-1086, March, 1993. [29] S. Ramprasad, N. R. Shanbhag, and I. N. Hajj, “ Decorrelating (DECOR) transformations for low-power adaptive filters,” Proc. of Intl. Symp. on Low-Power Electronics Design, August 1998, Monterey, CA. [30] H. Samueli, “An improved search algorithm for the design of multiplierless FIR filters with powers-of-two coefficients,” IEEE Trans. Circuits Syst., vol. 36, pp. 1044-1047. July, 1989. [31] Y. C. Lim et al, “Signed power-of-two term allocation scheme for the design of digital fitlers,” IEEE Trans. Circuits Syst., vol. 46, pp. 577-584. May, 1999. [32] M. T. Sun, T. C. Chen and A. M. Gottlieh, “VLSI implementation of a 16x16 discrete cosine transform,” IEEE Trans. Circuits, Systems, vol. CAS-36, pp.610-617, Apr. 1989. [33] S. Uramoto, Y. Inoue, A. Takabatake, J. Takeda, H. Yamashita, H. Terane, and M. Yoshimoto, ‘A 100Mhz 2-D discrete cosine transform core processor,’ IEEE J. Solid State Circuits, vol. 27, no. 4, pp. 492-499, Apr. 1992. [34] K. R. Rao and P. Yip, Discrete Cosine Transforms - Algorithms, Advantages, Applications, Boston, MA: Academic, 1990. [35] K. R. Rao, and J. J. Hwang, Techniques and Standards for image, Video and Audio Coding, New Jersey: Prentice Hall, 1996 [36] K. Nourji and N. Demassieux, "Optimization of real-time VLSI architectures for distributed arithmetic-based algorithms: application to HDTV filters," IEEE ISCAS, vol. 4. pp.223-226,1994. [37] K. Suzuki, M. Yamashina, J. Goto, T. Inoue, Y. Koseki, T. Horiuchi, N. Hamatake, K. Kumagai, T. Enomoto, and H. Yamada, “A 2.4ns, 16-bit, 0.5um, CMOS arithmetic logic unit for microprogrammable video sequence processor LSIs,” Proc. IEEE CICC., May, 1993, pp.12.4.1-12.4.4. [38] IEEE Std 1180-1990: ‘IEEE standard specifications for the Implementations of 8×8 inverse discrete cosine transform’. [39] T. S. Chang, C-S Kung and C.-W. Jen, “A simple processor core design for DCT/IDCT,” accepted for publication in IEEE Trans. Circuits and Systems for Video Technology, 1999. [40] M. Kovac and N. Ranganathan, "JAGUAR: A VLSI Architecture for JPEG Image Compression Standard," Proceedings of IEEE, vol. 83, no. 2, pp. 247-258, Feb 1995. [41] A. Madisetti and A. N. Wilson, Jr., “A 100 Mhz 2-D 8×8 DCT/IDCT Processor for HDTV Applications,” IEEE Trans. Video Technology, vol. 5, no. 2, pp. 158-164, Apr. 1995. [42] C.-Y. Hung, and P. Landman, “Compact inverse discrete cosine transform circuit for MPEG video decoding,” IEEE Workshop on Signal Processing Systems, pp. 364 — 373, 1997. [43] Y. Katayama, T. Kitsuki, and Y. Ooi, “A block processing unit in a single-chip MPEG-2 video encoder LSI,” IEEE Workshop on Signal Processing Systems, pp. 459-468, 1997. [44] R. Rambaldi, A. Ugazzoni, and R. Guerrieri, “A 35 W 1.1 V gate array 8×8 IDCT processor for video-telephony,” in Proc. IEEE ICASSP, vol. 5. pp.2993-2996, 1998. [45] K. Okamoto et al, “ A DSP for DCT-based and wavelet-based video codecs for consumer applications,” IEEE J. Solid-State Circuits, vol. 32, no. 3, pp. 460-467, Mar. 1997. [46] T. Xanthopoulos and A. Chandrakasan, “ A low-power IDCT macrocell for MPEG2 MP@ML exploiting data distribution properties for minimal activity,” in Proc. Symp. VLSI Circ., pp.38-39, 1998. [47] H. S. Hou, “A fast recursive algorithm for computing the discrete cosine transform,” IEEE Trans. on Acoust., Speech, Signal Processing, vol. ASSP-35, no. 10. pp. 1455-1461, Oct. 1987. [48] C. Loeffler, A. Ligtenberg and G. S. Moschytz, “Practical fast 1-D DCT algorithms with 11 multiplications,” in Proc. IEEE ICASSP, vol. 2, pp. 988-991.1989. [49] N. I. Cho and S. U. Lee, “Fast algorithm and implementations of 2-D DCT,” IEEE Trans. Circuits Sys., vol. 38, no. 3, pp. 297-305, Mar. 1991. [50] Y-P Lee, T-H Chen, L. G. Chen, M. J. Chen, and C. W Ku, “A cost-effective architecture for 8×8 two-dimensional DCT/IDCT using direct method,” IEEE Trans. Video Technology, vol. 7, no. 3, June 1997. [51] L.-G. Chen; J.-Y. Jiu, H.-C. Chang, Y.-P. Lee, and C.-W. Ku, “Low power 2D DCT chip design for wireless multimedia terminals,” Proc. ISCAS, vol.4, pp. 41-44, 1998 [52] S.-C. Hsia, B.-D. Liu, J.-F. Yang, and B.-L. Bai, “VLSI implementation of parallel coefficient-by-coefficient two-dimensional IDCT processor,” IEEE Trans. Video Technology, vol. 5, pp. 396-406, Oct. 1995 [53] J. Golston, “Single-chip H.324 videoconferencing, ” IEEE Micro, vol. 16, no. 4, pp. 21-33, Aug. 1996. [54] W. Houl, “An 8×8 discrete cosine transform implementation on the TMS320C25 or TMS320C30,” Application report: SPRA115, Texas Instruments., 1997. [55] “TMS320C62x assembly benchmarks”, http://www.ti.com/sc/docs/dsps/products/c6000/62xbench.htm, 1997. [56] M. Yoshida, H. Ohtomo, I. Kuroda, “A new generation 16-bit general purpose programmable DSP and its video rate application,” IEEE Workshop on VLSI Signal Processing, pp. 93 —101, 1993. [57] I. Kuroda, “Processor architecture driven algorithm optimization for fast 2-D-DCT,” IEEE Workshop on VLSI Signal Processing, VIII, pp. 481 —490, 1995. [58] Compass, PASSPORT library, 0.6 micron 5-volt high performance standard cell library, 1996. [59] T.-S. Chang and C.-W. Jen, “Hardware-efficient implementations for transforms in programmable logic device,” submitted to IEE Proceedings - Computers and Digital Techniques, 1999. [60] C.H, Dick, “FPGA based systolic array architectures for computing the discrete Fourier transform”, IEEE International Symposium on Circuits and Systems, 1996. vol.2, pp. 465 — 468. [61] J. Heron, D. Trainor, R. Woods, M. P Fargues, R.D. Hippenstiel, “Image compression algorithms using re-configurable logic”, Thirty-First Asilomar Conference on Signals, Systems & Computers, 1997 vol. 1 , pp. 399 —403. [62] N. W. Bergmann, and Y. Y. Chung, “Video compression with custom computers”, IEEE Transactions on Consumer Electronics, vol. 43, no. 3, pp. 925-933, Aug. 1997. [63] D. R. Bull and G. Wacey, “Bit-serial digital filter architecture using RAM-based delay operators”, IEE Proc.-Circuits Devices Syst., vol. 141, no. 5, Oct. 1994. [64] G. R. Goslin and B. Newgard, ?-Tap, 8-Bit FIR Filter Application Guide," Xilinx Publications, 1994. [65] Xilinx, The Programmable Logic Data Book, 1994. [66] T.-S. Chang, J.-I. Guo and C.-W. Jen, “Hardware efficient DFT designs with cyclic convolution and subexpression sharing” submitted to IEEE Trans on Circuits and System, ⅡAnalog and Digital Signal Processing, 1999. [67] T.-S. Chang, J.-I. Guo and C.-W. Jen, “A new hardware efficient algorithm and architecture for 1-d discrete Fourier transform” submitted to IEEE Trans. on Signal Processing, 1999. [68] H. T. Kung, “ Why systolic architectures?” Comput. Mag., vol. 15, pp.37-45, Jan. 1982. [69] J. A. Beraldin, T. Aboulnasr, and W. Steenart, “Efficient one-dimensional systolic array realization of discrete Fourier transform,” IEEE Trans. Acoust. Speech, Signal Processing, vol. 36, pp. 1665-1667, Oct, 1988. [70] J. A. Beraldin, T. Aboulnasr, and W. Steenaart, “Efficient one-dimensional systolic array realization of the discrete Fourier transform,” IEEE Trans. Circuits Syst., vol. 36, no. 1, Jan. 1989. [71] W. H. Fang and M. L. Wu, “An efficient unified systolic architecture for the computation of discrete trigonometric transforms,” Proc. ISCAS, vol.3, pp. 2092 —2095, 1997 [72] L. W. Chang and M. Y. Chen, “A new systolic array for discrete Fourier transform,” IEEE Trans. Acoust. Speech, Signal Processing, vol. 36, pp. 1665-1666, Oct. 1988. [73] N. R. Murthy and M. N. s. Swamy, “On the real-time computation of DFT and DCT through systolic architectures,” IEEE Trans. Signal Processing, vol. 42, no. 4, pp. 988-991, Apr. 1994. [74] E. Chan and S. Panchanathan, “A VLSI architecture for DFT, ” Proc. the 36th Midwest Symposium on Circuits and Systems, vol. 1, pp. 292-295, 1993 [75] S. G. Sedukhin, “A new systolic architecture for pipeline prime factor DFT algorithm,” Fourth Great Lakes Symposium on VLSI. pp. 40-45, 1994. [76] S. Gudvangen, and A. Holt, “Computation of prime factor DFT and DHT/DCT algorithms using cyclic and skew-cyclic bit-serial semisystolic IC convolvers,” IEE Proceedings-G, vol. 137, no. 5, pp. 373 — 389, Oct. 1990. [77] T. S. Chang, and C-W. Jen, “Hardware efficient transform designs with cyclic formulation and subexpression sharing,” Proc. ISCAS, vol. 2. pp. 398-401, 1998. [78] C.M. Radar, “Discrete Fourier transforms when the number of data samples is prime,” Proc. IEEE, vol. 56, pp. 1107-1108, 1968. [79] C.S. Burrus and T.W. Parks, DFT/FFT and convolution algorithms, John Wiley & Sons, 1985. [80] L. R. Rabiner and B. Gold, Theory and Application of Digital Signal Processing. Chap. 10, Prentice-Hall, 1975. [81] Y.-H. Chan, and W.-C. Siu, “A cyclic correlated structure for the realization of the discrete cosine transform,” IEEE Trans. Circuits Syst. II. vol. 39, pp. 109-113, Feb, 1992 [82] V. Boriakoff, “FFT computation with systolic arrays, a new architecture,” IEEE Trans. Circuits Syst. II. vol. 41, no. 4, pp. 278-284, Apr. 1994. [83] S.-F. Hsiao and C.-Y. Yen, “New unified VLSI architectures for computing DFT and other transforms,” Proc. ISCAS, 1997. [84] E. H. Wold and A. M. Despain, “Pipeline and parallel-pipeline FFT processors for VLSI implementations,” IEEE Trans. on computers, vol. C-33, no. 5, pp. 414-426, May, 1984. [85] M. Vergara et al, “A 195KFFT/s (256-points) high performance FFT/IFFT processor for OFDM applications,” 1998. [86] L. Jia et al, “A new VLSI-oriented FFT algorithm and implementation”, Proc. IEEE ASIC Conference, pp. 337 — 341, 1998 [87] S. He and M. Torkelson, “Design and implementation of a 1024-point pipeline FFT processor,” IEEE Proc. CICC, pp. 131-134, 1998. [88] H. K. Garg, Digital signal processing algorithms: number theory, convolution, fast Fourier transforms, and applications, CRC Press, 1998. [89] C. Joanblanq, F. Rothan, and P. Senn, “A video delay line compiler,” in Proc. ISCAS, May, 1990. [90] M. Mehendale, S. D. Sherlekar, and G. Venkatesh, “Coefficient optimization for low power realization of FIR filters,” in IEEE Workshop on VLSI Signal Processing, pp. 352-361, 1995. [91] Y. C. Lim, “Predictive coding for FIR filter wordlength reduction,” IEEE Tran. Circuits Syst, vol. 32, no. 4, pp. 365-372, Apr. 1985. [92] N Sankarayya, K. Roy, and D. Bhattacharya, “Algorithms for low Power and high speed FIR filter realization using differential coefficients,” IEEE Tran. Circuits Syst. II. vol. 44. No. 6. pp.488-497, June, 1997. [93] T. S. Chang, and C. W. Jen, “Low power FIR filter realizations with differential coefficients and inputs,” Proc. ICASSP, pp. 3009-3012, 1998. May, Seattle, WA. [94] T. S. Chang, Y-H. Chu, and C.-W. Jen, “Low power FIR filter realization with differential coefficients and inputs,” revised, IEEE Trans on Circuits and System, ⅡAnalog and Digital Signal Processing, 1999. [95] N Sankarayya, K. Roy, and D. Bhattacharya, “Optimizing computations in a transposed direct form realization of floating-point LTI-FIR systems,” Proc. ICCAD. pp. 120-125, Nov, 1997. [96] T.-S. Chang and C.-W. Jen, “High speed pipelined programmable FIR filter design,” submitted to IEE Proceedings: Circuits, Device and Systems, 1999. [97] B. Edwards, A. Corry, N. Weste, and C. Greenberg, “ A single chip ghost canceller,” in Proc. IEEE 1992 Custom Integrated Circuits Conf., pp.26.5.1-4, May 1992. [98] D. J. Pearson et al. "Digital FIR filters for high speed PRML disk read channels,", IEEE J. Solid-State Circuits, vo.30, no. 12, pp. 1517-1523, Dec. 1995. [99] M. Sid-ahmed, “A systolic realization for 2-D digital filters,” IEEE Trans. on Acoust., Speeach, Signal Processing, vol. 37, pp.560-565, Apr. 1989. [100] O. L.Mac Sorley, “ High speed arithmetic in binary computers,” IRE Proc., vol. 49, pp. 67-91, Jan. 1961. [101] D. Villeger and V. G. Oklobdzija, “Evaluation of Booth encoding techniques for parallel multiplier implementation.” Electronics Letters, vol. 29, no. 23, pp.2016-2017, Nov. 1993. [102] P. Song and G. De Michelli, “Circuits and architecture trade-offs for high speed multiplication,” IEEE J. Solid-State Circuits, vol. 26, no. 9, pp.1184-1198, Sept. 1991. [103] R. Hossain, L. D. Wronski, and A. Albicki “Low power design using double edge flips-flops,” IEEE Trans. VLSI systems, vol. 2. No. 2, pp.261-265, June 1994. [104] V. G. Oklobdzija, D. Villeger and S. S. Liu, “A method for speed optimized partial product reduction and generation of fast parallel multipliers using an algorithmic approach,” IEEE Trans. Computers, vol. 45, no. 3, pp. 294-306, Mar. 1996.
|