跳到主要內容

臺灣博碩士論文加值系統

(216.73.217.174) 您好!臺灣時間:2026/08/09 11:48
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:林宗良
研究生(外文):Tzung-Liang Lin
論文名稱:可變位元率之CELP語音編碼器
論文名稱(外文):A VARIABLE BIT RATE CELP SPEECH CODER
指導教授:李清坤
指導教授(外文):Ching-Kuen Lee
學位類別:碩士
校院名稱:大同大學
系所名稱:電機工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2002
畢業學年度:90
語文別:中文
論文頁數:56
中文關鍵詞:語音編碼線性預估向量量化線頻譜頻率
外文關鍵詞:speech codinglinear predictionvector quantizationline spectral frequency
相關次數:
  • 被引用被引用:0
  • 點閱點閱:267
  • 評分評分:
  • 下載下載:16
  • 收藏至我的研究室書目清單書目收藏:0
語音是人類最方便的溝通方式,在無線及網路通訊日益頻繁的現代,語音壓縮技術不斷的在追求低位元率與高語音品質。在語音編碼器的設計上使用可變位元率 (variable bit rate) 的方式就成為兼顧維持語音品質與降低位元率的極佳選擇。本篇論文的主要目標即是以1982年制訂的美國聯邦標準 (Federal Standard) FS-1016的4.8 kbps碼本激源式線性預估 (Code-Excited Linear Prediction, CELP) 語音編碼器為藍本,設計一個可變位元率語音編碼器。
我們的設計構想主要源自於一般語音編碼器的一個時框 (frame) 約為20-30 ms,然而語音是一種隨時間緩慢變化的訊號,例如一般母音通常會持續200-300 ms,在此期間聲道幾乎沒有變化。基於此一特性,語音在相鄰時框間的相似性非常高,而其代表參數例如線性預估編碼 (Linear Predictive Coding, LPC) 參數及線頻譜頻率 (Line Spectral Frequencies, LSF) 當然也就可能高度相似,也因此我們並不一定需要每個時框都傳送一組新的參數。
為了更有效的利用頻寬,本研究介紹一種適應性前向/後向量化方式 (adaptive forward/backward quantization, AFBQ)。具體而言,我們經由一個以實驗設定的臨界值來比較目前時框 (current frame) 與過去若干時框的LPC參數之間的差異以決定是要傳送新的LPC參數或是只需告知解碼端過去LPC參數的位置。利用此種AFBQ的機制,我們即可減少傳送LPC參數的需求而不致影響合成語音的聽覺品質。為了進一步降低位元率,我們也將FS-1016 CELP語音編碼器中的LSF參數量化器由34位元的純量量化器 (scalar quantizer, SQ) 改為10位元的向量量化器 (vector quantizer, VQ) 。此外為了提昇語音品質,我們在LSF向量量化器中使用了符合聽覺加權的距離量測,在語音合成部份則加入了線頻譜頻率內插法 (interpolation) 以增進合成語音頻譜變化的平滑化。LSF內插可在不增加傳輸位元的情況下改善語音品質,但會增加15 ms的編碼延遲。
從我們的實驗結果及非正式的聽覺測試中顯示,可變位元率語音編碼器使用AFBQ及LSF向量量化,可在低位元率的情況下保持語音品質,例如將位元率降至3.9 kbps 時,合成語音的平均段落信號雜訊比 (segmental signal-to-noise ratio, segSNR) 雖然會降低約0.6 dB,但在聽覺上並不易察覺有所不同。系統加入LSF內插後segSNR則可顯著提升約0.8 dB,其聽覺品質則與FS-1016 CELP coder相當。

In the era of mobile and network communication, speech is still the most natural and convenient manner for human to exchange information. Attempts are made continuously to pursuit speech coding techniques with lower bit rates and better synthetic speech quality. The use of variable bit rate (VBR) coders is undeniably an attractive approach for maintaining speech quality at lower average bit rate. The aim of this thesis is thus to design a VBR speech coder based on the Code Excited Linear Prediction (CELP) at 4.8 kbps, which was standardized as Federal Standard FS-1016 in 1982.
The basic idea of our system design is from the observation: the frame size of most speech coders is around 20-30 ms, while the speech signal is slowly time-varying, e.g., vowel sounds may last for 200-300 ms, during which the vocal tract remains nearly unchanged. This observation suggests that the speech parameters, such as the Linear Predictive Coding (LPC) parameters and the Line Spectral Frequencies (LSFs), may share high similarity between the current frame and some temporally closed previous frames. This means that it is not necessarily to transmit a set of new parameters for each frame. Instead, speech parameters of a previous frame may be used in the decoder to save the bit rate.
Based on the concept described above, we introduced an adaptive forward/backward quantization (AFBQ) [1] scheme to reduce the required for transmitting of LPC parameters. Specifically, the spectral distances between the current frame and some previous frames are calculated and an experimentally determined threshold is used to decide either the LPC parameters of the current frame should be transmitted or it is sufficient to transmit only a location index of a previous frame. The AFBQ scheme can reduce bit rate at a minimum cost of speech quality. To further reduce the bit rate, instead of using a 34-bit scalar quantizer, the proposed VBR coder utilizes a 10-bit vector quantizer (VQ) for the quantization of the LSF parameters. On the effort of speech quality improvement, we adopted a perceptual weighted distance measure in the LSF vector quantizer and incorporated an interpolation scheme for LSF parameters to smooth the spectral changes in the synthetic speech. The LSF interpolation scheme can improve the speech quality without the need of transmitting extra bits, but at the cost of 15-ms longer coding delay.
Our experimental results and informal listening test showed that, by using the AFBQ scheme and LSF vector quantizer, the proposed VBR speech coder could maintain speech quality at a lower average bit rate. For example, the VBR coder at 3.9 kbps can retain the average segmental signal-to-noise ratio (segSNR) with only 0.6 dB lower than that of the 4.8 kbps CELP coder, and their synthetic speech quality can hardly be differentiated. The experimental results also showed that the inclusion of the LSF interpolation scheme did improve the speech quality with a higher average segSNR of 0.8 dB.

ABSTRACT IN CHINESE I
ABSTRACT IN ENGLISH III
ACKNOWLEDGEMENTS V
CONTENTS VI
LIST OF FIGURES VIII
LIST OF TABLES IX
CHAPTER 1 INTRODUCTION 1
1.1 Introduction 1
1.2 Research Motivation 2
CHAPTER 2 THE FUNDAMENTALS OF THE CELP STRUCTURE 5
2.1 Introduction 5
2.2 CELP Coder Algorithm Description 6
2.2.1 Receiver 7
2.2.2 Transmitter 8
2.2.2.1 Linear Prediction Analysis 9
2.2.2.2 The Computation of LSF 12
2.2.2.3 Adaptive Codebook Search 14
2.2.2.4 Fixed Codebook Search 18
2.3 CELP Bit Allocation Format 19
CHAPTER 3 THE PROPOSED SYSTEM STRUCTURE 21
3.1 The Block Diagram of the Proposed System 21
3.2 Adaptive Forward/Backward Quantizer 24
3.2.1 Algorithm of Adaptive Forward/Backward Quantizer 26
3.2.2 Modification of Window Shape 28
3.3 Vector Quantization (VQ) of the LSFs 29
3.4 Interpolation of LSF 30
3.5 Combination of the Proposed System and LSF Interpolation33 CHAPTER 4 SIMULATION RESULTS AND DISCUSSIONS 36
4.1 Performance Evaluations 37
4.2 Discussions of the Proposed Systems 41
CHAPTER 5 CONCLUSIONSREFERENCES 44 REFERENCES 45

[1] J. Vass, Y. Zhao and X. Zhuang, “Adaptive forward-backward quantizer for low bit rate high quality speech coding,” IEEE Trans. Speech Audio Processing, vol. 1, pp. 552 —557, Nov. 1997.
[2] D. Kemp, R. Sueda and T. Tremain, “An Evaliation of 4800 bps Voice Coders,” in Proc. IEEE ICASSP, pp. 200-203, 1989.
[3] R. Fenichel, Federal Standard 1016, Telecommunications: Analog to Digital Conversion of Radio Voice by 4,800 bit/second Code Excited Linear Prediction (CELP), National Communication System, Office of Technology and Standards, Washington, DC, Feb. 1991.
[4] J. Campbell, V. Welch and T. Tremain, “An Expandable Error-Protected 4800 bps CELP Coder,” in Proc. IEEE ICASSP, pp. 735-738, 1989.
[5] D. Rahikka, T. Tremain, V. Welch and J. Campbell, “CELP Coding for Land Mobile Radio Applications,” in Proc. IEEE ICASSP, pp. 465-468, 1990.
[6] D. Lin, “New Approaches to Stochastic Coding of Speech Sources at Very Low Bit Rates,” Signal Processing III: Theories and Applications, pp. 445-448, 1986.
[7] B. Atal and M. Schroeder, “Adaptive predictive coding of speech signals.” The Bell System Technical Journal, pp. 1973-1987, 1970.
[8] J. Makhoul, “Linear prediction: A tutorial review,” Proc. IEEE, pp. 561-580, 1995.
[9] J. H. Chen, R. V. Cox, Y. C. Lin, N. Jayant and M. J. Melchner, “A low-delay CELP coder for the CCITT 16kb/s speech coding standard,” IEEE J. Select. Areas Commun., vol. 10, pp. 830-849, June 1992.
[10] N. Sugamura and N. Farvardin, “Quantizer Design in LSF Speech Analysis-Synthesis,” IEEE J. Select. Areas Commun., vol. 6. pp. 432 440, Feb 1988.
[11] K. K. Paliwal and B. S. Atal, “Efficient Vector Quantization of LPC Parameters at 24 Bits/Frame,” in Proc. IEEE ICASSP, pp. 661-664, 1991.
[12] J. A. Asenstorfer, “Source-Channel Coding for CELP Speech Coders,” Ph.D. dissertation, Adelaide Univ, Sept. 1994.

QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top