跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.226) 您好!臺灣時間:2026/08/08 10:29
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:廉凱成
研究生(外文):Kai-Cheng Lien
論文名稱:可變位元率MELP語音編碼器
論文名稱(外文):A VARIABLE BIT RATE MELP SPEECH CODER
指導教授:李清坤
指導教授(外文):Ching-Kuen Lee
學位類別:碩士
校院名稱:大同大學
系所名稱:通訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2002
畢業學年度:90
語文別:英文
論文頁數:66
中文關鍵詞:可變位元率語音編碼器
外文關鍵詞:MELP Speech Coder
相關次數:
  • 被引用被引用:0
  • 點閱點閱:242
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:1
這是一個第三代 (3G) 無線通訊標準的時代,儘管視訊、數據通信等多媒體應用日益普及,但語音通訊仍然是最重要的無線行動服務之一。利用編碼技術把語音訊號壓縮成少量參數以利於傳送的低位元率語音編碼器就變得日漸重要。而在語音編碼器的設計上使用可變位元率 (variable bit rate) 的方式也就成為兼顧維持語音品質與降低平均位元率的極佳選擇。本篇論文的目標就是以德州儀器所研發出來並且被美國國防部採用為 2.4 kbps新標準 Federal Standard 1017 (FS-1017) 的混合激源線性預估 (Mixed Excitation Linear Prediction, MELP) 技術為基礎的可變位元率的 MELP 語音編碼器。
在 2.4 kbps的FS-1017 MELP 語音編碼器中,其取樣頻率為 8 kHz ,解析度為 16 bit ,每個分析時框 (frame) 是 22.5 微秒 (ms) 且在每個分析時框中的輸出位元流 (bit stream) 為 54 個位元,其中的 25 個位元是用來代表線性預估參數 (Linear Predictive Coding, LPC) ,此參數係以一個多階向量量化器 (Multi-stage Vector Quantizer, MSVQ) 量化,因此,有超過 46% 的頻寬都在傳送 LPC 參數。所以,有效地量化 LPC 參數是減少整體位元率的可行途徑。
在 FS-1017 MELP 編碼器中,代表 LPC 參數被量化的是線頻譜係數 (Line Spectral Frequency LSF), 量化 LSF 的 MSVQ 為固定四階的向量量化器。我們對此四階向量量化器的實驗數據顯示平均有超過 20% 的情況不需要使用到完整四階的向量量化器就可以達成透明量化器 (transparent quantizer) 所要求的量化前後平均頻譜失真 (spectral distortion) 小於 1 分貝的要求。上述觀察形成了設計可變階數向量量化器 (Variable-Stage Vector Quantizer, VSVQ ) 的構想。具體而言,我們在 VSVQ 的各階向量量化器中插入了一個經實驗設定的臨界值,用以決定是否需要進行下一階的向量量化。
實驗結果顯示我們的可變位元率 MELP 編碼器所合成出來的語音品質非常接近FS-1017 標準的 2.4 kbps MELP 編碼器。在平均位元率降至 2.1 kbps 的情況下,兩者的語音品質沒有可察覺的差別,其平均的信號差異比 (Signal-to-Difference Ratio, SDR) 亦可達 85 分貝。除此之外,我們在本篇所提出來的可變階數向量量化器已經證實成功地應用在 MELP 語音編碼器,相信還可以應用在其他種類的語音編碼器來擴展其用途。

In the era of third-generation (3G) wireless personal communications, though applications of multimedia such as video and data communication have become more and more popular, speech communication is still one of the most important mobile radio services. Consequently, speech coding techniques that can compress speech information into as few parameters as possible is increasingly important. To achieve the goal, the use of variable bit rate speech coders is certainly an attractive approach to retain overall high voice quality at low average bit rate. The aim of this thesis is thus to develop a variable-rate speech coder based on the Federal Standard 1017 (FS-1017), a 2.4 kbps mixed excitation linear prediction (MELP) coder originally developed by the Texas Instruments and then standardized by the U.S. Department of Defense.
In the FS-1017 standard 2.4 kbps MELP coder, sampling rate is 8 kHz with 16 bit resolution and the frame size is 22.5 ms. Each frame has a bit stream of 54 bits, wherein 25 bits are LPC coefficients, which account for 46% of the required bandwidth. Therefore effectively quantization of the LPC coefficients is essential to reduce overall bit rate for the MELP coder.
In the FS-1017 MELP coder, LPC parameters are transformed to line spectral frequencies (LSFs) and then quantized by a fixed four-stage vector quantizer. However, our experimental results showed that, with only one- to three-stage VQ, more than 20% quantized LSFs could satisfy the requirement of “transparent quantization,” i.e., having an average spectral distortion (SD) less than 1 dB. Accordingly, we proposed to utilize a variable-stage vector quantizer (VSVQ) to design a variable-rate MELP speech coder. Specifically, we insert an experimentally determined threshold after each stage of the VSVQ to determine whether the SD requirement is satisfied. When the answer is yes, the procedure of the VSVQ is stop to save the bits for the following VQ stages.
Our experimental results showed that the speech quality of the proposed variable-rate MELP coder is very close to that of FS-1017 standard 2.4 kbps MELP coder. When the average bit rate is around 2.1 kbps, there is no audible difference between the FS-1017 and the proposed variable-rate MELP coders. The experimental results also showed that the Signal-to-Difference Ratio (SDR) between the synthetic speech of the FS-1017 and that of the proposed variable-rate MELP coder is as high as 85 dB. The structure of the variable-stage vector quantizer we proposed in this research has been proved to be a success in MELP coder. We believe that it also has high potential to be used in other types of speech coders to extend their usage.

ABSTRACT IN CHINESE
ABSTRACT IN ENGLISH
ACKNOWLEDGEMENTS
CONTENTS
LIST OF FIGURES
LIST OF TABLES
CHAPTER 1 INTRODUCTION
1.1 Motivation for Speech Coding
1.2 Speech Coding Standardizations
1.3 Motivation for Our Research
CHAPTER 2 FUNDAMENTALS OF THE MELP SPEECH CODER
2.1 Introduction
2.2 Linear Predictive Coding
2.3 LPC to LSF Transformation
2.4 Bandwidth Expansion
2.5 The Model of MELP Coder
2.5.1 Mixed Excitation and Bandpass Filters
2.5.2 Pitch and Gain Calculation
2.5.2.1 Integer Pitch Calculation
2.5.2.2 Fractional Pitch Refinement
2.5.2.3 Final Pitch Calculation
2.5.2.4 Gain Calculation
2.5.3 Aperiodic Pulses
2.5.4 Adaptive Spectral Enhancement
2.5.5 Pulse Dispersion Filter
2.5.6 Fourier Magnitude Modeling
2.6 Quantization of Parameters in MELP Coder
2.7 Efficient vector quantization of LPC Parameters
2.8 The MSVQ Algorithm in MELP Coder
2.8.1 The Concept of Multi-Stage Vector Quantizer
2.8.2 M-L Search Procedure
2.9 Conclusions
CHAPTER 3 A VARIABLE BIT RATE MELP SPEECH CODER
3.1 Introduction
3.2 The Effect of Different Number of Stage on Spectral Distortion
3.3 The Main Scheme in a Variable Bit Rate MELP Coder
3.3.1 The Structure of a Variable Rate MELP Coder
3.3.2 Effect and Selection of Threshold Value
3.4 Bit Allocation in a Variable Bit Rate MELP Coder
3.5 Conclusions
CHAPTER 4 THE SIMULATION AND PERFORMANCE
4.1 Implementation of Recorded Speech Data
4.2 Comparison of Original and a Variable Bit Rate MELP Coder
4.3 Comparison of Variable Bit Rate Coder and other Coders
4.4 Performance of a Variable Bit Rate Coder
4.5 Conclusions
CHAPTER 5 CONCLUSIONS
APPENDIX A LOOK-UP TABLE FOR PITCH QUANTIZATION OF FS-1017 MELP CODER
REFERENCES

[1] R. V. Cox, “Three new speech coders from the ITU cover a range of applications,”
IEEE Communications Magazine, pp. 40-47 Sept. 1997.
[2] A. V. McCree and T. P. Barnwell III, “A mixed excitation LPC vocoder model for
low bit rate speech coding,” IEEE Trans. Speech and Audio Processing, vol. 3, pp.
242-250, July 1995.
[3] B. S. Atal and S. L. Hanauer, “Speech analysis and synthesis by linear prediction of
the speech wave,” J. Acoust. Soc. Amer., vol. 50, pp. 637-655, Aug. 1971.
[4] T. E. Tremain, “The government standard linear predictor coding algorithm:
LPC-10,” Speech Technol., pp. 40-49, Apr. 1982.
[5] B. Atal, “Efficient coding of LPC parameters by temporal decomposition,” in Proc.,
IEEE ICASSP, 1983, pp. 81-85.
[6] P. Kabal and R. Ramachandran, “The computation of line spectral frequencies using
Chebyshev polynomials,” IEEE Trans. Acoustics Speech and Signal Processing, vol.
ASSP-34, pp. 1419-1426, Dec. 1986.
[7] A. V. McCree and T. P. Barnwell III, “Improving the performance of a mixed
excitation LPC vocoder in acoustic noise,” in Proc., IEEE ICASSP, Sept. 1992, pp.
II137-II140.
[8] Federal Information Processing Standards Publication (Draft), Specification for the
analog to digital conversion of voice by 2,400 bit/second mixed excitation linear
prediction (MELP), Jan. 1998.
[9] K. K, Paliwal and B. Atal, “Efficient vector quantization of LPC parameters at 24
bits/frame,” in Proc., IEEE ICASSP, Mar. 1991, pp. 661-664.
[10] F. K. Soong and B. H. Juang, “Line Spectral Pair (LSP) and speech data
compression,” in Proc., IEEE ICASSP, Oct. 1984, pp. 1-4.
[11] P. Ronald and S. John, “Incorporating perception into LSF quantization — some
experiments,” in Proc. IEEE ICASSP, Apr. 1997 pp. 1347-1350.
[12] L. Hanzo, F. Clare, A. Somerville and J. P. Woodard, Voice Compression and
Communications. New York: IEEE Press, 2001.
[13] P. LeBlanc, B. Bhattacharya, S. A. Mahmoud, and V. Cuperman, “Efficient search
and design procedures for robust multi-stage VQ of LPC parameters for 4 kb/s
speech coding,” IEEE Trans. Speech and Audio Processing, vol. 1, pp. 373-385, Oct.
1993.
[14] A. Buzo, A. H. Gray, R. M. Gray, and J. D. Markel, “Speech coding based upon
vector quantization,” IEEE Trans. Acoustics Speech and Signal Processing ASSP-28,
pp. 562-574, Oct. 1980.
[15] M. Tie, D. Wang, and C. Fan, “A novel variable-rate MELP speech coder,” in Proc.,
IEEE ICSP 2000, Nov. 2000, pp. 693-696.

QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top