跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.28) 您好!臺灣時間:2026/07/24 19:52
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:謝珮昕
研究生(外文):Pei-shin Hsieh
論文名稱:視訊中之字幕移除及其遺失資料回復
論文名稱(外文):Removal and Missing Data Recovery of Subtitle in Video
指導教授:曾建誠曾建誠引用關係
指導教授(外文):Chien-cheng Tseng
學位類別:碩士
校院名稱:國立高雄第一科技大學
系所名稱:電腦與通訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:2012
畢業學年度:100
語文別:中文
論文頁數:92
中文關鍵詞:遺失資料回復視訊字幕
外文關鍵詞:Video subtitleMissing data recovery
相關次數:
  • 被引用被引用:0
  • 點閱點閱:212
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:0
在本論文中,將探討如何將視訊中之解釋字幕移除,並修補這些字幕區域中的遺失影像資料。首先,利用字幕的一些特徵,如字幕的位置、字幕是白色的、字幕之文字為緊湊排列和字幕的幾何特性等,來完成字幕的偵測、以及字幕的定位和切割;其次,本論文探討如何回復字幕區域中遺失的影像資料;共有三種方法被提出,一是最近鄰居預測法,即利用最近的鄰居像素值,來預測和估計遺失的資料;二是基於範本影像修補,即利用影像上原有之未遺失資訊,以塊狀填補之方式回復遺失資訊;最後一個是時頻交替遞迴法,即假設影像信號是有限頻寬,然後利用時域和頻域之交互遞迴修正,來完成資料修補的工作。最後,本論文利用預先建立的電影片段之資料庫來測試所提方法之有效性,實驗結果顯示,修補後之電影片段,在人眼感官上是可接受的。
In this paper, the removal of subtitle in video is studied and the recovery of missing image data in subtitle region is also investigated. First, the features of video subtitle are used to develop the detecting, locating and segmenting algorithms of the subtitle region. The features are that the location of subtitle is in the middle-bottom region of image, the color of characters in subtitle is white, the subtitle letters and words are placed in a spatially compact arrangement, and the subtitle region has a rectangular geometric shape. Next, three methods are presented to recover the missing image data in the subtitle region of video. First is the nearest neighbor prediction method in which the pixel value of missing data is estimated by using the weighted average value of the nearest neighbor pixels, second is the exemplar-based image inpainting method in which the block region in the missing data of image is inpainted by using the confidence data of image, third is the time-frequency domains iterative modification method in which the image signal is assumed to be band-limited and an iterative method is used to estimate the missing data. Finally, a video database is established to evaluate the performance of the proposed method. The experimental results show that the quality of recovered video is acceptable for human visual perceptual system.
中文摘要 i
英文摘要 ii
誌謝 iv
目錄 v
表目錄 vii
圖目錄 viii
壹、 緒論 1
1.1. 前言 1
1.2. 研究動機 1
1.3. 研究方法 1
1.4. 論文架構 5
貳、 字幕偵測 6
2.1. 電影畫面分析 6
2.1.1. 無字幕之電影畫面 6
2.1.2. 有字幕之電影畫面 7
2.2. 字幕分析 7
2.3. 字幕偵測 9
2.3.1. 前處理 9
2.3.2. 偵測方法 17
2.3.3. 後處理 25
2.4. 結論 40
參、 字幕之定位和切割 41
3.1. 前言 41
3.2. 幾何 44
3.3. 緊湊排列 51
3.4. 結論 55
肆、 最近鄰居預測法 57
4.1. 由左至右,由上至下填補 57
4.2. 由右至左,由下至上填補 59
4.3. 分四區各自掃描填補 60
4.4. 由外而內填補 62
4.5. 結論 64
伍、 基於範本影像修補 66
5.1. 前言 66
5.2. 演算法 69
5.2.1. 計算優先權 70
5.2.2. 影像修補 71
5.3. 結論 74
陸、 時頻交替遞迴回復法 75
6.1. 時頻交替遞迴 75
6.2. 實驗結果 79
6.3. 結論 83
柒、 結論與未來展望 85
7.1. 結論 85
7.2. 未來展望 85
參考文獻 87
[1]R. Lienhart, and F. Stuber,“Automatic text recognition in digital videos,” Image and Video Processing IV, Proc. of SPIE, pp. 11–20, 1996.
[2]Y. Zhong, K. Karu, and A.K. Jain, “Locating text in complex color images,” Proc. of the Third Int. Conf. on Document Analysis and Recognition, IEEE Computer Society, vol. 1, pp. 1523-1535, 1995.
[3]C. Garcia, and X. Apostolidis, “Text detection and segmentation in complex color images,” Proc. of the Acoustics, Speech, and Signal Processing, IEEE Int. Conf. on, vol. 04, pp. 2326-2329, 2000.
[4]Y. Zhong, H. Zhang, and A.K. Jain, “Automatic Caption Localization in Compressed Video,” IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 22, pp. 385-392, 2000.
[5]Xueming Qian, and Guizhong Liu, “Text Detection, Localization and Segmentation in Compressed Videos,” Acoustics, Speech and Signal Processing, Proc. of IEEE Int. Conf. on, ICASSP ''06, vol. 2, pp. II-385-388, 2006.
[6]V. Wu, R. Manmatha, and E.M. Riseman, “TextFinder: An Automatic System to Detect and Recognize Text In Images,” IEEE Trans. on Pattern Anal. and Machine Intelligence, vol. 21, pp. 1224-1229, 1999.
[7]Congjie Mi, Yuan Xu, Hong Lu, and Xiangyang Xue, “A Novel Video Text Extraction Approach Based on Multiple Frames,” Fifth Int. Conf. on, Information, Communications and Signal Processing, pp. 678-682, 2005.
[8]C. Liu, C. Wang, and R. Dai, “Text Detection in Images Based on Unsupervised Classification of Edge-based Features,” Proc. of the 8th Int. Conf. on Document Analysis and Recognition, pp. 610-614, 2005.
[9]Xiaolan Wang,“The Research of Subtitles Regional Location Algorithm Based on Video Caption Frames,” Second International Symposium on Intelligent Information Technology Application, pp. 886-889, 2008.
[10]Liang Sang, Jingqi Yan, “Rolling and Non-Rolling Subtitle Detection with Temporal and Spatial Analysis for News Video,” Proceedings of 2011 International Conference on Modelling, Identification and Control, Shanghai, China, June 26-29, 2011.
[11]You-Li Chen, and Bor-Sen Chen, “Model-Based Multirate Representation of Speech Signals and Its Application to Recovery of Missing Speech Packets,” IEEE Transactions on Speech and Audio Processing. Vol. 5, No. 3, pp. 220-231, May 1997.
[12]A. Adler, V. Emiya, M. G. Jafari, M. Elad, R. Gribonval, and M. D. Plumbley, “Audio Inpainting,” IEEE Transactions on Audio, Speech, and Language Processing, Early Access, 2012.
[13]S.-C. Pei, Y.-C. Zeng, and C.-H. Chang, “Virtual restoration of ancient Chinese paintings using color contrast enhancement and lacuna texture synthesis,” IEEE Transactions on Image Processing, Vol. 13, No. 3, pp. 416-429, March 2004.
[14]S. C. Pei and Y.C. Zeng, “A novel image recovery algorithm for visible watermarked images,” IEEE Transactions on Information Forensics and Security, Vol. 1, No. 4, pp. 543-550, Dec. 2006.
[15]Xin Li, “Image Recovery Via Hybrid Sparse Representations: A Deterministic Annealing Approach,” IEEE Journal of Selected Topics in Signal Processing, Vol. 5, No. 5, pp. 953-962, 2011.
[16]A. Criminisi, P. Perez, and K. Toyama, “Region filling and object removal by exemplar-based image inpainting,” IEEE Transactions on Image Processing, Vol. 13, No. 9, pp. 1200-1212, 2004.
[17]O.G. Guleryuz, “Nonlinear approximation based image recovery using adaptive sparse reconstructions and iterated denoising-part I: theory,” IEEE Transactions on Image Processing, Vol. 15, No. 3, pp. 539-554, 2006.
[18]O.G. Guleryuz, “Nonlinear approximation based image recovery using adaptive sparse reconstructions and iterated denoising-part II: adaptive algorithms,” IEEE Transactions on Image Processing, Vol. 15, No. 3, pp. 555-571, 2006.
[19]Haricharan Lakshman, Martin Koppel, Patrick Ndjiki-Nya, and Thomas Wiegand, “Image Recovery Using Sparse Reconstruction Based Texture Refinement,” ICASSP 2010, pp. 786-789, 2010.
[20]V. M. Patel, R. Maleh, A. C. Gilbert, and R. Chellappa, “Gradient-based Image Recovery Methods from Incomplete Fourier Measurements,” IEEE Transactions on Image Processing, Vol. 21, No. 1, pp. 94-105, 2012.
[21]Julien Mairal, Michael Elad, and Guillermo Sapiro, “Sparse Representation for Color Image Restoration,” IEEE Transactions on Image Processing, Vol. 17, No. 1, pp. 53-69, January 2008.
[22]A Prieto-Guerrero, C Mailhes, and F Castanie, “Recovering Electrocardiogram Missing Samples in Wireless Transmission,” Computers in Cardiology 2009, pp. 845-848, 2009.
[23]H. Garudadri, Y. Chi, S. Baker, S. Majumdar, P. K. Baheti, and D. Ballard, “Diagnostic grade wireless ECG monitoring,” 2011 Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 850-855, 2011.
[24]Paulo Jorge S. G. Ferreira, “Iterative interpolation of ECG signals,” 14th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, pp. 2740-2741. 1992.
[25]R. Rodrigues, “Filling in the gap: A general method using neural networks,” 2010 Computing in Cardiology, pp. 453-456, 2010.
[26]P. Langley, S. King, K. Wang, D. Zheng, R. Giovannini, M. Bojarnejad, and A. Murray, “Estimation of missing data in multi-channel physiological time-series by average substitution with timing from a reference channel,” 2010 Computing in Cardiology, pp. 309-312, 2010.
[27]N. Theera-Umpon, P. Phiphatkhunarnon, and S. Auephanwiriyakul, “Data reconstruction for missing electrocardiogram using linear predictive coding,” IEEE International Conference on Mechatronics and Automation (ICMA 2008), pp. 638-643, 2008.
[28]P. Stoica, Jian Li, and Jun Ling, “Missing Data Recovery Via a Nonparametric Iterative Adaptive Approach,” IEEE Signal Processing Letters, Vol. 16, No. 4, pp. 241-244, 2009.
[29]P. S. Naidu and Bina Paramasivaiah, “Estimation of Sinusoids from Incomplete Time Series,” IEEE Transactions on Acoustics, Speech, and Signal Processing, Vol. ASSP-32, No. 3, pp. 559-562, 1984.
[30]Andor Bariska, “Recovering Periodically Spaced Missing Samples,” IEEE Signal Processing Magazine, pp. 127-129, Nov. 2007.
[31]A. Papoulis, “A new algorithm in spectral analysis and band-limited extrapolation,” IEEE Transactions on Circuits and Systems, Vol. 22, No. 9, pp. 735-742, 1975.
[32]Edited by Farokh Marvasti, Nonuniform Sampling Theory and Proactice, Kluwer Academic/Plenum Publishers, New York, 2001.
[33]M. Jones, “The discrete Gerchberg algorithm,” IEEE Transactions on Acoustics, Speech and Signal Processing, Vol. 34, No. 3, pp. 624-626, 1986.
[34] A. Criminisi, P. Perez, and K. Toyama, “Region Filling and Object Removal by Exemplar-Based Image Inpainting,” IEEE Transactions on Image Processing, Vol. 13, No. 9, pp. 1200-1212, 2004.
[35] Timothy K. Shin, Nick C. Tang and Wonjun Lee, “Video Inpainting and Implant via Diversified Temporal Continuations,” ACM SIGMULTIMEDIA, pp. 133-136, 2006.
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top