跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.249) 您好!臺灣時間:2026/10/08 07:39
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

我願授權國圖
: 
twitterline
研究生:董呈煌
研究生(外文):Cheng-Huang Tung
論文名稱:手寫中文文句辨識之研究
論文名稱(外文):A Study of Handwritten Chinese Text Recognition
指導教授:李錫堅李錫堅引用關係
指導教授(外文):Hsi-Jian Lee
學位類別:博士
校院名稱:國立交通大學
系所名稱:資訊工程研究所
學門:工程學門
學類:電資工程學類
論文種類:學術論文
論文出版年:1994
畢業學年度:82
語文別:英文
論文頁數:124
中文關鍵詞:手寫中文文句辨識、手寫中文產生器、大分類、辨識模組、語言模式、前後文後處理、未知詞
外文關鍵詞:Handwritten Chinese recognition、generator、candidate selection
相關次數:
  • 被引用被引用:1
  • 點閱點閱:260
  • 評分評分:
  • 下載下載:0
  • 收藏至我的研究室書目清單書目收藏:1
本論文提出一個手寫中文文句辨識系統,包含手寫中文文字辨識,前後文
處理及辭典的維護。首先,我們建立一個手寫中文產生器,用於產生手寫
中文字形。利用產生的手寫字形,可以計算切割文字影像成區塊的各種方
法之效能,亦可推導出對影像區塊的特徵抽取。在效能評估之後,即可建
立一個含有大分類及辨識模組的文字辨識系統。因為文字辨識仍然消耗多
數的執行時間,本論文提出了多階層的前大分類模組,以減少執行時間。
每一個階層利用從輸入字形抽取出的單一特徵,以除去資料庫中不被認同
的字。而在前大分類中所使用的特徵順序,是依據由訓練字庫計算出的消
減率大小排定。我們又提出以分類樹做前大分類的方法。實驗結果顯示所
提出的模式可以有效地降低執行時間,而又不降低文字辨識的精確度。為
了要增加文字辨識的準確度,本論文更提出一個新的方法,來偵測與改正
被辨識模組認錯的輸入字形。此方法包含兩個辨識模組,同時識別輸入字
形。如果這兩個辨識模組的辨識結果不同,輸入字形即被駁回。經由使被
接受訓練字形的準確度達到最大之過程,可建立第二個辨識模組。因為辨
識階段能正確的辨識大多數的輸入字形,同時對每一被駁回的輸入字形僅
輸出少量的候選字。對每一個被駁回的輸入字形,可根據前後文的訊息,
利用字的二元馬可夫語言模式,準確的選擇一個最好的候選字。因為語言
模式可以任意地與其他的辨識系統結合使用,本論文提出以詞集為本的新
語言模式,進一步提升語言模式的效能。辭典中的詞預期以語意來分詞集
,但是原本的辭典內並不含有語意訊息,必須藉著使用另一辭典—同義詞
詞林,訓練詞群的語意特徵。再依據語意的相近度,可以將所有詞群分為
m個詞集。查詢該m個詞集,即可將原辭典中的詞分成m個詞集。因此描述
二元前後文訊息的參數空間只需要 m*m。由實驗得知,這個新語言模式的
效能比字的二元語言模式的效能高出3.2%。在前後文後處理中,找出包
含於候選字集中的詞需用最多時間,為了使這項工作更有效率,辭典中的
詞序是依據詞首的兩個字來排序。因為包含所有詞首兩字的索引陣列是呈
稀疏狀,本論文使用了列位移方法壓縮該稀疏索引陣列。實驗值顯示,壓
縮率可達到 224,而且使用具有這種結構的辭典,可以很快地找出含在候
選字集中的詞,使前後文後處理的速度大幅加快。
In this thesis,we propose a Chinese text processing system for
handwritten Chinese character recognition,contextual
postprocessing and maintenance of the dictionary.A handwritten
Chinese character generator is created for generating
handwritten Chinese character images.By utilizing the generated
character images,we measure the performance of segmenting an
image into a number of meshes by different methods,and derive
the feature extraction for an image mesh. After the performance
measurement,a character recognition system consisting of a
candidate selection module and a matching module is
established. Because character recognition still takes much
execution time,we propose a multi-stage candidate pre-selection
module to reduce the execution time. In each stage,we use a
single feature computed from the input character image to
eliminate impossible character categories. The features used in
candidate pre- selection are ordered according to the reduction
rates evaluated from a set of training characters. We also
propose a method for organizing the character database as a
classification tree. The experimental results show that the
proposed model can reduce the total execution time
significantly without decreasing the precision of character
recognition. We present a new approach for detecting and
correcting characters erroneously identified by the matching
module. Two matching modules are applied at the recognition
stage to recognize an input character image simultaneously. If
the matching results of the two modules for a character image
are not the same,the character image is rejected at the
recognition stage. Here,we construct the second recognition
module by maximizing the accuracy of the accepted training
characters. Because the recognition stage recognizes most of
the input characters correctly and outputs a small number of
candidates for each rejected character,a character bigram
Markov language model can be applied to choose a candidate with
high recognition rate.
COVER
ABSTRACT(IN CHINESE)
ABSTRACT(IN ENGLISH)c99
ACKNOWLEDGEMENTS
TABLE OF CONTENTS
LIST OF FIGURES
LIST OF TABLES
CHAPTER 1 INTRODUCTION
1.1 Motivations
1.2 System Architecture
1.3 Thesis Organization
CHAPTER 2 SURVEY OF RELATED RESEARCHES
2.1 Character Recognition and Generation
2.2 Contextusl Postprocessing
CHAPTER 3 A CHARACTER GENERATOR AND ITS APPLICATIONS
3.1 Motivation
3.2 Line-vector Character Generation
3.3 Character Image Generation
3.4 Applications of a Character Generator
3.4.1 Stability Analysis for Segmentation of Images
3.4.2 Stability Analysis for Segmentation of Nonlinearly Normal ized Images
3.4.3 Derivation of a Matrix for Feature Extraction
3.4.4 Construction of a Recognition System
3.4.5 Recognition Rates for Recognizing CCL/HCCR1 Character Database
3.5 Discussions
CHAPTER 4 SPEED IMPROVEMENT OF CHARACTER RECOGNITION
4.1 Motivation
4.2 Multi-stage Candidate Pre-selection
4.3 Design of a Classification Tree for Candidate Pre-selection
4.4 Experimental Results
4.5 Dscussions
CHAPTER 5 DETECTION AND CORRECTION OF ERRONEOUSLY IDENTIFIED CHARATTERS
5.1 Motivation
5.2 Detection of Recognition Errors
5.3 The Language Model for Correcting Recognition Errors
5.4 Experimental Results
5.5 Discussions
CHAPTER 6 A LANGUAGE MODEL BASED ON CLUSTERED WORDS
6.1 Motivation
6.2 Construction of Semantic Attributes of Work Classes
6.3 Updating the Counts of Semantic Attributeds of Word Classes
6.4 Grouping Word Classes into m Groups
6.5 Markov Language Model Based on Clustered Words
6.6 Experimental Results
6.7 Discussions
CHAPTER 7 A NEW DICTIONARY STRUCTURE AND ITS APPLICATION
7.1 Motivation
7.2 Dictionary Construction and Compression
7.3 A Fast Algorithm for Finding Words in the Candidate Sets
7.4 Experimental Results
7.5 Discussions
CHAPTER 8 CORPUS-BASED DICTIONARY MAINTENANCE
8.1 Motivation
8.2 Basic Measurement for Identifying Unknown Words
8.3 Identification of Unknown Words
8.4 Interactive pattern Editing
8.5 Algorithm for Identifying Unknow Words
8.6 Experimental Results
8.7 Discussions
CHAPTER 9 CONCLUSIONS
9.1 Conclusions
9.2 Future Researches
BIBLIOGRAPHY
VITA
PUBLICATION LIST OF CHENG-HUANG TUNG
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top