跳到主要內容

臺灣博碩士論文加值系統

(216.73.216.141) 您好!臺灣時間:2026/08/23 11:45
字體大小: 字級放大   字級縮小   預設字形  
回查詢結果 :::

詳目顯示

: 
twitterline
研究生:黃啟銘
研究生(外文):Chi-Ming Huang
論文名稱:醫學異質資料庫整合系統
論文名稱(外文):Heterogeneous Databases Integration System In Medicine
指導教授:翁昭旼翁昭旼引用關係蔣以仁蔣以仁引用關係
指導教授(外文):JM WongJI Chiang
學位類別:碩士
校院名稱:國立臺灣大學
系所名稱:醫學工程學研究所
學門:工程學門
學類:綜合工程學類
論文種類:學術論文
論文出版年:2003
畢業學年度:91
語文別:中文
論文頁數:66
中文關鍵詞:綱目異質資料庫整合Join
外文關鍵詞:schemaheterogeneous databases integrationjoin
相關次數:
  • 被引用被引用:1
  • 點閱點閱:346
  • 評分評分:
  • 下載下載:62
  • 收藏至我的研究室書目清單書目收藏:2
由於組織內通常有許多的異質資料源-即資料散佈在各部門,甚至各部門使用相異的資料庫管理系統和作業系統,這些差異會阻礙內部資訊的取得,從而讓決策者不能及時製定決策或者降低決策品質,而使組織喪失競爭力。為了解決此一問題,必須發展一套資料存取中介軟體,讓使用者可輕易地整合異質資料源,此即為本研究之目的。
在本研究中開發系統主要分為三個階段:第一階段是建立一個半自動化的異質資料庫整合系統,讓使用者從中選擇欲整合的表格,並且建立表格關聯,而系統即依據此關聯圖將表格和資料轉移到共同資料庫。第二階段則是利用表格運算子將表格轉成符合需求的表格,以解決當選取表格未符合需求導致其無法整合在一起的問題。第三階段是利用StringMap來作字串相似性比對,以判別所欲結合之欄位是否具有高度關聯性或者是相同欄位,以提昇表格Join之品質。
為了讓使用者在使用上更方便,本系統係採用圖形化元件串接,讓使用者可以藉由簡單的滑鼠拖拉(drag-and-draw)動作輕易地完成異質資料庫整合之目的,不須撰寫SQL程式。而為了解決跨平台問題,本系統所有程式均係以JAVA開發,並且提供可擴充的元件化功能,讓使用者可隨需要增加系統的整合功能。在系統建置過程中,所採用之資料為台大醫院申報給健保局的住院醫療費用紀錄,包括:住院醫療費用清單明細檔及住院醫療費用(醫令清單)明細檔等二個檔案,系統驗證結果證實本系統確實可行。
A variety of application systems implemented by different kinds of database systems is presented in enterprises and hospitals. It is a hazard for enterprises and organizations for further usages, such as decision-making and data analysis. Hence, to solve the problem, it is necessary to provide a database middleware system, which can facilitate the users to integrate heterogeneous data sources conveniently.
This thesis proposes a heterogeneous database integration system performing extraction, transformation, loading. There are three phases for this: firstly, build up a semi-automatic heterogeneous database integration system to let users choose the tables which users want to join and align the relationships of tables. Then the system will transfer the related tables and data to the database containers according to the relationships. Finally, join data from heterogeneous database tables.
目 次
<摘要>................................................I
<Abstract>............................................II
<誌謝>.........................................III
<目次>...............................................IV
<圖次>................................................V
<表次>................................................VI
第 一 章 緒論...........................................1
1-1 前言...............................................1
1-2 開發系統動機與目的.................................2
1-3 開發系統階段內容...................................5
第 二 章 背景介紹與文獻探討...........................6
2-1分散式資料庫系統分類 ..............................6
2-2異質資料庫整合系統的綱目整合.......................7
2-3 資料庫逆向工程...................................10
2-3.1 判別一對一的二元關聯(Relationship)或為父/子類別關聯.12
2-3.2 判別一對多的二元關聯或為簡單弱實體類別............12
2-3.3判別多對多的二元關聯或為互動弱實體類別 ............13
2-3.4判別N元關聯...................................13
2-4 整合遠端資料庫系統...............................14
2-4.1 載入網要......................................14
2-4.2 處理實體內差異的運算............................14
2-4.3 處理實體間的運算子.............................15
2-4.4 處理查詢......................................16
2-5 DW資料不一致之研究...............................17
2-5.1 資料不一致的系統化分析架構......................17
2-5.2 來源資料的整合................................18
2-5.3 解決DW資料不一致的模型架構圖....................19
第 三 章 系統設計與實作..............................20
3-1 系統描述..........................................20
3-1.1 資料庫中介系統.................................20
3-1.2 系統架構介紹...................................21
3-2 建構半自動化異質資料庫整合系統 ...................21
3-2.1 連結資料庫....................................21
3-2.2 資料源集樹狀圖.................................23
3-2.3 建立表格關聯..................................24
3-2.4 轉移資料表....................................26
3-2.5 查詢資料......................................28
第 四 章 資料一致性之處理...........................29
4-1 實體識別..........................................29
4-2 StringMap.........................................31
4-2.1 k維超空間投影 .................................32
4-2.2 R-TREES.......................................35
4-2.2.1分割節點演算法............................38
4-2.2.2比較兩棵R-Trees ..........................40
4-3 表格運算子........................................42
4-3.1 資料過濾運算子.................................43
4-3.2 Sort-Merge Join表格運算子 .......................45
4-3.3拓撲排序......................................52
第 五 章 系統驗證....................................54
5-1 半自動化異質資料庫整合系統........................55
5-1.1測試資料-健保住院申報資料庫......................55
5-1.2查詢資料圖形化 .................................55
5-2 Sort/Merge Join測試結果 ..........................56
5-2.1資料一致性檢查 .................................58
5-3 StringMap測試 ....................................60
第 六 章 總結..........................................63
6-1 結論..............................................63
6-2 系統功能介紹......................................65
6-3 未來工作..........................................66
參考文獻................................................67
圖次
圖1-1: 資料處理流程圖....................................2
圖2-1: 多資料庫系統分類圖................................7
圖2-2: AutoMed Repositories API架構圖....................8
圖2-3: USM建構子和圖形符號..............................11
圖2-4: 綱目轉換流程圖....................................12
圖2-5: 整合資料庫的四層架構圖............................14
圖2-6: 解決DW資料不一致的模型架構圖.....................19
圖3-1: 資料庫中介系統架構圖..............................20
圖3-2: 異質資料庫整合系統架構圖..........................21
圖3-3: 利用JDBC連線資料庫的架構圖.......................22
圖3-4: 儲存資料庫連線資訊的階層圖........................23
圖3-5: 連線池............................................23
圖3-6: 資料源樹狀圖......................................24
圖3-7: 建立異質資料庫表格關聯圖..........................25
圖4-1: K維超空間投影.....................................32
圖4-2: K維超空間投影後之損失距離.........................32
圖4-3: 建構RTree圖......................................36
圖4-4: 區域分割圖........................................38
圖4-5: 兩棵RTrees圖.....................................40
圖4-6: 依據樣版字串走訪字串路徑圖........................44
圖4-7: 交叉比對圖........................................45
圖4-8: sort/merge join圖.................................45
圖4-9: 拓樸排序範例......................................52
圖4-10: 拓樸排序堆疊資料進出圖...........................52
圖4-11: 鄰接串列圖.......................................53圖5-1: 住院申報費用表格關聯圖............................54
圖5-2: 藥物費用表格關聯圖................................55
圖5-3: 藥物費用橫條圖....................................55
圖5-4: Sort/Merge Join表格運算子串接圖...................56
圖5-5: 增加區塊大小的排序時間............................56
圖5-6: 多次合併和一次合併的排序時間......................57
圖5-7: 欄位分析表格運算子串接圖..........................58
圖5-8: 欄位分析結果......................................59
圖5-9: 不同編輯距離計算方式的執行時間....................60
圖5-10: 不同投影維度的回收率和精確率.....................61
圖5-11: 三個字串編輯距離的三角關係.......................62
表次
表2-1: 資料不一致時之處理方式............................18
表3-1: 系統需求表........................................21
表3-2: 不同類JDBC驅動程式優缺點.........................22
表3-3: 全域表格名稱格式..................................24
表3-4: 資料完整性限制....................................25
表3-4: SQL資料型態轉至X資料型態.........................27
表3-5: X資料型態轉至SQL資料型態.........................27
表4-1: 淨相關試驗資料....................................29
表4-2: 編輯距離矩陣......................................33
表4-3: 損失距離矩陣......................................33
[Alva96] Alvaro Monge and Charles Elkan. The field matching problem: Algo-
rithms and applications, Knowledge Discovery and Data Mining, 1996.
[Alva97] Alvaro Monge and Charles Elkan. An efficient domain-independent algorithm for detecting approximately duplicate database records, Research Issues on Data Mining and Knowledge Discovery, 1997.
[Anto84] Antonin Guttman. R-trees: a dynamic index structure for spatial searching, In Proc. ACM SIGMOD Int. Conf. on Management of Data, 1984.
[Auto01] AutoMed: Automatic Generation of Mediator Tools for Heterogeneous Database Integration. Available at
http://www.doc.ic.ac.uk/%7Epjm/automed/
[Bati85] Carlo Batini and Maurizio Lenzerini. A Comparative Analysis of Methodologies for Database Schema Integration, ACM Computing Surveys, Vol. 18, No.4 1985, pp.323-364.
[Chia94] Roger H.L. Chiang, Terry Barron and Veda C. Storey. Reverse Engineering of Relational Databases: Extraction of an EER Model from a Relational Database. Data & Knowledge Engineering, Vol 12, No.2, 1994, pp. 107-142.
[Chia96] Roger H.L. Chiang, Terry Barron and Veda C. Storey. A Framework for the Design and Evaluation of Reverse Engineering Methods for Relational Databases. Data & Knowledge Engineering, Vol 21, No 1, 1996, pp. 57-77.
[Chri95] Christos Faloutsos, King-Ip Lin. FastMap: A fast algorithm for indexing, data-mining and visualization of traditional and multimedia datasets, Proceedings of 24th ACM SIGMOD International Conference on Management of Data, May 1995.
[Data95] DataJoiner. A Multidatabase Server, IBM ,1995.
[Edga01] Edgar Jaspe . Query Translation in Heterogeneous Database Environments, 2001. Available at http://www.dcs.bbk.ac.uk/~edgar/Report.pdf
[Eepe93] Ee-Peng Lim, Satya Prabhakar and etc.. Entity Identification in Database Integration, Information Sciences, 1993.
[Elma02] Ramez Elmasri, Yu-Chi Wu and etc.. Conceptual Modeling for Customized XML Schemas, ER 2002, LNCS 2503, 2002, pp. 429-443.
[Euge86] Eugene W. Myers. An O(ND) difference algorithm and its variations, Algorithmica, vol. 1, 1986, pp. 251-266.
[Gonz99] Gonzalo Navarro. A guided tour to approximate string matching, ACM Computing Surveys, 1999.
[Haas97] L.M. Haas, R.J. Miller and etc.. Transforming Heterogeneous Data with Database Middleware. Beyond Integration, IEEE Computer Society Technical Committee on Data Engineering, 1997.
[Lee01] Dongwon Lee and Yousub Hwang. Extracting Semantic Metadata and Its Visualization, ACM Crossroads, Vol. 7, No.3, 2001, pp.19-27.
[Leon86] Leonard Shapiro. Join processing in database systems with large main mermories, ACM Transactions on Database Systems, Vol. 11, No. 3, 1986, pp. 239-264.
[Lian02] Liang Jin, Chen Li and Sharad Mehrotra. Efficient similarity string joins in large data sets, UCI ICS technical report, 2002.
[Lim00] Ee-Peng Lim and Roger H.L. Chiang. The Integration of Relationship Instances From Heterogeneous Databases, Decision Support Systems, Vol.29, 2000, pp. 153-167.
[Lin00] Yaw-Jen Lin. The Personal Database Framework for Integrating Remote Database Systems on the Internet, NTU, CSIE, 2000.
[Lowr75] Roy Lowrance and Robert A. Wagner. An Extension of the String-to-String Correction Problem, Journal of the ACM, 1975.
[Pare98] Christine Parent and Stefano Spaccapietra. Issues and Approaches of Database Integration, Communications of the ACM, 41, 1998, pp.166-178.
[Park01] Jinsoo Park. Schema Integration Methodology and Toolkit for Heterogeneous and Distributed Geographic Databases, MISRC, 2001, pp.01-31
[Ramo02] Ramon Lawrence and Ken Barker. Using Unity to Semi-Automatically Integrate Relational Schema, Proceedings of the 18th Internatioinal Conference on Data Engineering, IEEE, 2002
[Ram99] Sudha Ram, Jinsoo Park and etc.: A Comprehensive Framework for Classifying Data- and Schema-Level Semantic Conflicts in Geographic and Non-Geographic Databases, In Proceedings of the 9th International Workshop on Information Technology and Systems, 1999, pp.185-190.
[Ram95] Sudha Ram. Intelligent Database Design Using the Unifying Semantic Model, Information and Management, Vol. 29, No. 4, 1995, pp. 191-206.
[Reze99] Fernando de Ferreira Rezende, Ulrich Hermsen and etc.. A Practical Approach to Access Heterogeneous and Distributed Database, Lecture Notes in Computer Science, Vol. 1626, 1999, pp.317-332.
[Rica96] Ricardo A. Baeze-Yates and Chris H. Perleberg. Fast and practical approximate string matching, Information Processing Letters 59, 1996, pp.21-27.
[Rich02] Richard Lenz and Klaus A Kuhn. Integration of Heterogeneous and Autonomous Systems in Hositals, Business briefing: global healthcare issue 3, 2002. Available at
http://www.wmrc.com/businessbriefing/pdf/health3_2002/reference/Ref13.pdf
[Shaf96] John Shafer, Rakesh Agrawal and Manish Mehta. SPRINT: A scalable parallel classier for data mining. In VLDB, 1996, pp.544-555.
[She90] Amit P. Sheth and James A. Larson. Federated database systems for managing distributed, heterogeneous and autonomous databases, ACM Computing Surveys 22, 1990, pp.183-236.
[Toby86] Toby J. Teorey, Dongoing Yang and James P. Fry. A Logical Design Methodology for Relational Databases Using the Extended Entity-Relationship Model, Computing Surveys, Vol. 18, No. 2,1986.
[Turk01] Can Türker and Michael Gertz. Semantic integrity support in SQL:1999 and commercial(Object-)relational database management systems, the VLDB journal, 10, 2001, pp.241-269.
[Weig01] Weiguo Fan, Hongjun Lu and etc.. Discovering and Reconciling Value Conflicts for Numerical Data Integration, Information Systems, 2001.
[Xueq99] Xuequn Wu. A CORBA-Based Architecture for Integrating Distributed and Heterogeneous Databases, 5th International Conference on Engineering of Complex Computer Systems, 1999, pp. 18-22.
[Ye02] D.Y. Ye, M.C. Lee and T.I. Wang. Mobile Agents for Distributed Transactions of a Distributed Heterogeneous Database System, Lecture Notes in Computer Science, Vol. 2453, 2002, pp.403-412.
[林91] 林克韋:DW資料不一致之研究,中央大學,資訊管理研究所,民91
QRCODE
 
 
 
 
 
                                                                                                                                                                                                                                                                                                                                                                                                               
第一頁 上一頁 下一頁 最後一頁 top