ETL Best Practices for Data Quality Checks in RIS Databases

Journal Title: Informatics - Year 2019, Vol 6, Issue 1

Abstract

The topic of data integration from external data sources or independent IT-systems has received increasing attention recently in IT departments as well as at management level, in particular concerning data integration in federated database systems. An example of the latter are commercial research information systems (RIS), which regularly import, cleanse, transform and prepare the analysis research information of the institutions of a variety of databases. In addition, all these so-called steps must be provided in a secured quality. As several internal and external data sources are loaded for integration into the RIS, ensuring information quality is becoming increasingly challenging for the research institutions. Before the research information is transferred to a RIS, it must be checked and cleaned up. An important factor for successful or competent data integration is therefore always the data quality. The removal of data errors (such as duplicates and harmonization of the data structure, inconsistent data and outdated data, etc.) are essential tasks of data integration using extract, transform, and load (ETL) processes. Data is extracted from the source systems, transformed and loaded into the RIS. At this point conflicts between different data sources are controlled and solved, as well as data quality issues during data integration are eliminated. Against this background, our paper presents the process of data transformation in the context of RIS which gains an overview of the quality of research information in an institution’s internal and external data sources during its integration into RIS. In addition, the question of how to control and improve the quality issues during the integration process in RIS will be addressed.

Authors and Affiliations

Otmane Azeroual, Gunter Saake and Mohammad Abuosba

Keywords

Related Articles

Applications of Blockchain Technology to Logistics Management in Integrated Casinos and Entertainment

The gaming industry has evolved into a multi-functional smart city that combines integrated casinos and entertainment (ICE). ICE logistics involve supply chains with various stages in geographically-distributed locatio...

Unstructured Text in EMR Improves Prediction of Death after Surgery in Children

Text fields in electronic medical records (EMR) contain information on important factors that influence health outcomes, however, they are underutilized in clinical decision making due to their unstructured nature. We...

Embracing First-Person Perspectives in Soma-Based Design

A set of prominent designers embarked on a research journey to explore aesthetics in movement-based design. Here we unpack one of the design sensitivities unique to our practice: a strong first person perspective—where...

Teaching HCI Skills in Higher Education through Game Design: A Study of Students’ Perceptions

Human-computer interaction (HCI) is an area with a wide range of concepts and knowledge. Therefore, a need to innovate in the teaching-learning processes to achieve an effective education arises. This article describes...

Creating a Multimodal Translation Tool and Testing Machine Translation Integration Using Touch and Voice

Commercial software tools for translation have, until now, been based on the traditional input modes of keyboard and mouse, latterly with a small amount of speech recognition input becoming popular. In order to test wh...

Download PDF file
  • EP ID EP44166
  • DOI https://doi.org/10.3390/informatics6010010
  • Views 275
  • Downloads 0

How To Cite

Otmane Azeroual, Gunter Saake and Mohammad Abuosba (2019). ETL Best Practices for Data Quality Checks in RIS Databases. Informatics, 6(1), -. https://www.europub.co.uk/articles/-A-44166