Data Preprocessing and Data Validations
Maria Cristina Mariani, Osei Kofi Tweneboah, Maria Pia Beccar-Varela
Abstract
Maria Cristina Mariani, Osei Kofi Tweneboah, Maria Pia Beccar-Varela
Abstract
This chapter discusses various techniques for preprocessing and validating data before applying any technique to analyze and describe it. The tasks in data preprocessing are as follows: data cleaning, data transformation, and data reduction. The steps and techniques for data cleaning will vary from dataset to dataset. There are three types of missing data according to the mechanisms of missingness namely, missing completely at random, missing at random, and missing not at random. The best solution to handling missing data is to prevent it from occuring by well-planning the study and using effective methods for the data collection. Many algorithms attempt to find trends in the data by comparing features of data points. Data transformation can be simple or complex based on the required changes to the data between the source data and the target data. The chapter describes two normalization techniques namely, min-max normalization and Z-score normalization.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This chapter discusses various techniques for preprocessing and validating data before applying any technique to analyze and describe it. The tasks in data preprocessing are as follows: data cleaning, data transformation, and data reduction. The steps and techniques for data cleaning will vary from dataset to dataset. There are three types of missing data according to the mechanisms of missingness namely, missing completely at random, missing at random, and missing not at random. The best solution to handling missing data is to prevent it from occuring by well-planning the study and using effective methods for the data collection. Many algorithms attempt to find trends in the data by comparing features of data points. Data transformation can be simple or complex based on the required changes to the data between the source data and the target data. The chapter describes two normalization techniques namely, min-max normalization and Z-score normalization.
Key concepts: Missing data, Normalization (sociology), Preprocessor, Data pre-processing, Database normalization, Computer science, Data mining, Data reduction