By: Milind Zodge

Overview

Data preparation is one of the early step in machine learning process. This step is typically followed right after the exploration of problem, identification of data sources and data descriptive exploration. Once you identified the data source you gather data from that source and conduct data exploration using various techniques like box-plot, co-relation matrix and scatter plot to understand the data. By doing this you may surface few challenges with the data which you need to handle before you use the data for training.

In this post I am focusing on few of such challenges and possible solutions.

Challenges

  1. Data can be incomplete:  For this case try to enrich data, get more data attributes from different sources and/or use external data sources
  2. Data can be missing: For this case you can use imputation logic e.g using mean values to fill in for missing data, update Null values to Nan or you may want to eliminate that instance all together
  3. Data can be untidy: What I mean by untidy is that you may have one column with multiple variables or variables in rows and columns for such case you will to use various technique like pivot/un-pivot the most common used technique is melt and cast process
  4. Data can be sparsed: For this case try to change representation of data e.g. COO matrix. If there are lot of zeros then you can normalize the data
  5. Data may have high cardinality variables: For this you can use binning to avoid using this column in training, e.g. primary key of the record
  6. Data with varying scales: You can use rescaling the attributes technique
  7. Data have outliers: Means out-of-range values unknown categorical values use binning also known as discretization or winsorizing basically assigning lesser  weight
  8. Data have lots of features: One of the way you can reduce data set is by eliminating unwanted features, you can use univariate selection technique which selects those features that have strong relationship with the target variable
  9. Data have lots of dimensionality:  For this case use dimension reduction techniques like PCA, principal Component Analysis to reduce dimensions

Discover more from Milind Zodge

Subscribe now to keep reading and get access to the full archive.

Continue reading