Pandas dataframe: Synthetic data generation

Viewed 240

I have a data frame df that contains 3 classes( classification Problem). The data contains most of the columns as categorical and the dataset is imbalanced. I am trying to generate a synthetic dataset that replicates the characteristics and features of the original data frame.

Q1. Does data.make_classification from scikit-learn can be used to generate synthetic data to balance the imbalanced df?

Q2. Does data.make_classification is used for random data generation only and not reproduce similar data with existing data df?

1 Answers

make_classification is mainly for generating synthetic data from scratch for simple tests. It basically just samples from Gaussians, which is probably not how you want to fill in your data. It also does not support categorical features.

There are many libraries for synthetic data generation (quick Google/Github search) which attempt to mimic properties of the original data.

You should also consider oversampling techniques, some of which involve synthetic data generation. Check out the Imbalanced Learn library for a nice introduction, including algorithms which support categorical and mixed data.

Related