Split train and test set according to categorical columns

Viewed 993

I have a dataframe containing around 25000 rows and 32 columns. I'd like to split this dataset into a train and test test (80/20). However, there are certain columns 1-hot encoded. Now when splitting the data I would like to get the same proportion of each 1-hot encoded column into the training set.

col_1     col_2   ..  col_31    col_32
  1          0         0         0
  1          0         0         0
...
  0          0         1         0
  0          0         1         0

So in the training set there should be 80% of the rows where each column equals 1. I've looked at different splitting methods from Sci-kit learn but was not able to find one that could accommodate my needs. Is there anyone with a solution or that is able to help me?

1 Answers

Initializing pandas frame

import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.DataFrame(np.random.randint(0,2, size=(100, 32)))

setting 1-hot-encoded = True/False if count(no.of.zero in a row) == 1

df['1-hot-encoded'] =  df.apply(lambda row: True if np.count_nonzero(row) == 1 else False, axis=1)

splitting while maintaining ratio of 1-hot-encoded

train, test = train_test_split(df, test_size=0.2, stratify=df['1-hot-encoded'])
Related