For our demonstration, we’ll just use the ten digits dataset from sklearn. Pendigits dataset consists of 10 classes from digit 0 to digit 9.
from sklearn.datasets import load_digits
digits = load_digits()
print(digits.data.shape)
print(digits.target.shape)
Output looks like -
(1797, 64)
(1797,)
So each digit consists of some sample dataset. I would like to have a subsample of each class from the dataset. For example from digit 0 to digit 9, I need 50 subsamples of each class present in the dataset.
print(digits.data.shape)
print(digits.target.shape)
Result should be(50 subsample * 10 class = 500 subsample) -
(500,64)
(500)
Result should consist of subsample of each class available in the dataset.