Sampling for large class and augmentation for small classes in each batch

Viewed 317

Let's say we have 2 classes one is small and the second is large.

I would like to use for data augmentation similar to ImageDataGenerator for the small class, and sampling from each batch, in such a way, that, that each batch would be balanced. (Fro minor class- augmentation for major class- sampling).

Also, I would like to continue using image_dataset_from_directory (since the dataset doesn't fit into RAM).

4 Answers

You can use tf.data.Dataset.from_generator that allows more control on your data generation without loading all your data into RAM.

def generator():
 i=0   
 while True :
   if i%2 == 0:
      elem = large_class_sample()
   else :
      elem =small_class_augmented()

   yield elem
   i=i+1
  

ds= tf.data.Dataset.from_generator(
         generator,
         output_signature=(
             tf.TensorSpec(shape=yourElem_shape , dtype=yourElem_ype))
    

This generator will alterate samples between the two classes,and you can add more dataset operations(batch , shuffle..)

What about sample_from_datasets function?

import tensorflow as tf
from tensorflow.python.data.experimental import sample_from_datasets

def augment(val):
    # Example of augmentation function
    return val - tf.random.uniform(shape=tf.shape(val), maxval=0.1)

big_dataset_size = 1000
small_dataset_size = 10

# Init some datasets
dataset_class_large_positive = tf.data.Dataset.from_tensor_slices(tf.range(100, 100 + big_dataset_size, dtype=tf.float32))
dataset_class_small_negative = tf.data.Dataset.from_tensor_slices(-tf.range(1, 1 + small_dataset_size, dtype=tf.float32))

# Upsample and augment small dataset
dataset_class_small_negative = dataset_class_small_negative \
    .repeat(big_dataset_size // small_dataset_size) \
    .map(augment)

dataset = sample_from_datasets(
    datasets=[dataset_class_large_positive, dataset_class_small_negative], 
    weights=[0.5, 0.5]
)

dataset = dataset.shuffle(100)
dataset = dataset.batch(6)

iterator = dataset.as_numpy_iterator()
for i in range(5):
    print(next(iterator))

# [109.        -10.044552  136.        140.         -1.0505208  -5.0829906]
# [122.        108.        141.         -4.0211563 126.        116.       ]
# [ -4.085523  111.         -7.0003924  -7.027302   -8.0362625  -4.0226436]
# [ -9.039093  118.         -1.0695585 110.        128.         -5.0553837]
# [100.        -2.004463  -9.032592  -8.041705 127.       149.      ]

Set up the desired balance between the classes in the weights parameter of sample_from_datasets.

As it was noticed by Yaoshiang, the last batches are imbalanced and the datasets length are different. This can be avoided by

# Repeat infinitely both datasets and augment the small one
dataset_class_large_positive = dataset_class_large_positive.repeat()
dataset_class_small_negative = dataset_class_small_negative.repeat().map(augment)

instead of

# Upsample and augment small dataset
dataset_class_small_negative = dataset_class_small_negative \
    .repeat(big_dataset_size // small_dataset_size) \
    .map(augment)

This case, however, the dataset is infinite and the number of batches in epoch has to be further controlled.

I didn't totally follow the problem. Would psuedo-code this work? Perhaps there are some operators on tf.data.Dataset that are sufficient to solve your problem.

ds = image_dataset_from_directory(...)

ds1=ds.filter(lambda image, label: label == MAJORITY)
ds2=ds.filter(lambda image, label: label != MAJORITY)

ds2 = ds2.map(lambda image, label: data_augment(image), label)

ds1.batch(int(10. / MAJORITY_RATIO))
ds2.batch(int(10. / MINORITY_RATIO))

ds3 = ds1.zip(ds2)

ds3 = ds3.map(lambda left, right: tf.concat(left, right, axis=0)

You can use the tf.data.Dataset.from_tensor_slices to load the images of two categories seperately and do data augmentation for the minority class. Now that you have two datasets combine them with tf.data.Dataset.sample_from_datasets.

# assume class1 is the minority class
files_class1 = glob('class1\\*.jpg')
files_class2 = glob('class2\\*.jpg')

def augment(filepath):
    class_name = tf.strings.split(filepath, os.sep)[0]
    image = tf.io.read_file(filepath)
    image = tf.expand_dims(image, 0)
    if tf.equal(class_name, 'class1'):
        # do all the data augmentation
        image_flip = tf.image.flip_left_right(image)
    return [[image, class_name],[image_flip, class_name]]

# apply data augmentation for class1
train_class1 = tf.data.Dataset.from_tensor_slices(files_class1).\
map(augment,num_parallel_calls=tf.data.AUTOTUNE)
train_class2 = tf.data.Dataset.from_tensor_slices(files_class2)

dataset = tf.python.data.experimental.sample_from_datasets(
datasets=[train_class1,train_class2], 
weights=[0.5, 0.5])

dataset = dataset.batch(BATCH_SIZE)
Related