How does tensorflow pipeline work with data that does not fit in memory?

Viewed 48

Hi I'm new to machine learning and I had a problem loading a big image dataset. I saw some video about pipeline in tensorflow and I think I quite understand the concept behind: you give some data, you work on it (map, filter, ecc...) and you get new data to give to the model.

My question was about how can this system handle a very large dataset?

images_filenames = tf.constant(image_list)
masks_filenames = tf.constant(mask_list)

dataset = tf.data.Dataset.from_tensor_slices((images_filenames,
                                              masks_filenames))

def process_path(image_path,mask_path):
    img = tf.io.read_file(image_path)
    img = tf.image.decode_png(img,channels=3)
    img = tf.image.convert_image_dtype(img,tf.float32) #this do the same as dividing by 255 to set the values between 0 and 1 (normalization)
    mask = tf.io.read_file(mask_path)
    mask = tf.image.decode_png(mask,channels=3)
    mask = tf.math.reduce_max(mask,axis=-1,keepdims=True)
    return img , mask

def preprocess(image,mask): 
    input_image = tf.image.resize(image,(96,128),method='nearest')
    input_mask = tf.image.resize(mask,(96,128),method='nearest')
    
    return input_image , input_mask

image_ds = dataset.map(process_path) # apply the preprocces_path function to our dataset
processed_image_ds = image_ds.map(preprocess) # apply the preprocess function to our dataset

In this case if I map my dataset I will load every image into the dataset. So what differences is it from a normal pandas dataframe?

0 Answers
Related