Hi I'm new to machine learning and I had a problem loading a big image dataset. I saw some video about pipeline in tensorflow and I think I quite understand the concept behind: you give some data, you work on it (map, filter, ecc...) and you get new data to give to the model.
My question was about how can this system handle a very large dataset?
images_filenames = tf.constant(image_list)
masks_filenames = tf.constant(mask_list)
dataset = tf.data.Dataset.from_tensor_slices((images_filenames,
masks_filenames))
def process_path(image_path,mask_path):
img = tf.io.read_file(image_path)
img = tf.image.decode_png(img,channels=3)
img = tf.image.convert_image_dtype(img,tf.float32) #this do the same as dividing by 255 to set the values between 0 and 1 (normalization)
mask = tf.io.read_file(mask_path)
mask = tf.image.decode_png(mask,channels=3)
mask = tf.math.reduce_max(mask,axis=-1,keepdims=True)
return img , mask
def preprocess(image,mask):
input_image = tf.image.resize(image,(96,128),method='nearest')
input_mask = tf.image.resize(mask,(96,128),method='nearest')
return input_image , input_mask
image_ds = dataset.map(process_path) # apply the preprocces_path function to our dataset
processed_image_ds = image_ds.map(preprocess) # apply the preprocess function to our dataset
In this case if I map my dataset I will load every image into the dataset. So what differences is it from a normal pandas dataframe?