I am working on a problem where I need to apply some transformation to my dataset using the map function that tf.data.Dataset provides. The idea is to apply this transformation that rely on some random number and then chain this transformation with another function.
The idea is something like that:
dataset = tf.data.Dataset.from_tensor_slices([1, 1, 1, 1, 1, 1])
dataset = dataset.map(lambda x: x + tf.random.uniform([], minval=0, maxval=9, dtype=tf.dtypes.int32)) #map function is done once
I thought that if I print dataset twice I should expect the same values, however, the result is the following.
ds = dataset.zip((dataset,dataset))
print(list(ds.as_numpy_iterator()))
#output -> [(8, 2), (2, 1), (8, 9), (2, 2), (6, 7), (2, 2)]
Any clues on how can I get exactly the same values after a .map transformation which relies on random numbers? It seems that the map function is done twice instead of once as I declared in the code snippet.
P.D: Using a random seed does the trick but its just hidding the problem.
[EDITED]
What I need is to perform an operation like this:
ds = tf.data.Dataset.range(1, 10)
y = ds.map(lambda x: my_random_operation(x))
x = ds.map(lambda x: another_random_operation(x))
dataset = tf.data.Dataset.zip((x, y))
As you can see, x depends on y, and it doesn't seem that this behavior is happening. That's why I asked how to apply the random operation first and then apply the map to illustrate this.