I've trained a (only-convs) network using small images patches, and now I am using it on larger images. Its working pretty well, except that when the image is much much larger, I end up getting out of memory.
Is there a way to predict the output of the network without saving intermediate states (the ones used for computing the backward pass?).
I just need the output of the forward pass, and I have no use at all for the gradients of the backward pass (the network is already trained).
Ps: I've already tried to split the greater image into patches, but doing so breaks the spatial correlation (particularly on the borders of each patch).
Edit: I am posting the function (code in python) I use for predicting a given img. It takes the img, adjust to the proper shape and value range, predicts, and returns:
def predictImg(img_name,autoencoder):
img = np.asarray(Image.open(img_name)).astype('float32')
img2 = np.expand_dims(np.transpose(img.astype('float32')[:,:,:3],(2,0,1))/255.0,0) # setting up for shape (sample,channels,x,y), ranging between [0,1]
prediction = autoencoder.predict(img2)
return np.transpose(np.round(np.clip(prediction[0]*255,0,255)),(1,2,0)) # getting a [0,255] image (discretized color values)