Google Colab not saving model or model checkpoints?

Viewed 1334

I've been using Colab to train my models, but it's quite infuriating that so far I have only been able to save the weights to my Google Drive, not the whole model, or even model checkpoints.

I mounted Google Drive with:

from google.colab import drive
drive.mount('/content/gdrive')

And I know that I can read files from the Drive as this code works:

import numpy as np
with np.load("/content/gdrive/MyDrive/trainingData.npz") as f:
    dataX = f["dataX"]
    dataY = f["dataY"]

And I set up the TPU using the following:

%tensorflow_version 2.x
import tensorflow as tf
print("Tensorflow version " + tf.__version__)

try:
  tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection
  print('Running on TPU ', tpu.cluster_spec().as_dict()['worker'])
except ValueError:
  raise BaseException('ERROR: Not connected to a TPU runtime; please see the previous cell in this notebook for instructions!')

tf.config.experimental_connect_to_cluster(tpu)
tf.tpu.experimental.initialize_tpu_system(tpu)
tpu_strategy = tf.distribute.experimental.TPUStrategy(tpu)

But when I run the following code, no model checkpoints get saved:

with tpu_strategy.scope():
  model = Sequential()
  model.add(LSTM(256, input_shape=(dataX.shape[1], dataX.shape[2])))
  model.add(Dropout(0.2))
  model.add(Dense(dataY.shape[1], activation="softmax"))
  model.compile(loss='categorical_crossentropy', optimizer='adam')

  filepath="/content/gdrive/MyDrive/weights-improvement-{epoch:02d}-{loss:.4f}.hdf5"
  checkpoint = ModelCheckpoint(filepath, monitor='loss', verbose=1, save_best_only=True, mode='min')
  callbacks_list = [checkpoint]

  model.fit(dataX, dataY, epochs=50, batch_size=128)

I can't even just save the model normally: model.save("/content/gdrive/MyDrive/model") gives:

UnimplementedError: File system scheme '[local]' not implemented (file: 'model/variables/variables_temp/part-00000-of-00001')
    Encountered when executing an operation using EagerExecutor. This error cancels all future operations and poisons their output tensors.

The interesting thing is that I can still save model weights, via model.save_weights("/content/gdrive/MyDrive/model.h5")

However, as I want to be able to save the whole model for future training, just saving the weights is not satisfactory.

What errors have I made and how can I save my model?

0 Answers
Related