Continue training style-gan 2 network after crash

Viewed 2205

I've been trying to train a style-gan2 network using a custom dataset. Unfortunately the server I'm currently running the computations on is somewhat unstable, causing it to crash after three days of training. Is there any way for me to continue training the network using the last snapshot of the network before it crashed? I have seen some references to continued training of a network, but neither the style-gan or style-gan2 github pages mention it.

3 Answers

After diggin through the code a bit I figured it out. Turns out there is a resume_pkl variable in training\training_loop. By setting that variable to the path of the snapshot I wanted to resume from I was able to restart the training process. The network has currently resumed training, I'll make another comment here if I encounter any further issues.

look in your stylegan2-master/results/ and find the most recent checkpoint, something like :

network-snapshot-005120.pkl

then you gotta edit a couple variables in training_loop.py

plug in the full path to that checkpoint pkl file (into variable "resume_pkl")

then convert the kimg value ("005120") to a float, and plug it into resume_kimg. resume_kimg important since it needs to know where to resume the learning rate curve thing.

heres what mien looks like:

resume_pkl = '/mnt/harddrive/stylegan2encoder-master/results/00012-stylegan2-testexperiment-1gpu-config-f/network-snapshot-005120.pkl',

resume_kimg  = 5120.0,

as for resume_time, i just leave it at zero because i know its training for like 100 days.

after that,

go back and run the same command you used to start the first session.

python run_training.py --num-gpus=1 --data-dir=/mnt/harddrive/stylegan2encoder-master/datasets/ --config=config-f --dataset=testexperiment
Related