How do I create pairs of images and captions in OpenAI CLIP format?

Viewed 56

I have the following problem. I want to make training data for OpenAIs dall-e in the CLIP format, i.e. make pairs of image and text (captions) with the same filename. I've tried to do that in the code below, but it doesn't work as I want it. The function "caption_image" creates 5 captions for each image in the folder, and prints that, but when I try to create a txt file with the 5 correct captions for each image, somehow all the text files end up with the just 1 caption, and it's the same caption for all the images in the folder.

Can anyone see what the mistake is in the code below?

Thanks a lot. I'm new to coding, and I hope you can help.

import os
from google.colab import drive
drive.mount('/content/drive')

images = []

path = '/content/drive/My Drive/Godard_imgs/Eloge_jpgs/eloge_sample/'
for filename in [filename for filename in os.listdir(path) if filename.endswith(".png") or filename.endswith(".jpg")]:
  image = os.path.join(path, filename)
  images.append(image)

for img in images:
  captionss = caption_image(img, args, net, preprocess)
  for caption in captionss:
    for filename in [filename for filename in os.listdir(path) if filename.endswith(".png") or filename.endswith(".jpg")]:
      with open("{}.txt".format(filename), "w") as f: 
        f.write(caption + '\n')
0 Answers
Related