Context
I'm trying to retrieve a large amount of data to train a CNN. More specifically, I'm looking for pictures of Swimming pools. I have found a lot of them in the open-images-v6 database made by Google. So now, I just want to download these particular images (I don't want 9 Millions images to end up in my download folder).
Problem
In order to do this, I followed carefully the instructions given on the Download page (see : https://storage.googleapis.com/openimages/web/download.html). So, I installed "fiftyone", tried out the "testing" procedure (which would be loading the "quickstart" dataset and navigating through the data) and have not encountered any issues so far.
But when I tried to retrieve the Swimming pool images with the following code, I went through a lot of issues :
import fiftyone as fo
import fiftyone.zoo as foz
dataset = foz.load_zoo_dataset(
"open-images-v6",
split="validation",
label_types="detections",
classes="Swimming pool"
)
session = fo.launch_app(dataset)
I will skip right to the problem I couldn't figure out : when I run the code, it properly downloads a bunch of .csv files, but when it tries to download the data (the images) it shows a pretty bad looking error :
botocore.exceptions.ClientError: An error occurred (404) when calling the HeadObject operation: Not Found
State of the art
After hours of searching the origin of the error, I eventually discovered that it was somehow linked with AWS, but I have absolutely no clue what I can do on this field.
I saw a random tutorial on internet that recommended to install "awscli" via PIP but nothing changed.
I tried to import other datasets with the same procedure (i.e foz.load_zoo_dataset("coco-2017")) and it seemed to work (at least the download started but I stopped it early).
Thank you for your time.