Download all files and subdirectories from a Google Drive directory from R

Viewed 669

There are some prior related questions (1, 2, 3), but nothing quite what I want, and I can't get the code example to work that Jenny Bryan posted in 2018.

I have a folder shared with me with some large files. The files are nested. So I want to recurse into the sub-directories and get all files from each. In my case, there are only two layers, but it would be nice with an approach that works for arbitrary number of layers.

The most obvious command to try is simply telling it to download the folder, hoping it will figure out the substructure:

#load the libraries
library(tidyverse)
library(googledrive)

#folder link to id
#hidden for privacy reasons
jp_folder = "https://drive.google.com/drive/folders/XXXXX" 
folder_id = drive_get(as_id(jp_folder))

#download in entirety
drive_download(folder_id)

Unfortunately, this doesn't work because it apparently cannot deal with folders:

> drive_download(folder_id)
Error: Not a recognized Google MIME type:
  * application/vnd.google-apps.folder

Here's my attempt at avoiding this issue by going into each subdir:

#load the libraries
library(tidyverse)
library(googledrive)

#folder link to id
#hidden for privacy reasons
jp_folder = "https://drive.google.com/drive/folders/XXXXX" 

#get the id data frame
folder_id = drive_get(as_id(jp_folder))

#find files in folder
files = drive_ls(folder_id)

#loop dirs and download files inside them
for (i in seq_along(files$name)) {
  i_dir = drive_ls(files$id[i])
  
  #download files
  walk(i_dir$id, ~ drive_download(as_id(.x)))
}

The files object seems fine (replacing the strings with fillers):

# A tibble: 6 x 3
  name   id                                drive_resource   
* <chr>  <chr>                             <list>           
1 A      AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA <named list [32]>
2 B      BBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBB <named list [32]>
3 C      CCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCC <named list [31]>
4 D      DDDDDDDDDDDDDDDDDDDDDDDDDDDDDDDDD <named list [31]>
5 E      EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE <named list [31]>
6 F      FFFFFFFFFFFFFFFFFFFFFFFFFFFFFFFFF <named list [31]>

However, when one attempts to get the contents of the subdir, it throws this error:

> i_dir = drive_ls(files$id[i])
Error: 'path' does not identify at least one Drive file.

What's wrong here?

1 Answers

Actually, it is simple: drive_ls() wants a data frame input with 1 row, not a character vector. The error message is misleading (it would be nice if it simply told the user to give it a data frame). If one changes the code to that, and adds the required loops, one can automatically download the contents of the sub-dirs. It will fail if there are sub-sub-dirs. A proper recursive function needs to be written and implemented in the package.

This code works for me:

#load the libraries
library(tidyverse)
library(googledrive)
  
#folder link to id
jp_folder = "https://drive.google.com/drive/folders/XXXXX"
folder_id = drive_get(as_id(jp_folder))

#find files in folder
files = drive_ls(folder_id)

#loop dirs and download files inside them
for (i in seq_along(files$name)) {
  #list files
  i_dir = drive_ls(files[i, ])
  
  #mkdir
  dir.create(files$name[i])
  
  #download files
  for (file_i in seq_along(i_dir$name)) {
    #fails if already exists
    try({
      drive_download(
        as_id(i_dir$id[file_i]),
        path = str_c(files$name[i], "/", i_dir$name[file_i])
      )
    })
  }
}

This version skips files already downloaded.

Related