write multiple parquet files by chunk_size

Viewed 398

I see the chunk_size argument in arrow::write_parquet(), but it doesn't seem to behave as expected. I would expect the code below to generate 3 separate parquet files, but only one is created, and nrow > chunk_size.

library(arrow)
# .parquet dir and file path
td <- tempdir()
tf <- tempfile("", td, ".parquet")
on.exit(unlink(tf))

# dataframe with 3e6 rows
n  <- 3e6 
df <- data.frame(x = rnorm(n))

# write with chunk_size 1e6, and view directory
write_parquet(df, tf, chunk_size = 1e6)
list.files(td)

Returns one file instead of 3:

[1] "25ff74854ba6.parquet"  
# read parquet and show all rows are there
nrow(read_parquet(tf))

Returns:

[1] 3000000

We can't pass multiple file name arguments to write_parquet(), and I don't want to partition, so write_dataset() also seem inapplicable.

2 Answers

The chunk_size parameter refers to how much data to write to disk at once, rather than the number of files produced. The write_parquet() function is designed to write individual files, whereas, as you said, write_dataset() allows partitioned file writing. I don't believe that splitting files on any other basis is supported at the moment, though it is a possibility in future releases. If you had a specific reason for wanting 3 separate files, I'd recommend separating the data into multiple datasets first and then writing each of those via write_parquet().

(Also, I am one of the devs on the R package, and can see that this isn't entirely clear from the docs, so I'm going to open a ticket to update those - thanks for flagging this up)

Short of an argument to write_parquet() like max_row that defaults to reasonable number (like 1e6), we can do something like this:

library(arrow)
library(uuid)
library(glue)
library(dplyr)

write_parquet_multi <- function(df, dir_out, max_row = 1e6){
  
  # Only one parquet file is needed
  if(nrow(df) <= max_row){
    cat("Saving", formatC(nrow(df), big.mark = ","), 
        "rows to 1 parquet file...")
    write_parquet(
      df, 
      glue("{dir_out}/{UUIDgenerate(use.time = FALSE)}.parquet"))
    cat("done.\n")
  }
  
  # Multiple parquet files are needed
  if(nrow(df) > max_row){
    count = ceiling(nrow(df)/max_row)
    start = seq(1, count*max_row, max_row)
    end   = c(seq(max_row, nrow(df), max_row), nrow(df))
    uuids = UUIDgenerate(n = count, use.time = FALSE)
    
    cat("Saving", formatC(nrow(df), big.mark = ","), 
        "rows to", count, "parquet files...")
    for(j in 1:count){
      write_parquet(
        dplyr::slice(df, start[j]:end[j]), 
        glue("{dir_out}/{uuids[j]}.parquet"))
    }
    cat("done.\n")
  }
  
}


# .parquet dir and file path
td <- tempdir()
tf <- tempfile("", td, ".parquet")
on.exit(unlink(tf))

# dataframe with 3e6 rows
n  <- 3e6 
df <- data.frame(x = rnorm(n))

# write parquet multi
write_parquet_multi(df, td)

list.files(td)

This returns:

[1] "7a1292f0-cf1e-4cae-b3c1-fe29dc4a1949.parquet"    
[2] "a61ac509-34bb-4aac-97fd-07f9f6b374f3.parquet"    
[3] "eb5a3f95-77bf-4606-bf36-c8de4843f44a.parquet"    
Related