I have a multiple datasets stored in a partitioned parquet format using the same partitioning file structure, e.g. the directory structure is like:
1/a/predictions.parquet
1/a/summary.parquet
2/a/predictions.parquet
2/a/summary.parquet
1/b/predictions.parquet
1/b/summary.parquet
2/b/predictions.parquet
2/b/summary.parquet
and I want to read the two datasets independently using arrow::open_dataset(). I know I can use list.files(pattern = "predictions.*parquet") to get just the files I want then read those in with open_dataset(), however in this case I loose the partitioning.
Here's an example of what I want to do:
library(arrow)
library(dplyr)
tf <- tempfile()
dir.create(tf)
predictions <- expand.grid(var1 = 1:2, var2 = c("a", "b")) %>%
mutate(prediction = rnorm(nrow(.)))
summary <- expand.grid(var1 = 1:2, var2 = c("a", "b")) %>%
mutate(var3 = runif(nrow(.)))
write_dataset(predictions, tf,
partitioning = c("var1", "var2"),
basename_template = "predictions-{i}.parquet",
hive_style = FALSE)
write_dataset(summary, tf,
partitioning = c("var1", "var2"),
basename_template = "summary-{i}.parquet",
hive_style = FALSE)
list.files(tf, recursive = TRUE)
# partitioning lost
list.files(tf, recursive = TRUE, pattern = "predictions.*parquet",
full.names = TRUE) %>%
open_dataset() %>%
collect()
# tries to read both datasets at once
open_dataset(tf, partitioning = c("var1", "var2")) %>%
collect()
# what i want to do
open_dataset(tf, pattern = "predictions.*parquet",
partitioning = c("var1", "var2")) %>%
collect()
unlink(tf)