I am learning Apache Arrow with R and I am trying to get a better understanding of the partitioning mechanisms.
I have a folder with more than 5 000 CSV files that have that
structure: cci_v4_2004106_ppz_takuvik_above_45n.csv
library(arrow)
#>
#> Attaching package: 'arrow'
#> The following object is masked from 'package:utils':
#>
#> timestamp
ds <- open_dataset("~/Desktop/ppz/", format = "csv")
ds
#> FileSystemDataset with 5824 csv files
#> longitude: double
#> latitude: double
#> primary_production: double
Is it possible to infer the partitioning from the filenames? For instance, I would like to use something like:
# ds <- open_dataset("~/Desktop/ppz/", format = "csv", partitioning = c("year", "yday"))
Can I define a schema such as partitioning = c("year", "yday") takes the values from the filename (here: year = 2004, yday = 106).
Created on 2022-03-25 by the reprex package (v2.0.1)