I am working with many csv files that are labelled with the month of year in brackets. For example:
files_names <- list.files("data/", recursive = TRUE, full.names = TRUE)
[1] "data/BOC_All_ATMImage_(Aug 2020).txt" "data/BOC_All_ATMImage_(Aug 2021).txt"
[3] "data/BOC_All_ATMImage_(Feb 2021).txt" "data/BOC_All_ATMImage_(Feb_2020).txt"
[5] "data/BOC_All_ATMImage_(May 2021).txt" "data/BOC_All_ATMImage_(Nov 2019).txt"
column_names <- files_names %>%
str_extract(., "(?<=\\().*?(?=\\))") %>%
str_to_lower() %>%
str_replace(., " ", "_")
"aug_2020" "aug_2021" "feb_2021" "feb_2020" "may_2021" "nov_2019"
I am using the map2 function in purrr to process the csv files and setting a column name using files_names and column_names in a loop.
data <-
map2(files_names, column_names,
~ read_csv(.x, guess_max = 50000) %>%
mutate(
day = 01,
month_year = str_extract(.x, "(?<=\\().*?(?=\\))"),
date_dmy = paste0(day, "-", month_year),
date = dmy(date_dmy),
"{.y}" := 1
),
.id = "group"
)
I need to figure out how to arrange this list so each data set is in chronological order. One approach is to arrange the initial character vectors (files_names and column_names) before feeding them into to loop. Or perhaps it would be easier to simply arrange the data list so the data frames are chronologically ordered? I have created a date variable in each data frame so this could be another approach, but I'm not sure how to reorder the list by a date variable.