Should you use S3 when methods would take different arguments?

Viewed 23

I'm debating whether to use S3 for a package I'm writing or whether to just create separate functions. Both functions do very similar things but one is intended for use on a single input and the other on a series of inputs and hence have different arguments and return different objects. At first I thought I should use S3 however now I'm leaning towards just creating two separate functions (kind of like how yardstick creates separate {metric}() and {metric}_vec() versions for each of the functions in the package.

Say I have a function like read_head() that expects a file_path then reads in the file and returns the head of it.

read_head <- function(file_path, n = 10) { 
  readr::read_csv(file_path, n_max = n)
}

And that I have another function read_head_files() that takes in a dataframe containing a column file_paths and creates a new list-column contents that contains the head of each file. For example

df_example <- tibble::tibble(file_paths = c("folder/file1.csv", "folder/file2.csv"))

read_head_files <- function(df, n = 10){
  dplyr::mutate(df, contents = purrr::map(file_paths, read_head, n = n) 
}

base R S3 generics like print(), str() etc. suggest that S3 can be used pretty liberally with objects that have reasonably different input and return objects. However then the yardstick package example mentioned above makes me think that perhaps in this case it is more appropriate to just create separate functions so that I can expose the argument names directly (note that yardstick though does use S3 so that matrices, dataframes, tables all have their own methods for dispatch).

I'm curious what the guidance is on when you should use S3 / OOP versus just creating individual functions (and specifically recommendations for the case above)?

0 Answers
Related