I am building a Python Class that works with PySpark dataframes or Hive tables. The input data can either be a str for table_name or a DataFrame. What is the Python best practice to do this? I wanted to do this so that it is flexible depending on the use case. The only difference is on how the data should be passed to the class (either as a table or a Spark DataFrame), everything else would be the same.
See example below:
class DataPipeline:
def __init__(self, data):
if isinstance(data, str):
self.df = spark.read.table(data)
elif isinstance(data, DataFrame):
self.df = data
else:
raise ValueError("some error")
def process_data(self):
# do something with the self.df here