I have many dataframes whose columns have the same order (the column name may differ for each dataframe). And there are 2 columns with a timestamp type but the problem is that in some dataframes, it has a date type. So I cannot merge it with union function.
I want to union all these dataframe but I don't want to cast to_timestamp for each dataframe.
My approach is to change the type of the first dataframe, then the remaining dataframe will follow the type of the first one but it does not work.
from pyspark.sql import functions as F
def change_type_timestamp(df):
df = df.withColumn("A", F.to_timestamp(F.col("A"))) \
.withColumn("B", F.to_timestamp(F.col("B")))
return df
dfs = [df1, df2, df3, ...]
dfs[0] = change_type_timestamp(dfs[0])
reduce(lambda a, b: a.union(b), dfs)
How can I union all the dataframe without changing the type of each dataframe one-by-one?