so I am trying to calculate the days between the date column and today. And filter table where the diff is more than 5. In spark, you could do something like
datediff(lit(today),df.date) > 5
In pyarrow what I am doing is following
dates = pa.compute.days_between(df['date'], today)
df = df.append_column('days_diff' , dates)
filtered = df.filter(pc.field('days_diff') > 5)
df = df.remove_column('days_diff')
But this creates a new column which is memory overhead. Is it possible to have calculated column for filter only?