I would like to drop columns in my pyarrow table that are null type. Basically NullType columns are columns where all the rows have null data. I would like to drop them since they are not used by me and they cause a conflict when I import them in Spark. Here is an exemple of how I do this right now:
import pyarrow as pa
def drop_na_columns(df):
null_columns = []
schema = df.schema
for (name_, type_) in zip(schema.names, schema.types):
if type_ == pa.null():
null_columns.append(name_)
return df.drop(null_columns)
def create_dataframe(list_dict: dict) -> pa.table:
fields = set()
for d in list_dict:
fields = fields.union(d.keys())
dataframe = pa.table({f: [row.get(f) for row in list_dict] for f in fields})
return drop_na_columns(dataframe)
This works fine for most of my use cases but I also have nested structures in my tables and sometimes one column in a nested structure is a null type column. My question is to know if there is a way to identify the type of every sub-column struct, and then to drop a column in a nested structure. If any of you have an idea or has already tackled this problem I would be delighted to hear about it. Thanks