I have a file in my directory. I read it in with Python and then extract certain information from the file to place inside a dataframe as a new row.
Next, I want to read in a second file ... do the same steps as above and append the extracted data as a new row below the original row.
This process will repeat for hundreds of files sitting in one directory.
The two files below are fed into the input_file parameter in my function.
File 1: frame_IceCat_Category_788.feather
File 2: frame_IceCat_Category_42.feather
Function:
def miss_val_df(input_file):
read_df = feather.read_feather(input_file)
read_df.shape
perc_missing = (read_df.isnull().sum()*100) / len(read_df)
missing_df = pd.DataFrame({'col_name': read_df.columns, 'percent_missing': perc_missing})
avg_missing = np.average(missing_df['percent_missing'])
missing_df_2 = pd.DataFrame({'filename': input_file, 'num_rows': read_df.shape[0], 'num_cols': read_df.shape[1], 'missing_data_%': avg_missing}, index=[0])
missing_df_2 = missing_df_2.append(missing_df_2)
return missing_df_2
Calling my function twice:
miss_val_df('frame_IceCat_Category_788.feather')
miss_val_df('frame_IceCat_Category_42.feather')
My output table is:
As you can see, the output table is wrong. It overwrote the data from my first file: "frame_IceCat_Category_788.feather"
What am I doing wrong? How do I fix the above code?
