how to append dataframe only with data to headers in pandas?

Viewed 403

I have a dataframe with only data . I have hardocded headers in my code. How do I append data to the headers with the data already in data frame. data.csv has random data with random columns and I have to pick only certain records with specific filter which I am doing by iloc and assigning to data frame df_NEW. below is my code:

import pandas as pd

df=pd.read.csv("C:\\users\\data.csv")
headers=['col1','col2','col3','col4','col5']
df_NEW=df[df.iloc[:,3]=='NEW']       
df_NEW_final=pd.DataFrame(df_NEW,columns=header)

I know how to append the headers to the data while reading csv. But here, I have to read csv and assign the data based on iloc filter to one data frame and then the result we get from df_NEW, we have to add the columns to this data frame.

df_NEW_final=pd.DataFrame(df_NEW,columns=header)

above line gives me only headers but without data.

Also, I have 46 headers and only few columns with values . If data is not present then that column should go as blank.

How do i do this ?

1 Answers

You could get the number of columns (as int) after your df.iloc operation.

df_NEW_col_length = len(df.columns)

Now you can keep the values you need from "headers"

headers[:df_NEW_col_length]

Here my code:

import pandas as pd
df = pd.read.csv("C:\\users\\data.csv")
df_NEW = df[df.iloc[:,3]=='NEW']  
# count how many columns there are in df_new
df_NEW_col_length = len(df.columns)
headers=['col1','col2','col3','col4','col5']
new_headers= headers[:df_NEW_col_length]
df_NEW.columns = new_headers

NOTE: This solution assumes a fixed/sorted headers

EDIT: You can dynamically map your new data model in this way:

df_NEW = pd.DataFrame({0:[20, 10],
                       1:[30, 20],
                       2:[40, 40]
                       })

df_NEW_col_length = len(df_NEW.columns)
headers = ['col1', 'col2', 'col3', 'col4', 'col5']
new_headers = headers[:df_NEW_col_length]
new_col_df = [col for col in df_NEW]

final = {}
for i,index in  zip(new_headers,new_col_df):
      final.update({i: df_NEW[index].tolist()})

final_df = pd.DataFrame(final)

output:

Out[1]: 

   col1  col2  col3
0    20    30    40
1    10    20    40
Related