I would like to read multiple data sets and combine them into a single Pandas dataframe with a year column.
My sample data sets include newyork2000.txt, newyork2001.txt, newyork2002.txt.
Each data set contains 'address' and 'price'.
Below is the newyork2000.txt:
253 XXX st, 150000
2567 YYY st, 200000
...
3896 ZZZ rd, 350000
My final single dataframe should look like this:
year address price
2000 253 XXX st 150000
2000 2567 YYY st 200000
...
2000 3896 ZZZ rd 350000
...
2002 789 XYZ ave 450000
So, I need to combine all data sets, create the year column, and name the columns.
Here is my code to create a single dataframe:
years=[2000,2001,2002]
df=[]
for i years:
df.append(pd.read_csv("newyork" + str(i) + ".txt", header=None))
dfs=pd.concat(df)
But, I could not create the year column and name the columns. Please help me solve this problem.