How to reshape 'year' from a single feature?

Viewed 92

I have a df containing columns 'year' and 'per capita income (US$)'.

plt.scatter(df.year, df['per capita income (US$)'], color='red')
plt.xlabel('Year')
plt.ylabel('Per Capita Income (US$)')
plt.show()

reg = linear_model.LinearRegression()
reg.fit(df[['year']], df['per capita income (US$)'])

reg.predict(2011) 

Error message received:

ValueError: Expected 2D array, got scalar array instead:
array=2011.
Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.

I adjusted the call of reg.predict() to:

reg.predict([[2011]])

Code executed without error, however, the .predict() function didn't return the desired output.

print(df.columns)

Index(['year', 'per capita income (US$)'], dtype='object')
1 Answers

You should reshape your X to be a 2D array not 1D array. Fitting a model requires requires a 2D array. i.e (n_samples, n_features).
When you use .reshape(-1,1) it adds one dimension to the data.

X = df['year'].values.reshape(-1,1)
y = df['per capita income (US$)'].values
reg = linear_model.LinearRegression()
reg.fit(X,y)
print(reg.predict([[2011]]))
Related