I generate a pandas dataframe from read_sql_query. It has three columns, "results, speed, weight"
I want to use scikit-learn LinearRegression to fit results = f(speed, weight)
I haven't been able to find the correct syntax that would allow me to pass this dataframe, or column slices of it, to LinearRegression.fit(y, X).
print df['result'].shape
print df[['speed', 'weight']].shape
(8L,)
(8, 2)
but I cannot pass that to fit
lm.fit(df['result'], df[['speed', 'weight']])
It throws a deprecation warning and a ValueError
DeprecationWarning: Passing 1d arrays as data is deprecated in 0.17 and willraise ValueError in 0.19.
ValueError: Found arrays with inconsistent numbers of samples: [1 8]
What is the efficient, clean way to take dataframes of targets and features, and pass them to fit operations?
This is how I generated the example:
import pandas as pd
import numpy as np
from datetime import datetime, timedelta
date_today = datetime.now()
days = pd.date_range(date_today, date_today + timedelta(7), freq='D')
np.random.seed(seed=1111)
data = np.random.randint(1, high=100, size=len(days))
data2 = np.random.randint(1, high=100, size=len(days))
data3 = np.random.randint(1, high=100, size=len(days))
df = pd.DataFrame({'test': days, 'result': data,'speed': data2,'weight': data3})
df = df.set_index('test')
print(df)