Python RandomForestRegressor gives identical values for all predictions

Viewed 336

I'm running a RandomForestRegressor model and all my test data returns the same values despite having different inputs. I asked a friend to run the identical code and there it seems to be working. I tried running it locally and on Google cloud, but I still have the problem...

Here is the problematic code and the result. This used to be working, by the way. Any ideas? Edit: link to the input file here: https://drive.google.com/open?id=1PHKHJHiDUrTC93Wn7O_6JkaUPFE0-AfJ

df = pd.read_excel('sportsref-qbdata-raw-v2.xlsx', header=1)

cols_to_keep = ['Player', 'QB_score', 'Height-in', 'Weight', 'BMI', 'BCS School', 'Cmp', 'Att', 'Pct', 'Yds', 'AY/A', 'TD', 'Int', 'Rate', 'Rush_Att', 'Rush_Yds', 'Rush_Avg', 'Rush_TD']

df = df[cols_to_keep]
df['inter_rate'] = df['Int'] / df['Att']
df['td_attpt'] = df['TD'] / df['Att']

X = df.drop(['Player', 'QB_score'], axis=1)
y = df['QB_score']

# Fit random forest regression model
model_1 = RandomForestRegressor(n_estimators=10, random_state=0)
model_2 = RandomForestRegressor(n_estimators=40, random_state=0)
model_1.fit(X, y)
model_2.fit(X, y)

y_1 = model_1.predict(X)
y_2 = model_2.predict(X)

df['rf_1'], df['rf_2'] = y_1, y_2

print('Random Forest with 10 trees %s' %model_1.score(X, y))
print('Random Forest with 40 trees %s' %model_2.score(X, y))

mahomes = {'BMI': 29.48,
'BCS school': 1,
'Cmp': 857,
'Att': 1349,
'Pct': 63.5,
'Yds': 11252,
'AY/A': 8.8,
'TD': 93,
'Int': 29,
'Rate': 152.0,
'inter_rate': 0.02149740548,
'td_attmpt': 0.06893995552,
'Height-in': 74.13,
'Weight': 225,
'Rush_Att': 308 ,
'Rush_Yds': 845 ,
'Rush_Avg': 2.7 ,
'Rush_TD': 22}

trubisky = {'BMI': 29.09,
'BCS school': 1,
'Cmp': 386,
'Att': 572,
'Pct': 67.5,
'Yds': 4762,
'AY/A': 9.0,
'TD': 41,
'Int': 10,
'Rate': 157.6,
'inter_rate': 0.01748251748,
'td_attmpt': 0.07167832167,
'Height-in': 74.13,
'Weight': 222,
'Rush_Att': 120 ,
'Rush_Yds': 439 ,
'Rush_Avg': 3.7 ,
'Rush_TD': 8}

watson = {'BMI': 28.96,
'BCS school': 1,
'Cmp': 814,
'Att': 1207,
'Pct': 67.4,
'Yds': 10168,
'AY/A': 8.7,
'TD': 90,
'Int': 32,
'Rate': 157.5,
'inter_rate': 0.02651201325,
'td_attmpt': 0.07456503728,
'Height-in': 74.13,
'Weight': 221,
'Rush_Att': 435 ,
'Rush_Yds': 1934 ,
'Rush_Avg': 4.4 ,
'Rush_TD': 26}

mh = pd.DataFrame.from_dict(mahomes, orient='index')
mh = mh.T
tb = pd.DataFrame.from_dict(trubisky, orient='index')
tb = tb.T

wa = pd.DataFrame.from_dict(watson, orient='index')
wa = wa.T

print("Mahomes ", model_1.predict(mh))
print("Mahomes ", model_2.predict(mh))

print("Trubisky ", model_1.predict(tb))
print("Trubisky ", model_2.predict(tb))

print("Watson ", model_1.predict(wa))
print("Watson ", model_2.predict(wa))

OUTPUT:

Random Forest with 10 trees 0.7632960480975086

Random Forest with 40 trees 0.8421489310137844

Mahomes [9.96887821]

Mahomes [7.9205464]

Trubisky [9.96887821]

Trubisky [7.9205464]

Watson [9.96887821]

Watson [7.9205464]

0 Answers
Related