There will be multiple steps to creating an ensemble model.
Start by creating the two models individually. For the first model, split the data by group and train two individual models, then join the two models together in a function. For the second model, the data can be left in its entirety (aside from removing testing data). Then, create another method to join the other two models into one ensemble model.
To demonstrate, I'll start by importing the necessary modules and loading in the dataframe:
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
data_str = """group,feature_1,feature_2,year,dependent_variable
group_a,12,19,2010,0.4
group_a,11,13,2011,0.9
group_a,10,5,2012,1.2
group_a,16,9,2013,3.2
group_b,8,29,2010,0.6
group_b,9,33,2011,0.1
group_b,111,15,2012,2.1
group_b,16,19,2013,12.2"""
data_list = [row.split(",") for row in data_str.split("\n")]
data = pd.DataFrame(data_list[1:], columns = data_list[0])
train = data.loc[data["year"] != "2013"]
test = data.loc[data["year"] == "2013"]
This will be using a RandomForestRegressor ensemble model, but any regression model can be used. In addition, it should be noted that the dataframe used here differs from the given dataframe in that this dataframe has its rows indexed from 0 rather than being indexed by group, and group is instead a column within the dataframe.
To construct the first model:
- split the data into data for group a and for group b
- train two independent models
- join the models
The first two steps are done below:
# Splitting Data
train_a = train.loc[train["group"] == "group_a"]
train_b = train.loc[train["group"] == "group_b"]
test_a = test.loc[test["group"] == "group_a"]
test_b = test.loc[test["group"] == "group_b"]
# Training Two Models
model_a = RandomForestRegressor()
model_a.fit(train_a.drop(["dependent_variable", "year", "group"], axis = "columns"), train_a.dependent_variable)
model_b = RandomForestRegressor()
model_b.fit(train_b.drop(["dependent_variable", "year", "group"], axis = "columns"), train_b.dependent_variable)
Then, their predict methods can be joined together:
def individual_predictor(group, feature_1, feature_2):
if group == "group_a": return model_a.predict([[feature_1, feature_2]])[0]
elif group == "group_b": return model_b.predict([[feature_1, feature_2]])[0]
This will take in a group and two features individually and return the prediction. This can be adapted to whatever input and output type is necessary.
To create the second model, leave the data as whole and only train one model, which also removes the necessity to join the models:
model = RandomForestRegressor()
model.fit(train.drop(["dependent_variable", "year", "group"], axis = "columns"), train.dependent_variable)
Finally, you can join the models together into an ensemble model by averaging the result of their predict methods:
def ensemble_predict(group, feature_1, feature_2):
return (individual_predictor(group, feature_1, feature_2) + model.predict([[feature_1, feature_2]])[0]) / 2
Again, this takes in a group and two features then returns the result. This will likely need to be adapted into another format, such as taking in a list of list of inputs and outputting a list of predictions.