I am working on a project where I need to take groups of data and predict the next value for that group using a time series model. In my data, I have a grouping variable and a numeric variable.
Here is an example of my data:
import pandas as pd
data = [
["A", 10],
["B", 10],
["C", 15],
["D", 12],
["A", 18],
["B", 19],
["C", 14],
["D", 22],
["A", 20],
["B", 25],
["C", 12],
["D", 30],
["A", 36],
["B", 27],
["C", 10],
["D", 45]
]
data = pd.DataFrame(
data,
columns=[
"group",
"value"
],
)
What I want to do is to create a for loop that iterates over the groups and predicts the next value for A, B, C, and D. Essentially, my end result would be a new data frame with 4 rows, one for each new predicted value. It would look something like this:
group pred_value
A 40
B 36
C 8
D 42
Here is my attempt at that so far:
from statsmodels.tsa.ar_model import AutoReg
final=pd.DataFrame()
for i in data['group']:
group = data[data['group']==i]
model = AutoReg(group['value'], lags=1)
model_fit = model.fit()
yhat = model_fit.predict(len(group), len(group))
final = final.append(yhat,ignore_index=True)
Unfortunately, this produces a data frame with 15 rows and I'm not sure how to get the end result that I described above.
Can anyone help point me in the right direction? Any help would be appreciated! Thank you!