I build an support/resistance like feature for my deep learning model (cryptocurrency prediction). The newly trained model's results are very good, with the normal evaluation the results are also very good. But with an evaluation that only predicts on the last sample in the dataframe the results are not nearly as good. The evaluation is done over the same period with the same amount of samples.
So I am concerned that the support/resistance feature uses future data somehow. I have had this issue before when I made the code for the first time but I already fixed that error. So does anyone have any ideas if this code does indeed use future data which would not be possible in real-life.
This problem happened since this new feature so the rest of the code is good.
The code works like the following, it retrieves the high & low per slices of 12 samples. Then it fills in the value of the high/low from the point itself till just before the next high/low.
Code:
# find lows & highs in window.
lows, highs = [], []
max = int(len(df) / window)
diff = len(df) - (max * window)
for index in range(max):
if index == 0:
sliced = df.iloc[index*window: ((index+1)*window)+diff]
else:
sliced = df.iloc[(index*window)+diff: ((index+1)*window)+diff]
high = sliced["high"].max()
index = sliced.index[sliced['high'] == high].tolist()[0]
highs.append([index, high])
low = sliced["low"].min()
index = sliced.index[sliced['low'] == low].tolist()[0]
lows.append([index, low])
# fill in highs.
max = len(highs)
filled = []
for index in range(max):
if index == 0: # this does fill in future data but the first rows are always dropped past this because of other features so this is not the problem.
for i in range(0, highs[index][0]):
filled.append([i, highs[index][1]])
if index < max-1:
for i in range(highs[index][0], highs[index+1][0]):
filled.append([i, highs[index][1]])
elif index == max-1:
for i in range(highs[index][0], len(df)):
filled.append([i, highs[index][1]])
highs = filled
# fill in lows.
max = len(lows)
filled = []
for index in range(max):
if index == 0: # this does fill in future data but the first rows are always dropped past this because of other features so this is not the problem.
for i in range(0, lows[index][0]):
filled.append([i, lows[index][1]])
if index < max-1:
for i in range(lows[index][0], lows[index+1][0]):
filled.append([i, lows[index][1]])
elif index == max-1:
for i in range(lows[index][0], len(df)):
filled.append([i, lows[index][1]])
lows = filled
# fill support & resistance into df.
for index, high in highs:
df.at[index, "resistance"] = high
for index, low in lows:
df.at[index, "support"] = low
return df[["support", "resistance"]]
As a candlestick graph it would look like this. The purple points are the new highs and lows.
Does anyone see some kind of error which could make the feature use future data that is not available in real-life and therefore could explain the differences between the evaluations?
