So, the objective I am chasing here is to do Time Series Forecasting using Tensorflow. My data is structured like this:
| timestamp | float64 | numerical |
|---|---|---|
| feature a | float64 | numerical |
| feature b | float64 | numerical |
| feature c | string | categorical |
| target | string | categorical |
The issue I am facing is to properly input the categorical data columns I have.
Currently, my pipeline looks roughly like this:
- Split the dataframe into training, validation and test
n = len(data)
train_df = data[0:int(n*0.7)]
val_df = data[int(n*0.7):int(n*0.9)]
test_df = data[int(n*0.9):]
- Use
tf.keras.preprocessing.timeseries_dataset_from_arrayto preprocess the data
data = np.array(data, dtype=np.float32)
ds = tf.keras.preprocessing.timeseries_dataset_from_array(
data=data,
targets=None,
sequence_length=self.total_window_size,
sequence_stride=1,
shuffle=False,
batch_size=32,)
- Split the resulting
tf.datasetinto windows (closely following the tensorflow documentation for time series forecasting) - I have this simple model
lstm_model = tf.keras.models.Sequential([
tf.keras.layers.LSTM(32, return_sequences=True),
tf.keras.layers.Dense(units=1)
])
- Then I compile and fit the model using my training and validation data
This does compile if I map the categorical data columns to int by myself (low cardinality). But that leaves the column in an ordinal state. If I don't do that, step 2 does not work.
Now, I looked into One Hot Encoding but could not figure out a way to implement this into my current pipeline. Maybe I completely misunderstood the process but the problem is that I need a dataframe with only float32 values as input in step 2. Therefore, I would have to encode the data before that step. Every example for One Hot Encoding in Tensorflow I found performed the encoding within the model as an additional layer; not during the preprocessing stage.
I may be completely off, but I cannot figure out a way to fix that. I would be over the moon if somebody could point me towards the right direction.