Python, how to cluster multivariate time series?

Viewed 931

I have a toy time series dataframe in this format:

>>> df
                 dtime dev        sw1        sw2
0  2020-01-01 00:00:00  A1   5.496714   5.593792
1  2020-01-01 00:15:00  A1   5.417291   6.385936
2  2020-01-01 00:30:00  A1   6.758800   6.056747
3  2020-01-01 00:45:00  A1   8.189697   7.862034
4  2020-01-01 01:00:00  A1   6.988069   6.595961
5  2020-01-01 01:15:00  A1   7.543641   6.080126
6  2020-01-01 01:30:00  A1   9.912546  10.208666
7  2020-01-01 01:45:00  A1   9.656324   9.917379
8  2020-01-01 02:00:00  A1   8.974970   8.980084
9  2020-01-01 02:15:00  A1  10.542560  10.307973
0  2020-01-01 00:00:00  B1   4.536582   3.121212
1  2020-01-01 00:15:00  B1   5.089826   4.669180
2  2020-01-01 00:30:00  B1   6.353073   6.010359
3  2020-01-01 00:45:00  B1   4.753386   3.951109
4  2020-01-01 01:00:00  B1   5.497304   5.336019
..                 ...  ..        ...        ...
5  2020-01-01 01:15:00  H3   3.044125   3.456906
6  2020-01-01 01:30:00  H3   1.753714   2.575774
7  2020-01-01 01:45:00  H3   0.812104   2.708897
8  2020-01-01 02:00:00  H3   0.647316   0.401928
9  2020-01-01 02:15:00  H3  -1.987569  -2.741305
0  2020-01-01 00:00:00  I3   4.780328   3.890814
1  2020-01-01 00:15:00  I3   4.801557   3.985747
2  2020-01-01 00:30:00  I3   5.366783   5.289681
3  2020-01-01 00:45:00  I3   2.815063   3.156215
4  2020-01-01 01:00:00  I3   1.969284   2.245975
5  2020-01-01 01:15:00  I3   1.720465   2.547648
6  2020-01-01 01:30:00  I3   2.582069   2.595071
7  2020-01-01 01:45:00  I3   1.439862   2.893396
8  2020-01-01 02:00:00  I3   0.025795  -0.238861
9  2020-01-01 02:15:00  I3   0.513267   3.233437

[90 rows x 4 columns]

Each row presents d(ate)time, device, and position of device's two switches. I have to cluster this time series by the position of the switches. So, similarly set switches (with minimal distance) should form a cluster.

Dataframe is created in a way that devices A1, B1, and C1 should form cluster 1, devices D2, E2, F2, should form cluster 2, and devices G3, H3, I3 should form cluster 3.

The code how time series are build is as follows:

import pandas as pd
import numpy as np

nper = 10 # number of periods
dtime_fr = '2020-01-01' # datetime from
freq = '15T'
dtime_range = pd.date_range(dtime_fr, periods=nper, freq=freq) # timestamps creation

df = pd.DataFrame()

# reising baseline
baseline_start = 5
baseline_stop = 10
baseline_linspace = np.linspace(baseline_start, baseline_stop, nper)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'A1', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'B1', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'C1', 'sw1': ts, 'city':'C1'})], sort=False)

# steady baseline
baseline_start = 5
baseline_stop = 5
baseline_linspace = np.linspace(baseline_start, baseline_stop, nper)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'D2', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'E2', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'F2', 'sw1': ts, 'city':'C1'})], sort=False)


# falling baseline
baseline_start = 5
baseline_stop = 0
baseline_linspace = np.linspace(baseline_start, baseline_stop, nper)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'G3', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'H3', 'sw1': ts, 'city':'C1'})], sort=False)
ts = baseline_linspace + np.random.normal(size=nper)
df = pd.concat([df, pd.DataFrame({'dtime' : dtime_range, 'dev': 'I3', 'sw1': ts, 'city':'C1'})], sort=False)

df.insert(3, 'sw2', df.sw1 + np.random.normal(size=len(df)))

df

If I have only one switch (sw1), I would do the clustering in the following way:

df_sw1 = df.pivot(index='dev', values='sw1', columns='dtime')
dist = pdist(df_sw1, metric='euclidean')
Z = scipy.cluster.hierarchy.linkage(dist)

fig, ax = plt.subplots(figsize=(20,15))
scipy.cluster.hierarchy.dendrogram(Z, labels=df_sw1.index, orientation='top');

from sklearn.cluster import AgglomerativeClustering

hclust = AgglomerativeClustering(n_clusters=3)
hclust.fit(df_sw1)

hclust.labels_

array([0, 0, 0, 2, 2, 2, 1, 1, 1], dtype=int64)

enter image description here

In case when I have to consider both switches, if two devices are alike, sw1 will have (almost) the same value as other device's sw1, and at the same time sw2 will have (almost) the same value as other device's sw2.

I am not sure how to do that.

What format dataframe should have in order to calculate pdist? How to do the clustering if I have to consider both switches?

0 Answers
Related