To find the right aggregation level for my data, I have to split the day into frames of different sizes.
Example data:
da = data.frame(timestamp = c("2016-01-24 01:17:37 GMT" ,"2016-01-24 02:09:41 GMT", "2016-01-24 13:34:35 GMT", "2016-01-24 15:17:56 GMT", "2016-01-24 18:14:55 GMT"))
da
timestamp
1 2016-01-24 01:17:37 GMT
2 2016-01-24 02:09:41 GMT
3 2016-01-24 13:34:35 GMT
4 2016-01-24 15:17:56 GMT
5 2016-01-24 18:14:55 GMT
For example, I could start cutting the day in 24 parts. Then 0:00 to 1:00 is part 1, 1:00 to 2:00 is part 2 etc.
da2 = data.frame(timestamp = c("2016-01-24 01:17:37 GMT" ,"2016-01-24 02:09:41 GMT", "2016-01-24 13:34:35 GMT", "2016-01-24 15:17:56 GMT", "2016-01-24 18:14:55 GMT"),
daypart = c(2, 3, 14, 16, 19))
da2
timestamp daypart
1 2016-01-24 01:17:37 GMT 2
2 2016-01-24 02:09:41 GMT 3
3 2016-01-24 13:34:35 GMT 14
4 2016-01-24 15:17:56 GMT 16
5 2016-01-24 18:14:55 GMT 19
Or into 48 parts. Then 0:00 to 0:30 is part 1, 0:30 to 1:00 part 2 etc:
da48 = data.frame(timestamp = c("2016-01-24 01:17:37 GMT" ,"2016-01-24 02:09:41 GMT", "2016-01-24 13:34:35 GMT", "2016-01-24 15:17:56 GMT", "2016-01-24 18:14:55 GMT"),
+ daypart = c(3, 5, 28, 31, 37))
da48
timestamp daypart
1 2016-01-24 01:17:37 GMT 3
2 2016-01-24 02:09:41 GMT 5
3 2016-01-24 13:34:35 GMT 28
4 2016-01-24 15:17:56 GMT 31
5 2016-01-24 18:14:55 GMT 37
I found this post Pos on how to convert time to categorical variable, which has already helpes, but how can I code this in such a way that I only have to change the number of parts I want to cut the day into?