Supposed I have some financial data of minutes as below, I would like to write a user-defined function(below code is ugly and complicated), how do I get the 5-minute/10-minute/30-minute/1 hour/8 hour/24 hours data with rows summary using Python/pandas out of CSV?
TIME OPEN HIGH LOW CLOSE VOLUME
----------------------------------------------
0 1592194620 3046.00 3048.50 3046.00 3047.50 505
1 1592194630 3047.00 3048.00 3046.00 3047.00 162
2 1592194640 3047.50 3048.00 3047.00 3047.50 98
3 1592194650 3047.50 3047.50 3047.00 3047.50 228
4 1592194660 3048.00 3048.00 3047.50 3048.00 136
5 1592194670 3048.00 3048.00 3046.50 3046.50 174
6 1592194680 3046.50 3046.50 3045.00 3045.00 134
7 1592194690 3045.50 3046.00 3044.00 3045.00 43
8 1592194700 3045.00 3045.50 3045.00 3045.00 214
9 1592194710 3045.50 3045.50 3045.50 3045.50 8
10 1592194720 3045.50 3046.00 3044.50 3044.50 152
.......
.......
19999 1591594660 3048.00 3048.00 3047.50 3048.00 136
The sample output as below:
3048.50 2140 2020-06-13 04:34:00
3050.50 67 2020-06-13 04:35:00
3049.50 1489 2020-06-13 04:36:00
3047.50 987 2020-06-13 04:37:00
......
3099.50 2 2020-06-14 04:34:00
Below is my stupid code:
import pandas as pd
import pymysql
conn = pymysql.connect( host = "localhost",
user="root",
passwd="root",
db="demo")
sql = "SELECT TIME, OPEN, HIGH, LOW, CLOSE, VOLUME FROM demo_table;"
df = pd.read_sql(sql, conn)
# 12 hours for 1000 records
for i in range(1000, 20000-1000,1):
high_price = df.loc[i,['high']][0]
df_1000 = df.loc[i-1000:i]
df_high = df_1000[df_1000['high']>high_price]
high_count = df_high.shape[0]
df_last = df_high.tail(1)
time_dt = pd.Timestamp(df_last['TIME'], unit='s')
print(high_price, high_count, time_dt )