I have a huge amount of data stored in parquet files with a timestep frequency of 5 minutes. I would like to resample in hourly frequency. I have the query doing that in sql and I use the integration between Arrow and Duckdb.
The problem is the following I am exporting the resampled data in parquet files and when I re-import them in a arrow dataset I have the old problem of timezone in arrow-duckdb.
How to be sure to not export timezone in duckdb? or equivalently how to set
Set TimeZone = NULL;
in duckdb?
I receive Not implemented Error: Unknown TimeZone setting.
This is the code importing and exporting parquet files.
import duckdb
import pyarrow as pa
import pyarrow.dataset as ds
import glob
# Open dataset
db = ds.dataset('database_source/filtered_parquet/')
# Create duckdb connection and cursor
con = duckdb.connect()
c = con.cursor()
for file in glob.glob("database_source/PF-*"): # directories stored by data
day = file[26:36] # e.g. 2021-01-01
c.execute(f""" COPY
(SELECT date_trunc('hour', casedate) as timeindex,
country,
CAST(nominalv as USMALLINT) as nominalv,
MEDIAN(value) as value
FROM db
WHERE date_trunc('day', casedate) = '{day}'
GROUP BY timeindex,
country,
nominalv)
TO 'out/resampled/{day}.parquet';""").fetchall()
Files are correctly exported, then I import again my arrow dataset as
new_db = ds.dataset('out/resampled/')
# Create duckdb connection and cursor
con = duckdb.connect()
c = con.cursor()
c.execute("""SELECT count(value) from new_db""").fetchall()
and I got RuntimeError: Not implemented Error: Unsupported Internal Arrow Type tsu:UTC