How to use dask.dataframe to read sql query from oracle?

Viewed 38

I am using sqlalchemy to fetch records, but there are over millions of rows, I am trying to use dask.dataframe:

I am using following codes:

import pandas as pd
import sqlalchemy
import time
import dask.dataframe as dd

start_time = time.time()

conn = "oracle+cx_oracle://{}:{}@{}:{}/?service_name={}".format(username, password, hostname, port, service_name)
engine = sqlalchemy.create_engine(conn)
query = """SELECT * FROM {table} WHERE {condition}"""

ddf = dd.from_pandas(pd.read_sql(query, engine), npartitions=10)

print("--- %s seconds ---" % (time.time() - start_time))

Since it use pandas.read_sql() first then transferred to dask.dataframe, how to use directly dask.dataframe.read_sql_query() or dask.dataframe.read_sql_table() to achieve this? I guess it can save some time for large-scale query.

0 Answers
Related