I have a pandas df which is over 10 million in rows. I'm trying to convert this pandas df to spark df using the below method.
spark_session = SparkSession.builder.appName('pandasToSparkDF').getOrCreate()
# Pandas to Spark
spark_df = spark_session.createDataFrame(pandas_df)
This process is taking ~9 minutes to convert pandas df to spark df of 10 million rows on Databricks. Which is too long.
Is there any other way where I can convert it faster?
Thanks. Appreciate the help.