Error loading spark sql context for redshift jdbc url in glue

Viewed 118

Hello I am trying to fetch month-wise data from a bunch of heavy redshift table(s) in glue job.

As far as I know glue documentation on this is very limited. The query works fine in SQL Workbench which I have connected using the same jdbc connection being used in glue 'myjdbc_url'.

Below is what I have tried and seeing error -

from pyspark.context import SparkContext
sc = SparkContext()
sql_context = SQLContext(sc)
df1 = sql_context.read \
            .format("jdbc") \
            .option("url", myjdbc_url) \
            .option("query", mnth_query) \
            .option("forward_spark_s3_credentials","true") \
            .option("tempdir", "s3://my-bucket/sprk") \
            .load()
print("Total recs for month :"+str(mnthval)+" df1 -> "+str(df1.count()))

However it shows me driver error in the logs as below -

: java.sql.SQLException: No suitable driver at java.sql.DriverManager.getDriver(DriverManager.java:315) at org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions$$anonfun$6.apply(JDBCOptions.scala:105) at org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions$$anonfun$6.apply(JDBCOptions.scala:105) at scala.Option.getOrElse(Option.scala:121)

I have used following too but to no avail. Ends up in Connection refused error.

sql_context.read \
                .format("com.databricks.spark.redshift")
                .option("url", myjdbc_url) \
                .option("query", mnth_query) \
                .option("forward_spark_s3_credentials","true") \
                .option("tempdir", "s3://my-bucket/sprk") \
                .load()

What is the correct driver to use. As I am using glue which is a managed service with transient cluster in the background. Not sure what am I missing. Please help what is the right driver?

0 Answers
Related