Spark JDBC connection to Google Cloud Spanner failure

Viewed 72

I am trying to read data from Google Cloud Spanner using PySpark JDBC connection. My Spark application is running on a Dataproc cluster. I am using the official Google Cloud Spanner JDBC driver as found here. Below is a snippet of the PySpark code:

project = <<PROJECT_ID>>
instance = <<INSTANCE_ID>>
databases = <<DATABASE_ID>>
spanner_connection_url = 'jdbc:cloudspanner:/projects/' + project + '/instances/' + instance + '/databases/' + databases
df = spark.read \
    .format("jdbc") \
    .option("url", spanner_connection_url) \
    .option("driver", "com.google.cloud.spanner.jdbc.JdbcDriver") \
    .option("dbtable", "test_employee") \
    .load()

However, my Spark job fails with below error: enter image description here

This is usually the standard way of setting up JDBC connection from Spark, or for Spanner is there something that needs to be done differently.

2 Answers

The error seems to indicate that the JDBC driver cannot find a class in one of its dependencies. Could it be that you are only adding the .jar file of the JDBC driver, but not any of its dependencies?

Or put another way: How are you making sure that not only the JDBC driver, but also all its dependencies are being added to the classpath?

The missing class is from com.google.api:gax-grpc https://mvnrepository.com/artifact/com.google.api/gax-grpc. I downloaded and checked google-cloud-spanner-jdbc-2.7.6.jar from https://search.maven.org/artifact/com.google.cloud/google-cloud-spanner-jdbc/2.7.6/jar, but it doesn't seem to include dependencies, only Spanner JDBC classes are found. So you might need to add the missing dependencies.

$ jar tf google-cloud-spanner-jdbc-2.7.6.jar

...
com/google/cloud/spanner/jdbc/JdbcClob.class
com/google/cloud/spanner/jdbc/JdbcDataSource.class
com/google/cloud/spanner/jdbc/JdbcBlob$1.class
...
Related