Spark on Kubernetes (non EMR) with AWS Glue Data Catalog

Viewed 431

I am running spark jobs on EKS and these jobs are submitted from Jupyter notebooks.

We have all our tables in an S3 bucket and their metadata sits in Glue Data Catalog.

I want to use the Glue Data Catalog as the Hive metastore for these Spark jobs. I see that it's possible to do when Spark is run in EMR: https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-hive-metastore-glue.html

but is it possible from Spark running on EKS?

I have seen this code released by aws: https://github.com/awslabs/aws-glue-data-catalog-client-for-apache-hive-metastore but I can't understand if patching of the Hive jar is necessary for what I'm trying to do. Also I need the hive-site.xml file for connecting Spark to the metastore, how can I get this file from Glue Data Catalog?

1 Answers

I found a solution for that.

I created a new spark image with this instructions: https://github.com/viaduct-ai/docker-spark-k8s-aws

and finally at my job yaml file, I added some configurations

sparkConf:
   ...
    spark.hadoop.fs.s3a.impl: "org.apache.hadoop.fs.s3a.S3AFileSystem"
    spark.hadoop.fs.s3.impl: "org.apache.hadoop.fs.s3a.S3AFileSystem"
Related