How to read from one s3 account and write to another in Databricks

Viewed 273

My cluster is configured to use ROLE_B which gives me access to BUCKET_IN_ACCOUNT_B but not to BUCKET_IN_ACCOUNT_A. So I assume XACCOUNT_ROLE to access BUCKET_IN_ACCOUNT_A. The following code works just fine.

sc._jsc.hadoopConfiguration().set("fs.s3a.credentialsType", "AssumeRole")
sc._jsc.hadoopConfiguration().set("fs.s3a.stsAssumeRole.arn", XACCOUNT_ROLE)
sc._jsc.hadoopConfiguration().set("fs.s3a.acl.default", "BucketOwnerFullControl")

df = spark.read.option("multiline","true").json(BUCKET_IN_ACCOUNT_A)

But, when I try to write this dataframe back to BUCKET_IN_ACCOUNT_B like below, I get java.nio.file.AccessDeniedException.

df.write \
  .format("delta") \
  .mode("append") \
  .save(BUCKET_IN_ACCOUNT_B)

I assume this is the case because my spark cluster is still configured to use XACCOUNT_ROLE. My question is, how do I switch back to ROLE_B?

sc._jsc.hadoopConfiguration().set("fs.s3a.credentialsType", "AssumeRole")
sc._jsc.hadoopConfiguration().set("fs.s3a.stsAssumeRole.arn", ROLE_B)
sc._jsc.hadoopConfiguration().set("fs.s3a.acl.default", "BucketOwnerFullControl")

did not work.

1 Answers

I never ended up discovering a way to switch the cluster role again. But I did figure out an alternative way of accomplishing reading from one s3 bucket and writing to another. Basically, I mounted the XACCOUNT bucket and was able to write to BUCKET_IN_ACCOUNT_B without having to interfere with the cluster's role.

dbutils.fs.unmount("/mnt/MOUNT_LOCATION")
dbutils.fs.mount(BUCKET_IN_ACCOUNT_A, "/mnt/MOUNT_LOCATION",
  extra_configs = {
    "fs.s3a.credentialsType": "AssumeRole",
    "fs.s3a.stsAssumeRole.arn": XACCOUNT_ROLE,
    "fs.s3a.acl.default": "BucketOwnerFullControl"
  }
)

df = spark.read.option("multiline","true").json("/mnt/MOUNT_LOCATION")
df.write \
  .format("delta") \
  .mode("append") \
  .save(BUCKET_IN_ACCOUNT_B)
Related