How can I bucketing with s3 using aws glue?

Viewed 894

I tried partitioning and bucketing using AWS Glue on S3. But the bucketing did not work. Only the partitioning did work. How can I bucketing with AWS Glue?

datasink4 = glueContext.write_dynamic_frame.from_options(
    frame = dropnullfields3,
    connection_type = "s3",
    connection_options = {"path": s3_output_full,
                          "partitionKeys": ["PARTITIONKEY"],
                          "bucketColumns": ["ROW_ID"],
                          "numberOfBuckets": 12},
    format = "parquet",
    transformation_ctx = "datasink4")

job.commit()
1 Answers

I think they're not supported yet

My script is using bucketBy function instead; but it'd replace existing data in defined path

df_name, job_df = (str(transform_name), df)
datasink_path = "s3://sink-bucket/job-data/"
writing = job_df.write.format('parquet').mode("append") \
                          .partitionBy('event_day') \
                          .bucketBy(3, 'bucketed_field') \
                          .saveAsTable(df_name, path = datasink_path)
Related