When I run a spark job on EMR via Airflow I can specify the log directory on s3 in the spark configuration file. But the logs are then written into a directory named after the cluster ID. The same applies to the logs of the individual steps.
This results in a path like this:
s3://<bucket>/<my_logging_dir>/j-4I93KLS332Ps/steps/s-48503RJF32K/...
If possible, I would like to tell the EMR cluster to save the logs with custom prefixes to have something like this:
s3://<bucket>/<my_logging_dir>/my-spark-cluster-2022-03-07-15-34/steps/job1/...
How could I achieve this?