I am running pyspark locally on my mac 16GB with 6 cores and I am reading a json file and converting it to parquet.I am receiving this error only for one file, rest of files pyspark is able to read and write to s3.
Error:
org.apache.hadoop.fs.s3a.auth.NoAuthWithAWSException: Credentials requested after provider list was closed
It is a big file 500 mb and I have repartitioned it to provide parallelism
try:
investment_df=readRawJson(spark, config, "investments",schema=schema.Table_DDL.investments_ddl())
investment_df=investment_df.repartition(1000)
writeParquet(investment_df, config, "investments")
except (pyspark.sql.utils.CapturedException, TypeError) as error:
print("Exception found in raw to processed file at investments function, " + str(error))
raise
command used to run pyspark:
spark-submit --conf spark.network.timeout=800 --conf spark.driver.memoryOverhead=10g --packages com.amazonaws:aws-java-sdk:1.11.901,org.apache.hadoop:hadoop-aws:3.3.1,net.snowflake:snowflake-jdbc:3.13.10,net.snowflake:spark-snowflake_2.12:2.9.2-spark_3.1 main.py