I'm getting the following series of errors, specifically the one that says "Input path does not exist", in my AWS Glue Job (written in pyspark) that is supposed to crawl a specific bucket location. The problem is that I've specified the folder location (and not individual filenames), and the fact that it is somehow able to determine the actual filenames before showing this error is what is intriguing. The following error is what is gleaned from the CloudWatch logs:
21/11/11 12:36:15 ERROR ProcessLauncher: Error from Python:Traceback (most recent call last): File "/tmp/CloudTrailLogBatchProcessing", line 11, in df = spark.read.option("multiline", "true").json('s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/') File "/opt/amazon/spark/python/lib/pyspark.zip/pyspark/sql/readwriter.py", line 274, in json return self._df(self._jreader.json(self._spark._sc._jvm.PythonUtils.toSeq(path))) File "/opt/amazon/spark/python/lib/py4j-0.10.7-src.zip/py4j/java_gateway.py", line 1257, in call answer, self.gateway_client, self.target_id, self.name) File "/opt/amazon/spark/python/lib/pyspark.zip/pyspark/sql/utils.py", line 63, in deco return f(*a, **kw) File "/opt/amazon/spark/python/lib/py4j-0.10.7-src.zip/py4j/protocol.py", line 328, in get_return_value format(target_id, ".", name), value) py4j.protocol.Py4JJavaError: An error occurred while calling o74.json. : org.apache.hadoop.mapreduce.lib.input.InvalidInputException: Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0000Z_GGxAlmYsJSPcVsAA.json.gz Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0000Z_XaZZwCZ0WbyAXlIh.json.gz Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0005Z_02o2hkXQbNMs8H1c.json.gz Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0005Z_Nl4Z2awclBvau922.json.gz Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0010Z_3imeLnOAlakY8Exn.json.gz Input path does not exist: s3:///AWSLogs//CloudTrail/us-east-2/2021/11/11/_CloudTrail_us-east-2_20211111T0010Z_V5l4nU9tIat4zYqg.json.gz
The following is part of the pyspark script saved as a job in AWS Glue:
import boto3
from pyspark.context import SparkContext
from pyspark.sql.session import SparkSession
import pyspark.sql.types as T
spark = SparkSession.builder.master("local[1]").appName("S3Crawler").getOrCreate()
# Read from multiple JSON or JSON.gz files in a directory
df = spark.read.option("multiline", "true").json('s3a://<foldername>/AWSLogs/<accountID>/CloudTrail/us-east-2/2021/11/11/')
print(df)
from pyspark.sql import functions as F
Records = df.withColumn("Records", F.explode(F.col("Records")))