We have data stored in s3 partitioned in the following structure:
bucket/directory/table/aaaa/bb/cc/dd/
where aaaa is the year, bb is the month, cc is the day and dd is the hour.
As you can see, there are no partition keys in the path (year=aaaa, month=bb, day=cc, hour=dd).
As a result, when I read the table into Spark, there is no year, month, day or hour columns.
Is there anyway I can read the table into Spark and include the partitioned column without:
- changing the path names in s3
- iterating over each partition value in a loop and reading each partition one by one into Spark (it is a huge table and this takes far too long and is obviously sub-optimal).