Apache Iceberg to index AWS S3

Viewed 388

I have a usecase where there are about 100M files stored on S3. I have a manifest file separately for the location of these files based on my data model. I want to understand if Apache Iceberg is a good fit to provide indexing of my S3 files.

Reading through Iceberg Documentation, seems like it talks about creating a table with one column being target S3 location. If this is the case how is it different from a normal relational table where I store the S3 path and other column being relevant to my model.

Any pointers or examples where people have used S3 indexing with Iceberg would be helpful.

1 Answers

This is a bit different than what's going on. What Iceberg does is create a secondary level of metadata separate from the actual table data. This metadata is what actually has the field of "path" for the particular row.

If you take a look at this diagram enter image description here

The Path information is stored in the "manifest file" along with any metrics for that specific file. This means when you perform a read, the query can determine which files to actually read without touching the underlying data files.

For example if you have 3 files with column A, the manifest file will look approximately like

path/to/file1 | Max Value of ColA in File1 | Min Value of ColA in File1
path/to/file2 | Max Value of ColA in File2 | Min Value of ColA in File2
path/to/file3 | Max Value of ColA in File3 | Min Value of ColA in File3

So when I read that file I evaluate any predicates on ColA without touching file1, file2 or file3.

This "metadata" can be read from a "metadata table" in Iceberg but this is just a SQL View of the metadata files for a particular snapshot.

Related