Ingest Blockchain Data into AWS Pipeline

Viewed 41

I need to ingest data from Solana blockchain (~100M blocks) into an AWS pipeline in order to do some sort of blockchain analysis and data extraction. This data ingestion process will have 2 phases:

  • Catching up phase - Data ingestion needs fetch data from the block 0 all the way to the latest blocks
  • New blocks ingestion phase - Newly added blocks need to be ingested periodically

The solution that I came up with is as follows:

I first set up a periodic lambda with concurrency 1 which finds ranges of ~500k valid block numbers per lambda invocation, starting from block 0. This lambda sends messages which contain these block numbers to SQS. In order to accomplish this, I have to keep track of the highest block number in between lambda invocations.

Another lambda, with concurrency N, takes items in batches of ~450 block numbers from the SQS and sends a HTTP request to get the actual data from these 450 blocks from the blockchain. It only takes one HTTP request to accomplish this since Solana allows for batched RPC requests. This block data is then sent to another SQS queue where it will be picked up by workers for data extraction.

Since I'm fairly new to the AWS ecosystem, I'm wondering if this is the right approach to do this type of thing. I'm especially concerned with the first lambda, since keeping state between lambdas seems kinda hacky to me. But then again, lambdas are preferred to EC2 since it's easier to maintain them.

0 Answers
Related