Batch data via repetitive api calls and ingestion to bigquery via apache beam

Viewed 109

I am working on a use case. In this use case, I have to ingest data from a rest endpoint associated to Adobe AEM query builder. Now, because of the size of the data, I will have to pull data in batches. Let me elaborate the last statement. Consider an endpoint x. Let us assume that it returns back more than 200k points if I send the data limit in the query as no limit or to be very specific set the p.limit parameter to -1. Now, to systematically pull data from this endpoint, I will need to use pagination functionality of adobe aem by setting the p.guessTotal and p.limit parameter. This, will help me pull only a specific number of points in a given call. eg. 1000 points, 2000 points, etc. Also, this will not timeout the AEM, as well as, provide all the data. Now, I want to pull this data using Apache Beam with Google Cloud Dataflow runner because of its stability and scalability. If this was a one time pull, the implementation is straightforward. But, in the above scenario, we have a recursive data pull. That's where I am not able to figure out, how to implement this in Apache Beam. I cannot hardcode values in a config file as a dynamic configuration is preferred. I request some guidance int this context. Please let me know, if more elaboration is required. I am happy to explain more. Thank you.

0 Answers
Related