I have a requirement to read 40 million records from database, process them in parallel(by making REST api calls) and write the status back to the database. At any point in time, I can only make 4000 parallel calls to REST api (as the REST api cannot scale up more than 4000 tps), i.e. Read 4000 records, make 4000 REST api calls, write their status back to DB and fetch next 4000 and so on.
I am considering two options:
Spring batch (remote partitioning)on AWS fargate
2 different spring batch modules. One module (Driver) calculates the total number of records in the client table and intelligently updates another table with the total number of records each Worker instance will be processing. If there are 10 worker instances then Worker Instance 1 processes Record 1 through 4 million. W2 processes 4 million 1 through 8 million etc. The workers will keep processing in batches of 400 records from each task to maintain the throttling limit. The driver can poll the DB to see if any record is still not processed or errored out and marks the batch job complete.
Depending on the answers to following questions I can decide which architecture to follow:
I have a requirement to replay the whole batch including partitoning the whole dataset, thrice. This is to make sure that whatever records retuned with error status during the initial run (REST API failure/timeout) will be sent for re-processing again. Is this possible with Remote partitioning?
In remote partitioning how is a worker instance removing the entry from the queue upon successful completion of the job? Is it handled by Spring internally?
If a worker task comes down by any chance will all the records be picked up and reprocessed again? i.e. will the records again be sent to the REST APIS even through some of them were sent the first time and a response was received. But the task crashed mid process?
How can we stop a Spring batch running on Cloud.
How can we restart a Spring batch running on cloud if we had to stop it mid-process?