Horizontal scaling with Spring Batch

Viewed 780

I have an application which utilizes Spring Batch to read and process records from a common database. The job is triggered at a fixed time from a scheduler and works fine on a single app instance. I would like to horizontally scale this application to improve processing time whilst using the same database. Is there anything within Spring Batch (a semaphore) to manage the data being accessed by multiple instances, so as to prevent them accessing and modifying the same records?

I've done a search and have only managed to find multi-threading within the same app instance.

Many Thanks

6 Answers

You can solve it in architecture level,

By using load balancer, you can separate the requests to chunks and it will process it parallelly by sending these requests to different nodes / instances.

For example: Mysql has 1m records, you can fetch the data in chunks 0-100k .. 900k-1m, send the data to processor micro-service over Ribbon or other load balancer.

Automatically it will send it every time to a different node in order.

Good luck

Is there anything within Spring Batch (a semaphore) to manage the data being accessed by multiple instances, so as to prevent them accessing and modifying the same records?

No, not across instances. It is up to you to make sure each instance works on a distinct data set.

I've done a search and have only managed to find multi-threading within the same app instance.

In addition to multi-threaded steps, Spring Batch also offers partitioned steps where each worker is assigned a distinct partition. Workers could be local threads or remote JVMs. You can create as many remote workers as you want, so this approach allows you to horizontally scale you job.

You can use a separate db table (accessible to all instances) to save the last unique-sort key processed in every page (assuming you are using PagingReader). Then read the highest key from the db table, and use it in your SQL query (like, WHERE value > key) for the page-read, so that when the next page in any instance runs the SQL query, it picks up the records with key-value greater than the ones processed in other pages.

Is there anything within Spring Batch (a semaphore) to manage the data being accessed by multiple instances, so as to prevent them accessing and modifying the same records?

No, not across instances. It is up to you to make sure each instance works on a distinct data set.

With respect to this, I have done POC on distributing dataset across instances by using a distributed lock (used dynamodb lock client).

Approach I have used here is to spin up multiple instances of an application and these instances would fight for the dataset lock (or jobs in spring batch world) that it has to run for at the start-up. With enough monitoring, you can make sure all datasets are assigned to some instance.

More details at my repo : https://github.com/vinodhinic/scale-spring-batch

Related