Why data is getting duplicated while migrating data from Mysql to S3 using multiple trigger for same AWS Lambda function

Viewed 57

I am using Springboot in AWS Lambda to migrate data from MySQL DB to S3 on a regular interval. My lambda calls the functions apply() which I set in environment variable FUNCTION_NAME in AWS lambda console.

It is working fine when I invoke my lambda with one trigger of cloudwatch event Scheduler.

But to achieve concurrency and complete the task faster as there will be millions of records on prod, I added three trigger at the same time but due to this there is a problem that the same data is moved to s3 three times, once by each lambda.

Although I have annotated my apply() method with @Transactional. Below is the overridden apply() method of Function<T, R> interface of java.util.function.

@Transactional
@Override
public String apply(Map map) {
Supplier<Stream<Object>> streamSupplier = () -> repo.getAllRecords();
        try (Stream<Object> responseStream = streamSupplier.get()) {
            List<Object> dataList = new LinkedList<>();
            responseStream.forEach(item -> {
                if(dataList.size() != 30000) {
                    dataList.add(item);
                } else {
                    uploadToS3(dataList);
                    dataList.clear();
                }
                entityManager.detach(item);
            });

       if(dataList.size() > 0) {
            uploadToS3(dataList);
            dataList.clear();
        }
    }

This is how I am fetching result from DB

@Transactional
@QueryHints(value = {
            @QueryHint(name = HINT_FETCH_SIZE, value = "10000"),
            @QueryHint(name = HINT_CACHEABLE, value = "false"),
            @QueryHint(name = READ_ONLY, value = "true")
    })
    @Query(value = "select * from table_x",nativeQuery = true)
    public Stream<Object> getAllRecords();

I am new in this field. Please help me in understanding why the data is migrated three times even using Transactional?

0 Answers
Related