Should I use ALLOW FILTERING in Cassandra to delete associated entities in a multi-tenant app?

Viewed 51

I have a spring-boot project and I am using Cassandra as database. My application is a tenant based application and all my tables include the tenantId. It is always part of the partition key of all tables but I have also other columns which are part of the partition keys.

So, the problem is; I want to remove a specific tenant from my database but I can't do it directly. Because I need the other part of the partition key.

I have two solutions for it in mind.

  1. I will allow filtering and select all the tenant specific entities and then remove them one by one in the application.
  2. I will use the findAll() method and fetch all the data and then filter in the application and delete all the tenant specific data.

Example:

        public class DeleteTenant{

         @Autowired MyRepository myRepo;
      
         public void cleanTenantWithoutDbFiltering(String tenantId){
         myRepo.findAll()
          .stream()
          .filter(entity -> entity.getTenantId().equals(tenantId)) // ??
          .forEach(MyRepository::remove);
         }

         public void cleanTenantWithDbFiltering(String tenantId){
         myRepo.getTenantSpecificData(tenantId)
          .forEach(MyRepository::remove);
        }
    }

My getTenantSpecificData(String tenantId) query would look like that:

@AllowFiltering
@Query("Select * from myTable where tenantId = ?1 ALLOW FILTERING")
public List<MyEntity> getTenantSpecificData(String tenantId);

Do you have any other idea about it? If not which one do you think would be more efficient? Filtering in the application itself or in the cassandra.

Thanks in advance for your answers!

1 Answers

It isn't clear to me how you've modelled your data because you haven't provided examples of your schema but in any case, the use of ALLOW FILTERING is never going to be a good idea because it means that your query has to do a full table scan of all the relevant tables unless the tenant ID is the partition key.

You will need to come up with a different approach such as writing a Spark app that will efficiently go through the tables to identify partitions/rows to delete. Cheers!

Related