My program generates large amount time-series data into the following table:
CREATE TABLE AccountData
(
PartitionKey text,
RowKey text,
AccountId uuid,
UnitId uuid,
ContractId uuid,
Id uuid,
LocationId uuid,
ValuesJson text,
PRIMARY KEY (PartitionKey, RowKey)
)
WITH CLUSTERING ORDER BY (RowKey ASC)
The PartitionKey is a dictionary value (one of 10) and the RowKey is DateTime converted to long.
Now due to the crazy amount of data that is being generated by the program, every ContractId has a different retention policy in the code. The code goes and deletes old data based on the retention for the specific ContractId.
I am now running into problems where during a SELECT statement it picks up too many Tombstones and I get an error.
What Table Compaction strategy should I use to solve this Tombstone problem?