How to achieve Tiered Time series storage , last hour in memory, rest on disk with MongoDB. any alternatives

Viewed 20

I am working on building a video analytics platform. We are using deepstream to process some videos and push frame-by-frame metadata and deep learning model outputs to a Kafka broker. We need to push this data to a database - MongoDB. We are then planning to use pandas to load from the DB and perform analysis of data in certain time windows.

The problem is that we may be processing about 100 cameras at 30fps each, with upto 20 objects of interest in the frame. That will make it 60k messages/second. We will use a kafka connector for MongoDB. From mongodb, there are going to be many services that will query recent data so I am looking for a way to keep last 1 hour data in memory for fast querying.

Is there a possibility to achieve this?

I want tiered time series storage. Past hour's data needs to be in-memory for quick querying and lookup. everything else can be in disk. It has to be nosql as data format and entries will keep changing based on objects present in the frame. It can't be a seperate DB as I wan't to keep this detail behind a common interface. The querying scripts shouldn't know where data is coming from - memory or disk.

I looked into sharding but it looks like I can't get the shard key to keep rolling with current_time - 1 hour. I also checked TTL and expiryAfterSeconds but they are only meant to delete and not possible to do custom callbacks.

I will also be interested to know if there are any alternatives

0 Answers
Related