Python RocksDB Combine / Merge DBs

Viewed 99

I find myself in a situation where I have my rather large (100s of gbs) dataset split up in many 'chunks', where each chunk is rocksdb database with the exact same format of keys / values. I would need to make fast queries across the entire dataset, so it would make sense for me to combine these chunks into a single rocksdb database rather than make the same query on chunk 1, chunk 2, ... chunk n and combine the results, as combining would fully utilize the optimized rocksdb methods like multi_get ect. Here is a code example to demonstrate what I mean:

import rocksdb

chunk1 = rocksdb.DB('chunk1.db', rocksdb.Options(create_if_missing=True))
chunk2 = rocksdb.DB('chunk2.db', rocksdb.Options(create_if_missing=True))
chunk3 = rocksdb.DB('chunk3.db', rocksdb.Options(create_if_missing=True))

# Populate dbs here

chunk1.multi_get(keys)
# 1, 2, None, None, None, None
chunk2.multi_get(keys)
# None, None, 3, 4, None, None
chunk3.multi_get(keys)
# None, None, None, None, 5, 6

# I am looking for how to implement this combine method
full_set = combine(chunk1, chunk2, chunk3)

full_set.multi_get(keys)
# 1, 2, 3, 4, 5, 6

Is there a way to implement the combine method above without opening each database chunk and then reading every line and writing it to combined database? The reason I don't want to just fully reread and recreate the combined database is that the chunks are populated dynamically onto my machine and can change every few days, so frequently waiting for a long combine operation would not be ideal. I have attempted the solution discussed here, however I was not able to get a combined database after running repair, as it would discard all the new .sst files and only keep the original set.

Alternatively, if anyone knows any other database similar to rocksdb that easily achieves this, I am also open to suggestions.

0 Answers
Related