I have a cluster of 4 raspberries Pi4 8 GB RAM (r1, r2, r3 and r4) and each of them has a SSD attached storage.
The cluster is configured with MongoDB Shard, so, I structured them as follows:
- mongod-config daemon running on r1
- mongod-shard daemon running on r1, r2, r3, r4
- mongos (router) daemon running on r1, r2 and r3
From my laptop I am loading a very large test database into a collection named "testcoll". The "testcoll" is shared among r1, r2, r3 and r4. I configured range shard key as follows:
sh.shardCollection("mydb.testcoll",{_id: 1},{unique:true})
where _id is the ranged key, since I am uploading data indicating in my csv, that the key is the first column in _id,col1,col2,col3.
So I have only one index: _id. I loaded first 50M of records very quickly (around 2 hours).
After 100M records loaded I notices it becomes slower (6 hours passed).
Now, at 200M of records seems to be extremely slow to insert new data.
I verified that amount data are almost equally distributed over the four SSDs.
However, queries are extremely quick and they are not affected by the numbers of records loaded into the collection.
Query time remains constant and independent by the numbers f records (this is very good).
Insert time increases exponentially as I insert new records, why?
If needed I can post here logs/informations eventually required in comments.