Mongodb Bulk write Updateone or Updatemany

Viewed 836

I want to know if its faster(importing) using updateone or updatemany with bulk write.My code for importing the data into the collection with pymongo look is this:

for file in sorted_files:
    df = process_file(file)
    for row, item in df.iterrows():
        data_dict = item.to_dict()
        bulk_request.append(UpdateOne(
            {"nsamples": {"$lt": 12}},
            {
                "$push": {"samples": data_dict},
                "$inc": {"nsamples": 1}
            },
            upsert=True
        ))
    result = mycol1.bulk_write(bulk_request)

When i tried update many the only thing i change is this:

...
...
bulk_request.append(UpdateMany(..
..
..

I didnt see any major difference in insertion time.Shouldnt updateMany be way faster? Maybe i am doing something wrong.Any advice would be helpful! Thanks in advance!

Note:My data consist of 1.2m rows .I need each document to contain 12 subdocuments.

2 Answers

updateOne updates only the first matching document for nsamples: {$lt: 12}. So updateOne should be faster.

However, why do you insert them one by one? Put all in one document and make a single update. Similar to this one:

sample_data = [];
for row, item in df.iterrows():
    data_dict = item.to_dict();
    sample_data.append(data_dict);
db.mycol1.updateOne(
  {"nsamples": {"$lt": 12}},
  { 
     "$push": { samples: { $each: sample_data } },
     "$inc": {"nsamples": len(sample_data) }
  }
)

@Wernfried Domscheit's answer is correct.

This answer is specific to your scenario.

If you don't mind not updating records to existing documents and insert new documents altogether, use the below code which is the best optimized for your use case.

sorted_files = []
process_file = None
for file in sorted_files:
    df = process_file(file)
    sample_data = []
    for row, item in df.iterrows():
        sample_data.append(item.to_dict())
        if len(sample_data) == 12:
            mycol1.insertOne({
                "samples": sample_data,
                "nsamples": len(sample_data),
            })
            sample_data = []
    mycol1.insertOne({
        "samples": sample_data,
        "nsamples": len(sample_data),
    })

If you want to fill up your existing records with 12 objects and then, create new records, use the below code logic.

Note: I have not tested the code in my local, its just to understand the flow for you to use.

for file in sorted_files:
    df = process_file(file)
    sample_data = []
    continuity_flag = False
    for row, item in df.iterrows():
        sample_data.append(item.to_dict())
        if not continuity_flag:
            sample_rec = mycol1.find_one({"nsamples": {"$lt": 12}}, {"nsamples": 1})
            if sample_rec is None:
                continuity_flag = True
            elif sample_rec["nsamples"] + len(sample_data) == 12:
                mycol1.update_one({
                    "_id": sample_rec["_id"]
                }, {
                    "$push": {"samples": {"$each": sample_data}},
                    "$inc": {"nsamples": len(sample_data)}
                })
        if len(sample_data) == 12:
            mycol1.insert_one({
                "samples": sample_data,
                "nsamples": len(sample_data),
            })
            sample_data = []
    if sample_data:
        mycol1.insert_one({
            "samples": sample_data,
            "nsamples": len(sample_data),
        })
Related