Extract files from MongoDB

Viewed 934

Office documents (Word, Excel, PDF) have been uploaded to a website over the last 10 year. The website does not have a way to download all files, only individual files one at a time. This would take days to complete so I got in touch with the website and asked them to provide all the files. They provided a Mongo database dump that included several JSON and BSON files and they stated this was the only way they could provide the files.

I would like to extract the original office documents from the BSON file to my Windows computer, keeping the folder structure and metadata (when the file was created, etc.), if possible.

I have installed a local version of Mongo on my Windows 10 computer and imported the JSON and BSON files. Using MongoDB Compass, I can see these files have been imported as collections including a 2.73GB fs.chunks.bson file that I am assuming contains the office documents. I have Googled what the next step should be, but I am unsure how to proceed. Any help would be appreciated.

2 Answers

What you need to do is restore the dumps into your database, this can be done using the mongorestore command, some GUI interfaces like robo3T can also provide a way to do it. Make sure your mongo version is the same as the website Mongo's version, otherwise you are risking data corruption which would be a pain to handle.

Now let's talk about Mongo's file system GridFS, it has 2 collections: fs.files collection contains file metadata while the fs.chunks contains the actual file data. In theory every file will have multiple chunks, this storing method purpose was to make streaming data more efficient.

To actually read the file from GridFS you'll have to fetch eacg of the file documents from fs.files collection first, then fetch the matching chunks from fs.chunks collection for each of them. Once you fetched all the chunks you can "create" your file and do whatever you want with it.

Here is a sudo sample of what needs to be done:

files = db.fs.files.find({});

... for each file ....
chunks = db.fs.chunks.find( { files_id: file._id } ).sort( { n: 1 } )
data = chuncks[0].data + .... + chunks[n].data;

...
do whatever you want with the data. remember to check the file type from the file metadata, different types will require different actions.

I had to do something similar.

First I restored the files and chunks BSON backups into my MongoDB.

mongorestore -d db_name -c fs.chunks chunks.bson
mongorestore -d db_name -c fs.files files.bson

(note that you need to replace db_name with your database name)

This was enough for GridFS to function.

Next I wrote up a script to extract the files out of the database. I used PHP to do this as it was already setup where I was working. Note that I had to install the MongoDB driver and library (with composer). If you are on Windows it is easy to install the drive, you just need to download the dll from here and place it in the php/ext folder. Then add the following to the php.ini:

extension=mongodb

Below is a simple version of a script that will dump all files, it can be easily extended to customise folders, prevent overlapping names etc.

include('vendor/autoload.php');
$client = new MongoDB\Client("mongodb://localhost:27017");

$bucket = $client->local->selectGridFSBucket();
$files = $bucket->find();

foreach($files as $file){
    $fileId = $file['_id'];
    $filename = explode('.',$file['filename']);
    $ext = $filename[1];
    $filename = $filename[0];

    $output = fopen('files/'.$filename.".".$ext, 'wb');

    $bucket->downloadToStream($fileId, $output);
}
Related