How to resume parsing a large JSON stream at a specific point in Node.js with stream-json package?

Viewed 270

I'm so happy to have been shown how to actually parse JSON objects out of an extremely large JSON file using the stream-json module. It was a pretty simple and straightforward built-in method it turns out!

However, looking at the Wiki of stream-json, I don't see a way to resume processing at a specific object index. How can I do that? Or more generally if there's no quick way to do that, how can I resume at a specific place in the plain JSON stream or just plain gzipped string stream?

For background, I am parsing the Wikidata JSON dump which is about 90GB. I was doing an initial round of parsing through the dump items to collect the classes of each entity/item, but then my parser crashed because of some unexpected data structure 760,000 items into the process!!! Who knows how many items there actually are haha. But I have restarted the process, this time with a try catch to skip over bad data. I forgot to time it initially, but timing it this time around it is taking on average 4 seconds per 1000 items to parse this JSON stream. So 4s * 760 = 3040s is about 50 minutes to get back to where I left off. That is a long time to wait, especially since I am going to have to do 5 or 10 rounds of parsing as I uncover new things I want to retrieve from this data. So I would like to know how to jump to the point where it crashes or where I left off, if at all possible.

How can I do that with stream-json? Perhaps there is a way to rework the StreamArray code to support jumping to a specific item. Or perhaps I can somehow associate the raw wikidata/latest-all.json.gz gzip string stream to the stream-json, so I can point to some location in the latest-all.json.gz string where the fs.createReadStream should start reading from? Something like that.

Would love to know how to resume at some specific place. How can I do that in this case?

My current code is essentially this:

const fs = require('fs')
const zlib = require('zlib')
const { parser } = require('stream-json')
const { streamArray } = require('stream-json/streamers/StreamArray')

let input = fs.createReadStream(`./wikidata/latest-all.json.gz`)

input
  .pipe(zlib.createGunzip())
  .pipe(parser())
  .pipe(streamArray())
  .on('data', d => handle(d.value))

function handle(object) {
  // each entity in the json stream
}

For context on where I'm coming from, I am thinking along the same lines as resuming a wget download with the -c option. If your internet cuts out with wget downloading a 100GB file, or a set of 10,000 files, you can just do wget -c <url> essentially and continue where you left off, rather than having to re-download everything. I would like to do that equivalent sort of thing with the JSON stream parser. To skip all the entities I have already parsed before it crashed or the program exited.

If there is no straightforward way to do this, what sort of hack could I use?

Actually, the first few thousand items took 4 seconds each, now it looks like they are taking 0.5-1s each as we move into the later data in the stream, so after 30 minutes I am at item 1,271,000. Slightly better, but still a long wait.

0 Answers
Related