I'm so happy to have been shown how to actually parse JSON objects out of an extremely large JSON file using the stream-json module. It was a pretty simple and straightforward built-in method it turns out!
However, looking at the Wiki of stream-json, I don't see a way to resume processing at a specific object index. How can I do that? Or more generally if there's no quick way to do that, how can I resume at a specific place in the plain JSON stream or just plain gzipped string stream?
For background, I am parsing the Wikidata JSON dump which is about 90GB. I was doing an initial round of parsing through the dump items to collect the classes of each entity/item, but then my parser crashed because of some unexpected data structure 760,000 items into the process!!! Who knows how many items there actually are haha. But I have restarted the process, this time with a try catch to skip over bad data. I forgot to time it initially, but timing it this time around it is taking on average 4 seconds per 1000 items to parse this JSON stream. So 4s * 760 = 3040s is about 50 minutes to get back to where I left off. That is a long time to wait, especially since I am going to have to do 5 or 10 rounds of parsing as I uncover new things I want to retrieve from this data. So I would like to know how to jump to the point where it crashes or where I left off, if at all possible.
How can I do that with stream-json? Perhaps there is a way to rework the StreamArray code to support jumping to a specific item. Or perhaps I can somehow associate the raw wikidata/latest-all.json.gz gzip string stream to the stream-json, so I can point to some location in the latest-all.json.gz string where the fs.createReadStream should start reading from? Something like that.
Would love to know how to resume at some specific place. How can I do that in this case?
My current code is essentially this:
const fs = require('fs')
const zlib = require('zlib')
const { parser } = require('stream-json')
const { streamArray } = require('stream-json/streamers/StreamArray')
let input = fs.createReadStream(`./wikidata/latest-all.json.gz`)
input
.pipe(zlib.createGunzip())
.pipe(parser())
.pipe(streamArray())
.on('data', d => handle(d.value))
function handle(object) {
// each entity in the json stream
}
For context on where I'm coming from, I am thinking along the same lines as resuming a wget download with the -c option. If your internet cuts out with wget downloading a 100GB file, or a set of 10,000 files, you can just do wget -c <url> essentially and continue where you left off, rather than having to re-download everything. I would like to do that equivalent sort of thing with the JSON stream parser. To skip all the entities I have already parsed before it crashed or the program exited.
If there is no straightforward way to do this, what sort of hack could I use?
Actually, the first few thousand items took 4 seconds each, now it looks like they are taking 0.5-1s each as we move into the later data in the stream, so after 30 minutes I am at item 1,271,000. Slightly better, but still a long wait.