Download entire history of a Wikipedia page

Viewed 854

I'd like to download the entire revision history of a single article on Wikipedia, but am running into a roadblock.

It is very easy to download an entire Wikipedia article, or to grab pieces of its history using the Special:Export URL parameters:

curl -d "" 'https://en.wikipedia.org/w/index.php?title=Special:Export&pages=Stack_Overflow&limit=1000&offset=1' -o "StackOverflow.xml"

And of course I can download the entire site including all versions of every article from here, but that's many terabytes and way more data than I need.

Is there a pre-built method for doing this? (Seems like there must be.)

2 Answers

The example above only gets information about the revisions, not the actual contents themselves. Here's a short python script that downloads the full content and metadata history data of a page into individual json files:

import mwclient
import json
import time

site = mwclient.Site('en.wikipedia.org')
page = site.pages['Wikipedia']

for i, (info, content) in enumerate(zip(page.revisions(), page.revisions(prop='content'))):
    info['timestamp'] = time.strftime("%Y-%m-%dT%H:%M:%S", info['timestamp'])
    print(i, info['timestamp'])
    open("%s.json" % info['timestamp'], "w").write(json.dumps(
        { 'info': info,
            'content': content}, indent=4))
Related