How to extract data points in dynamic graphs to text file? (Datawrapper/html/json)

Viewed 102

I want to extract the points in a text file for the linked graph (datawrapper) that appear when you move the cursor over the various plotted lines (different days, different proportions).

This graph (looked at the code with "view-source:")

So far, I have only extracted single tables, for example, with BeautifulSoup or in R on simple HTML pages. In this case, I do not know what the most elegant approach would be.

What are common ways or solutions to solve this problem that I can dig into? Since I want to extract several such plots as tables, a solution that automates this would be desirable.

Thanks for your suggestions

1 Answers

I'm a little late on this, but I hope this might be helpful. The graphics are generated by javascript--we can dig into the scripts and find the data that they use in JSON format.

How I've solved this is to, using python:

  1. parse the webpage with BeautifulSoup:
url = "https://datawrapper.dwcdn.net/RE9Rq/1/"
raw = requests.get(url)
soup = bs(raw.text, "lxml")
  1. Locate the script that has all the data that we want, and remove some unnecessary characters:
raw_data = soup.find_all("script")[1]
string_data = str(raw_data).replace("\\", "")
  1. Nested inside of this mess of a string is the JSON that contains our data. So, we can use regex to find and extract it:
raw_json = (
    "{" + re.findall(r"\"data\":\{.*?\]", string_data, flags=re.MULTILINE)[1] + "}}"
)
  1. This is still technically in string form, so we need to parse it using the json library:
data = json.loads(raw_json)["data"]

Now, you should have all the data, so you can look through it, find whatever pieces you want, and transform it into a dataframe (pandas.DataFrame.from_records() is quite helpful)


full code:

import requests
import re
import json
from bs4 import BeautifulSoup as bs

url = "https://datawrapper.dwcdn.net/RE9Rq/1/"
raw = requests.get(url)
soup = bs(raw.text, "lxml")
raw_data = soup.find_all("script")[1]
string_data = str(string_data).replace("\\", "")
raw_json = (
    "{" + re.findall(r"\"data\":\{.*?\]", raw_data, flags=re.MULTILINE)[1] + "}}"
)

data = json.loads(raw_json)["data"]
Related