Scrapy Splash, How to deal with onclick?

Viewed 388

I'm trying to scrape the following site

I'm able to receive a response but i don't know how can i access the inner data of the below items in order to scrape it:

I noticed that accessing the items is actually handled by JavaScript and also the pagination.

What should i do in such case?

enter image description here

Below is my code:

import scrapy
from scrapy_splash import SplashRequest


class NmpaSpider(scrapy.Spider):
    name = 'nmpa'
    http_user = 'hidden' # as am using Cloud Splash
    allowed_domains = ['nmpa.gov.cn']

    def start_requests(self):
        yield SplashRequest('http://app1.nmpa.gov.cn/data_nmpa/face3/base.jsp?tableId=27&tableName=TABLE27&title=%E8%BF%9B%E5%8F%A3%E5%8C%BB%E7%96%97%E5%99%A8%E6%A2%B0%E4%BA%A7%E5%93%81%EF%BC%88%E6%B3%A8%E5%86%8C&bcId=152904442584853439006654836900', args={
            'wait': 5}
        )

    def parse(self, response):
        goal = response.xpath("//*[@id='content']//a/@href").getall()
        print(goal)
1 Answers

If you use some breakpoints you'll see its a frustrating job that I explain what I understood from my research. when you are working with this kind of situation you have two ways:

1 - [Easy Way] use selenium and open a browser and click on each link and get the returned contents easily, you can run multiple browsers and get link contents simultaneously.

2 - [Hard Way] simulate what website does (by making similar functions inside python) and do exactly what website does in JS but in the end instead of showing the results, just save it in a variable and use it the way you want.

Now if you choose the HARD WAY this is what I found:

the link JS is like this:

commitForECMA(callbackC,'content.jsp?tableId=26&tableName=TABLE26&tableView=国产医疗器械产品(注册&Id=138150',null)

it calls a function named commitForECMA and get what this function returns and pass it to callBackC function.

well this was obvious, but its important to know what these functions do and how to replicate it.

  • commitForECMA:

this is the function:

function commitForECMA($_8, $_10, $_12) {
    request = createXMLHttp();
    request.onreadystatechange = $_8;
    if ($_12 == null) {
        _$du(request, _$Fe('uM6r2MG'), _$Fe("jp0YV"), $_10);
        request.setRequestHeader(_$Fe("XACeXwDYXwcTV8Ur2"), _$Fe("YwDYgwceLwDT7iCYX3Ce9FKyvHKwPFa"));
    } else {
        var $_9 = "";
        var $_19 = $_12.elements;
        var $_0 = $_19.length;
        for (var $_18 = 0; $_18 < $_0; $_18++) {
            var $_14 = _$3P($_19, $_18);
            if ($_14.type != _$Fe("yQ6YPMK20") && _$3P($_14, _$Fe('uwbm7wKV')) != "") {
                if ($_9.length > 0) {
                    $_9 += "&" + $_14.name + "=" + _$3P($_14, _$Fe('kwbm7wKV'));
                } else {
                    $_9 += $_14.name + "=" + _$3P($_14, _$Fe('Ewbm7wKV'));
                }
                $_9 += _$Fe("Jx2J03Up2Hsl");
            }
        }
        _$du(request, _$Fe('uM6r2MG'), _$Fe("HVlesYq"), $_10);
        $_9 = encodeURI($_9);
        $_9 = encodeURI($_9);
        request.setRequestHeader(_$Fe("d3CmOFDVz3CeXwoxBMq"), _$Fe("FMbZz3CmOFDV"));
        request.setRequestHeader(_$Fe("yACeXwDYXwcTV8Ur2"), _$Fe("g3UraMD2O3UpNMCgB8cT6w6QzRbenM1TTQbS2MbJBRDY9"));
    }
    request.send($_9);
    if ($_12 != null) {
        $_12.reset();
    }
}

yes as you can see it just creates a XMLHTTP request which (for the links in question) Posts the $_10 content to the server and get the results in callBackC function which is now in $_8. but the trick here is the $_10 contents goes through ~13000 lines of code to create links like this:

http://app1.nmpa.gov.cn/data_nmpa/face3/content.jsp?6SQk6G2z=GBK-56.it.xmhx8IaDT25ZyaSxljrwULe8AkNw8QjmeNqdT0YqZYbMZ2P6Jgn3ZUIgh3ibPI81bjA6xUCKJmzy1LD.4AZnk4g4G_iMO4tdiebiVDoPPtdVDIkDWw0OnDHek.d_2r.PfBtuIoxDvrbGDL.Lv2AuD6lxiObz_lldDHq6HnEw_irAP1hCH.Dr3KdW33DN2w0X1R75N3f8GXdHinmxXLtYbZNYZEE9K7lk9AGmBWgcTds.XgGVW3gDS5OEwoRat44Ecke8k7ZXoY_2revEbUrD8UpOrGprlPEwVYuAvLoTSZX8WJEWQ_QT2CDjNw0FOwAECzsFJa4hGgUtjCPzG&c1SoYK0a=GBK-4aeKAo74EouxLY.stFwdwvXQQG_hXMGG8gB0Hhe6V2Il9k9c8yiTLqduIXpv2RNt.H.weYXeF5XhV0CR2lATieRmk.cs8.fPhNpfGx7JkG1uacp75kDcmXsNtuKgbzRUHZh8vkj4UEYbPcwIYIOw5gFG_cMi9n1GYq0AXXK9UQn9IsmjCBuI7AOFw.pk91OgjvkJCcg2y0y3yDkGwZPcg5EktfAXi.PjmfaecWg8hodU87q6B3ZuPxhel9K9I3EDBxzCHtZqt_0YFlkJCcK4hLq

the problem is with obfuscation and also the nested variables and functions that can keep you out of track for hours if you try to debug it line by line (which I did) and the code makes the characters after content.jsp? part one by one and that explains why its about 13000 lines!!

this part request.send($_9); should have a body for request because its a POST request and $_9 was always null! it seems there are more protection levels to it as it seems.

  • callBackC:

well the callbackC is apparently a simple function to get responseText and show it to user:

function callbackC() {
    if (request.readyState == 1) {
        _$c2(document.getElementById(_$Fe("Y3CeXwDYXwq")), '=', _$Fe('vFKyXRUxEYlTW'), _$Fe("kHDxnHOaB3vE5HDxnHOSNMKQGQ6xOHK2z3Kw2Qne7MCm9FKyvhbwNROg"));
    }
    if (request.readyState == 4) {
        if (request.status == 200) {
            oldContent[oldContent.length] = request.responseText;
            _$c2(document.getElementById(_$Fe("b3CeXwDYXwq")), '=', _$Fe('EFKyXRUxEYlTW'), request.responseText);
            request = null;
        } else {
            _$c2(document.getElementById(_$Fe("u3CeXwDYXwq")), '=', _$Fe('BFKyXRUxEYlTW'), "<br><br><br><span style=font-size:x-large;color:#215add>服务器未返回数据</span>");
        }
    }
}

I didn't quiet get what those _$Xx functions do (because it goes so deep that its out of my patient!) but it seems they simply replaced the document.getElementById("someThing").innerText="Contents"; with multi layered functions so we can't understand the code easily, and the request.responseText is what you need which is HTML code for the table of results.

there is also a 3rd way which I don't know if you can implement it in your code, but since these functions are in a public scope you can simply override them by redefining these two functions (or replace the functions in the link with your own functions and run them). I tried to get the URL for the request which gave me the link I used in middle of this post, but it didn't worked (I just override the callBackC function and get request.responseURL) and the link gave me 404 error.

I don't think I said all I got from my observations but I think it's enough for you to know what you are up against if you are not already aware, and I hope I was helpful.

Reference:

XMLHttpRequest: Living Standard — Last Updated 16 August 2021

Related