List all media and document files loaded with webpage with requests python

Viewed 1062

I'm looking for a way to list all loaded files with the requests module. Like there is in chrome's Inspector Network tab, you can see all kinds of files that have been loaded by the webpage.

enter image description here

The problem is the file(in this case .pdf file) I want to fetch does not have a specific tab, and the webpage loads it by javascript and AJAX I guess, because even after the page loaded completely, I couldn't find a tag that has a link to the .pdf file or something like that, so every time I should goto Networks tab and reload the page and find the file in the loaded resources list. Is there any way to catch all the loaded files and list them using the Requests module?

1 Answers

When a browser loads an HTML file it then interprets the contents of that file. It may discover that there is a tag referencing an external JavaScript URL. The browser will then issue a GET request to retrieve that file. When said file is received, it hen interprets the JavaScript file by executing the code within. That code might contain AJAX code that in turn fetches more files. Or the HTML file may reference an extern CSS file with a tag or image file with an tag. These files will also be loaded by the browser and can be seen when you run the browser's inspector.

In contrast, when you do a get request with the requests module for a particular URL, only that one page is fetched. There is no logic to interpret the contents of the returned page and fetch those images, style sheets, JavaScript files, etc. that are referenced within the page.

You can, however, use Python to automate a browser using a tool such as Selenium WebDriver, which can be used to fully download a page.

Related