I am looking for framework in Python to extract key information from thousands of different websites - such as "office location", "CEO", etc. Ideally, the script would read in a website url, identify some "key terms" such as "location", "offices", "team members", etc and print the corresponding metrics.
My only related experience is in using Scrapy is in extracting information that follows a pattern on one specific webpage (i.e. extracting tables from Wikipedia), but not sure if Scrapy or BeautifulSoup would work for this sort of project. Was wondering if Scrapy would be my best bet, and if so, what would be the correct syntax to use for this type of project. I've already tried some variations of
import scrapy
from bs4 import BeautifulSoup
import urllib
class OurfirstbotSpider(scrapy.Spider):
name = 'ourfirstbot'
start_urls = [
'https://en.wikipedia.org/wiki/List_of_common_misconceptions',
]
def parse(self, response):
#yield response
headings = response.css('.mw-headline').extract()
datas = response.css('ul').extract()
for item in zip(headings, datas):
all_items = {
'headings' : BeautifulSoup(item[0]).text,
'datas' : BeautifulSoup(item[1]).text,
}
yield all_items
with no avail, due to each site having a different layout and none of them following a specific pattern. Any help would be appreciated.