I'm developing some code to scrape text from websites. I'm not interested to scrape the entire page, but just in sections of the page that contain certain words. Ideally, I want to scrape the entire paragraph that contains the word. I've seen examples that use the .find_all("p") line, however I found that many websites do not use HTML-defined paragraphs ("p"). Therefore I would like to refrain from that.
Right now, I'm using an approach were the text before and after a certain word is searched. However, the problem here is that the same sentences can be mentioned multiple times over. For example in the code below, the sentence "Drought is pushing food prices up sharply in East Africa" is mentioned 3 times. Here is the code:
from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
import re
url = "https://www.un.org/africarenewal/news/drought-pushing-food-prices-sharply-east-africa"
req = Request(url, headers={"User-Agent": 'Mozilla/5.0'})
page = urlopen(req, timeout = 5) # Open page within 5 seconds. This line skips 'empty' websites
htmlParse = BeautifulSoup(page.read(), 'lxml') #html5lib
SearchWords = ["drought", "water", "food"] # text must contain these words
textP = ""
text = ""
for word in SearchWords:
print(word)
for r in re.findall(re.compile('.{0,100}'+word+'.{0,100}'), htmlParse.text):
textP = textP + r
text= text + textP
print(text)
As mentioned, I would ideally get all the paragraphs that contain a certain word, without duplicates. Has anyone any experience with this? Much much appreciated!