I am trying to grab the text on a webpage which is organized into paragraphs with <p> tags including a subclass that also contains <p> tags. When I use find_all, it grabs the tags I want as well as the ones within the subclasses. I want to exclude the subclass tags.
Here is the website's html:
> div id="storytext" class="class_id"
> div id="imageid" class="imagewrap"
> # here is the subclass <p> tags and are being scraped when I use find_all()
# the following that is the only <p> tag I want to include
> <p>...</p>
> <p>...</p>
> <p>...</p>
> <p>...</p>
> <p>...</p>
My code:
story_text = soup.find('div', {'id':'storytext'})
paragraphs = story_text.find_all('p')
for p in paragraphs:
story.append(p.text)
This pulls everything under the div, {'class':'imagewrap'}