I am trying to scrape a page using python and beautiful soup bs4
I want to keep the text in the <p> element in the page along with the emojis in this text.
The first attempt was:
import urllib
import urllib.request
from bs4 import BeautifulSoup
urlobject = urllib.request.urlopen("https://example.com")
soup = BeautifulSoup(urlobject, "lxml")
result= list(map(lambda e: e.getText(), soup.find_all("p", {"class": "text"})))
But this doesn't include emojis. I then tried to remove .getText() and just keep :
result= list(map(lambda e: e, soup.find_all("p", {"class": "text"})))
Which made me realize the emojis in this website are in the alt of img tags:
<p class="text">I love the night<img alt="" class="emoji" src="etc"/><span>!</span></p>
So what I want to do is :
- getText() for
pwith classtext - But also get
altforimgwithclass=emoji
And keep the text and the emojis as one sentence.
Is there any way to do this?
Any help would be appreciated.