Given a paragraph on a Wikipedia page (i.e. Lana Del Ray) using BeautifulSoup I need to extract the text from the paragraph and mark the word or set of words that have enclosed a link.
For example, consider the first sentences of the first paragraph:
Elizabeth Woolridge Grant (born June 21, 1985), known professionally as Lana Del Rey, is an American singer and songwriter. Her music is noted for its cinematic quality and exploration of tragic romance, glamour, and melancholia, containing references to contemporary pop culture and 1950s–1960s Americana.
I need to have it in this format:
Elizabeth Woolridge Grant (born June 21, 1985), known professionally as Lana Del Rey, is an American singer and songwriter. Her music is noted for its cinematic quality and exploration of tragic romance, START_A glamour END_A, and START_A melancholia END_A, containing references to contemporary START_A pop culture END_A and START_A 1950s–1960s Americana END_A.
Until now I could extract either one paragraph or link separately using something like this:
from urllib import request
from bs4 import BeautifulSoup
soup = BeautifulSoup(request.urlopen("https://en.wikipedia.org/wiki/Lana_Del_Rey").read())
for tag in soup.select('p a[href]'):
if tag['href'].startswith('/wiki/'):
text = tag.text.strip()
print(text)
for tag in soup.select('p'):
text = tag.text.strip()
print(text)
Is it possible to extract the text of the paragraph
tag and identify which words have a link attached to them?