I have an HTML string, with several <em>...</em> tags in it. I need to find all the indexes of these tags relatively to the string, where all the tags are removed.
For example:
from bs4 import BeautifulSoup
string = "<em>This</em> is <em>a sample</em> string"
string_without_tags = BeautifulSoup(string, "lxml").text
# [(0, 4), (8, 16)] <=> "This" and "a sample"
print(string_without_tags[:4], ", ", string_without_tags[8:16], sep="")
I think I could just use a loop, but maybe there is more efficient way to do what I need?