I want to iterate on every "p" tags in a XML document and be able to get the current element's xpath but I don't find anything that does it.
The kind of code I tried:
from bs4 import BeautifulSoup
xml_file = open("./data.xml", "rb")
soup = BeautifulSoup(xml_file, "lxml")
for i in soup.find_all("p"):
print(i.xpath) # xpath doesn't work here (None)
print("\n")
Here is a sample XML file that I try to parse:
<?xml version="1.0" encoding="UTF-8"?>
<article>
<title>Sample document</title>
<body>
<p>This is a <b>sample document.</b></p>
<p>And there is another paragraph.</p>
</body>
</article>
I would like my code to output:
/article/body/p[0]
/article/body/p[1]