Can't get <a> tag from LXML

Viewed 32

I am scraping urban dictionary with Python for the top definition, example, author, and like/dislike of a word/phrase. I am using lxml to access the site and pull the xml data. I proceed to call request for the XPath data and store it in tree. However when it returns, it doesn't return what is expected.

import os
from lxml import html
import requests

page = requests.get("https://www.urbandictionary.com/define.php?term=Food")
tree = html.fromstring(page.content)

# relative XPath to the definition
example = tree.xpath('//*[@id="content"]/div[1]/div[3]')

print(example)

out >> [' that has ever ', ' to ', '.']

It skips over some words, namely the words that have an tag. I'm looking for it to return: The best thing that has ever happened to earth., or maybe ['The best thing ', 'that has ever ', 'happened ', 'to ', 'earth', '.']

I dont really care if it is in an array/list form, or a string form, all I want is that lxml includes the words under a tag in the return, however it would do that. How would I go about getting the content as well?

Thanks in advance

1 Answers

Try it this way:

example = tree.xpath('//*[@id="content"]//div[@class="meaning"]')
print(example[0].text_content())

Output:

The best thing that has ever happened to earth.

To get all definitions, change it to:

example = tree.xpath('//div[@class="meaning"]')
for ex in example:
    print(ex.text_content())

Output:

The best thing that has ever happened to earth.
when you've done something so cringe you can't stop replaying it in your head and it stops you from getting on with your every day life
the solution to all of women's problems.
The best thing ever
a substance you eat,then poop out.usually followed my a nap.
A basic human right that is restricted to stores and restaurants. If you try to steal food for your starving family you'll be locked away.
Food: as in what models dont eat
Related