Using XPath to retrieve elements inside a <script> tag

Viewed 3429

I am trying to use XPath to get elements on a page which reside within a <script> tag. For example:

<div id="foo">
    <script>
        <p>You can't get me.</p>
    </script>
</div>

If I try response.xpath('//div[@id="foo"]//p') or response.xpath('//div[@id="foo"]/script/p'), both return an empty array.

How can I get the elements within the <script> tag using XPath?

1 Answers

eLRuLL provided an even more elegant and better answer to my question. His solutions is as follows:

from scrapy import Selector

#First, retrieve the content within the <script> tag:
text = response.xpath('//script/text()').extract_first()
#Then, create a Selector
sel = Selector(text=text)
#Now we can use XPath normally as if the text was a common HTML response
sel.xpath(//p/text()).extract_first()

Old answer: The <script> node has only text type children. That's why XPath don't get deeper down a <script> tag. But, I found a way around it.

#First, retrieve the content within the <script> tag:
text = response.xpath('//script/text()').extract_first()
#Then, encode it
text_encoded = text.encode('utf-8')
#Now, convert it to a HtmlResponse object
text_in_html = HtmlResponse(url='some url', body=text_encoded, encoding='utf-8')
#Now we can use XPath normally as if the text was a common HTML response
text_in_html.xpath(//p/text()).extract_first()
Related