XPath select following until some condition?

Viewed 88

I'm having a trouble selecting product from following node. Here's the html:

 <div>
      <p>Order ID 1</p>
      <p style="display:none"></p>
      <p>product 1</p>

      <p>Order ID 2</p>
      <p style="display:none"></p>
      <p>product 1</p>
      <p>product 2</p>
      
      <p>Order ID 3</p>
      <p style="display:none"></p>
      <p>product 1</p>
      <p>product 2</p>
    
      <p>Order ID 4</p>
      <p style="display:none"></p>
      <p>product 1</p>
      <p>product 2</p>
      <p>product 3</p>
     
      <p>Order ID 5</p>
      <p style="display:none"></p>
      <p>product 1</p>  
      
    </div>

I selected Order ID with following code:

//div/p[@style="display:none"]/preceding-sibling::p[1]

Is there any way to select product? code I tried :

//div/p[@style="display:none"]/following::p[not(@style="display:none" )]

result :

<p>product 1</p>
<p>Order ID 2</p>
<p>product 1</p>
<p>product 2</p>
<p>Order ID 3</p>
<p>product 1</p>
<p>product 2</p>
<p>Order ID 4</p>
<p>product 1</p>
<p>product 2</p>
<p>product 3</p>
<p>Order ID 5</p>
<p>product 1</p>

How to deselect order ID

3 Answers

You can try using the text() content, as follow:

//div/p[contains(text(), 'product')]/text()

or

//div/p[not(contains(text(), 'Order'))]/text()

Using Scrapy in python the output using extract() function is:

['product 1', 'product 1', 'product 2', 'product 1', 'product 2', 'product 1', 'product 2', 'product 3', 'product 1']

I. Use:

/div/p[@style='display:none']
      /following-sibling::p[not(@style)]
                             [not(following-sibling::p[1][@style='display:none'])]

II. Explanation

In simple words this XPath expression instructs the XPath engine to do the following:

Get all following siblings of all p elements that are children of the top element div and have a style attribute with value the string "display:none", such that (these following siblings) don't have a style attribute themselves, and are not an immediate preceding sibling of a p element that has a style attribute with value the string "display:none"

III. XSLT - based verification:

<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
 <xsl:output omit-xml-declaration="yes" indent="yes"/>

  <xsl:template match="/">
    <xsl:copy-of select=
    "/div/p[@style='display:none']
            /following-sibling::p[not(@style)]
                                   [not(following-sibling::p[1][@style='display:none'])]
    "/>
  </xsl:template>
</xsl:stylesheet>

When this transformation is applied on the provided XML document:

<div>
    <p>Order ID 1</p>
    <p style="display:none"></p>
    <p>product 1</p>
    <p>Order ID 2</p>
    <p style="display:none"></p>
    <p>product 1</p>
    <p>product 2</p>
    <p>Order ID 3</p>
    <p style="display:none"></p>
    <p>product 1</p>
    <p>product 2</p>
    <p>Order ID 4</p>
    <p style="display:none"></p>
    <p>product 1</p>
    <p>product 2</p>
    <p>product 3</p>
    <p>Order ID 5</p>
    <p style="display:none"></p>
    <p>product 1</p>
</div>

The XPath expression is evaluated and its (wanted, correct) result is copied to the output:

<p>product 1</p>
<p>product 1</p>
<p>product 2</p>
<p>product 1</p>
<p>product 2</p>
<p>product 1</p>
<p>product 2</p>
<p>product 3</p>
<p>product 1</p>

Here is a screenshot of evaluating this XPath expression with the XPath Visualizer:

enter image description here

You can check for p tags that their following sibling doesn't have that style (this won't apply to Order ID i).

scrapy shell

In [1]: from scrapy import Selector

In [2]: html=""" <div>
   ...:       <p>Order ID 1</p>
   ...:       <p style="display:none"></p>
   ...:       <p>product 1</p>
   ...: 
   ...:       <p>Order ID 2</p>
   ...:       <p style="display:none"></p>
   ...:       <p>product 1</p>
   ...:       <p>product 2</p>
   ...:       
   ...:       <p>Order ID 3</p>
   ...:       <p style="display:none"></p>
   ...:       <p>product 1</p>
   ...:       <p>product 2</p>
   ...:     
   ...:       <p>Order ID 4</p>
   ...:       <p style="display:none"></p>
   ...:       <p>product 1</p>
   ...:       <p>product 2</p>
   ...:       <p>product 3</p>
   ...:      
   ...:       <p>Order ID 5</p>
   ...:       <p style="display:none"></p>
   ...:       <p>product 1</p>  
   ...:       
   ...:     </div>"""

In [3]: sel = Selector(text=html)

In [4]: sel.xpath('//div/p[@style="display:none"]/following::p[not(following::p[1][@style="display:none"])]/text()').ge
   ...: tall()
Out[4]:
['product 1',
 'product 1',
 'product 2',
 'product 1',
 'product 2',
 'product 1',
 'product 2',
 'product 3',
 'product 1']
Related