How to select elements containing special characters in XPath?

Viewed 517

I am trying to exclude three <td> elements from a result set:

<td>
    &#x1F947;
</td>
<td>
    &#x1F948;
</td>
<td>
    &#x1F949;
</td>

I've tried using:

td[not(contains(., '&#x1F948;'))]

For example, but the element I don't want still comes back...

2 Answers

In the xpath expression, you need to use the escape conventions of the host language. Using &-escaping is fine if the host is XSLT, but if it’s JavaScript, for example, you’ll need to use backslash escaping.

To avoid the labyrinth of escaping conventions, just use literal Unicode characters themselves, which can be searched and then copy-and-pasted from sites such as Compart:

Char Entity Ref Literal Unicode XPath
&#x1F947; //td[not(contains(.,''))]
&#x1F948; //td[not(contains(.,''))]
&#x1F949; //td[not(contains(.,''))]

Here's a single XPath 2.0+ expression that will select all td elements in the document except those consisting of only the targeted special characters:

//td[not(normalize-space() = ('', '',''))]

In XPath 1.0, you'd have to write out the clauses separately:

//td[not(normalize-space() = '') and 
     not(normalize-space() = '') and 
     not(normalize-space() = '')]

Rearrange via DeMorgan's per taste. Go back to contains() if you truly want to test via substring containment rather than string value equality.

Related