Take the following HTML snippet, read in with rvest:
page <- read_html('<div class="author">A name</div>
<div id="tag">A tag</div>
<div>A number</div>
<div>A date</div>
<p class="articleParagraph dearticleParagraph">A text</p>
')
I can select all the <div> nodes like this:
page %>% html_nodes("div")
How do I select only those <div> nodes that do not have any class or id, i.e. not the first two?
Building on the first comment, I have made some progress:
page %>% html_nodes("div:not([class])")
excludes the first div, as desired.
page %>% html_nodes("div:not([id])")
excludes the second div, as desired.
However, I cannot combine the two:
page %>% html_nodes("div:not([class])") %>% html_nodes("div:not([id])")
returns an empty node set.